Video super-resolution processing method, video super-resolution processing device and storage medium

By using a deep learning model to perform regional super-resolution processing and deblurring optimization on video frames, the problems of high computational complexity and insufficient real-time performance in low-resolution video processing are solved, achieving a balance between efficient video super-resolution and real-time performance.

CN112950465BActive Publication Date: 2025-09-19BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110104274.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-26
Publication Date
2025-09-19
Estimated Expiration
2041-01-26

AI Technical Summary

Technical Problem

In the fields of medical imaging, video surveillance, and remote sensing imaging, existing technologies have low video resolution due to factors such as imaging equipment limitations, atmospheric disturbances, and scene motion changes, making it difficult to meet the needs of high-resolution video processing and analysis.

Method used

A deep learning model is used to perform super-resolution on part of the video frame to be processed, including determining the first processing area and performing super-resolution processing, combining deblurring and image smoothing methods to optimize the computational complexity to meet real-time requirements.

Benefits of technology

While ensuring the video super-resolution processing effect, the computational complexity of high-resolution video is reduced, achieving a balance between video super-resolution effect and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112950465B_ABST
    Figure CN112950465B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video super-resolution processing method, a video super-resolution processing device, and a storage medium. The video super-resolution processing method includes: obtaining a video to be processed, and determining a frame to be processed in the video to be processed; if the resolution of the frame to be processed is greater than a resolution threshold, determining a first processing area in the frame to be processed, and super-resolution processing is performed on the first processing area based on a deep learning model, wherein the first processing area is a partial area in the frame to be processed. Through the embodiments of the present disclosure, while ensuring the video super-resolution processing effect of the video to be processed, the real-time requirements of video super-resolution are met, and a balance is achieved between the video super-resolution effect and the real-time performance of video super-resolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a video super-resolution processing method, a video super-resolution processing device, and a storage medium. Background Art

[0002] With the development of technology, high-resolution video provides users with a clearer and more comfortable visual experience in fields such as medical imaging, video surveillance, and remote sensing imaging, and related technical research has attracted widespread attention. However, in real life, due to factors such as imaging equipment limitations, atmospheric disturbances, and scene motion changes, the actual video obtained is often of low resolution, which makes subsequent video processing and analysis difficult and fails to meet people's needs.

[0003] The resulting super-resolution reconstruction technology, whose basic task is to reconstruct the corresponding high-resolution image or video from the original low-resolution image or video, has become a research hotspot in the field of image processing. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides a video super-resolution processing method, a video super-resolution processing device and a storage medium.

[0005] According to one aspect of an embodiment of the present disclosure, a video super-resolution processing method is provided, comprising: obtaining a video to be processed, and determining a frame to be processed in the video to be processed; if the resolution of the frame to be processed is greater than a resolution threshold, determining a first processing area in the frame to be processed, and performing super-resolution processing on the first processing area based on a deep learning model, wherein the first processing area is a partial area in the frame to be processed.

[0006] In one embodiment, the video super-resolution processing method further includes: if the resolution of the frame to be processed is less than the resolution threshold, super-resolution processing is performed on the entire area of ​​the frame to be processed based on the deep learning model.

[0007] In one embodiment, determining the first processing area in the frame to be processed includes: determining an area within a set range centered on the geometric center of the frame to be processed as the first processing area; and / or determining an area in the frame to be processed where pixels at the same position in adjacent frames to be processed and whose pixel value difference is greater than a set pixel difference threshold are located as the first processing area.

[0008] In one embodiment, before performing super-resolution processing, the video super-resolution processing method further includes: performing deblurring processing on the frame to be processed.

[0009] In one embodiment, performing deblurring processing on the frame to be processed includes: performing deblurring processing on the frame to be processed based on the deep learning model.

[0010] In one embodiment, the video super-resolution processing method further includes: performing super-resolution processing on a second processing area of ​​the frame to be processed based on an image smoothing processing method, where the second processing area is other areas in the frame to be processed except the first processing area.

[0011] According to another aspect of the embodiments of the present disclosure, a video super-resolution processing device is provided, which includes: an acquisition module for acquiring a video to be processed; a determination module for determining a frame to be processed in the video to be processed, and if the resolution of the frame to be processed is greater than a resolution threshold, determining a first processing area in the frame to be processed, wherein the first processing area is a partial area in the video frame; and a super-resolution processing module for performing super-resolution processing on the first processing area based on a deep learning model.

[0012] In one embodiment, the processing module is further configured to: when the resolution of the frame to be processed is less than the resolution threshold, perform super-resolution processing on the frame to be processed based on the deep learning model.

[0013] In one embodiment, the first processing area is the motion area and / or central area in the video frame, and the determination module is further used to: determine the area where the geometric center of the video frame is located as the first processing area; and if the difference in pixel values ​​at the same pixel position in adjacent video frames of the video to be processed is greater than a set threshold, determine the area where the pixel position is located as the first processing area.

[0014] In one embodiment, the video super-resolution processing device further includes: a deblurring processing module, configured to perform deblurring processing on the frame to be processed.

[0015] In one embodiment, the deblurring module performs deblurring on the frame to be processed in the following manner: performing deblurring on the frame to be processed based on the deep learning model.

[0016] In one embodiment, the super-resolution processing module is further used to: perform super-resolution processing on a second processing area of ​​the frame to be processed based on an image smoothing processing method, where the second processing area is other areas in the frame to be processed except the first processing area.

[0017] According to another aspect of an embodiment of the present disclosure, a video super-resolution processing device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to: execute any one of the aforementioned video super-resolution processing methods.

[0018] According to another aspect of the embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided. When the instructions in the storage medium are executed by the processor of the mobile terminal, the mobile terminal can execute any one of the aforementioned video super-resolution processing methods.

[0019] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: when the resolution of the frame to be processed is greater than the resolution threshold, the first processing area in the frame to be processed is super-resolved based on the deep learning model, which can ensure the video super-resolution processing effect of the video to be processed while meeting the real-time requirements of video super-resolution, thereby achieving a balance between the video super-resolution effect and the real-time performance of video super-resolution.

[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0022] Figure 1 The figure is a flowchart of a video super-resolution processing method according to an exemplary embodiment of the present disclosure.

[0023] Figure 2 The figure is a flowchart of a video super-resolution processing method according to another exemplary embodiment of the present disclosure.

[0024] Figure 3 The figure is a flowchart of a video super-resolution processing method according to another exemplary embodiment of the present disclosure.

[0025] Figure 4 The figure is a flowchart of a video super-resolution processing method according to another exemplary embodiment of the present disclosure.

[0026] Figure 5 The figure is a flowchart of a video super-resolution processing method according to another exemplary embodiment of the present disclosure.

[0027] Figure 6 The figure is a flowchart of a video super-resolution processing method according to another exemplary embodiment of the present disclosure.

[0028] Figure 7The figure is a flowchart of a video super-resolution processing method according to an exemplary embodiment of the present disclosure.

[0029] Figure 8 The figure shows an application scenario and effect diagram of a video super-resolution processing method according to an exemplary embodiment of the present disclosure.

[0030] Figure 9 The figure is a block diagram of a video super-resolution processing device according to an exemplary embodiment of the present disclosure.

[0031] Figure 10 The figure is a block diagram of a video super-resolution processing device according to another exemplary embodiment of the present disclosure.

[0032] Figure 11 The present invention is a block diagram showing a device for video super-resolution processing according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0033] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0034] With the advent of high-definition display devices and ultra-high-definition video resolution formats, high-resolution videos offer users a clearer and more comfortable visual experience. The demand for reconstructing high-resolution videos from low-resolution video is increasing. Video super-resolution, a technology used to reconstruct high-resolution video from a given low-resolution video, is widely used in fields such as HDTV, satellite imagery, and video surveillance.

[0035] The super-resolution method calculates the unknown pixel values ​​in a high-resolution image using a given low-resolution image input. This method requires only a small amount of calculation and is very fast. However, the reconstruction effect is poor, especially for images with more high-frequency information. Single image super-resolution refers to the use of a low-resolution image to reconstruct its corresponding high-resolution image. In contrast, video super-resolution uses multiple related low-resolution video frames to reconstruct their corresponding high-resolution video frames. When the format of the video to be processed is large, for example, the number of frames per second is 24, 30 or more, the video super-resolution algorithm has a large amount of computation and complex steps. The played video will have obvious pauses and slowness, which brings a very bad experience to the user.

[0036] Based on this, the present disclosure provides a video super-resolution processing method. When the resolution of the frame to be processed is greater than the resolution threshold, the video frame is processed in regions, and video super-resolution processing is performed based on the depth model only in the determined area. While ensuring the video super-resolution processing effect of the video to be processed, the real-time requirements of video super-resolution are met.

[0037] Figure 1 This is a flow chart of a video super-resolution processing method according to an exemplary embodiment of the present disclosure. Figure 1 As shown, the video super-resolution processing method includes the following steps.

[0038] In step S101, a video to be processed is obtained, and frames to be processed in the video to be processed are determined.

[0039] In step S102, if the resolution of the frame to be processed is greater than the resolution threshold, a first processing area in the frame to be processed is determined, and super-resolution processing is performed on the first processing area based on the deep learning model.

[0040] In the embodiment of the present disclosure, the video to be processed can be the original video within the field of view captured by an image acquisition device such as a camera, or a video obtained after preprocessing the original video, or a video downloaded from the network. The frame to be processed is determined in the video to be processed, and the video frame obtained by parsing the video to be processed is determined as the frame to be processed. The frame to be processed can be a plurality of consecutive frames that can be used as a reference. For example, the frame to be processed can be the i-th frame in the video to be processed, and the i-1-th frame and the i+1-th frame adjacent to it. Using multiple related frames to be processed, their corresponding high-resolution frames are reconstructed. In addition to utilizing the spatial correlation within a single image, the temporal correlation between adjacent low-resolution video frames is also utilized. The frame to be processed can also be the i-th frame in the video to be processed, and the i-1-th frame, i-2-th frame, i+1-th frame and i+2-th frame adjacent to it.

[0041] The resolution threshold can be pre-set, and the resolution of the frame to be processed is compared with the pre-set resolution threshold. It can be understood that the resolution threshold can be set to different values ​​according to the different video super-resolution processing capabilities. The specific value of the resolution threshold is not limited in the embodiment of the present disclosure. In the embodiment of the present disclosure, the first processing area is a partial area in the video frame. When the resolution of the frame to be processed is greater than the resolution threshold, the first processing area in the frame to be processed is determined, and the first processing area is super-resolved based on the deep learning model. The time it takes to generate a super-resolved image based on the deep learning super-resolution network model is related to the resolution of the input frame to be processed. For the same super-resolution network model, the time it takes to generate a super-resolved image for a video frame with a resolution of 360P is approximately 1 / 3 to 1 / 4 of the time it takes to generate a super-resolved image for a video frame with a resolution of 720P. Therefore, when the resolution of the frame to be processed is large, that is, greater than the set resolution threshold, only a partial area in the video frame is super-resolved using the deep learning-based video super-resolution model.

[0042] In the embodiment of the present disclosure, it can be a deep learning model pre-trained for video super-resolution processing, and the network structure of the deep model can be a basic structure such as a super-resolution generative adversarial network (SRGAN). When performing video super-resolution processing, in order to achieve a better processing effect, the frame to be processed can be taken as a plurality of adjacent video frames, and the plurality of video frames of the video to be processed are used as the input of the neural network. After the deep learning model for video super-resolution processing extracts features from the plurality of video frames through multiple convolution operations, it expands the feature size of the input plurality of video frames to twice the original size through upsampling operations, and then performs deconvolution through a multi-layer convolutional network to generate a super-resolution image. The network structure of the deep model can also be other structures used for super-resolution generative adversarial networks (SRGAN).

[0043] According to an embodiment of the present disclosure, when the resolution of the frame to be processed is greater than a resolution threshold, the frame to be processed is determined in the acquired video to be processed, and a partial area in the frame to be processed, i.e., a first processing area, is determined, and super-resolution processing is performed on the first processing area based on a deep learning model. This reduces the computational complexity of the super-resolution algorithm for high-resolution video, while ensuring the video super-resolution processing effect of the video to be processed, and meets the real-time requirements of video super-resolution, thus achieving a balance between the video super-resolution effect and the real-time performance of video super-resolution.

[0044] Figure 2 is a flow chart of a video super-resolution processing method according to an exemplary embodiment. Figure 2 As shown, the video super-resolution processing method includes the following steps.

[0045] In step S201, a video to be processed is obtained, and frames to be processed in the video to be processed are determined.

[0046] In step S202, if the resolution of the frame to be processed is greater than the resolution threshold, a first processing area in the frame to be processed is determined, and super-resolution processing is performed on the first processing area based on the deep learning model.

[0047] In step S203 , if the resolution of the frame to be processed is less than the resolution threshold, super-resolution processing is performed on the entire area of ​​the frame to be processed based on the deep learning model.

[0048] In an embodiment of the present disclosure, a video to be processed is obtained, and a frame to be processed in the video to be processed is determined. The first processing area is a portion of the frame to be processed. When the resolution of the frame to be processed is greater than a resolution threshold, the first processing area in the frame to be processed is determined, and the first processing area is super-resolved based on a deep learning model. That is, when the resolution of the frame to be processed is greater than the resolution threshold, only the first processing area in the frame to be processed is super-resolved using the deep learning-based video super-resolution model.

[0049] The time it takes for a deep learning-based super-resolution network model to generate a super-resolution image is related to the resolution of the input frame to be processed. For the same super-resolution network model, the time it takes to generate a super-resolution image for a 360P frame to be processed is approximately one-third to one-quarter of the time it takes for a 720P frame to be processed. Therefore, when the resolution of the frame to be processed is small, that is, when the resolution of the frame to be processed is less than the set resolution threshold, super-resolutioning the entire area of ​​the video frame using a deep learning-based video super-resolution model does not incur excessive computational overhead or affect the video super-resolution processing speed.

[0050] According to an embodiment of the present disclosure, when the resolution of the frame to be processed is greater than the resolution threshold, the frame to be processed is determined in the acquired video to be processed, and a partial area in the frame to be processed, i.e., the first processing area, is determined, and super-resolution processing is performed on the first processing area based on a deep learning model. When the resolution of the frame to be processed is less than the set resolution threshold, video super-resolution processing is performed on all areas in the frame to be processed using a video super-resolution model based on deep learning. This reduces the computational complexity of the super-resolution algorithm under high-resolution video, and while ensuring the video super-resolution processing effect of the video to be processed, the real-time requirements of video super-resolution are met.

[0051] Figure 3 is a flow chart of a video super-resolution processing method according to an exemplary embodiment. Figure 3As shown, the video super-resolution processing method includes the following steps.

[0052] In step S301, a video to be processed is obtained, and frames to be processed in the video to be processed are determined.

[0053] In step S302, if the resolution of the frame to be processed is greater than the resolution threshold, the area within the set range centered on the geometric center of the frame to be processed is determined as the first processing area; and / or, the area within the frame to be processed where the pixel points at the same position in adjacent frames to be processed and whose pixel value difference is greater than the set pixel difference threshold are located is determined as the first processing area.

[0054] In step S303, super-resolution processing is performed on the first processing area based on the deep learning model.

[0055] In the embodiment of the present disclosure, the first processing area is the motion area in the frame to be processed and / or the central area of ​​the frame to be processed. The central area, as distinguished from the background area in the frame to be processed, can be an area within a set range with the geometric center of the frame to be processed as the center. For example, the area within the set range with the geometric center of the frame to be processed as the center can be a regular graphic area such as a circle, a square, a rectangle, or other areas within a set range. The area where the pixels at the same position in adjacent frames of the video to be processed and whose pixel value difference is greater than the set pixel difference threshold are located is the motion area, that is, the area where the moving object in the video is located. In video super-resolution processing, the central area and / or the motion area include a large amount of information and are the focus of attention. For the same frame to be processed, the central area and the motion area may overlap, partially overlap, or not overlap.

[0056] Obtain a video to be processed and determine a frame to be processed in the video to be processed. The first processing area is a portion of the frame to be processed. When the resolution of the frame to be processed is greater than a resolution threshold, the central area and / or motion area in the frame to be processed is determined, and the central area and / or motion area are super-resolved based on a deep learning model. That is, when the resolution of the frame to be processed is greater than the resolution threshold, only the key areas of interest in the frame to be processed are super-resolved using a deep learning-based video super-resolution model.

[0057] According to an embodiment of the present disclosure, when the resolution of a frame to be processed is greater than a resolution threshold, the frame to be processed is determined in the acquired video to be processed, and the central region and / or motion region in the frame to be processed, i.e., the first processing region, is determined. The first processing region is then super-resolved based on a deep learning model. This reduces the computational complexity of the super-resolution algorithm for high-resolution video, while ensuring the effectiveness of super-resolution processing of the video to be processed and meeting the real-time requirements of video super-resolution.

[0058] Figure 4 is a flow chart of a video super-resolution processing method according to an exemplary embodiment. Figure 4 As shown, the video super-resolution processing method includes the following steps.

[0059] In step S401, a video to be processed is obtained, and a frame to be processed in the video to be processed is determined.

[0060] In step S402 , if the resolution of the frame to be processed is greater than the resolution threshold, a first processing area in the frame to be processed is determined.

[0061] In step S403, the frame to be processed is deblurred, and the first processing area is super-resolved based on the deep learning model.

[0062] In an embodiment of the present disclosure, a video to be processed is obtained, and a frame to be processed in the video to be processed is determined. When the resolution of the frame to be processed is greater than a resolution threshold, a first processing area in the frame to be processed is determined, where the first processing area is a portion of the frame to be processed, and the first processing area in the frame to be processed is subjected to video super-resolution processing based on a deep learning model. That is, when the resolution of the frame to be processed is greater than the resolution threshold, only the first area in the frame to be processed is subjected to video super-resolution processing using a deep learning-based video super-resolution model.

[0063] Before the frame to be processed is super-resolved, the frame to be processed is deblurred. There may be motion blur caused by moving objects or scene switching problems in the video. The frame to be processed is deblurred to remove the blur caused by motion, improve the quality of the video frame to be processed as the model input, effectively remove interference information, and improve the quality of video super-resolution processing. It can be understood that the processing method for deblurring the frame to be processed in the embodiment of the present disclosure can be a method based on deep learning or traditional image optimization used in the current technology, and the embodiment of the present disclosure does not limit the processing method of deblurring.

[0064] According to an embodiment of the present disclosure, when the resolution of the frame to be processed is greater than the resolution threshold, the frame to be processed is determined in the acquired video to be processed, and a partial area in the frame to be processed, i.e., the first processing area, is determined, and the first processing area is super-resolved based on a deep learning model. Before the frame to be processed is super-resolved, the frame to be processed is deblurred. The quality of the frame to be processed is improved, and the computational complexity of the super-resolution algorithm under high-resolution video is reduced. While ensuring the effect of super-resolution processing of the video to be processed, the effect of super-resolution processing of the video is further improved, and the information contained in the super-resolved video is more effectively obtained, while meeting the real-time requirements of video super-resolution.

[0065] Figure 5 is a flow chart of a video super-resolution processing method according to an exemplary embodiment. Figure 5 As shown, the video super-resolution processing method includes the following steps.

[0066] In step S501, a video to be processed is obtained, and frames to be processed in the video to be processed are determined.

[0067] In step S502 , if the resolution of the frame to be processed is greater than the resolution threshold, a first processing area in the frame to be processed is determined.

[0068] In step S503 , the frame to be processed is deblurred based on the deep learning model, and the first processing area is super-resolution processed.

[0069] In an embodiment of the present disclosure, a video to be processed is obtained, and a frame to be processed in the video to be processed is determined. When the resolution of the frame to be processed is greater than a resolution threshold, a first processing region is determined in the frame to be processed, where the first processing region is a portion of the frame to be processed. That is, when the resolution of the frame to be processed is greater than the resolution threshold, only the first region in the frame to be processed is subjected to video super-resolution processing using a deep learning-based video super-resolution model.

[0070] Before super-resolution processing, the frames to be processed are deblurred. Videos may contain motion blur caused by moving objects or scene transitions. Deblurring the frames to be processed removes motion blur, improves the quality of the frames to be processed as model input, effectively removes interference information, and improves the quality of video super-resolution.

[0071] In the disclosed embodiment, to further improve the operating speed of the video overclocking model, the neural network used to deblur the frames to be processed can be combined with the neural network used to super-resolution the frames to be processed. That is, the first half of the deep learning model is used to deblur the frames to be processed, and the second half is used to super-resolution the frames to be processed. When determining the overall loss function of the model, the weighted sum of the deblurring loss function and the super-resolution loss function is taken.

[0072] According to an embodiment of the present disclosure, when the resolution of the frame to be processed is greater than the resolution threshold, the frame to be processed is determined in the acquired video to be processed, and a partial area in the frame to be processed, i.e., the first processing area, is determined, and super-resolution processing is performed on the first processing area based on a deep learning model. Before super-resolution processing is performed on the frame to be processed, deblurring processing is performed on the frame to be processed. The neural network for deblurring the frame to be processed is combined with the neural network for super-resolution processing of the frame to be processed, and the partial output data of the network for super-resolution processing of the frame to be processed in the deep learning model is used for deblurring processing, thereby reducing the amount of model calculation, increasing the model operation speed, and further improving the video super-resolution effect.

[0073] Figure 6 is a flow chart of a video super-resolution processing method according to an exemplary embodiment. Figure 6 As shown, the video super-resolution processing method includes the following steps.

[0074] In step S601, a video to be processed is obtained, and frames to be processed in the video to be processed are determined.

[0075] In step S602, if the resolution of the frame to be processed is greater than the resolution threshold, a first processing area in the frame to be processed is determined, and super-resolution processing is performed on the first processing area based on the deep learning model.

[0076] In step S603, super-resolution processing is performed on the second processing area of ​​the frame to be processed based on the image smoothing processing method, where the second processing area is other areas of the frame to be processed except the first processing area.

[0077] In an embodiment of the present disclosure, a video to be processed is obtained, and a frame to be processed in the video to be processed is determined. The first processing area is a partial area in the frame to be processed. When the resolution of the frame to be processed is greater than the resolution threshold, the first processing area in the frame to be processed is determined, and the first processing area is subjected to super-resolution processing based on a deep learning model. That is, when the resolution of the frame to be processed is greater than the resolution threshold, only the first area in the frame to be processed is subjected to video super-resolution processing using a video super-resolution model based on deep learning. In the frame to be processed, the area considered by the first processing area is the second processing area. The second processing area can be an area with less user attention, such as a background area or a non-motion area in the frame to be processed. The second processing area includes a small amount of information in the video super-resolution processing. In the second processing area, super-resolution processing is performed based on an image smoothing processing method (such as a bicubic interpolation method). Compared with the processing method based on the deep learning model, the processing speed based on the image smoothing processing method is faster, which significantly reduces the computational complexity of super-resolution processing under high-resolution video and does not cause excessive computational overhead.

[0078] According to an embodiment of the present disclosure, when the resolution of the frame to be processed is greater than the resolution threshold, the frame to be processed is determined in the acquired video to be processed, and a partial area in the frame to be processed, i.e., a first processing area, is determined, and super-resolution processing is performed on the first processing area based on a deep learning model. In the second processing area, super-resolution processing of the video to be processed is performed based on an image smoothing processing method. This reduces the computational complexity of the super-resolution algorithm for high-resolution video, while ensuring the effect of super-resolution processing of the video to be processed and meeting the real-time requirements of video super-resolution.

[0079] Figure 7 This is a flow chart of a video super-resolution processing method according to an exemplary embodiment of the present disclosure. Figure 7 As shown, the video super-resolution processing method includes the following steps.

[0080] In step S701, a video to be processed is obtained, and frames to be processed in the video to be processed are determined.

[0081] In step S702 , the frame to be processed is deblurred based on the deep learning model.

[0082] In step S703 , it is determined whether the resolution of the frame to be processed is greater than a resolution threshold.

[0083] When it is determined that the resolution of the frame to be processed is greater than the resolution threshold, step S704 is executed.

[0084] When it is determined that the resolution of the frame to be processed is less than the resolution threshold, step S705 is executed.

[0085] In step S704, a first processing area in the frame to be processed is determined, and super-resolution processing is performed on the first processing area based on the deep learning model.

[0086] In step S705 , super-resolution processing is performed on the entire area of ​​the frame to be processed based on the deep learning model.

[0087] In step S706, super-resolution processing is performed on the second processing area of ​​the frame to be processed based on an image smoothing processing method.

[0088] In step S707 , the processing results of the first processing area and the second processing area are merged to generate a super-resolution video.

[0089] In the disclosed embodiment, a video to be processed is obtained, and a frame to be processed in the video to be processed is determined. Before super-resolution processing is performed on the frame to be processed, the frame to be processed is deblurred. To further improve the operating speed of the video overclocking processing model, the neural network used for deblurring the frame to be processed can be combined with the neural network used for super-resolution processing of the frame to be processed.

[0090] In the embodiment of the present disclosure, the first processing area is a partial area in the frame to be processed. When the resolution of the frame to be processed is greater than the resolution threshold, the first processing area in the frame to be processed is determined, and the first processing area is super-resolved based on the deep learning model. That is, when the resolution of the frame to be processed is greater than the resolution threshold, only the first area in the video frame is super-resolved using the video super-resolution model based on deep learning. In the frame to be processed, the area considered as the first processing area is the second processing area. In the second processing area, super-resolution processing is performed based on image smoothing processing methods, such as bicubic interpolation, and the processing speed is compared with super-resolution processing based on the deep learning model, and the processing effect is not much different. The processing results of the first processing area and the second processing area are merged, and the overlapping parts in the frame to be processed are removed to generate a super-resolved video of the video to be processed.

[0091] When the resolution of the frame to be processed is less than the set resolution threshold, the entire area in the video frame is super-resolved using a deep learning-based video super-resolution model, which does not cause excessive computational overhead and does not affect the video super-resolution processing speed.

[0092] According to an embodiment of the present disclosure, when the resolution of the frame to be processed is greater than a resolution threshold, the frame to be processed is determined in the acquired video to be processed, and a partial area within the frame to be processed, i.e., a first processing area, is determined, and super-resolution processing is performed on the first processing area based on a deep learning model. This reduces the computational complexity of the super-resolution algorithm for high-resolution video, while ensuring the effect of super-resolution processing on the video to be processed, while meeting the real-time requirements of video super-resolution, thus achieving a balance between the effect of video super-resolution and the real-time performance of video super-resolution.

[0093] Figure 8 The figure shows an application scenario and effect diagram of a video super-resolution processing method according to an exemplary embodiment of the present disclosure. Figure 8 The process of super-resolution processing of the frames to be processed based on the deep learning model is shown. The frames to be processed are used as model input and input into the deep learning network model. After multiple convolutions are performed on multiple frames to be processed to extract features, the feature sizes of the input frames to be processed are expanded to twice the original size through upsampling operations. After that, deconvolution is performed through a multi-layer convolutional network. The model outputs super-resolution images of multiple frames to be processed, realizing video super-resolution processing of the video to be processed.

[0094] Based on the same concept, an embodiment of the present disclosure also provides a video super-resolution processing device.

[0095] It is understandable that the video super-resolution processing device provided by the embodiment of the present disclosure includes hardware structures and / or software modules corresponding to the execution of each function in order to realize the above functions. In combination with the units and algorithm steps of each example disclosed in the embodiment of the present disclosure, the embodiment of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiment of the present disclosure.

[0096] Figure 9 FIG. 1 is a block diagram of a video super-resolution processing device according to an exemplary embodiment of the present disclosure. Figure 9 The video super-resolution processing device 100 includes an acquisition module 101, a determination module 102 and a super-resolution processing module 103.

[0097] The acquisition module 101 is used to acquire the video to be processed.

[0098] The determination module 102 is used to determine a frame to be processed in the video to be processed, and if the resolution of the frame to be processed is greater than a resolution threshold, determine a first processing area in the frame to be processed, where the first processing area is a partial area in the frame to be processed.

[0099] The super-resolution processing module 103 is used to perform super-resolution processing on the first processing area based on a deep learning model.

[0100] In one embodiment, the super-resolution processing module 103 is further configured to: when the resolution of the frame to be processed is less than a resolution threshold, perform super-resolution processing on the entire area of ​​the frame to be processed based on the deep learning model.

[0101] In one embodiment, the determination module 102 is also used to: determine an area within a set range centered on the geometric center of the frame to be processed in the frame to be processed as a first processing area; and / or determine an area in the frame to be processed where pixels at the same position in adjacent frames to be processed and whose pixel value difference is greater than a set pixel difference threshold are located as the first processing area.

[0102] Figure 10 FIG. 1 is a block diagram of a video super-resolution processing device according to an exemplary embodiment of the present disclosure. Figure 10 The video super-resolution processing device 100 further includes: a deblurring processing module 104.

[0103] The deblurring processing module 104 is configured to perform deblurring processing on the frame to be processed.

[0104] In one embodiment, the deblurring module 104 performs deblurring on the frame to be processed in the following manner: performing deblurring on the frame to be processed based on a deep learning model.

[0105] In one embodiment, the super-resolution processing module 103 is further configured to perform super-resolution processing on a second processing area of ​​the frame to be processed based on an image smoothing processing method, where the second processing area is an area other than the first processing area in the frame to be processed.

[0106] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0107] Figure 11 FIG2 is a block diagram illustrating an apparatus 200 for video super-resolution processing according to an exemplary embodiment of the present disclosure. For example, apparatus 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0108] Reference Figure 11, apparatus 200 may include one or more of the following components: a processing component 202 , a memory 204 , a power component 206 , a multimedia component 208 , an audio component 210 , an input / output (I / O) interface 212 , a sensor component 214 , and a communication component 216 .

[0109] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 202 may include one or more modules to facilitate interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate interaction between the multimedia component 208 and the processing component 202.

[0110] The memory 204 is configured to store various types of data to support operations on the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0111] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 200.

[0112] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0113] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.

[0114] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0115] The sensor assembly 214 includes one or more sensors for providing various aspects of the status assessment of the device 200. For example, the sensor assembly 214 can detect the open / closed state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor assembly 214 can also detect changes in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and temperature changes of the device 200. The sensor assembly 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0116] The communication component 216 is configured to facilitate wired or wireless communication between the device 200 and other devices. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0117] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0118] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 204 including instructions, which can be executed by the processor 220 of the apparatus 200 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0119] It is understood that in this disclosure, "plurality" refers to two or more than two, and other quantifiers are similar. "And / or" describes the association relationship of related objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects before and after are in an "or" relationship. The singular forms "a", "the" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0120] It will be further understood that the terms "first," "second," and the like are used to describe various types of information, but such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another and do not indicate a particular order or level of importance. In fact, the terms "first," "second," and the like are fully interchangeable. For example, first information could be referred to as second information, and similarly, second information could be referred to as first information without departing from the scope of this disclosure.

[0121] It is further understood that although operations are described in a particular order in the drawings in the embodiments of the present disclosure, this should not be construed as requiring that the operations be performed in the particular order shown or in a serial order, or that all of the operations shown be performed to obtain the desired results. In certain circumstances, multitasking and parallel processing may be advantageous.

[0122] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0123] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video super-resolution processing method, characterized in that: The method comprises: Acquire a video to be processed, and determine relevant frames to be processed in the video to be processed; If the resolution of the frame to be processed is greater than the resolution threshold, determining a first processing area in the frame to be processed, and performing super-resolution processing on the first processing area based on the deep learning model, where the first processing area is a partial area in the frame to be processed; Performing super-resolution processing on a second processing area of ​​the frame to be processed based on an image smoothing processing method, where the second processing area is other areas of the frame to be processed except the first processing area; The determining of the first processing area in the frame to be processed includes: Determining, in the frame to be processed, an area within a set range centered around the geometric center of the frame to be processed as the first processing area; and / or An area where pixels at the same position in adjacent frames to be processed and having a pixel value difference greater than a set pixel difference threshold are located is determined as the first processing area.

2. The video super-resolution processing method according to claim 1, wherein: The method further comprises: If the resolution of the frame to be processed is less than the resolution threshold, super-resolution processing is performed on the entire area of ​​the frame to be processed based on the deep learning model.

3. The video super-resolution processing method according to claim 1 or 2, characterized in that: Before performing the super-resolution processing, the method further includes: Deblurring is performed on the frame to be processed.

4. The video super-resolution processing method according to claim 3, wherein: Deblurring the frame to be processed includes: Deblurring is performed on the frame to be processed based on the deep learning model.

5. A video super-resolution processing device, characterized in that: The device comprises: An acquisition module is used to obtain the video to be processed; a determination module, configured to determine a correlated frame to be processed in the video to be processed, and if the resolution of the frame to be processed is greater than a resolution threshold, determine a first processing area in the frame to be processed, the first processing area being a partial area in the frame to be processed; determine an area within a set range centered on the geometric center of the frame to be processed as the first processing area; and / or determine an area in the frame to be processed where pixels at the same position in adjacent frames to be processed and whose pixel value difference is greater than a set pixel difference threshold are located as the first processing area; The super-resolution processing module is used to perform super-resolution processing on the first processing area based on a deep learning model; and to perform super-resolution processing on the second processing area of ​​the frame to be processed based on an image smoothing processing method, where the second processing area is other areas in the frame to be processed except the first processing area.

6. The video super-resolution processing device according to claim 5, characterized in that: The processing module is further configured to: When the resolution of the frame to be processed is less than the resolution threshold, super-resolution processing is performed on the entire area of ​​the frame to be processed based on the deep learning model.

7. The video super-resolution processing device according to claim 5 or 6, characterized in that: The video super-resolution processing device also includes: The deblurring processing module is used to perform deblurring processing on the frame to be processed.

8. The video super-resolution processing device according to claim 7, characterized in that: The deblurring module performs deblurring on the frame to be processed in the following manner: Deblurring is performed on the frame to be processed based on the deep learning model.

9. A video super-resolution processing device, characterized in that: include: processor; memory for storing processor-executable instructions; The processor is configured to execute the video super-resolution processing method according to any one of claims 1 to 4.

10. A non-transitory computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to execute the video super-resolution processing method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Real-time super-resolution method and system based on FPGA

    CN108765282A

  • Method and system of reconstructing super-resolution image

    US20120051667A1