Video processing method, device, storage medium and electronic device

By dividing video into keyframes and non-keyframes, and using efficient super-resolution models for processing, the problems of low efficiency and poor subjective effects of video in the prior art are solved, and more efficient processing and better video quality are achieved.

CN114332709BActive Publication Date: 2025-05-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111638616.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-05-06
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The existing video super-resolution technology has low computational efficiency and poor subjective video effects, especially when processing non-I-frames, the algorithm effect is not ideal.

Method used

Keyframes and non-keyframes are processed by dividing videos into keyframes and non-keyframes, and using efficient super-resolution models, respectively. Specifically, single-frame image inference is performed on the keyframe using the first super-resolution model, and the super-resolution frame and non-keyframe features of the keyframe are fused based on the second super-resolution model to obtain the super-resolution frame of the non-keyframe.

Benefits of technology

The calculation efficiency of video super-resolution processing is improved, and while maintaining efficient processing speed, the subjective effect of the video is improved, ensuring a good balance between processing speed and subjective effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332709B_ABST
    Figure CN114332709B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video processing method, device, storage medium and electronic device. The method includes: dividing a video into groups including a predetermined number of video frames, and classifying the video frames in each group into key frames and non-key frames; performing super-resolution processing on the key frames and non-key frames of each group respectively through a super-resolution model to obtain super-resolution frames of key frames and super-resolution frames of non-key frames of each group; encoding the super-resolution frames of key frames and super-resolution frames of non-key frames of each group into super-resolution videos, wherein the super-resolution model is a model configured to use the super-resolution frames of key frames as a reference for super-resolution processing of non-key frames. The video super-resolution method according to the present disclosure can make full use of the results after super-resolution of key frames to guide the super-resolution processing of non-key frames, and has good benefits in terms of processing speed and subjective effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video technology, and in particular to a video processing method and device, a corresponding video super-resolution model training method, an electronic device, and a computer-readable storage medium. Background Art

[0002] Super Resolution technology is an image / video processing technology that can process LR (Low Resolution) images or videos into HR (High Resolution) images or videos, thereby improving the resolution and quality of images or videos.

[0003] Video super-resolution technology is mainly divided into the following methods. One method is to decode the video to be processed and extract frames, and then serially call the super-resolution model for each frame for processing. After processing each frame, the resulting frame is encoded to obtain the final output video. This method performs the same amount of computation on all video frames, without considering the characteristics of various types of video frames in the actual video (for example, I frames, P frames, B frames). This method of processing all video frames equally will cause a lot of computational redundancy, thereby reducing computational efficiency.

[0004] Another way is to decode the video into video frames, and then only perform model processing on I frames, and perform lightweight model inference on non-I frames or directly abandon inference, thereby improving the calculation speed. However, this model algorithm does not handle non-I frames well. Although this method has a fast inference speed, the subjective effect of the video after algorithm processing is not satisfactory. Summary of the invention

[0005] The present disclosure provides a video processing method and device, as well as a corresponding video super-resolution model training method, electronic device and computer-readable storage medium, so as to at least solve the problems of low super-resolution calculation efficiency and poor subjective effect of super-resolution video in the related art, and may not solve any of the above problems.

[0006] According to a first aspect of the present disclosure, a video processing method is provided, characterized in that it includes: dividing a video into groups including a predetermined number of video frames, and classifying the video frames in each group into key frames and non-key frames; performing super-resolution processing on the key frames and non-key frames of each group respectively through a super-resolution model to obtain super-resolution frames of key frames and super-resolution frames of non-key frames of each group; encoding the super-resolution frames of key frames and super-resolution frames of non-key frames of each group into super-resolution videos, wherein the super-resolution model is a model configured to use the super-resolution frames of key frames as a reference for super-resolution processing of non-key frames.

[0007] According to a first aspect of the present disclosure, the super-resolution model includes a first super-resolution model and a second super-resolution model, and the super-resolution model is trained in the following manner: performing single-frame image inference on a key frame through the first super-resolution model to obtain a super-resolution frame of the key frame; fusing the features of the super-resolution frame of the key frame and the non-key frame based on the second super-resolution model, and performing inference based on the fused features to obtain a super-resolution frame of the non-key frame; determining a first loss function for the first super-resolution model according to the key frame and the super-resolution frame of the key frame, and adjusting the parameters of the first super-resolution model based on the value of the first loss function; determining a second loss function for the second super-resolution model according to the non-key frame and the super-resolution frame of the non-key frame, and adjusting the parameters of the second super-resolution model based on the second loss function.

[0008] According to a first aspect of the present disclosure, classifying video frames in each group into key frames and non-key frames includes: classifying a predetermined type of video frames in the group as key frames, and classifying the remaining video frames as non-key frames, or classifying the first frame in the group as a key frame, and classifying the remaining video frames as non-key frames.

[0009] According to a first aspect of the present disclosure, each of a first super-resolution model and a second super-resolution model includes an inference main network, which includes multiple residual convolution layers and fast upsampling layers for performing inference operations, wherein the number of channels and the number of residual convolution layers included in the first super-resolution model are higher than the number of channels and the number of residual convolution layers included in the second super-resolution model.

[0010] According to a first aspect of the present disclosure, performing single-frame image inference on a key frame through a first super-resolution model to obtain a super-resolution frame of the key frame includes: calculating the depth features of the key frame layer by layer through the multiple residual convolution layers; and upsampling the depth features of the key frame into a super-resolution frame of the key frame with high resolution through a fast upsampling layer.

[0011] According to the first aspect of the present disclosure, the second super-resolution model also includes a first feature extractor, a second feature extractor and a stitching unit, wherein the features of the super-resolution frame of the key frame and the non-key frame are fused based on the second super-resolution model, and reasoning is performed based on the fused features to obtain the super-resolution frame of the non-key frame, including: extracting the features of the super-resolution frame of the key frame by the first feature extractor and the features of the non-key frame by the second feature extractor; performing feature stitching and convolution on the extracted features of the super-resolution frame of the key frame and the features of the non-key frame by the stitching unit to obtain the fused features; performing calculations on the fused features layer by layer through multiple residual convolution layers of the inference main network to extract deep features of the non-key frame; and upsampling the deep features of the non-key frame into a super-resolution frame of the non-key frame with high resolution through a fast upsampling layer.

[0012] According to the first aspect of the present disclosure, each residual convolution layer of the first super-resolution model and the second super-resolution model includes a plurality of residual convolution blocks connected in series, and each residual convolution block includes a plurality of basic convolution operation units, which are configured to perform convolution operations on input features to extract deeper features.

[0013] According to a first aspect of the present disclosure, in a first super-resolution model and a second super-resolution model, input features and output features of a main inference network are connected across layers; input features and output features of each residual convolution layer are connected across layers; and output features of the first basic convolution operation unit and the last convolution operation unit in each residual block of the residual convolution layer are connected across layers.

[0014] According to a second aspect of the present disclosure, a video processing device is provided, including: a grouping unit, configured to divide a video into groups including a predetermined number of video frames, and classify the video frames in each group into key frames and non-key frames; a video processing unit, configured to perform super-resolution processing on the key frames and non-key frames of each group respectively through a super-resolution model to obtain super-resolution frames of key frames and super-resolution frames of non-key frames of each group; an encoding unit, configured to encode the super-resolution frames of key frames and super-resolution frames of non-key frames of each group into super-resolution videos, wherein the super-resolution model is a model configured to use the super-resolution frames of key frames as a reference for super-resolution processing of non-key frames.

[0015] According to a second aspect of the present disclosure, the super-resolution model includes a first super-resolution model and a second super-resolution model, and the super-resolution model is trained in the following manner: performing single-frame image inference on a key frame through the first super-resolution model to obtain a super-resolution frame of the key frame; fusing the features of the super-resolution frame of the key frame and the non-key frame based on the second super-resolution model, and performing inference based on the fused features to obtain a super-resolution frame of the non-key frame; determining a first loss function for the first super-resolution model according to the key frame and the super-resolution frame of the key frame, and adjusting the parameters of the first super-resolution model based on the value of the first loss function; determining a second loss function for the second super-resolution model according to the non-key frame and the super-resolution frame of the non-key frame, and adjusting the parameters of the second super-resolution model based on the second loss function.

[0016] According to a second aspect of the present disclosure, the grouping unit is configured to: classify a predetermined type of video frames in the group as key frames and classify the remaining video frames as non-key frames, or classify the first frame in the group as a key frame and classify the remaining video frames as non-key frames.

[0017] According to a second aspect of the present disclosure, each of the first super-resolution model and the second super-resolution model includes an inference main network, which includes multiple residual convolution layers and fast upsampling layers for performing inference operations, wherein the number of channels and the number of residual convolution layers included in the first super-resolution model are higher than the number of channels and the number of residual convolution layers included in the second super-resolution model.

[0018] According to the second aspect of the present disclosure, the first super-resolution model is configured to: calculate the depth features of the key frame layer by layer through the multiple residual convolution layers; upsample the depth features of the key frame into a super-resolution frame of the key frame with high resolution through a fast upsampling layer.

[0019] According to a second aspect of the present disclosure, the second super-resolution model also includes a first feature extractor, a second feature extractor and a stitching unit, wherein the second super-resolution model is configured to: extract features of the super-resolution frame of the key frame through the first feature extractor and features of the non-key frame through the second feature extractor; perform feature stitching and convolution on the extracted features of the super-resolution frame of the key frame and the features of the non-key frame through the stitching unit to obtain the fused features; perform calculations on the fused features layer by layer through multiple residual convolution layers of the inference subject network to extract deep features of the non-key frame; and upsample the deep features of the non-key frame to a super-resolution frame of the non-key frame with high resolution through a fast upsampling layer.

[0020] According to the second aspect of the present disclosure, each residual convolution layer of the first super-resolution model and the second super-resolution model includes a plurality of residual convolution blocks connected in series, and each residual convolution block includes a plurality of basic convolution operation units, which are configured to perform convolution operations on input features to extract deeper features.

[0021] According to a second aspect of the present disclosure, in the first super-resolution model and the second super-resolution model, the input features and output features of the main inference network are connected across layers; the input features and output features of each residual convolution layer are connected across layers; the output features of the first basic convolution operation unit and the last convolution operation unit in each residual block of the residual convolution layer are connected across layers.

[0022] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and at least one memory storing computer executable instructions, wherein when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to execute the video processing method as described above.

[0023] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the video processing method as described above.

[0024] According to a fifth aspect of the present disclosure, a computer program product is provided, wherein instructions in the computer program product are executed by at least one processor in an electronic device to perform the video processing method as described above.

[0025] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects: fully utilizing the results after key frame super-resolution to guide the super-resolution processing of non-key frames, and using a lightweight processing model for non-key frames can also ensure good subjective effects, with good benefits in terms of processing speed and subjective effects.

[0026] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0028] Figure 1 is a flowchart illustrating a method for training a video super-resolution model according to an exemplary embodiment of the present disclosure.

[0029] Figure 2 is a schematic diagram illustrating a process of performing key frame inference by a video super-resolution model according to an exemplary embodiment of the present disclosure.

[0030] Figure 3 is a schematic diagram illustrating a process of performing non-keyframe inference by a video super-resolution model according to an exemplary embodiment of the present disclosure.

[0031] Figure 4 is a schematic diagram illustrating a model network structure for performing key frame inference according to an exemplary embodiment of the present disclosure.

[0032] Figure 5 is a schematic diagram illustrating a model network structure for performing non-keyframe reasoning according to an exemplary embodiment of the present disclosure.

[0033] Figure 6 is a block diagram illustrating a training apparatus for a video super-resolution model according to an exemplary embodiment of the present disclosure.

[0034] Figure 7 is a flowchart illustrating a video processing method according to an exemplary embodiment of the present disclosure.

[0035] Figure 8 is a block diagram illustrating a video processing apparatus according to an exemplary embodiment of the present disclosure.

[0036] Fig. 9 is a block diagram illustrating an electronic device for performing a video processing method according to an exemplary embodiment of the present disclosure.

[0037] Fig.10 is a block diagram illustrating an electronic device for performing a video processing method according to another exemplary embodiment. DETAILED DESCRIPTION

[0038] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation methods described in the following examples do not represent all implementation methods consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the attached claims.

[0040] It should be noted that the phrase "at least one of the items" in the present disclosure includes three types of parallel situations: "any one of the items", "a combination of any number of the items", and "all of the items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step 1 and step 2" which means the following three parallel situations: (1) executing step 1; (2) executing step 2; (3) executing step 1 and step 2.

[0041] Figure 1 is a flowchart illustrating a method for training a video super-resolution model according to an exemplary embodiment of the present disclosure.

[0042] like Figure 1 As shown, first, in step S110, the video is divided into groups including a predetermined number of video frames, and the video frames in each group are classified into key frames and non-key frames. According to an exemplary embodiment of the present disclosure, the video may be first decoded into a video frame sequence, and the video frames may be grouped into a plurality of groups in sequence. For example, a plurality of groups may be obtained according to a predetermined number, and each group may have the same number (e.g., 10 frames), that is, the video is divided at intervals of 10 frames.

[0043] According to an exemplary embodiment of the present disclosure, the first frame in the divided group may be classified as a key frame, and the remaining video frames may be classified as non-key frames. Alternatively, a specific type of frame (e.g., I frame) in the divided group may be classified as a key frame, and the remaining video frames may be classified as non-key frames. The key frame here may be, for example, a video frame with a large amount of information and which may generally be used as a reference frame in the encoding and decoding process of other video frames.

[0044] Next, in step S120, a single-frame image inference is performed on the key frame by the first super-resolution model to obtain a super-resolution frame of the key frame. According to an exemplary embodiment of the present disclosure, the first super-resolution model according to an exemplary embodiment of the present disclosure may be implemented by a convolutional neural network. The first super-resolution model may perform layer-by-layer convolution calculations on the input key frame to obtain the features of the key frame and finally obtain a super-resolution frame of the key frame. This process is called a super-resolution inference process of the key frame. Figure 2 Let’s illustrate the super-resolution inference process of key frames.

[0045] like Figure 2As shown in the figure, assuming that the I frame (reference frame i) is used as the key frame in each group, the multiple convolutional layers included in the first super-resolution model (SR model) extract features from the reference frame i layer by layer. In other words, the first convolutional layer extracts the first layer features from the key frame, the second convolutional layer further extracts the second layer features from the first layer features..., and so on, the output result of the last nth convolutional layer can be used as the super-resolution frame of the key frame. The output result is saved and used in the subsequent super-resolution inference process of non-key frames.

[0046] Then, in step S130, the features of the super-resolution frame of the key frame and the non-key frame are fused by a second super-resolution model, and reasoning is performed based on the fused features to obtain the super-resolution frame of the non-key frame. According to an exemplary embodiment of the present disclosure, the second super-resolution model may also be implemented using a structure having multiple convolutional layers similar to the first super-resolution model. Figure 3 As shown, the second super-resolution model (SR model) according to an exemplary embodiment of the present disclosure can use the super-resolution result (key frame SR result) of the key frame obtained in step S120 and the remaining non-key frames (e.g., Figure 3 The i+1th frame, i+2th frame, ..., i+nth frame (excluding the i-th frame) in the video group shown is used to obtain super-resolution frames of non-key frames (i+1th frame result ..., i+nth frame result).

[0047] According to an exemplary embodiment of the present disclosure, each of the first super-resolution model and the second super-resolution model includes an inference main network, which includes a plurality of cascaded residual convolution layers and fast upsampling layers, wherein the number of channels and the number of residual convolution layers included in the first super-resolution model are higher than the number of channels and the number of residual convolution layers included in the second super-resolution model.

[0048] For example, the number of channels of the first super-resolution model can be greater than the number of channels of the second super-resolution model, and more convolutional layers are used, so that the convolutional neural network of the first super-resolution model is wider and deeper, and its computational workload can be, for example, twice that of the second super-resolution model. In this way, the effect of super-resolution processing of key frames can be guaranteed.

[0049] According to an exemplary embodiment of the present disclosure, the first super-resolution model may calculate the depth features of the key frames layer by layer through the multiple residual convolution layers, and upsample the depth features of the key frames into super-resolution frames with high resolution through the fast upsampling layer.

[0050] According to an exemplary embodiment of the present disclosure, the features of the super-resolution frame of the key frame and the non-key frame are fused through a second super-resolution model, and reasoning is performed based on the fused features to obtain the super-resolution frame of the non-key frame, including: extracting the features of the super-resolution frame of the key frame through a first feature extractor and the features of the non-key frame through a second feature extractor; performing feature splicing and convolution on the extracted features of the super-resolution frame of the key frame and the features of the non-key frame through a splicing unit to obtain the fused features; performing calculations on the fused features layer by layer through multiple residual convolution layers of an inference subject network to extract deep features of non-key frames; and upsampling the deep features of non-key frames into super-resolution frames of non-key frames with high resolution through a fast upsampling layer.

[0051] That is, the second super-resolution model may have two input branches to respectively input the super-resolution frame of the key frame and the non-key frame saved previously, and after extracting the respective depth features of the same dimension through the feature extractor, the two features are fused through the splicing unit, thereby helping the main reasoning network of the second super-resolution model to better process the non-key frame. Here, the feature extractor may be a convolutional module with a predetermined number of layers and dimensions.

[0052] As described above, since the second super-resolution model can be a lightweight inference model, it can obtain super-resolution results of non-key frames with less computational effort. At the same time, because it refers to the super-resolution results of key frames to perform inference operations, it can obtain good super-resolution results of non-key frames.

[0053] The following will refer to Figure 4 and Figure 5 An example of a specific structure of a super-resolution model according to an exemplary embodiment of the present disclosure is described.

[0054] Figure 4 and Figure 5 2 are schematic diagrams respectively illustrating network structures of a first super-resolution model and a second super-resolution model according to an exemplary embodiment of the present disclosure.

[0055] like Figure 4 As shown, the first super-resolution model includes a main inference network, which is composed of a plurality of cascaded residual convolution layers and fast upsampling layers. Each residual convolution layer includes a plurality of residual convolution blocks connected in series, and each residual convolution block includes a plurality of basic convolution operation units, which are configured to perform convolution operations on input features to extract deeper features. The fast upsampling layer performs fast upsampling on the output of the last residual convolution layer to obtain the final super-resolution result. Figure 4 In the keyframe LR IAfter being input, a super-resolution frame SR is obtained after passing through multiple residual convolution layers and fast upsampling layers I .

[0056] According to an exemplary embodiment of the present disclosure, cross-layer connections can be set between residual convolution layers, between residual convolution blocks, and between basic convolution operation units to add features of different depths, thereby helping the gradient of the convolutional neural network to be effectively back-propagated and obtaining better network optimization results.

[0057] For example, Figure 4 As shown, the input features and output features of the main inference network of the first super-resolution model are connected across layers (such as Figure 4 The global cross-layer connection shown in FIG. 1 ), the input features and output features of each residual convolutional layer of the main inference network of the first super-resolution model are cross-layer connected (as shown in FIG. Figure 4 The long cross-layer connection shown in the figure), and the output features of the first basic convolution operation unit and the last convolution operation unit in each residual block of the residual convolution layer are cross-layer connected (as shown in the figure). Figure 4 Here, the cross-layer connection of features refers to adding the features of different layers.

[0058] Figure 5 The second super-resolution model shown includes super-resolution frames SR respectively used to extract key frames I Features I and non-keyframe LR P Features p Two feature extractors. Super-resolution frames SR of extracted key frames I Features I and non-keyframe LR P Features p The concatenation unit is concatenated (i.e., the two feature matrices are merged into one feature matrix), and then the fused feature Feat is obtained after convolution. I-p =Conv(concat{Feat I ,Feat p}, where concat represents the feature concatenation operation and conv represents the convolution operation through the convolution layer of the convolutional neural network. Then, the fused features are passed through the main inference network composed of multiple cascaded residual convolution layers and fast upsampling layers to perform super-resolution inference to obtain super-resolution frames of each non-key frame.

[0059] Similar to the first super-resolution model, the input features and output features of the main inference network of the second super-resolution model are connected across layers (e.g. Figure 5The input features and output features of each residual convolution layer of the main inference network of the first super-resolution model are cross-layer connected (as shown in the long cross-layer connection in Figure 5), and the output features of the first basic convolution operation unit and the last convolution operation unit in each residual block of the residual convolution layer are cross-layer connected (as shown in Figure 6). Figure 5 short cross-layer connections shown).

[0060] It should be understood that the network structures, cross-layer connection methods, etc. of the above-mentioned first super-resolution model and the second super-resolution model are only for illustration. Those skilled in the art may adopt other network structures suitable for super-resolution image processing to obtain super-resolution frames of key frames, and obtain super-resolution frames of non-key frames based on the super-resolution frames of key frames and non-key frames.

[0061] After obtaining super-resolution frames of key frames and super-resolution frames of non-key frames, a first loss function for a first super-resolution model can be determined according to the key frames and the super-resolution frames of the key frames in step S140, and parameters of the first super-resolution model can be adjusted based on the value of the first loss function, and a second loss function for a second super-resolution model can be determined according to the super-resolution frames of non-key frames and non-key frames in step S150, and parameters of the second super-resolution model can be adjusted based on the second loss function.

[0062] For example, the high-resolution frame HR of the key frame I And super-resolution frames SR of key frames I The absolute value difference between them determines the first loss function Loss1 of the first super-resolution model = |HR I –SR I |, and based on the high-resolution frame HR of the non-keyframe p And non-keyframe super-resolution frame SR p The absolute value difference between them determines the second loss function Loss1 of the second super-resolution model = |HR p –SR p |, and then adjust the parameters of the main inference network according to the first loss function, and adjust the parameters of the feature extractor, splicing unit and main inference network of the second super-resolution model according to the second loss function, until the first loss function and the second loss function converge to the predetermined target value, thereby completing the model training process.

[0063] As mentioned above, the video super-resolution model trained by the above method utilizes the characteristics of the key frames of the video and uses the super-resolution processing results of the key frames as a reference for the super-resolution processing of non-key frames, thereby reducing the amount of model calculation while ensuring the processing effect of the model.

[0064] Figure 6is a block diagram illustrating a training apparatus for a video super-resolution model according to an exemplary embodiment of the present disclosure.

[0065] like Figure 6 As shown, the training device 600 of the video super-resolution model according to an exemplary embodiment of the present disclosure may include a grouping unit 610, a first reasoning unit 620, a second reasoning unit 630, a first parameter adjustment unit 640 and a second parameter adjustment unit 650.

[0066] The grouping unit 610 is configured to divide the video into groups including a predetermined number of video frames, and classify the video frames in each group into key frames and non-key frames.

[0067] According to an exemplary embodiment of the present disclosure, the grouping unit 610 classifies a predetermined type of video frames in the group as key frames and the remaining video frames as non-key frames, or classifies the first frame in the group as a key frame and the remaining video frames as non-key frames.

[0068] The first reasoning unit 620 is configured to perform single-frame image reasoning on the key frame through the first super-resolution model to obtain a super-resolution frame of the key frame. The second reasoning unit 630 is configured to fuse the super-resolution frame of the key frame and the features of the non-key frame based on the second super-resolution model, and perform reasoning based on the fused features to obtain the super-resolution frame of the non-key frame.

[0069] The first parameter adjustment unit 640 is configured to determine a first loss function for a first super-resolution model according to the key frame and the super-resolution frame of the key frame, and adjust the parameters of the first super-resolution model based on the value of the first loss function. The second parameter adjustment unit 650 is configured to determine a second loss function for a second super-resolution model according to the non-key frame and the super-resolution frame of the non-key frame, and adjust the parameters of the second super-resolution model based on the second loss function.

[0070] According to an exemplary embodiment of the present disclosure, each of the first super-resolution model and the second super-resolution model includes an inference main network, which includes multiple residual convolution layers and fast upsampling layers for performing inference operations, wherein the number of channels and the number of residual convolution layers included in the first super-resolution model are higher than the number of channels and the number of residual convolution layers included in the second super-resolution model.

[0071] According to an exemplary embodiment of the present disclosure, the first reasoning unit 630 is configured to: calculate the depth features of the key frame layer by layer through the multiple residual convolution layers; upsample the depth features of the key frame into a super-resolution frame of the key frame with high resolution through a fast upsampling layer.

[0072] According to an exemplary embodiment of the present disclosure, the second super-resolution model also includes a first feature extractor, a second feature extractor and a stitching unit, wherein the second inference unit 640 is configured to: extract features of the super-resolution frame of the key frame through the first feature extractor and features of the non-key frame through the second feature extractor; perform feature stitching and convolution on the extracted features of the super-resolution frame of the key frame and the features of the non-key frame through the stitching unit to obtain the fused features; perform calculations on the fused features layer by layer through multiple residual convolution layers of the inference body network to extract deep features of non-key frames; and upsample the depth features of the non-key frames to super-resolution frames of non-key frames with high resolution through a fast upsampling layer.

[0073] According to an exemplary embodiment of the present disclosure, each residual convolution layer of the first super-resolution model and the second super-resolution model includes a plurality of residual blocks connected in series, and each residual block includes a plurality of basic convolution operation units.

[0074] According to an exemplary embodiment of the present disclosure, the input features and output features of the first super-resolution model and the second super-resolution model are connected across layers; the input features and output features of each residual convolution layer of the first super-resolution model and the second super-resolution model are connected across layers; the output features of the first basic convolution operation unit and the last convolution operation unit in each residual block of the residual convolution layer are connected across layers. The input features and output features of the first super-resolution model and the second super-resolution model are connected across layers; the input features and output features of each residual convolution layer of the first super-resolution model and the second super-resolution model are connected across layers; the output features of the first basic convolution operation unit and the last convolution operation unit in each residual block of the residual convolution layer are connected across layers.

[0075] Figure 7 is a flowchart illustrating a video processing method according to an exemplary embodiment of the present disclosure.

[0076] First, in step S710, a video is divided into groups including a predetermined number of video frames, and the video frames in each group are classified into key frames and non-key frames.

[0077] As described above, according to an exemplary embodiment of the present disclosure, a frame sequence of a video may be sequentially divided into a plurality of groups according to a predetermined number, and key frames and non-key frames may be determined in each group. According to an exemplary embodiment of the present disclosure, the first frame in each group may be determined as a key frame, and the remaining video frames may be determined as non-key frames. Alternatively, a predetermined type of video frame in each group may be determined as a key frame, and the remaining video frames may be determined as non-key frames.

[0078] Next, in step S720, super-resolution processing is performed on the key frames and non-key frames of each group respectively through a super-resolution model to obtain super-resolution frames of the key frames and super-resolution frames of the non-key frames of each group.

[0079] Then, in step S730, super-resolution frames of key frames and super-resolution frames of non-key frames of each group are encoded into a super-resolution video, wherein the super-resolution model is a model configured to use the super-resolution frames of key frames as a reference for super-resolution processing of non-key frames.

[0080] According to an exemplary embodiment of the present disclosure, the super-resolution model includes a first super-resolution model and a second super-resolution model, and the super-resolution model can be trained in the following manner, that is, performing single-frame image reasoning on a key frame through the first super-resolution model to obtain a super-resolution frame of the key frame, fusing the super-resolution frame of the key frame and the features of the non-key frame based on the second super-resolution model, and performing reasoning based on the fused features to obtain a super-resolution frame of the non-key frame. Here, the first super-resolution model and the second super-resolution model are based on the above reference Figure 1-Figure 5 The method described above is trained and has the same Figure 1-Figure 5 The first super-resolution model and the second super-resolution model have the same structure, so the description of the first super-resolution model and the second super-resolution model will not be repeated here.

[0081] Figure 8 is a block diagram illustrating a video super-resolution apparatus according to an exemplary embodiment of the present disclosure.

[0082] like Figure 8 As shown, a video super-resolution apparatus 800 according to an exemplary embodiment of the present disclosure may include a grouping unit 810 , a video processing unit 820 , and an encoding unit 830 .

[0083] The grouping unit 810 is configured to divide the video into groups including a predetermined number of video frames, and classify the video frames in each group into key frames and non-key frames. As described above, according to an exemplary embodiment of the present disclosure, the frame sequence of the video can be sequentially divided into a plurality of groups according to a predetermined number, and key frames and non-key frames can be determined in each group. According to an exemplary embodiment of the present disclosure, the first frame in each group can be determined as a key frame, and the remaining video frames can be determined as non-key frames. Alternatively, a predetermined type of video frame in each group can be determined as a key frame, and the remaining video frames can be determined as non-key frames.

[0084] The video processing unit 820 is configured to perform super-resolution processing on the key frames and non-key frames of each group respectively through a super-resolution model to obtain super-resolution frames of the key frames and super-resolution frames of the non-key frames of each group, wherein the super-resolution model is a model configured to use the super-resolution frames of the key frames as a reference for super-resolution processing of the non-key frames.

[0085] The encoding unit 830 is configured to encode the super-resolution frames of the key frames and the super-resolution frames of the non-key frames of each group into a super-resolution video.

[0086] According to an exemplary embodiment of the present disclosure, the super-resolution model includes a first super-resolution model and a second super-resolution model, and the super-resolution model is trained in the following manner: performing single-frame image reasoning on a key frame through the first super-resolution model to obtain a super-resolution frame of the key frame, fusing the super-resolution frame of the key frame and the features of the non-key frame based on the second super-resolution model, and performing reasoning based on the fused features to obtain the super-resolution frame of the non-key frame, wherein the first super-resolution model and the second super-resolution model are based on the above reference Figure 1-Figure 5 The method described above is trained and has the same Figure 1-Figure 5 The first super-resolution model and the second super-resolution model have the same structure, so the description of the first super-resolution model and the second super-resolution model will not be repeated here.

[0087] Fig. 9 900 is a block diagram showing a structure of an electronic device for video super-resolution processing and / or video super-resolution model training according to an exemplary embodiment of the present disclosure. The electronic device 900 may be, for example, a smart phone, a tablet computer, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, or a desktop computer. The electronic device 900 may also be referred to as a user device, a portable terminal, a laptop terminal, a desktop terminal, or other names.

[0088] Typically, the electronic device 900 includes a processor 901 and a memory 902 .

[0089] The processor 901 may include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 901 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 901 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0090] The memory 902 may include one or more computer-readable storage media, which may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 902 is used to store at least one instruction, which is used to be executed by the processor 901 to implement the present disclosure. Figure 2-Figure 7 The method embodiments shown provide a video super-resolution model training method and / or a video super-resolution method.

[0091] In some embodiments, the electronic device 900 may further optionally include: a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902 and the peripheral device interface 903 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 903 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 904, a touch display screen 905, a camera 906, an audio circuit 907, a positioning component 908 and a power supply 909.

[0092] The peripheral device interface 903 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902, and the peripheral device interface 903 may be implemented on a separate chip or circuit board, which is not limited in this embodiment.

[0093] The radio frequency circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with communication networks and other communication devices through electromagnetic signals. The radio frequency circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The radio frequency circuit 904 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes, but is not limited to: a metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 904 may also include circuits related to NFC (Near Field Communication), which is not limited in the present disclosure.

[0094] The display screen 905 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos and any combination thereof. When the display screen 905 is a touch display screen, the display screen 905 also has the ability to collect touch signals on the surface or above the surface of the display screen 905. The touch signal can be input to the processor 901 as a control signal for processing. At this time, the display screen 905 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 905 can be one, set on the front panel of the electronic device 900; in other embodiments, the display screen 905 can be at least two, respectively set on different surfaces of the terminal 900 or in a folding design; in some further embodiments, the display screen 905 can be a flexible display screen, set on the curved surface or folding surface of the terminal 900. Even, the display screen 905 can also be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 905 can be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode) and the like.

[0095] The camera assembly 906 is used to capture images or videos. Optionally, the camera assembly 906 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 906 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0096] The audio circuit 907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 901 for processing, or input them into the radio frequency circuit 904 to achieve voice communication. For the purpose of stereo acquisition or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 900. The microphone may also be an array microphone or an omnidirectional acquisition microphone. The speaker is used to convert the electrical signal from the processor 901 or the radio frequency circuit 904 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 907 may also include a headphone jack.

[0097] Positioning component 908 is used to locate the current geographic location of electronic device 900 to implement navigation or LBS (Location Based Service). Positioning component 908 can be a positioning component based on the US GPS (Global Positioning System), China's Beidou system, Russia's Grenas system or the European Union's Galileo system.

[0098] The power supply 909 is used to power various components in the electronic device 900. The power supply 909 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 909 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0099] In some embodiments, the electronic device 900 further includes one or more sensors 910 , including but not limited to: an acceleration sensor 911 , a gyroscope sensor 912 , a pressure sensor 913 , a fingerprint sensor 914 , an optical sensor 915 , and a proximity sensor 916 .

[0100] The acceleration sensor 911 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal 900. For example, the acceleration sensor 911 can be used to detect the components of gravity acceleration on the three coordinate axes. The processor 901 can control the touch display screen 905 to display the user interface in a horizontal view or a vertical view according to the gravity acceleration signal collected by the acceleration sensor 911. The acceleration sensor 911 can also be used for collecting game or user motion data.

[0101] The gyro sensor 912 can detect the body direction and rotation angle of the terminal 900, and the gyro sensor 912 can cooperate with the acceleration sensor 911 to collect the user's 3D actions on the terminal 900. The processor 901 can implement the following functions based on the data collected by the gyro sensor 912: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0102] The pressure sensor 913 can be set on the side frame of the terminal 900 and / or the lower layer of the touch display screen 905. When the pressure sensor 913 is set on the side frame of the terminal 900, the user's holding signal of the terminal 900 can be detected, and the processor 901 performs left and right hand recognition or shortcut operation according to the holding signal collected by the pressure sensor 913. When the pressure sensor 913 is set on the lower layer of the touch display screen 905, the processor 901 controls the operability controls on the UI according to the user's pressure operation on the touch display screen 905. The operability controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.

[0103] The fingerprint sensor 914 is used to collect the user's fingerprint, and the processor 901 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 914, or the fingerprint sensor 914 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as a trusted identity, the processor 901 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, paying, and changing settings. The fingerprint sensor 914 can be set on the front, back, or side of the electronic device 900. When a physical button or a manufacturer logo is set on the electronic device 900, the fingerprint sensor 914 can be integrated with the physical button or the manufacturer logo.

[0104] The optical sensor 915 is used to collect the ambient light intensity. In one embodiment, the processor 901 can control the display brightness of the touch display screen 905 according to the ambient light intensity collected by the optical sensor 915. Specifically, when the ambient light intensity is high, the display brightness of the touch display screen 905 is increased; when the ambient light intensity is low, the display brightness of the touch display screen 905 is reduced. In another embodiment, the processor 901 can also dynamically adjust the shooting parameters of the camera component 906 according to the ambient light intensity collected by the optical sensor 915.

[0105] The proximity sensor 916, also called a distance sensor, is usually arranged on the front panel of the electronic device 900. The proximity sensor 916 is used to collect the distance between the user and the front of the electronic device 900. In one embodiment, when the proximity sensor 916 detects that the distance between the user and the front of the terminal 900 is gradually decreasing, the processor 901 controls the touch display screen 905 to switch from the screen-on state to the screen-off state; when the proximity sensor 916 detects that the distance between the user and the front of the electronic device 900 is gradually increasing, the processor 901 controls the touch display screen 905 to switch from the screen-off state to the screen-on state.

[0106] Those skilled in the art will understand that Fig. 9 The structure shown in the figure does not constitute a limitation on the electronic device 900, and may include more or less components than those shown in the figure, or combine certain components, or adopt a different component arrangement.

[0107] Fig.10 FIG. 1 is a block diagram of another electronic device 1000. For example, the electronic device 1000 may be provided as a server. Fig.10 , the electronic device 1000 includes one or more processing processors 1110 and a memory 1120. The memory 1120 may include one or more programs for executing the above video super-resolution method and / or video super-resolution model training method. The electronic device 1100 may also include a power supply component 1130 configured to perform power management of the electronic device 1100, a wired or wireless network interface 1140 configured to connect the electronic device 1100 to the network, and an input / output (I / O) interface 1150. The electronic device 1100 may operate based on an operating system stored in the memory 1120, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.

[0108] According to an embodiment of the present disclosure, a computer-readable storage medium storing instructions may also be provided, wherein when the instructions are executed by at least one processor, the at least one processor is prompted to execute the video super-resolution model training method and / or the video super-resolution method according to the present disclosure. Examples of computer-readable storage media here include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device is configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide the computer programs and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system, so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.

[0109] According to an embodiment of the present disclosure, a computer program product may also be provided, and instructions in the computer program product may be executed by a processor of a computer device to complete a video super-resolution model training method and / or a video super-resolution method.

[0110] According to the video super-resolution model training method and / or video super-resolution method, device, electronic device, and computer-readable storage medium disclosed in the present invention, video key frames can be fully processed, and the results after key frame super-resolution can be fully utilized to guide the super-resolution processing of non-key frames. When using a lightweight processing model for non-key frames, good subjective effects can also be guaranteed, and good benefits can be achieved in terms of processing speed and subjective effects.

[0111] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0112] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A video processing method, characterized in that: include: dividing the video into groups including a predetermined number of video frames, and classifying the video frames in each group into key frames and non-key frames; Performing super-resolution processing on the key frames and non-key frames of each group respectively through a super-resolution model to obtain super-resolution frames of the key frames and super-resolution frames of the non-key frames of each group; Encode the super-resolution frames of the key frames and the super-resolution frames of the non-key frames of each group into a super-resolution video, wherein the super-resolution model is a model configured to use a super-resolution frame of a key frame as a reference for super-resolution processing of a non-key frame, The super-resolution model includes a first super-resolution model and a second super-resolution model, and the super-resolution frames of the key frames and the non-key frames of each group obtained by the first super-resolution model are input into the second super-resolution model to obtain the super-resolution frames of the non-key frames. The first super-resolution model is configured to perform single-frame image reasoning on the key frame to obtain a super-resolution frame of the key frame, and the second super-resolution model is configured to fuse the super-resolution frame of the key frame and the features of the non-key frame, and perform reasoning based on the fused features to obtain the super-resolution frame of the non-key frame. Among them, the second super-resolution model also includes a first feature extractor, a second feature extractor and a stitching unit. The second super-resolution model is configured to extract the features of the super-resolution frame of the key frame through the first feature extractor and the features of the non-key frame through the second feature extractor, and perform feature stitching and convolution on the extracted features of the super-resolution frame of the key frame and the features of the non-key frame through the stitching unit to obtain the fused features.

2. The method according to claim 1, characterized in that The super-resolution model is trained in the following way: Determining a first loss function for a first super-resolution model according to the key frame and the super-resolution frame of the key frame, and adjusting parameters of the first super-resolution model based on a value of the first loss function; A second loss function for a second super-resolution model is determined according to the non-key frame and the super-resolution frame of the non-key frame, and parameters of the second super-resolution model are adjusted based on the second loss function.

3. The method according to claim 1, characterized in that The video frames in each group are classified as key frames and non-key frames including: Classifying a predetermined type of video frames in the group as key frames, and classifying the remaining video frames as non-key frames, or The first frame in the group is classified as a key frame, and the remaining video frames are classified as non-key frames.

4. The method according to claim 2, characterized in that Each of the first super-resolution model and the second super-resolution model includes an inference main network, which includes multiple residual convolution layers and fast upsampling layers for performing inference operations, wherein the number of channels and the number of residual convolution layers included in the first super-resolution model are higher than the number of channels and the number of residual convolution layers included in the second super-resolution model.

5. The method according to claim 4, characterized in that Performing single-frame image inference on the key frame by the first super-resolution model to obtain a super-resolution frame of the key frame includes: Calculating the depth features of the key frames layer by layer through the multiple residual convolutional layers; The deep features of the keyframes are upsampled to super-resolution frames with high-resolution keyframes through a fast upsampling layer.

6. The method according to claim 4, characterized in that The second super-resolution model is further configured as: Performing calculations on the fused features layer by layer through a plurality of residual convolutional layers of an inference subject network to extract deep features of non-key frames; The deep features of non-key frames are upsampled to super-resolution frames with high resolution through a fast upsampling layer.

7. The method according to claim 6, characterized in that Each residual convolution layer of the first super-resolution model and the second super-resolution model includes a plurality of residual convolution blocks connected in series, and each residual convolution block includes a plurality of basic convolution operation units, which are configured to perform convolution operations on input features to extract deeper features.

8. The method according to claim 7, characterized in that In the first super-resolution model and the second super-resolution model: The input features and output features of the subject reasoning network are connected across layers; The input features and output features of each residual convolution layer are connected across layers; The output features of the first basic convolutional operation unit and the last convolutional operation unit in each residual block of the residual convolutional layer are connected across layers.

9. A video processing device, characterized in that: include: a grouping unit configured to divide the video into groups including a predetermined number of video frames, and classify the video frames in each group into key frames and non-key frames; A video processing unit is configured to perform super-resolution processing on the key frames and non-key frames of each group respectively through a super-resolution model to obtain super-resolution frames of the key frames and super-resolution frames of the non-key frames of each group; An encoding unit configured to encode the super-resolution frames of the key frames and the super-resolution frames of the non-key frames of each group into a super-resolution video, wherein the super-resolution model is a model configured to use a super-resolution frame of a key frame as a reference for super-resolution processing of a non-key frame, The super-resolution model includes a first super-resolution model and a second super-resolution model, and the super-resolution frames of the key frames of the group obtained by the first super-resolution model and the non-key frames of the group are input into the second super-resolution model to obtain the super-resolution frames of the non-key frames. The first super-resolution model is configured to perform single-frame image reasoning on the key frame to obtain a super-resolution frame of the key frame, and the second super-resolution model is configured to fuse the super-resolution frame of the key frame and the features of the non-key frame, and perform reasoning based on the fused features to obtain the super-resolution frame of the non-key frame. Among them, the second super-resolution model also includes a first feature extractor, a second feature extractor and a stitching unit. The second super-resolution model is configured to extract the features of the super-resolution frame of the key frame through the first feature extractor and the features of the non-key frame through the second feature extractor, and perform feature stitching and convolution on the extracted features of the super-resolution frame of the key frame and the features of the non-key frame through the stitching unit to obtain the fused features.

10. The device according to claim 9, characterized in that The super-resolution model is trained in the following way: Determining a first loss function for a first super-resolution model according to the key frame and the super-resolution frame of the key frame, and adjusting parameters of the first super-resolution model based on a value of the first loss function; A second loss function for a second super-resolution model is determined according to the non-key frame and the super-resolution frame of the non-key frame, and parameters of the second super-resolution model are adjusted based on the second loss function.

11. The device according to claim 9, characterized in that The grouping unit is configured as: Classifying a predetermined type of video frames in the group as key frames, and classifying the remaining video frames as non-key frames, or The first frame in the group is classified as a key frame, and the remaining video frames are classified as non-key frames.

12. The device according to claim 10, characterized in that Each of the first super-resolution model and the second super-resolution model includes an inference main network, which includes multiple residual convolution layers and fast upsampling layers for performing inference operations, wherein the number of channels and the number of residual convolution layers included in the first super-resolution model are higher than the number of channels and the number of residual convolution layers included in the second super-resolution model.

13. The device according to claim 12, characterized in that The first super-resolution model is configured as: Calculating the depth features of the key frames layer by layer through the multiple residual convolutional layers; The deep features of the keyframes are upsampled to super-resolution frames with high-resolution keyframes through a fast upsampling layer.

14. The device according to claim 12, characterized in that The second super-resolution model is further configured as: Performing calculations on the fused features layer by layer through a plurality of residual convolutional layers of an inference subject network to extract deep features of non-key frames; The deep features of non-key frames are upsampled to super-resolution frames with high resolution through a fast upsampling layer.

15. The device according to claim 14, characterized in that Each residual convolution layer of the first super-resolution model and the second super-resolution model includes a plurality of residual convolution blocks connected in series, and each residual convolution block includes a plurality of basic convolution operation units, which are configured to perform convolution operations on input features to extract deeper features.

16. The device according to claim 15, characterized in that In the first super-resolution model and the second super-resolution model: The input features and output features of the subject reasoning network are connected across layers; The input features and output features of each residual convolution layer are connected across layers; The output features of the first basic convolutional operation unit and the last convolutional operation unit in each residual block of the residual convolutional layer are connected across layers.

17. An electronic device, characterized in that: include: at least one processor; at least one memory storing computer executable instructions, Wherein, when the computer executable instructions are executed by the at least one processor, the at least one processor is prompted to perform the method of any one of claims 1 to 8.

18. A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 8.

19. A computer program product, characterized in that The instructions in the computer program product are executed by at least one processor in an electronic device to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for enhancing video image quality

    CN111031346A