Video Processing Method, Apparatus, and Storage Medium Based on Super-Resolution Network

By introducing spatial attention blocks and global context modules into the residual network structure of the super-resolution network, the depth characteristics of the video are extracted and reconstruction are carried out, and the problem of slow processing speed of super-resolution networks in the prior art is solved, and efficient video super-resolution processing is achieved.

CN114463182BActive Publication Date: 2025-06-03SHENZHEN KANDAO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210128987.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-11
Publication Date
2025-06-03
Estimated Expiration
2042-02-11

AI Technical Summary

Technical Problem

Existing super-resolution networks need to train deep network structures when ensuring the ultra-clear quality of images, resulting in very slow data processing speed and are not suitable for video super-resolution processing.

Method used

The spatial attention block SA module and the global context module GCB module are set up in the residual network structure RIR module. The depth features of shallow features are extracted through these modules and reconstructed by reconstructing the convolutional layer to obtain super-resolution image frames.

Benefits of technology

On the basis of not increasing the depth of the network layer, the super-resolution processing effect is improved, the processing efficiency is improved, the smoothness of video data is ensured, and the excessive requirements for graphics card quality are reduced, so as to achieve a balance between high resolution effect and rapid processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463182B_ABST
    Figure CN114463182B_ABST
Patent Text Reader

Abstract

The present invention discloses a video processing method, including: inputting each low-resolution image frame of a low-resolution video into a trained super-resolution network for processing to obtain a super-resolution image frame; wherein, extracting shallow features of each low-resolution image frame through a shallow feature extraction convolutional layer; extracting deep features of the shallow features through a residual network structure RIR module; extracting deep features of the shallow features through a preset residual group RG; in RG, extracting deep features through stacked residual channel attention blocks RCAB and a spatial attention module SA module; in RCAB, extracting channel features through a global context module GCB; reconstructing based on the deep features through a reconstruction convolutional layer to obtain super-resolution image frames of each low-resolution image frame; and synthesizing all the super-resolution image frames into a super-resolution video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and particularly to a video processing method, device and storage medium based on a super-resolution network. Background Art

[0002] High-definition videos such as 8K videos captured by 8K cameras are not suitable for real-time transmission due to their large size. Therefore, before transmitting an 8K video, it will be compressed. The specific processing method can be interlaced compression (for example, only the pixels of odd or even rows), so that the 8K video will be compressed into a 4K video. At this time, the video viewing terminal needs to restore the 4K video. In addition, in real life, there are often scenarios where low-resolution videos are converted into high-resolution videos, such as restoring old photo videos, etc.

[0003] Existing super-resolution networks need to train very deep network structures to ensure the ultra-clear quality of images, and the data processing speed is very slow, which is not suitable for video super-resolution processing.

[0004] Therefore, a video processing method based on a super-resolution network is needed to solve the above technical problems. Summary of the Invention

[0005] The main purpose of the present invention is to provide a video processing method, device and computer-readable storage medium based on a super-resolution network, aiming to solve the problem that existing super-resolution networks need to train very deep network structures to ensure the ultra-clear quality of images, and the data processing speed is very slow, which is not suitable for video super-resolution processing.

[0006] To achieve the above purpose, the present invention provides a video processing method based on a super-resolution network, and the method includes:

[0007] Step S10, obtaining a low-resolution video to be processed;

[0008] Step S20, inputting each low-resolution image frame of the low-resolution video into a trained super-resolution network for processing to obtain a super-resolution image frame corresponding to each low-resolution image frame;

[0009] Wherein, the step S20 includes:

[0010] Step S21, extracting the shallow features of each low-resolution image frame through a shallow feature extraction convolutional layer;

[0011] Step S22, extracting the deep features of the shallow features through a residual network structure RIR module;

[0012] Wherein, the step S22 includes:

[0013] Step S221: Extract the deep features of the shallow features through a preset residual group RG;

[0014] Among them, the step S221 includes:

[0015] Step S101: In the residual group RG, extract deep features through stacked residual channel attention blocks RCAB and a spatial attention module SA module;

[0016] Among them, in the RCAB, extract channel features through a global context module GCB;

[0017] Step S23: Reconstruct based on the deep features through a reconstruction convolutional layer to obtain super-resolution image frames of each low-resolution image frame;

[0018] Step S30: Synthesize all the super-resolution image frames into a super-resolution video.

[0019] Optionally, the step S221, extracting the deep features of the shallow features through a preset residual group RG further includes:

[0020] Step S102: Extract the deep features of the shallow features through two stacked residual groups RG;

[0021] The step S101, extracting deep features through stacked channel attention blocks RCAB and a spatial attention module SA module in the residual group RG includes:

[0022] Step S201: In each residual group RG, extract deep features through two stacked RCAB modules, one SA module, and one convolutional layer.

[0023] Optionally, the step S22, extracting the deep features of the shallow features through a residual network structure RIR module includes:

[0024] Step S222: Forward the first low-frequency information at the input end of the RIR module to the output end of the RIR module through a long skip connection LSC;

[0025] The step S221, extracting the deep features of the shallow features through a preset residual group RG includes:

[0026] Step S103: Forward the second low-frequency information in the input end of the RG to the output end of the RG through a short skip connection SSC.

[0027] Optionally, the spatial attention module SA module includes a spatial attention convolutional layer and a spatial attention activation layer, wherein the number of output channels of the spatial attention convolutional layer is 1; after the SA module extracts the spatial feature quantity, it is added to the channel feature quantity in a superposition manner.

[0028] Optionally, in step S201, in each residual group RG, the deep feature extraction by two stacked RCAB modules, one SA module, and one convolutional layer includes:

[0029] Step S301, performing deep feature extraction on the shallow feature through the first RCAB module to obtain a first deep feature quantity;

[0030] Step S302, performing deep feature extraction on the first deep feature quantity through the second RCAB module to obtain a second deep feature quantity;

[0031] Step S303, performing spatial feature extraction based on the first deep feature quantity and the second deep feature quantity through the SA module to obtain a spatial feature quantity;

[0032] Step S304, performing superposition processing on the spatial feature quantity, the first deep feature quantity, and the second deep feature quantity, and using the obtained result as the output of the residual group RG.

[0033] Optionally, in step S303, performing spatial feature extraction based on the first deep feature quantity and the second deep feature quantity through the SA module to obtain a spatial feature quantity includes:

[0034] Step S401, extracting intermediate spatial feature information based on the first deep feature quantity and the second deep feature quantity through the spatial attention convolutional layer, wherein the number of channels of the spatial attention convolutional layer is 1;

[0035] Step S402, extracting a spatial feature quantity based on the intermediate spatial feature information through the activation function ReLU.

[0036] Optionally, in step S301, performing deep feature extraction on the shallow feature through the first RCAB module to obtain a first deep feature quantity includes:

[0037] Step S403, inputting the shallow feature into the first convolutional layer in the first RCAB module for processing to obtain a first intermediate channel feature;

[0038] Step S404, inputting the first intermediate channel feature into the activation function ReLU in the first RCAB module for processing to obtain a second intermediate channel feature;

[0039] Step S405: Input the second intermediate channel feature into the second convolutional layer in the first RCAB module for processing to obtain a third intermediate channel feature;

[0040] Step S406: Input the third intermediate channel feature into the global context module GCB module in the first RCAB module for processing to obtain a fourth intermediate channel feature;

[0041] Step S407: Superimpose the shallow feature, the first intermediate channel feature, the second intermediate channel feature, the third intermediate channel feature, and the fourth intermediate channel feature to obtain the first depth feature quantity.

[0042] Optionally, in step S406, inputting the third intermediate channel feature into the global context module GCB module in the first RCAB module for processing to obtain a fourth intermediate channel feature includes:

[0043] Step S501: Input the third intermediate channel feature into the first convolutional layer and the softmax layer in the GCB module for processing, and perform dot product processing on the obtained intermediate result and the third intermediate channel feature to obtain a compressed feature;

[0044] Step S502: Input the compressed feature into the second convolutional layer, the activation function ReLU, and the third convolutional layer in the GCB module for processing, and superimpose the generated intermediate result and the compressed feature to obtain the fourth intermediate channel feature

[0045] To achieve the above object, the present invention further provides a video processing device based on a super-resolution network. The video processing device includes:

[0046] A video acquisition module for acquiring a low-resolution video to be processed;

[0047] A super-resolution processing module for inputting each low-resolution image frame of the low-resolution video into a trained super-resolution network for processing to obtain a super-resolution image frame corresponding to each low-resolution image frame;

[0048] Wherein, the super-resolution processing module includes:

[0049] A shallow feature extraction sub-module for extracting the shallow features of each low-resolution image frame through a shallow feature extraction convolutional layer;

[0050] A depth feature extraction sub-module for extracting the depth features of the shallow features through a residual network structure RIR module;

[0051] The deep feature extraction module is further configured to extract the deep features of the shallow features through a preset residual group RG;

[0052] The deep feature extraction module is further configured to perform deep feature extraction through stacked residual channel attention blocks RCAB and a spatial attention module SA module in the residual group RG;

[0053] Among them, in the RCAB, channel features are extracted through a global context module GCB;

[0054] The reconstruction sub-module is configured to perform reconstruction based on the deep features through a reconstruction convolutional layer to obtain super-resolution image frames of each low-resolution image frame;

[0055] The synthesis module is configured to synthesize all the super-resolution image frames into a super-resolution video.

[0056] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described above are implemented.

[0057] The video processing method, device and computer-readable storage medium provided by the present invention first obtain a low-resolution video to be processed, and then input each low-resolution image frame of the low-resolution video into a trained super-resolution network for processing to obtain super-resolution image frames corresponding to each low-resolution image frame. During the processing, shallow features of each low-resolution image frame are extracted through a shallow feature extraction convolutional layer, and deep features of the shallow features are extracted through a residual network structure RIR module. In the residual network structure RIR module, channel features are extracted through a residual channel attention block RCAB, a spatial attention module SA module and a global context module GCB, and then reconstruction is performed based on the deep features through a reconstruction convolutional layer to obtain super-resolution image frames of each low-resolution image frame. All the super-resolution image frames are synthesized into a super-resolution video. By the above method, compared with the prior art, the spatial attention block SA module and the global context module GCB module set in the residual network structure RIR module can improve the super-resolution processing effect without increasing the depth of the network layer, make up for the lack of depth, and at the same time, the reduction of the depth of the network layer is more conducive to improving the processing efficiency, ensuring that the smoothness of video data is not affected during the super-resolution processing, reducing the excessive requirements for the quality of the graphics card, and while taking into account the high-resolution effect, the processing speed can be closer to real-time processing, reducing resource occupancy and improving the user experience. Description of the Drawings

[0058] Figure 1It is a schematic diagram of the system structure of the hardware operating environment involved in the solution of the embodiment of the present invention;

[0059] Figure 2 It is a schematic flowchart of an embodiment of the method for video processing based on super resolution of the present invention;

[0060] Figure 3 It is an example diagram of the super resolution network structure in the embodiment of the present invention;

[0061] Figure 4 It is an example diagram of the composition structure of the residual group RG in the embodiment of the present invention;

[0062] Figure 5 It is an example diagram of the structure of the spatial attention module SA module in the embodiment of the present invention;

[0063] Figure 6 It is a schematic flowchart of the refinement of step S201 in the embodiment of the present invention;

[0064] Figure 7 It is a schematic flowchart of the refinement of step S303 in the embodiment of the present invention;

[0065] Figure 8 It is a schematic flowchart of the refinement of step S301 in the embodiment of the present invention;

[0066] Figure 9 It is an example diagram of the structure of the residual channel attention block RCAB in the embodiment of the present invention;

[0067] Figure 10 It is a schematic flowchart of the refinement of step S406 in the embodiment of the present invention;

[0068] Figure 11 It is an example diagram of the structure of the global context module GCB module in the embodiment of the present invention. Detailed implementation manners

[0069] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0071] In the prior art, the super resolution network needs to train a very deep network structure to ensure the ultra - clear quality of the image, and the data processing speed is very slow, which is not suitable for video super resolution processing.

[0072] To solve the above technical problems, the present invention provides a video processing method based on a super-resolution network. In this method, a spatial attention block SA module and a global context block GCB module set in the residual network structure RIR module can improve the super-resolution processing effect without increasing the network layer depth, making up for the lack of depth. At the same time, the reduction of the network layer depth is more conducive to improving the processing efficiency, ensuring that the video data fluency is not affected during the super-resolution processing, reducing the excessive requirements for the graphics card quality, and making the processing speed closer to real-time processing while taking into account the high-resolution effect, reducing resource occupancy, and improving the user experience.

[0073] As Figure 1 shown, Figure 1 is a schematic diagram of the system structure of the hardware operating environment involved in the embodiment solution of the present invention.

[0074] The terminal in the embodiment of the present invention can be a terminal device with computing capabilities, a PC, a smart phone, a tablet computer, an e-book reader, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a portable computer, or other movable terminal devices with a display function.

[0075] As Figure 1 shown, the terminal may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the foregoing processor 1001.

[0076] Optionally, the terminal may further include a camera, an RF (Radio Frequency) circuit, sensors, an audio circuit, a WiFi module, etc. Among them, the sensors such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display screen according to the brightness of the ambient light, and the proximity sensor can turn off the display screen and / or the backlight when the mobile terminal is moved to the ear. As a kind of motion sensor, the gravity acceleration sensor can detect the magnitude of the acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile terminal (such as horizontal and vertical screen switching, related games, magnetometer attitude calibration), vibration recognition related functions (such as pedometer, tapping), etc.; of course, the mobile terminal can also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be elaborated here.

[0077] Those skilled in the art can understand that Figure 1 the terminal structure shown in

[0078] does not constitute a limitation on the terminal, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Figure 1 As shown in

[0079] In the Figure 1 shown terminal, the network interface 1004 is mainly used to connect to the background server and communicate with the background server for data; the user interface 1003 is mainly used to connect to the client (user side) and communicate with the client for data; and the processor 1001 can be used to call the video processing program stored in the memory 1005 and perform the following operations:

[0080] Step S10, obtaining a low-resolution video to be processed;

[0081] Step S20, inputting each low-resolution image frame of the low-resolution video into a trained super-resolution network for processing to obtain a super-resolution image frame corresponding to each low-resolution image frame;

[0082] Among them, the step S20 includes:

[0083] Step S21, extracting the shallow features of each low-resolution image frame through a shallow feature extraction convolutional layer;

[0084] Step S22, extracting the deep features of the shallow features through a residual network structure RIR module;

[0085] Among them, the step S22 includes:

[0086] Step S221, extracting the depth features of the shallow features through a preset residual group RG;

[0087] Among them, the step S221 includes:

[0088] Step S101, in the residual group RG, extracting depth features through stacked residual channel attention blocks RCAB and a spatial attention module SA module;

[0089] Among them, in the RCAB, extracting channel features through a global context module GCB;

[0090] Step S23, reconstructing based on the depth features through a reconstruction convolutional layer to obtain super-resolution image frames of each low-resolution image frame;

[0091] Step S30, synthesizing all the super-resolution image frames into a super-resolution video.

[0092] Further, the processor 1001 can call the video processing program stored in the memory 1005 and also perform the following operations:

[0093] Step S102, extracting the depth features of the shallow features through two stacked residual groups RG;

[0094] Step S201, in each residual group RG, extracting depth features through two stacked RCAB modules, one SA module, and one convolutional layer.

[0095] Further, the processor 1001 can call the video processing program stored in the memory 1005 and also perform the following operations:

[0096] Step S222, forwarding the first low-frequency information at the input end of the RIR module to the output end of the RIR module through a long skip connection LSC;

[0097] Step S103, forwarding the second low-frequency information in the input end of the RG to the output end of the RG through a short skip connection SSC.

[0098] Further, the processor 1001 can call the video processing program stored in the memory 1005 and also perform the following operations:

[0099] Step S301, extracting depth features of the shallow features through a first RCAB module to obtain a first depth feature quantity;

[0100] Step S302, extracting depth features of the first depth feature quantity through a second RCAB module to obtain a second depth feature quantity;

[0101] Step S303: Based on the first depth feature quantity and the second depth feature quantity, perform spatial feature extraction through the SA module to obtain a spatial feature quantity.

[0102] Step S304: Superimpose the spatial feature quantity, the first depth feature quantity, and the second depth feature quantity, and use the obtained result as the output of the residual group RG.

[0103] Further, the processor 1001 can call the video processing program stored in the memory 1005 and further perform the following operations:

[0104] Step S401: Extract intermediate spatial feature information based on the first depth feature quantity and the second depth feature quantity through a spatial attention convolutional layer, where the number of channels of the spatial attention convolutional layer is 1.

[0105] Step S402: Extract a spatial feature quantity based on the intermediate spatial feature information through the activation function ReLU.

[0106] Further, the processor 1001 can call the video processing program stored in the memory 1005 and further perform the following operations:

[0107] Step S403: Input the shallow feature into the first convolutional layer in the first RCAB module for processing to obtain a first intermediate channel feature.

[0108] Step S404: Input the first intermediate channel feature into the activation function ReLU in the first RCAB module for processing to obtain a second intermediate channel feature.

[0109] Step S405: Input the second intermediate channel feature into the second convolutional layer in the first RCAB module for processing to obtain a third intermediate channel feature.

[0110] Step S406: Input the third intermediate channel feature into the global context module GCB module in the first RCAB module for processing to obtain a fourth intermediate channel feature.

[0111] Step S407: Superimpose the shallow feature, the first intermediate channel feature, the second intermediate channel feature, the third intermediate channel feature, and the fourth intermediate channel feature to obtain the first depth feature quantity.

[0112] Further, the processor 1001 can call the video processing program stored in the memory 1005 and further perform the following operations:

[0113] Step S501: Input the third intermediate channel feature into the first convolutional layer and the softmax layer in the GCB module for processing in sequence, perform dot product processing on the obtained intermediate result and the third intermediate channel feature to obtain a compressed feature;

[0114] Step S502: Input the compressed feature into the second convolutional layer, the activation function ReLU, and the third convolutional layer in the GCB module for processing in sequence, superimpose the generated intermediate result and the compressed feature to obtain the fourth intermediate channel feature.

[0115] Refer to Figure 2 , Figure 2 FIG. [X] is a schematic flowchart of an embodiment of the video processing method based on super-resolution of the present invention. The video processing method of the present invention can be applied in the technical field of restoring low-resolution video data into super-resolution video data. For example, in some specific technical scenarios, the 8K video captured by an 8K camera is not suitable for real-time transmission due to its large volume. Therefore, before transmitting the 8K video, compression processing will be performed on the 8K video. The specific compression method can be interlaced compression, such as only compressing the pixels of odd rows or even rows, so that the 8K video will be compressed into a 4K video.

[0116] Subsequently, the video playback end or the receiving end needs to restore the 4K video, and the video processing method of the present invention can be applied for restoration to restore the low-resolution video into a super-resolution video. The method of the present invention includes:

[0117] Step S10: Obtain the low-resolution video to be processed.

[0118] First, obtain the low-resolution video data to be processed. The low-resolution video data can be the above-mentioned 4K video data after compression processing, or other types of low-resolution video data, which is not limited herein. Among them, the low-resolution video is formed by multiple low-resolution image frames.

[0119] In some embodiments, the low-resolution video can be obtained based on a video processing instruction triggered by the user, or the steps of the video processing method of the present invention can be automatically executed starting from step S10 when receiving the low-resolution video. Correspondingly, the low-resolution video to be processed can be stored locally, determined based on the instruction selected by the user, or received through the network. The low-resolution video (or image frame) and the super-resolution video (or image frame) of the present invention are resolutions corresponding to two different high-definition specifications, and can be set and distinguished based on the definitions of existing technologies or actual application scenarios, which are not specifically limited herein.

[0120] Step S20: Input each low-resolution image frame of the low-resolution video into the pre-trained super-resolution network for processing to obtain the super-resolution image frame corresponding to each low-resolution image frame.

[0121] Based on Step S10, after obtaining the low-resolution video, extract each low-resolution image frame in the video, and input the low-resolution image frame into the trained super-resolution network for processing. After the super-resolution processes each low-resolution image frame, it outputs the super-resolution image frame corresponding to each low-resolution image frame.

[0122] Among them, the super-resolution network can be a model pre-trained and saved locally, or a model saved on other terminals can be called. During the process of training the super-resolution network, a preset number of low-resolution images and corresponding super-resolution images can be obtained as the training set. The low-resolution images in the training set are used as the network input, and the corresponding super-resolution images are used as the network output, and iterative processing is performed in the network to adjust the structural parameters in the network to obtain.

[0123] Specifically, in some embodiments, Step S20 may include:

[0124] Step S21: Extract the shallow features of each low-resolution image frame through the shallow feature extraction convolutional layer;

[0125] Step S22: Extract the deep features of the shallow features through the residual network structure RIR module.

[0126] Specifically, as Figure 3 shown, Figure 3 is a schematic diagram of the super-resolution network structure in an embodiment of the present invention. As shown in the figure, LR is the low-resolution image frame to be processed, and HR stack is the super-resolution image frame obtained after being processed by the super-resolution network. The super-resolution network of the present invention consists of a shallow feature extraction convolutional layer CONV① and a residual network structure RIR (Residual in Residual) module. Among them, the residual network structure RIR module includes at least one residual group Rg 1 and / or Rg2, and a reconstruction convolutional layer CONV②.

[0127] During the implementation process, the low-resolution image frame LR can be first input into the shallow feature extraction convolutional layer CONV① for shallow feature extraction to obtain the shallow features of the low-resolution image frame LR. Among them, the shallow feature extraction convolutional layer refers to the convolutional layer used to extract shallow features.

[0128] Then, the depth features are extracted from the shallow features through the residual network structure RIR module. The residual network structure RIR module of the present invention is based on the improved residual network structure, which can ensure the image clarity processing effect without increasing the number of network layers, thereby improving the efficiency of image frame super-resolution processing and meeting the actual application scenarios of video super-resolution processing. The specific implementation method is described as follows. Specifically, in some embodiments, step S22 includes:

[0129] Step S221, extracting the depth features of the shallow features through a preset residual group RG;

[0130] Among them, the step S221 includes:

[0131] Step S101, in the residual group RG, extracting depth features through stacked residual channel attention blocks RCAB and a spatial attention module SA module;

[0132] Among them, in the RCAB, channel features are extracted through a global context module GCB.

[0133] As Figure 3 and Figure 4 shown, where Figure 4 is an example diagram of the composition structure of the residual group RG in the embodiment of the present invention. The residual network structure RIR module of the present invention is composed of one or more residual groups RG (i.e., Residual Group). Among them, the residual group is composed of one or more residual channel attention blocks RCAB (i.e., Residual Channel Attention Block), a spatial attention module SA (i.e., Spatial Attention) module, and a convolutional layer. And, in the residual channel attention block RCAB, a global context module GCB (i.e., Global Context Block) is set to extract depth features. Through the above structural settings, thus, in the residual network structure RIR module, the depth features are extracted from the shallow features through the residual group RG. When the residual group extracts the depth features, the depth features are extracted through the residual channel attention block RCAB and the spatial attention module SA module, and the channel features are extracted through the global context module GCB. The above structural settings and extraction methods are further improvements on the basis of the existing residual network structure RIR module. By adding the spatial attention block SA module, the network's perception of spatial features can be improved; by adding the global context module GCB, the network's perception of channel features can be increased, and channel features independent of space can be better obtained, improving super-resolution.

[0134] Step S23, reconstructing based on the depth features through a reconstruction convolutional layer to obtain super-resolution image frames of each low-resolution image frame.

[0135] As Figure 3 shown, after extracting the depth features through the residual group RG, the extracted depth features are reconstructed through the reconstruction convolutional layer CONV② to obtain the super-resolution image frames corresponding to the low-resolution image frames.

[0136] Step S30, synthesize all the super-resolution image frames into a super-resolution video.

[0137] After obtaining the super-resolution image frames, determine the arrangement order of the super-resolution image frames based on the arrangement order of the low-resolution image frames in the low-resolution video in step S10, and synthesize the super-resolution video based on the arrangement order of the super-resolution image frames.

[0138] In the above video processing method based on the super-resolution network, first obtain the low-resolution video to be processed, and then input each low-resolution image frame of the low-resolution video into the trained super-resolution network for processing to obtain the super-resolution image frames corresponding to each low-resolution image frame. During the processing, extract the shallow features of each low-resolution image frame through the shallow feature extraction convolutional layer, and extract the depth features of the shallow features through the residual network structure RIR module. In the residual network structure RIR module, extract the channel features through the residual channel attention block RCAB, the spatial attention module SA module, and the global context module GCB, and then reconstruct based on the depth features through the reconstruction convolutional layer to obtain the super-resolution image frames of each low-resolution image frame. Synthesize all the super-resolution image frames into a super-resolution video. In this way, compared with the prior art, the spatial attention block SA module and the global context module GCB module set in the residual network structure RIR module can improve the super-resolution processing effect without increasing the depth of the network layers, make up for the lack of depth, and at the same time, the reduction of the depth of the network layers is more conducive to improving the processing efficiency, ensuring that the smoothness of the video data is not affected during the super-resolution processing, reducing the excessive requirements for the quality of the graphics card, and while taking into account the high-resolution effect, the processing speed can be closer to real-time processing, reducing resource occupancy and improving the user experience.

[0139] In some embodiments, the step S221 of extracting the depth features of the shallow features through the preset residual group RG further includes:

[0140] Step S102, extract the depth features of the shallow features through two stacked residual groups RG;

[0141] The step S101 of extracting depth features through the stacked residual channel attention block RCAB and the spatial attention module SA module in the residual group RG includes:

[0142] Step S201, in each residual group RG, perform deep feature extraction through two stacked RCAB modules, one SA module, and one convolutional layer.

[0143] Please refer to Figure 3 and Figure 4 , in some embodiments, two residual groups may be set in the residual network structure RIR to sequentially extract the deep features of the shallow features. Among them, the residual group RG-1 is used to extract the deep features of the shallow features, and the residual group RG-2 is used to further extract the deep features from the deep features extracted by RG-1, and then the deep features extracted by the residual group RG-1 and the residual group RG-2 are superimposed.

[0144] In each residual group RG, the shallow features or the deep features extracted based on the shallow features can be gradually processed through the combination of two residual channel attention blocks RCAB, one spatial attention block SA module, and one convolutional layer (such as Figure 4 the convolutional layer ③ in ), and the deep features extracted by each component are superimposed and then output.

[0145] In the above processing method, the deep features of the shallow features are extracted through two stacked residual groups RG; in each residual group RG, deep feature extraction is performed through two stacked RCAB modules, one SA module, and one convolutional layer. Through the above network structure setting and feature extraction method, the network structure can be simplified, the depth of the network layer can be reduced, and the feature extraction efficiency can be improved.

[0146] It can be understood that, in some embodiments, the residual group RG, which is a component of the residual network structure RIR module of the present invention, and each component of the residual group RG can be adjusted according to the actual application scenario or application field as needed, and is not limited to the number or components mentioned in the above embodiments.

[0147] Such as Figure 5 shown, where Figure 5This is a structural example diagram of the Spatial Attention module SA module in the embodiments of the present invention. In some embodiments, the Spatial Attention module SA module includes a spatial attention convolutional layer CONV④ and a spatial attention activation layer ReLU①. Among them, the number of output channels of the spatial attention convolutional layer is 1; after the SA module extracts the spatial feature quantity, it is added to the channel feature quantity in a superposition manner. In the above way, the Spatial Attention module SA module can be implemented in a simple way, that is, a structure composed of a convolutional layer and an activation layer. The convolutional layer and the activation layer are used to extract spatial features. Here, the convolutional layer CONV④ is different from an ordinary convolutional layer, and its number of output channels is 1. That is to say, the spatial features it extracts are independent of the channels. The subsequent network structure will superimpose this spatial feature quantity on the original feature quantity of the channel to enhance / weaken the feature quantity at the corresponding spatial position, achieving the adjustment effect on the spatial features.

[0148] In some embodiments, in step S22, extracting the depth features of the shallow features through the Residual Network Structure RIR module includes:

[0149] Step S222, forwarding the first low-frequency information at the input end of the RIR module to the output end of the RIR module through the Long Skip Connection LSC;

[0150] Step S221, extracting the depth features of the shallow features through the preset Residual Group RG includes:

[0151] Step S103, forwarding the second low-frequency information in the RG input end to the output end of the RG through the Short Skip Connection SSC.

[0152] Specifically, in the super-resolution processing process of low-frequency images, in order to achieve better processing effects, generally, an attempt is made to restore as much high-frequency information as possible. The low-frequency image contains relatively more low-frequency information, and this part of the information can be directly output to the output end of the network structure. Subtracting this part of the low-frequency information from the output signal at the output end can accurately obtain the output high-frequency information.

[0153] Such as Figure 3 and Figure 4As shown, in some embodiments, a long skip connection LSC (i.e., Long Skip Connection) can be set in the residual network structure RIR module, and a short skip connection SSC (i.e., Short Skip Connection) can be set in the residual group RG. Among them, the long skip connection LSC is used to identify the low-frequency information part in the input information of the residual network structure RIR module, that is, the first low-frequency information, and directly forward the first low-frequency information to the output end of the residual network structure RIR module for superposition processing. Thus, the residual network structure RIR module focuses on processing the more valuable high-frequency information part, improving the information processing efficiency of the residual network structure RIR module. The short skip connection SSC is used to identify the low-frequency information part in the input information of the residual group RG, that is, the second low-frequency information, and directly forward the second low-frequency information to the output end of the residual group RG for superposition processing, without performing complex feature extraction and other processing on the second low-frequency information through the internal components in the residual group RG. The residual group RG can concentrate resources on processing the more valuable high-frequency information part, improving the information processing efficiency of the residual group RG. Among them, the above-mentioned first low-frequency information and second low-frequency information respectively refer to the low-frequency information part of the input information at the input end of the residual network structure RIR module and the residual group RG, and are not limited to specific information.

[0154] Based on the above structure setting and processing method, thus, the first low-frequency information at the input end of the RIR module can be forwarded to the output end of the RIR module through the long skip connection LSC, enabling the network to only predict the high-frequency part of the signal; the second low-frequency information in the input end of the RG can be forwarded to the output end of the RG through the short skip connection SSC, thus skipping the rich low-frequency information, simplifying the information flow, concentrating resources on the feature extraction processing of the high-frequency information part, and improving the processing efficiency.

[0155] Please refer to Figure 6 , Figure 6 which is a detailed flowchart of step S201 in the embodiment of the present invention. In some embodiments, step S201 includes:

[0156] Step S301, performing deep feature extraction on the shallow features through the first RCAB module to obtain the first deep feature quantity.

[0157] As Figure 4As shown, each component in the residual group RG processes the input in a certain order. The first RCAB module, i.e., the first residual channel attention block, is the RCAB module in the residual group RG that first receives the input information. Based on the above video processing method, the input received by the first RCAB module is the shallow feature extracted by the convolutional layer CONV①. When the first RCAB module receives the shallow feature, it first performs the first deep feature extraction on the shallow feature to obtain the first deep feature quantity. Among them, the first deep feature refers to the intermediate deep feature obtained by the first RCAB module performing the first deep feature extraction on the shallow feature.

[0158] Step S302, perform deep feature extraction on the first deep feature quantity through the second RCAB module to obtain the second deep feature quantity.

[0159] As Figure 4 shown, the second RCAB module is the adjacent component sorted after the first RCAB module. After the first RCAB module extracts the first deep feature quantity, it sends the first deep feature quantity to the second RCAB module (i.e., the second residual channel attention block). After the second RCAB module receives the first deep feature quantity, it performs further deep feature extraction on the first deep feature quantity to obtain the second deep feature quantity. Among them, the second deep feature quantity refers to the intermediate deep feature obtained by the second RCAB module performing further deep feature extraction on the first deep feature quantity.

[0160] Step S303, perform spatial feature extraction based on the first deep feature quantity and the second deep feature quantity through the SA module to obtain the spatial feature quantity.

[0161] As Figure 4 shown, based on the above video processing method, the spatial attention block SA module can be set at the adjacent position after the second RCAB module. After the second RCAB module extracts the second deep feature quantity, it sends the first deep feature quantity and the second deep feature quantity to the spatial attention block SA module. After the spatial attention block SA module receives the first deep feature quantity and the second deep feature quantity, it performs spatial feature extraction based on the first deep feature quantity and the second deep feature quantity to obtain the spatial feature quantity.

[0162] Step S304, perform superposition processing on the spatial feature quantity, the first deep feature quantity, and the second deep feature quantity, and use the obtained result as the output of the residual group RG.

[0163] As Figure 4 shown, based on the above video processing method, the spatial feature quantity, the first deep feature quantity, and the second deep feature quantity can be subjected to superposition processing, and the obtained result of the superposition processing is used as the output of the residual group RG. In some embodiments, as Figure 4As shown, one or more of the first depth feature quantity and the second depth feature quantity can also be subjected to channel scaling through the channel downscaling convolutional layer CONV③ and then superimposed, and then used as the output of the residual group RG.

[0164] It can be understood that in some embodiments, there may be multiple residual group RG modules in a residual network structure, such as Figure 3 As shown, the above process is Figure 3 the data processing process of the residual group RG-1 shown. After obtaining the output of the above residual group RG-1, the output can be input into the residual group RG-2 for further feature extraction processing. The specific processing process can refer to the processing process of the residual group RG-1 and will not be elaborated here.

[0165] Please refer to Figure 7 , Figure 7 which is a detailed flowchart of step S303 in the embodiment of the present invention. In some embodiments, step S301 includes:

[0166] Step S401, extracting intermediate spatial feature information based on the first depth feature quantity and the second depth feature quantity through a spatial attention convolutional layer, where the number of channels of the spatial attention convolutional layer is 1;

[0167] Step S402, extracting a spatial feature quantity based on the intermediate spatial feature information through an activation function ReLU.

[0168] Specifically, in combination with Figure 5 , the spatial attention module SA module includes a spatial attention convolutional layer CONV④ and a spatial attention activation layer ReLU①, where the number of output channels of the spatial attention convolutional layer is 1; after the SA module extracts the spatial feature quantity, it is added to the channel feature quantity in a superimposed manner. In the above manner, the spatial attention module SA module can be implemented in a simple way, that is, a structure composed of a convolutional layer and an activation layer. The convolutional layer and the activation layer are used to extract spatial features. Here, the convolutional layer CONV④ is different from an ordinary convolutional layer, and its number of output channels is 1. That is to say, the spatial features it extracts are independent of channels. The subsequent network structure will superimpose this spatial feature quantity on the original channel feature quantity to enhance / weaken the feature quantity at the corresponding spatial position, achieving the adjustment effect on the spatial features.

[0169] Through the above structural settings, in the spatial attention module SA module, the intermediate spatial feature information can be extracted based on the first depth feature quantity and the second depth feature quantity through the spatial attention convolutional layer CONV⑤, and then the activation function ReLU extracts the spatial feature quantity based on the intermediate spatial feature information, thereby enhancing / weakening the feature quantity at the corresponding spatial position and achieving the adjustment effect on the spatial features. Among them, the intermediate spatial feature information refers to the intermediate spatial information that has not been activated by the activation function.

[0170] It can be understood that the process of extracting the spatial feature quantity in the above steps S401 to S402 is the process of extracting the spatial feature quantity of the spatial attention module in the residual group RG-1. The spatial attention module SA module in the residual group RG-2 extracts the spatial feature quantity based on the depth feature quantity further extracted by other components in the residual group RG-1 and the residual group RG-2. The specific extraction process can refer to the above steps and will not be elaborated here.

[0171] Please refer to Figure 8 and Figure 9 , Figure 8 is a schematic diagram of the refinement process of step S301 in the embodiment of the present invention;

[0172] Figure 9 is a structural example diagram of the residual channel attention block RCAB in the embodiment of the present invention. Specifically, step S301 includes:

[0173] Step S403, input the shallow feature into the first convolutional layer in the first RCAB module for processing to obtain the first intermediate channel feature.

[0174] As Figure 9 shown, in the first RCAB module (the first residual channel attention block), in the direction from the input end to the output end, each component can be sequentially set to perform intermediate channel feature extraction in turn. Starting from the input end, the first convolutional layer CONV⑤ can be set first as the convolutional layer that first receives the input information in the first RCAB module. Based on the above video processing method, after the shallow feature input value enters the first RCAB module, it is first received by the first convolutional layer CONV⑤ and undergoes channel feature extraction processing to obtain the first intermediate channel feature. Among them, the first intermediate channel feature refers to the intermediate channel feature obtained by the first convolutional layer CONV⑤ performing the first channel feature extraction processing on the shallow feature.

[0175] Step S404, input the first intermediate channel feature into the activation function ReLU in the first RCAB module for processing to obtain the second intermediate channel feature.

[0176] As Figure 9As shown, after the first convolutional layer CONV⑤, the activation function ReLU② can be set to further process the output of the first convolutional layer CONV⑤. After the first convolutional layer CONV⑤ outputs the first intermediate channel features, the first intermediate channel features are sent to the activation function ReLU②, and the activation function ReLU② performs activation processing on the first intermediate channel features to obtain the activated second intermediate channel features. Among them, the second intermediate channel features refer to the intermediate channel features after being processed by the activation function ReLU②.

[0177] Step S405: Input the second intermediate channel features into the second convolutional layer in the first RCAB module for processing to obtain third intermediate channel features.

[0178] As Figure 9 shown, after the activation function ReLU②, the second convolutional layer CONV⑥ can be set to further process the output of the activation function ReLU②. After the activation function ReLU② outputs the second intermediate channel features, the second intermediate channel features are sent to the second convolutional layer CONV⑥, and the second convolutional layer CONV⑥ processes the second intermediate channel features to obtain third intermediate channel features. Among them, the third intermediate channel features refer to the intermediate channel features obtained after the second convolutional layer CONV⑥ further processes the second intermediate channel features.

[0179] Step S406: Input the third intermediate channel features into the global context module GCB module in the first RCAB module for processing to obtain fourth intermediate channel features.

[0180] As Figure 9 shown, after the second convolutional layer CONV⑥, the global context module GUB (Global Context Block) module can be set to further process the output of the second convolutional layer CONV⑥. After the second convolutional layer CONV⑥ outputs the third intermediate channel features, the third intermediate channel features are sent to the global context module GCB module, and the global context module GCB module processes the third intermediate channel features to obtain fourth intermediate channel features. Among them, the fourth intermediate channel features refer to the intermediate channel features obtained after the global context module GCB module further processes the third intermediate channel features.

[0181] Step S407: Superimpose the shallow features, the first intermediate channel features, the second intermediate channel features, the third intermediate channel features, and the fourth intermediate channel features to obtain the first depth feature quantity.

[0182] As Figure 9As shown, based on the above video processing method, the first intermediate channel feature, the second intermediate channel feature, the third intermediate channel feature, and the fourth intermediate channel feature can be superimposed, and the result obtained after the superimposition processing is used as the output of the first RCAB module.

[0183] It can be understood that in some embodiments, there may be multiple residual channel attention block RCAB modules in a residual group RG, such as Figure 4 As shown, the above process is Figure 4 As shown in the sorting process of the first RCAB. After obtaining the output of the first RCAB, the output can be input into the second RACB for further channel feature extraction processing. The specific processing process can refer to the above processing process of the first RCAB, which will not be elaborated here.

[0184] In the above video processing method, the shallow feature is input into the first convolutional layer in the first RCAB module for processing to obtain the first intermediate channel feature; the first intermediate channel feature is input into the activation function ReLU in the first RCAB module for processing to obtain the second intermediate channel feature; the second intermediate channel feature is input into the second convolutional layer in the first RCAB module for processing to obtain the third intermediate channel feature; the third intermediate channel feature is input into the global context module GCB module in the first RCAB module for processing to obtain the fourth intermediate channel feature; the shallow feature, the first intermediate channel feature, the second intermediate channel feature, the third intermediate channel feature, and the fourth intermediate channel feature are superimposed to obtain the first depth feature quantity. Through the above processing method, compared with the prior art, the network structure can be simplified, the network's perception of channel features can be increased, and better channel features can be obtained. When setting the global context module GCB, for the spatial compression part, dot multiplication can be used instead of averaging to better obtain channel features independent of space. In the embedded channel feature part, superimposition can be used instead of multiplication to improve the gradient disappearance problem in the network training process and facilitate network convergence.

[0185] Please refer to Figure 10 and Figure 11 , Figure 10 which is the detailed flowchart of step S406 in the embodiment of the present invention, Figure 11 which is the structural example diagram of the global context module GCB module in the embodiment of the present invention. In some embodiments, step S406 includes:

[0186] Step S501, input the third intermediate channel feature into the first convolutional layer and the softmax layer in the GCB module in sequence for processing, and perform dot multiplication processing on the obtained intermediate result and the third intermediate channel feature to obtain the compressed feature;

[0187] In step S502, the compressed features are sequentially input into the second convolutional layer, the activation function ReLU, and the third convolutional layer of the GCB module for processing, and the generated intermediate result and the compressed features are superimposed to obtain the fourth intermediate channel features.

[0188] As Figure 11 shown, in the global context module (GCB module), in the direction from the input end to the output end, a first convolutional layer (CONV⑦), a softmax function, a multiplication operator, a second convolutional layer (CONV⑧), an activation function ReLU②, a third convolutional layer (CONV⑨), and a superimposition operator can be sequentially set to process the channel features. Based on the above structural settings and the above video processing algorithm, after the convolutional layer CONV⑥ outputs the third intermediate channel features, the third intermediate channel features can be sequentially input into the first convolutional layer CONV⑦ and the softmax layer in the GCB module for processing, and the intermediate results obtained by their respective processing are dot-multiplied with the third intermediate channel features to obtain compressed features. Then, the compressed features are sequentially input into the second convolutional layer CONV⑧, the activation function ReLU②, and the third convolutional layer CONV⑨ of the GCB module for processing, and the intermediate results generated by them and the compressed features are superimposed to obtain the fourth intermediate channel features and output them.

[0189] Through the processing method of the above global context module (GCB module), for the spatial compression part, multiplication can be used instead of averaging to better obtain channel features independent of space. In the part of embedding channel features, superimposition can be used instead of multiplication to improve the problem of gradient disappearance in the network training process, which is beneficial to network convergence.

[0190] In addition, the present invention also provides a video processing device based on a super-resolution network.

[0191] Among them, the video processing device based on the super-resolution network of the present invention includes:

[0192] A video acquisition module, configured to acquire a low-resolution video to be processed;

[0193] A super-resolution processing module, configured to input each low-resolution image frame of the low-resolution video into a trained super-resolution network for processing to obtain a super-resolution image frame corresponding to each low-resolution image frame;

[0194] Among them, the super-resolution processing module includes:

[0195] A shallow feature extraction sub-module, configured to extract shallow features of each low-resolution image frame through a shallow feature extraction convolutional layer;

[0196] A deep feature extraction sub-module for extracting the depth features of the shallow features through a residual network structure RIR module;

[0197] The deep feature extraction module is also used to extract the depth features of the shallow features through a preset residual group RG;

[0198] The deep feature extraction module is also used to perform deep feature extraction through stacked residual channel attention blocks RCAB and a spatial attention module SA module in the residual group RG;

[0199] Among them, in the RCAB, channel features are extracted through a global context module GCB;

[0200] A reconstruction sub-module for reconstructing based on the depth features through a reconstruction convolutional layer to obtain super-resolution image frames of each low-resolution image frame;

[0201] A synthesis module for synthesizing all the super-resolution image frames into a super-resolution video.

[0202] Among them, the specific implementation manners of the video processing device based on the super-resolution network of the present invention can refer to the various embodiments of the video processing method based on the super-resolution network of the present invention, which will not be elaborated here.

[0203] In addition, an embodiment of the present invention also proposes a computer-readable storage medium.

[0204] A computer program is stored on the computer-readable storage medium of the present invention, and when the program is executed by a processor, the steps of the video processing method described above are implemented.

[0205] Among them, the method implemented when the video processing program running on the processor is executed can refer to the various embodiments of the video processing method of the present invention, which will not be elaborated here.

[0206] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or system. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or system including that element.

[0207] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0208] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0209] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A video processing method based on a super-resolution network, characterized in that, the method includes: Step S10, obtaining a low-resolution video to be processed; Step S20, inputting each low-resolution image frame of the low-resolution video into a trained super-resolution network for processing to obtain a super-resolution image frame corresponding to each low-resolution image frame; wherein, the step S20 includes: Step S21, extracting shallow features of each low-resolution image frame through a shallow feature extraction convolutional layer; Step S22, extracting deep features of the shallow features through a residual network structure RIR module; wherein, the step S22 includes: forwarding the first low-frequency information at the input end of the RIR module to the output end of the RIR module through a long skip connection LSC; forwarding the second low-frequency information in the input end of the residual group RG to the output end of the RG through a short skip connection SSC; Step S221, extracting deep features of the shallow features through two stacked residual groups RG; wherein in each residual group RG, deep feature extraction is performed through two stacked RCAB modules, one SA module, and one convolutional layer; wherein, the step S221 includes: Step S301, inputting the shallow features into the first convolutional layer in the first RCAB module for processing to obtain a first intermediate channel feature; inputting the first intermediate channel feature into the activation function ReLU in the first RCAB module for processing to obtain a second intermediate channel feature; inputting the second intermediate channel feature into the second convolutional layer in the first RCAB module for processing to obtain a third intermediate channel feature; inputting the third intermediate channel feature into the global context module GCB module in the first RCAB module for processing to obtain a fourth intermediate channel feature; stacking the shallow features, the first intermediate channel feature, the second intermediate channel feature, the third intermediate channel feature, and the fourth intermediate channel feature to obtain a first deep feature quantity; Step S302, performing deep feature extraction on the first deep feature quantity through a second RCAB module to obtain a second deep feature quantity; Step S303, performing spatial feature extraction on the first deep feature quantity and the second deep feature quantity through the SA module to obtain a spatial feature quantity; Step S304, stacking and processing the spatial feature quantity, the first deep feature quantity, and the second deep feature quantity, and taking the obtained result as the output of the residual group RG; wherein, in the RCAB module, channel features are extracted through the global context module GCB; Step S23, reconstructing based on the deep features through a reconstruction convolutional layer to obtain a super-resolution image frame for each low-resolution image frame; Step S30, synthesizing all the super-resolution image frames into a super-resolution video; wherein inputting the third intermediate channel feature into the global context module GCB module in the first RCAB module for processing to obtain a fourth intermediate channel feature includes: Step S501: Input the third intermediate channel feature into the first convolutional layer and the softmax layer in the GCB module for processing in sequence, perform dot product processing on the obtained intermediate result and the third intermediate channel feature to obtain a compressed feature; Step S502: Input the compressed feature into the second convolutional layer, the activation function ReLU, and the third convolutional layer in the GCB module for processing in sequence, and superimpose the generated intermediate result and the compressed feature to obtain the fourth intermediate channel feature.

2. The video processing method based on a super-resolution network according to claim 1, characterized in that, the spatial attention module SA module includes a spatial attention convolutional layer and a spatial attention activation layer, wherein the output channel number of the spatial attention convolutional layer is 1; after the SA module extracts the spatial feature quantity, it is added to the channel feature quantity in a superimposed manner.

3. The video processing method based on a super-resolution network according to claim 1, characterized in that, in step S303, the SA module performs spatial feature extraction based on the first depth feature quantity and the second depth feature quantity, and the obtained spatial feature quantity includes: Step S401: Extract intermediate spatial feature information based on the first depth feature quantity and the second depth feature quantity through the spatial attention convolutional layer, wherein the channel number of the spatial attention convolutional layer is 1; Step S402: Extract the spatial feature quantity based on the intermediate spatial feature information through the activation function ReLU.

4. A video processing device based on a super-resolution network, characterized in that, the device includes: a video acquisition module for acquiring a low-resolution video to be processed; a super-resolution processing module for inputting each low-resolution image frame of the low-resolution video into a trained super-resolution network for processing to obtain a super-resolution image frame corresponding to each low-resolution image frame; wherein, the super-resolution processing module includes: a shallow feature extraction sub-module for extracting shallow features of each low-resolution image frame through a shallow feature extraction convolutional layer; a deep feature extraction sub-module for extracting deep features of the shallow features through a residual network structure RIR module; the deep feature extraction sub-module is further configured to forward the first low-frequency information at the input end of the RIR module to the output end of the RIR module through a long skip connection LSC; forward the second low-frequency information in the input end of the residual group RG to the output end of the RG through a short skip connection SSC; extract deep features of the shallow features through two superimposed residual groups RG; wherein in each residual group RG, deep feature extraction is performed through two superimposed RCAB modules, one SA module, and one convolutional layer. The steps of extracting the depth features of the shallow features through two stacked residual groups RG include: inputting the shallow features into the first convolutional layer in the first RCAB module for processing to obtain first intermediate channel features; inputting the first intermediate channel features into the activation function ReLU in the first RCAB module for processing to obtain second intermediate channel features; inputting the second intermediate channel features into the second convolutional layer in the first RCAB module for processing to obtain third intermediate channel features; inputting the third intermediate channel features into the global context module GCB module in the first RCAB module for processing to obtain fourth intermediate channel features; stacking the shallow features, the first intermediate channel features, the second intermediate channel features, the third intermediate channel features, and the fourth intermediate channel features to obtain a first depth feature quantity; performing depth feature extraction on the first depth feature quantity through a second RCAB module to obtain a second depth feature quantity; performing spatial feature extraction on the first depth feature quantity and the second depth feature quantity through the SA module to obtain a spatial feature quantity; performing a stacking process on the spatial feature quantity, the first depth feature quantity, and the second depth feature quantity, and using the obtained result as the output of the residual group RG; wherein, in the RCAB module, channel features are extracted through the global context module GCB; A reconstruction sub-module, configured to perform reconstruction based on the depth features through a reconstruction convolutional layer to obtain super-resolution image frames of each low-resolution image frame; A synthesis module, configured to synthesize all the super-resolution image frames into a super-resolution video; Wherein, inputting the third intermediate channel features into the global context module GCB module in the first RCAB module for processing to obtain fourth intermediate channel features includes: Sequentially inputting the third intermediate channel features into the first convolutional layer and the softmax layer in the GCB module for processing, performing a dot product process on the obtained intermediate result and the third intermediate channel features to obtain compressed features; Sequentially inputting the compressed features into the second convolutional layer, the activation function ReLU, and the third convolutional layer in the GCB module for processing, and stacking the generated intermediate result and the compressed features to obtain the fourth intermediate channel features.

5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Image super-resolution reconstruction method and system

    CN112862689A

  • Mine image super-resolution reconstruction method and system based on multi-scale residual network

    CN113592718A