Method, Enhancement Device, Electronic Device, and Medium for Compressed Video Frame Enhancement
Through the combination of multi-layer structure processing and parallel bidirectional network, video frames are enhanced, which solves the problem of distortion during video compression and achieves higher quality video frame recovery.
Patent Information
- Application Number
- CN202210375275.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-11
AI Technical Summary
Distortion will occur during video compression, such as block effect, ringing effect and blurred image quality, affecting the user's viewing experience.
The compressed video frame is processed using a multi-layer structure. By obtaining the frame features of the previous layer, the frame features of the current layer, the forward frame features and the backward frame features of the current layer, the multi-layer processing is carried out according to the preset parallel bidirectional network, and finally, the details and quality of the video frame are enhanced based on the residual block fusion screening results.
Provide a more sufficient basis through the recovery of complementary information between video frames, use information of different contribution degrees to integrate, effectively utilize frame information, alleviate the impact of distortion during compression, and improve the details and quality of video frames.
Smart Images

Figure CN114979664B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video technology, and in particular to a method for enhancing compressed video frames, an enhancement device, an electronic device, and a medium. Background Art
[0002] Generally, with the improvement of various demands for the use of product devices by people, Internet video streams are becoming increasingly popular, and the expectations and demands for high-quality and high-resolution videos are also growing. When using a product device to play a video, users often hope to maintain both the timeliness of the video played by the product device and the quality of the video played by the product device.
[0003] Generally, the amount of data of digital videos is extremely large, requiring high storage space and being inconvenient to transmit in a network with limited bandwidth. Usually, there is a high correlation between video data, that is, redundant information. To break through the limitations of storage space and transmission bandwidth, video compression technology has been proposed and developed rapidly.
[0004] Currently, the core of this technology is to remove redundant information in video data and extract a compact video data representation to achieve the purpose of data compression, thereby reducing the bit rate and facilitating the storage and transmission of high-quality and high-resolution videos. However, videos will inevitably be distorted during the compression process, such as block effects, ringing effects, and blurred image quality, seriously affecting the user's viewing experience. Summary of the Invention
[0005] To solve the above technical problems, the technical solution adopted in the first aspect of the present application is to provide a method for enhancing compressed video frames, the method including: processing a compressed video frame using a multi-layer structure to obtain at least the frame features of the previous layer and the frame features of the current layer; obtaining the frame features of the previous layer, the frame features of the current layer, the forward frame features of the current layer, and the backward frame features of the current layer to obtain a combined input; performing multi-layer processing on the combined input according to a preset parallel bidirectional network to obtain multiple screening results that meet the fusion conditions; and fusing the multiple screening results based on a residual block to obtain the enhanced features corresponding to the current layer.
[0006] To solve the above technical problems, the technical solution adopted in the second aspect of the present application is to provide an enhancement device, the enhancement device including: a processing module for processing a compressed video frame using a multi-layer structure to obtain at least the frame features of the previous layer and the frame features of the current layer; an obtaining module for obtaining the frame features of the previous layer, the frame features of the current layer, the forward frame features of the current layer, and the backward frame features of the current layer to obtain a combined input; the processing module is further configured to perform multi-layer processing on the combined input according to a preset parallel bidirectional network to obtain multiple screening results that meet the fusion conditions; and a fusion module for fusing the multiple screening results based on a residual block to obtain the enhanced features corresponding to the current layer.
[0007] To solve the above technical problems, the technical solution adopted in the third aspect of the present application is to provide an electronic device, which includes: a processor and a memory. A computer program is stored in the memory, and the processor is configured to execute the computer program to implement the method of the first aspect of the present application.
[0008] To solve the above technical problems, the technical solution adopted in the fourth aspect of the present application is to provide a computer-readable storage medium, which stores a computer program, and the computer program can implement the method of the first aspect of the present application when executed by a processor.
[0009] The beneficial effects of the present application are as follows: In order to achieve better video frame enhancement, the present application performs multi-layer processing on the combined input according to a preset parallel bidirectional network, so that the complementary information between video frames can provide a more sufficient basis for the restoration of the current frame information. And by fusing the screening results that meet the fusion conditions, and using the differences in different contribution degrees to enhance the details and quality of the video frames, it promotes the effective utilization of the frame information and alleviates the influence brought by other adverse factors. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0011] Figure 1 It is a schematic flowchart of an embodiment of the method for enhancing compressed video frames of the present application;
[0012] Figure 2 It is of the present application Figure 1 A schematic flowchart of a specific embodiment of step S12 in;
[0013] Figure 3 It is of the present application Figure 1 A schematic flowchart of a specific embodiment of step S13 in;
[0014] Figure 4 It is of the present application Figure 1 A schematic diagram of the overall framework structure of a specific embodiment of;
[0015] Figure 5 It is of the present application Figure 3 A schematic flowchart of a specific embodiment of step S31 in;
[0016] Figure 6 It is of the present application Figure 5Flow chart of a specific embodiment of step S42;
[0017] Figure 7 This application Figure 5 Flow chart of a specific embodiment of step S43;
[0018] Figure 8 This application Figure 5 Flow chart of a specific embodiment of step S44;
[0019] Figure 9 Flow chart of a specific embodiment of the dual attention fusion strategy of this application;
[0020] Figure 10 This application Figure 9 Flow chart of a specific embodiment;
[0021] Figure 11 Structural schematic block diagram of an enhanced device embodiment of this application;
[0022] Figure 12 Structural schematic block diagram of an electronic device embodiment of this application;
[0023] Figure 13 Circuit schematic block diagram of a computer-readable storage medium embodiment of this application. Detailed implementation manners
[0024] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.
[0025] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0026] It should also be understood that the terms used in this specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in this specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0027] It should be further understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0028] As used in this specification and the appended claims, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0029] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0030] To illustrate the technical solutions of the present application, the following specific embodiments are used to illustrate that the present application provides a method for enhancing compressed video frames. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the method for enhancing compressed video frames of the present application. The method specifically includes the following steps:
[0031] S11: Process the compressed video frame using a multi-layer structure to obtain at least the frame features of the previous layer and the frame features of the current layer;
[0032] Generally, for the processing of compressed video, a recurrent neural network is often used. The recurrent neural network of the present application includes at least a multi-layer structure, so a multi-layer structure can be used to perform preliminary processing on the compressed video frame.
[0033] Generally speaking, a compressed video is composed of video frames, and the video frame sequence can be synthesized into a video. In order to make full use of the information between video frames, when processing the compressed video frame using a multi-layer structure, at least the frame features of the previous layer and the frame features of the current layer can be obtained.
[0034] Since each layer of the multi-layer structure corresponds to its own frame feature respectively, if the multi-layer structure has N1 layers and the N1-layer structure is connected in sequence, after processing the compressed video frames, there are the frame features corresponding to the i-th layer and the frame features corresponding to the i+1-th layer in the N1 layer. If the frame features corresponding to the i+1-th layer are used as the frame features of the current layer, then the frame features corresponding to the i-th layer are the frame features of the previous layer; where N1 is an integer not less than 2, and i is an integer greater than 0 and not greater than N1-1; for example, when N1 is 2 as above, after processing the compressed video frames, there are at least the frame features corresponding to the first layer and the frame features corresponding to the second layer, where if the frame features corresponding to the second layer are used as the frame features of the current layer, then the frame features corresponding to the first layer are used as the frame features of the previous layer.
[0035] S12: Obtain the frame features of the previous layer, the frame features of the current layer, the forward frame features of the current layer, and the backward frame features of the current layer to obtain a combined input;
[0036] For the current layer, there are many and different sources of video frame information that can be obtained. For the first layer, at least the features of the current frame layer and the backward frame features of the current layer can be obtained; for the last layer, at least the frame features of the current layer and the forward frame features of the current layer can be obtained.
[0037] Here, the forward frame features of the current layer and the backward frame features of the current layer are described. When the multi-layer structure includes N2 layers and the N2-layer structure is connected in sequence, after processing the compressed video frames, there are frame features corresponding to each layer in the N2 layer. If the j+1-th layer is used as the current frame, then the frame features corresponding to the j+1-th layer are used as the frame features of the current layer, and the frame features corresponding to the j-th layer are the previous layer frame features of the j+1-th layer. The frame features corresponding to the first layer to the j-th layer in the N2 layer are the forward frame features of the j+1-th layer, and the frame features corresponding to the j+2-th layer to the N-th layer can be used as the backward frame features; where N2 is an integer not less than 3, and i is an integer greater than 1 and not greater than N2-1.
[0038] For easy understanding, a specific example is given here. If N2 is 4, the multi-layer structure can be the first layer, the second layer, the third layer, and the fourth layer connected in sequence. After processing the compressed video frames, there are at least the frame features corresponding to the first layer, the frame features corresponding to the second layer, the frame features corresponding to the third layer, and the frame features corresponding to the fourth layer; if the frame features corresponding to the third layer are used as the frame features of the current layer, then the frame features corresponding to the second layer can be used as the frame features of the previous layer, and the frame features corresponding to the first layer and the second layer can be used as the forward frame features, and the frame features corresponding to the fourth layer can be used as the backward frame features.
[0039] In order to make better use of the information resources between videos, for the video frame features in the middle layer, specifically, the frame features of the previous layer, the frame features of the current layer, the forward frame features of the current layer, and the backward frame features of the current layer can be obtained, and then these features can be combined to form a joint input.
[0040] S13: performing multi-layer processing on the joint input according to a preset parallel bidirectional network to obtain multiple screening results that meet the fusion conditions;
[0041] Specifically, in order to process the joint input, the present application also designs a preset parallel bidirectional network. The preset parallel bidirectional network has a sufficient processing layer structure, and the parameters between its video frames are also shared. Therefore, according to the preset parallel bidirectional network, the joint input is processed in multiple layers so that each video frame can make full use of the information flow of the remaining video frames, thereby achieving the purpose of enhancing frame details and quality.
[0042] Furthermore, the results obtained after multi-layer processing can be screened again, and different positions with different contributions can be selected. This can promote the effective use of information, integrate the characteristics of all stages in the preset bidirectional parallel network, and obtain multiple screening results that meet the fusion conditions.
[0043] S14: Based on the residual block, multiple screening results are fused to obtain the enhanced features corresponding to the current layer.
[0044] Usually, the added new network layer will at least not make the effect worse than the original one, so the application effect of the model can be improved more stably by deepening the number of layers. Residual blocks are usually set so that those with good effects are retained and those with poor effects are skipped to maintain the effective gradient of the newly added network layer.
[0045] Specifically, based on the residual block, multiple screening results are fused to obtain the enhanced features corresponding to the current layer. Finally, these enhanced features are reconstructed to obtain a clear image, thereby preserving the fidelity of the original video frame as much as possible.
[0046] Therefore, in order to achieve better video frame enhancement, the present application performs multi-layer processing on the joint input according to a preset parallel bidirectional network, so that the complementary information between video frames can provide a more sufficient basis for the recovery of the current frame information, and by fusing the screening results that meet the fusion conditions, the different contribution degrees are used to enhance the details and quality of the video frames, thereby promoting the effective use of frame information and alleviating the impact of other unfavorable factors.
[0047] Furthermore, the steps of obtaining the previous layer frame features, the current layer frame features, the current layer forward frame features, and the current layer backward frame features to obtain the joint input are as follows: Figure 2 ,Figure 2 This is an application Figure 1 and is a schematic flowchart of a specific embodiment of step S12 in the application, which specifically includes the following steps:
[0048] S21: Extract the features of the previous layer frame, the features of the current layer frame, the forward frame features of the current layer, and the backward frame features of the current layer;
[0049] Specifically, for the extraction of video frames, residual blocks can often be used. For example, this application can use a series of residual blocks with stride convolutions to achieve feature extraction. Since residual blocks can extract feature information at different gradients, the extracted features are often relatively comprehensive and their classification is also relatively reasonable.
[0050] Therefore, through a series of residual blocks with stride convolutions, the features of the previous layer frame, the features of the current layer frame, the forward frame features of the current layer, and the backward frame features of the current layer can be extracted to retain the details and quality of the video frame in different channel dimensions.
[0051] S22: Concatenate the features of the previous layer frame, the features of the current layer frame, the forward frame features of the current layer, and the backward frame features of the current layer according to their respective corresponding channel dimensions to obtain a joint input.
[0052] Among them, the preset parallel bidirectional network includes at least four layers of structures, where each layer of structure includes at least forward and backward branches that propagate independently in parallel and a preset attention strategy. The forward and backward branches that propagate independently in parallel include at least a backward propagation structure and a forward propagation structure.
[0053] Specifically, the four layers of structures include at least four recurrent networks connected in sequence, such as the first recurrent network, the second recurrent network, the third recurrent network, and the fourth recurrent network connected in sequence.
[0054] Among them, for the first recurrent network, the adjacent front and rear frames are used as the reference frame sources of the current frame at the same time; for the second recurrent network, the second frames in the front and rear directions are used as the reference frame sources of the current frame at the same time; for the third recurrent network, the third frames in the front and rear directions are used as the reference frame sources of the current frame at the same time; for the fourth recurrent network, the fourth frames in the front and rear directions are used as the reference frame sources of the current frame at the same time.
[0055] In this way, the reference frames of different layer recurrent network structures in the four-layer recurrent network are all video frames with different frame rates, that is, the reference frames of each layer of recurrent network are all video frames with different frame rates, and the reference frames of the forward and backward propagation of each layer of recurrent network are video frames with the same frame rate in opposite directions. Specific embodiments of each layer of recurrent network will be described in detail later.
[0056] Further, for the step of performing multi-layer processing on the combined input according to a preset parallel bidirectional network to obtain multiple screening results that meet the fusion conditions, please refer to Figure 3 , Figure 3 is a schematic flowchart of a specific embodiment of step S13 in this application Figure 1 , which specifically includes the following steps:
[0057] S31: Process the combined input according to the backpropagation structure to obtain a first result;
[0058] Since the forward and backward branches of parallel independent propagation contain two key points: one is the number of forward and backward reference frames. Assume the number of forward reference frames is M and the number of backward reference frames is N; the other is the source of the reference frame, that is, adjacent frames or cross-multiple-frame references, and the number of cross frames can be different. When cross-frame reference is set, the number of forward-propagation cross frames is R1 and the number of backward-propagation cross frames is R2, where forward propagation is also called backward propagation.
[0059] For example, there are three frames, the first frame, the second frame, and the third frame, and each frame corresponds to four layers. Based on the combined input of the second frame in the second layer, first process the frame features of the current layer (that is, the frame features of the second layer of the second frame), and then process the backward propagation of the previous layer (that is, the frame features of the first layer of the second frame). Specifically, it can also be like M forward reference frames and R1 forward-propagation cross frames based on cross-frame reference. Finally, process the frame features of the third layer and the fourth layer.
[0060] That is, according to the backpropagation structure, process the combined input in a similar manner as the second layer, the first layer, the third layer, and the fourth layer to obtain a first result. Specifically, according to the backpropagation structure, based on R1 forward-propagation cross frames of cross-frame reference, perform forward processing on M forward reference frames, and then a first result can be obtained.
[0061] S32: Process the combined input according to the forward propagation structure to obtain a second result;
[0062] That is, according to the normal processing mode, its forward propagation is processed in sequence as the first layer, the second layer, the third layer, and the fourth layer. Process the combined input according to the forward propagation structure to obtain a first result. Specifically, according to the forward propagation structure, based on R2 forward-backward propagation cross frames of cross-frame reference, perform backward processing on N forward reference frames, and then a second result can be obtained.
[0063] S33: Based on a preset attention strategy, screen the first result and the second result to obtain multiple screening results that meet the fusion conditions.
[0064] For each layer of the parallel loop structure, mainly for each frame of the video sequence, forward and backward feature propagation is respectively performed to utilize complementary information from the remaining videos. In order to screen and fuse the first result and the second result, the present application also designs a preset attention strategy, and then fuses the forward and backward propagation results, and the result will be passed to the next layer structure to make more full use of information.
[0065] Since both the first result and the second result may include multiple frame features, based on the preset attention strategy, by screening the first result and the second result, specifically, for example, by selecting multiple frame features through the frame features with weights that meet the preset gradient or preset threshold range, multiple screening results that meet the fusion conditions can be obtained, thereby promoting the effective utilization of information and alleviating the problem that the enhancement effect is affected due to insufficient accuracy in the processing process.
[0066] Among them, please refer to Figure 4 , Figure 4 which is Figure 1 the overall framework structure schematic diagram of a specific embodiment of the present application.
[0067] The network mainly includes optical flow estimation and an enhancement network implemented using a loop structure. Among them, the optical flow estimation is implemented using the SPyNet network of existing technologies. The enhancement network mainly includes three parts: feature extraction, a bidirectional parallel network based on an attention aggregation strategy, and frame reconstruction. The feature extraction stage is implemented using a series of residual blocks with stride convolutions, and the frame reconstruction stage will reconstruct the enhanced video frames by integrating the features of all stages in the bidirectional parallel network.
[0068] Denote the currently to-be-enhanced video frame as X i , and its forward and backward reference frames are respectively denoted as X f and X b , with the spatial size of H×W and the number of channels of C. Since the front and back adjacent frames are strongly correlated with the current frame, the front and back adjacent frames can be used as the reference frame sources of the current frame simultaneously, that is, M = N = 1, R1 = R2 = 0, that is, X f = X i-1 and X b = X i+1 . The overall structure is as Figure 4 shown, where the parallel bidirectional network is the network designed by the present application and includes four layers of structure.
[0069] Among them, each layer of structure can be divided into two parts: one is the forward and backward branches that propagate independently in parallel, as shown by the white dotted-line module B and the dark gray dotted-line module F in the figure and the corresponding color arrows connected to them. The other is to adopt an attention strategy to fuse the forward and backward propagation results, and the result will be passed to the next layer structure.
[0070] As Figure 4As shown, taking the backpropagation structure of the second-layer loop structure as an example, its specific structure is shown in the backpropagation module in the upper right corner of the above figure. The forward propagation is similar, except for the reference frame. The backpropagation structure at least includes: a pre-alignment stage, a deformation alignment stage, and a feature enhancement stage.
[0071] For video processing tasks, the utilization of inter-frame information is very important. On the one hand, the complementary information between video frames can provide a more sufficient basis for the restoration of the current frame information. On the other hand, it improves the temporal coherence of the processed video. Therefore, this application uses a loop structure to process video sequences, enabling each video frame to make full use of the information flow of the entire video frame, making up for the problem of insufficient self-information, and thus enhancing the details and quality of the video frame.
[0072] Furthermore, for the step of processing the joint input according to the backpropagation structure to obtain the first result, please refer to Figure 5 , Figure 5 is a schematic flowchart of a specific embodiment of step S31 in this application Figure 3 and specifically includes the following steps:
[0073] S41: Obtain the frame features of the current layer and the reference frame features;
[0074] S42: Through a warping operation, pre-align the reference frame features to the frame features of the current layer to obtain a pre-alignment result;
[0075] Specifically, in the pre-alignment stage, through the warping operation, the reference frame features can be pre-aligned to the frame features of the current layer, and the reference frame features are initially aligned to the frame features of the current layer to obtain a pre-alignment result.
[0076] S43: Combine the pre-alignment result and the frame features of the current layer, and perform deformable convolution on the reference frame features to obtain a deformation alignment result from the reference frame features to the frame features of the current layer;
[0077] Specifically, in the deformation alignment stage, based on the input of the current layer, combine the pre-alignment result and the frame features of the current layer, and perform deformable convolution on the reference frame features to obtain a deformation alignment result from the reference frame features to the frame features of the current layer.
[0078] S44: Based on a preset residual block, fuse and enhance the enhanced features of the previous layer, the frame features of the current layer, and the deformation alignment result to obtain the enhanced features of the current layer in the backpropagation as the first result.
[0079] Furthermore, for the step of pre-aligning the reference frame features to the frame features of the current layer through a warping operation to obtain a pre-alignment result, please refer to Figure 6 , Figure 6 is this application Figure 5Schematic diagram of the process of a specific embodiment in step S42, specifically including the following steps:
[0080] S51: Obtain the frame of the current layer and the reference frame;
[0081] In order to obtain the pre-aligned reference frame features with the frame features of the current layer, it is first necessary to obtain the frame of the current frame in the current layer and the reference frame, which are used as parameters for optical flow calculation, so as to calculate the optical flow information from the frame of the current layer to the reference frame, and then process to obtain the pre-aligned reference frame features with the frame features of the current layer.
[0082] S52: According to the optical flow information from the frame of the current layer to the reference frame, pre-align the reference frame features with the frame features of the current layer through a warping operation to obtain a pre-alignment result.
[0083] Specifically, according to the optical flow information from the frame of the current layer to the reference frame through a warping operation, the reference frame features are initially aligned to the frame features of the current layer to obtain a pre-alignment result.
[0084] Furthermore, for the step of performing deformable convolution on the reference frame features by combining the pre-alignment result and the current frame features to obtain the deformable alignment result from the reference frame features to the current frame features, please refer to Figure 7 , Figure 7 This is an application Figure 5 Schematic diagram of the process of a specific embodiment in step S43, specifically including the following steps:
[0085] S61: Based on the input of the current frame in the current layer, through multiple convolutional layers, combine the pre-alignment result and the frame of the current layer to generate a residual bias and a mask;
[0086] Specifically, for deformable alignment, a residual bias and a mask are generated by combining the pre-alignment result and the input of the current frame in the current layer through multiple convolutional layers.
[0087] S62: Add the residual bias and the optical flow information to obtain a first bias;
[0088] Specifically, add the residual bias to the above-mentioned optical flow information to obtain a first bias, where the above-mentioned optical flow information at least includes the optical flow information from the frame of the current layer to the reference frame.
[0089] S63: Use the first bias and the mask as the parameters of deformable convolution to perform deformable convolution on the reference frame features to obtain the deformable alignment result from the reference frame features to the frame features of the current layer.
[0090] Specifically, use the first bias and the mask as the parameters of deformable convolution to perform deformable convolution on the reference frame features to obtain the alignment result from the reference frame features to the current frame.
[0091] Further, based on a preset residual block, the enhanced feature of the previous layer, the frame feature of the current layer, and the deformation alignment result are fused and enhanced to obtain the enhanced feature of the current layer in backpropagation as the first result. Please refer to Figure 8 , Figure 8 is a schematic flowchart of a specific embodiment of step S44 in this application Figure 5 , which specifically includes the following steps:
[0092] S71: Based on the input of the current frame in the current layer, the enhanced feature of the previous layer, the frame feature of the current layer, and the deformation alignment result are input into a preset residual block;
[0093] Specifically, in the feature enhancement stage, combining the enhanced result of the backpropagation (opposite to the forward propagation) in the previous parallel bidirectional network layer, the input of the current frame in the current layer, and the alignment result of the reference frame feature, the enhanced feature of the previous layer, the frame feature of the current layer, and the deformation alignment result are input into a preset residual block, where the preset residual block can be multiple or a series.
[0094] S72: Through the preset residual block, the enhanced feature of the previous layer, the frame feature of the current layer, and the deformation alignment result are fused and enhanced to obtain the enhanced feature of the current layer in backpropagation, and the enhanced feature of the current layer is propagated to the next frame.
[0095] Specifically, through the preset residual block, the enhanced feature of the previous layer, the frame feature of the current layer, and the deformation alignment result are fused and enhanced to obtain the enhanced result of the current layer in this direction, and it is propagated to the next frame.
[0096] In addition, in addition to directly referring to the information of the front and rear adjacent frames, the information of the remaining frames can also provide additional supplementary information. Therefore, M = N = 1 can be set, and the R value is set according to the number of layers. Denote the current video frame to be enhanced as X i , and its forward and backward reference frames are respectively denoted as X f and X b , with a spatial size of H×W and a channel number of C.
[0097] For the first layer of the recurrent network, both the front and rear adjacent frames are used as the reference frame sources of the current frame, R1 = R2 = 0, that is, X f = X i-1 and X b = X i+1 . For the second layer of the recurrent network, the second frames in the front and rear directions are used as the reference frame sources of the current frame, R1 = R2 = 1, that is, X f = X i-2 and X b = X i+2。For the third - layer recurrent network, the third frames in the front - and - back directions are simultaneously used as the source of reference frames for the current frame, R1 = R2 = 2, that is, X f = X i-3 and X b = X i+3 。For the fourth - layer recurrent network, the fourth frames in the front - and - back directions are simultaneously used as the source of reference frames for the current frame, R1 = R2 = 3, that is, X f= X i-4 and X b = X i+5 。
[0098] That is to say, the reference frames propagated by each layer of the recurrent structure are video frames with different frame rates. And for each layer, the forward - and - backward propagated reference frames are video frames with the same frame rate in opposite directions.
[0099] Among them, the step of processing the joint input according to the backward - propagation structure to obtain the second result may specifically include:
[0100] Performing backward processing on the enhanced features of the previous layer to obtain one of the second results; performing backward processing on the frame features of the current layer to obtain the other of the second results; where the second result at least includes one of the second results and the other of the second results.
[0101] Furthermore, each layer of the parallel recurrent structure mentioned above mainly performs forward - and - backward feature propagation for each frame of the video sequence respectively to utilize complementary information from the remaining videos. To fully integrate this information into the current - frame features, the preset attention strategy at least includes a channel - attention branch and a spatial - attention branch.
[0102] Among them, for the step of screening the first result and the second result based on the preset attention strategy to obtain multiple screening results that meet the fusion conditions, please refer to Figure 9 and Figure 10 , Figure 9 is a schematic flowchart of a specific embodiment of the dual - attention fusion strategy of this application, Figure 10 is a schematic flowchart of a specific embodiment in this application Figure 9 . As shown in Figure 9 , taking the shallow - layer features, backward - propagation features, forward - propagation features in the feature - extraction stage, and the enhanced results of the previous parallel - recurrent - structure layer as inputs, first splice all the inputs along the channel dimension to obtain a joint input, which specifically includes the following steps:
[0103] S81: Based on the channel attention branch, perform average pooling on the joint input to obtain the average value of each feature channel; use two convolutions and a preset layer to train and learn the average value to obtain the weight of each feature channel; perform dimension expansion through the weight of each feature channel to obtain the first weight with the same dimension as the joint input.
[0104] Specifically, perform average pooling on the joint input to obtain the average value of each feature channel; then use two convolutional layers + sigmoid layer to learn this to obtain the weight of each feature channel; finally, through dimension expansion operation, obtain the first weight with the same dimension as the joint input.
[0105] S82: Based on the spatial attention branch, use convolutional layers of 1×1 and 3×3 respectively to train and learn the joint input to obtain the second weight of each spatial position of the joint input.
[0106] S83: Concatenate the first weight and the second weight along the channel dimension; call a convolution to fuse the first weight and the second weight to obtain a weight set with the same dimension as the joint input.
[0107] S84: According to the weight set, screen the first result and the second result in the way of point-by-point multiplication of the joint input to obtain multiple screening results that meet the fusion conditions.
[0108] In addition, the present application also provides an enhancement device. Please refer to Figure 11 , Figure 11 which is the structural schematic block diagram of the enhancement device embodiment of the present application. The enhancement device 60 includes: an acquisition module 61, a processing module 62, and a fusion module 63.
[0109] The processing module 62 is used to process the compressed video frame by using a multi-layer structure to obtain at least the previous layer frame feature and the current layer frame feature.
[0110] The acquisition module 61 is used to acquire the previous layer frame feature, the current layer frame feature, the forward frame feature of the current layer, and the backward frame feature of the current layer to obtain a joint input.
[0111] The processing module 62 is further used to perform multi-layer processing on the joint input according to a preset parallel bidirectional network to obtain multiple screening results that meet the fusion conditions.
[0112] The fusion module 63 is used to fuse multiple screening results based on a residual block to obtain the enhanced feature corresponding to the current layer.
[0113] Therefore, in order to achieve better video frame enhancement, the present application performs multi-layer processing on the combined input according to a preset parallel bidirectional network, enabling more sufficient basis to be provided for the restoration of the current frame information through the complementary information between video frames. Moreover, by fusing the screening results that meet the fusion conditions and utilizing the differences in different contribution degrees, the details and quality of the video frames are enhanced, promoting the effective utilization of frame information and alleviating the impact brought by other adverse factors.
[0114] To illustrate the technical solution of the present application, the present application also provides an electronic device, which can be a computer, a mobile phone, etc., and is not specifically limited. Please refer to Figure 12 , Figure 12 FIG. is a schematic block diagram of the structure of an embodiment of the electronic device of the present application. The electronic device 7 includes: a processor 71 and a memory 72. A computer program 721 is stored in the memory 72. The processor 71 is configured to execute the computer program 721 to implement the method as in the embodiment of the present application, which will not be elaborated herein.
[0115] In addition, the present application also provides a computer-readable storage medium. Please refer to Figure 13 , Figure 13 FIG. is a schematic circuit diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 80 stores a computer program 81. When the computer program 81 can be executed by a processor, it can implement the method as in the embodiment of the present application, which will not be elaborated herein.
[0116] If it is implemented in the form of a software functional unit and sold or used as an independent product, it can also be stored in a device with a storage function. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage device and includes several instructions (program data) for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage device includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, as well as electronic devices such as computers, mobile phones, laptop computers, tablet computers, cameras, etc. having the above storage media.
[0117] The description of the execution process of the program data in the device with a storage function can be referred to the description in the method embodiment of the present application above, which will not be elaborated herein.
[0118] The above are only embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present application.
Claims
1. A method for enhancing compressed video frames, characterized in that The method includes: Processing the compressed video frames using a multi-layer structure to obtain at least the frame features of the previous layer and the frame features of the current layer; Obtaining the frame features of the previous layer, the frame features of the current layer, the forward frame features of the current layer, and the backward frame features of the current layer to obtain a combined input; Among them, extracting the frame features of the previous layer, the frame features of the current layer, the forward frame features of the current layer, and the backward frame features of the current layer; Concatenating the frame features of the previous layer, the frame features of the current layer, the forward frame features of the current layer, and the backward frame features of the current layer according to their respective corresponding channel dimensions to obtain the combined input; Performing multi-layer processing on the combined input according to a preset parallel bidirectional network to obtain multiple screening results that meet the fusion conditions; among them, the preset parallel bidirectional network includes at least four layers of structures, and each layer of structure includes at least forward and backward branches that propagate independently in parallel and a preset attention strategy, and the forward and backward branches that propagate independently in parallel include at least a backward propagation structure and a forward propagation structure; Based on a residual block, fusing multiple screening results to obtain the enhanced features corresponding to the current layer; The performing multi-layer processing on the combined input according to a preset parallel bidirectional network to obtain multiple screening results that meet the fusion conditions includes: Processing the combined input according to the backward propagation structure to obtain a first result; Processing the combined input according to the forward propagation structure to obtain a second result; Based on the preset attention strategy, screening the first result and the second result to obtain multiple screening results that meet the fusion conditions.
2. The method according to claim 1, wherein The backward propagation structure at least includes: a pre-alignment stage, a deformation alignment stage, and a feature enhancement stage; The processing the combined input according to the backward propagation structure to obtain a first result includes: Obtaining the frame features of the current layer and the reference frame features; Through a warping operation, pre-aligning the reference frame features to the frame features of the current layer to obtain a pre-alignment result; Combining the pre-alignment result and the frame features of the current layer, and performing deformable convolution on the reference frame features to obtain a deformation alignment result from the reference frame features to the frame features of the current layer; Based on a preset residual block, fusing and enhancing the enhanced features of the previous layer, the frame features of the current layer, and the deformation alignment result to obtain the enhanced features of the current layer in the backward propagation as the first result.
3. The method according to claim 2, wherein The through a warping operation, pre-aligning the reference frame features to the frame features of the current layer to obtain a pre-alignment result includes: Obtaining the frame of the current layer and the reference frame; According to the optical flow information from the frame of the current layer to the reference frame, pre-aligning the reference frame features to the frame features of the current layer through a warping operation to obtain a pre-alignment result.
4. The method according to claim 3, wherein Combining the pre-alignment result and the frame features of the current layer, performing deformable convolution on the reference frame features to obtain a deformable alignment result from the reference frame features to the frame features of the current layer, including: Based on the input of the current frame in the current layer, through multiple convolutional layers, combining the pre-alignment result and the frames of the current layer to generate a residual bias and a mask; Adding the residual bias and the optical flow information to obtain a first bias; Using the first bias and the mask as parameters of deformable convolution, performing deformable convolution on the reference frame features to obtain the deformable alignment result from the reference frame features to the frame features of the current layer.
5. The method according to claim 4, wherein Based on a preset residual block, fusing and enhancing the enhanced features of the previous layer, the frame features of the current layer, and the deformable alignment result to obtain the enhanced features of the current frame in backpropagation as the first result, including: Based on the input of the current frame in the current layer, inputting the enhanced features of the previous layer, the frame features of the current layer, and the deformable alignment result into the preset residual block; Through the preset residual block, fusing and enhancing the enhanced features of the previous layer, the frame features of the current layer, and the deformable alignment result to obtain the enhanced features of the current layer in backpropagation, and propagating the enhanced features of the current layer to the next frame.
6. The method according to claim 1, wherein The four-layer structure at least includes four recurrent neural networks connected in sequence; Among them, the reference frames of different-layer recurrent neural network structures in the four-layer recurrent neural network are video frames with different frame rates, and the reference frames for forward and backward propagation of each layer of the recurrent neural network are video frames with the same frame rate in opposite directions.
7. The method according to claim 1, wherein Processing the combined input according to the backward propagation structure to obtain a second result, including: Performing backward processing on the enhanced features of the previous layer to obtain one of the second results; Performing backward processing on the frame features of the current layer to obtain the other of the second results; Wherein the second result at least includes one of the second results and the other of the second results.
8. The method according to claim 1, wherein The preset attention strategy at least includes a channel attention branch and a spatial attention branch; Based on the preset attention strategy, screening the first result and the second result to obtain multiple screening results that meet the fusion conditions, including: Based on the channel attention branch, performing average pooling operation on the combined input to obtain the average value of each feature channel; using two convolutions and a preset layer to train and learn the average value to obtain the weight of each feature channel; performing a dimension expansion operation through the weight of each feature channel to obtain a first weight with the same dimension as the combined input; Based on the spatial attention branch, using convolutional layers of 1×1 and 3×3 respectively to train and learn the combined input to obtain a second weight at each spatial position of the combined input; Concatenate the first weight and the second weight along the channel dimension; call a convolution to fuse the first weight and the second weight to obtain a weight set with the same dimension as the joint input; According to the weight set, screen the first result and the second result in a point-by-point multiplication manner with the joint input to obtain multiple screened results that meet the fusion conditions.
9. An enhancement device, characterized in that, The enhancement device includes: A processing module, configured to process the compressed video frame by using a multi-layer structure to obtain at least the frame feature of the previous layer and the frame feature of the current layer; An acquisition module, configured to acquire the frame feature of the previous layer, the frame feature of the current layer, the forward frame feature of the current layer, and the backward frame feature of the current layer to obtain a joint input; wherein, extract the frame feature of the previous layer, the frame feature of the current layer, the forward frame feature of the current layer, and the backward frame feature of the current layer; concatenate the frame feature of the previous layer, the frame feature of the current layer, the forward frame feature of the current layer, and the backward frame feature of the current layer according to their respective corresponding channel dimensions to obtain the joint input; The processing module is further configured to perform multi-layer processing on the joint input according to a preset parallel bidirectional network to obtain multiple screened results that meet the fusion conditions; wherein, the preset parallel bidirectional network includes at least four-layer structures, and each layer structure includes at least forward and backward branches that propagate independently in parallel and a preset attention strategy, and the forward and backward branches that propagate independently in parallel include at least a backward propagation structure and a forward propagation structure; the performing multi-layer processing on the joint input according to the preset parallel bidirectional network to obtain multiple screened results that meet the fusion conditions includes: processing the joint input according to the backward propagation structure to obtain a first result; processing the joint input according to the forward propagation structure to obtain a second result; based on the preset attention strategy, screening the first result and the second result to obtain multiple screened results that meet the fusion conditions; A fusion module, configured to fuse multiple screened results based on a residual block to obtain the enhanced feature corresponding to the current layer.
10. An electronic device, characterized in that, Including: A processor and a memory, wherein a computer program is stored in the memory, and the processor is configured to execute the computer program to implement the method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program can be executed by a processor, it implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Video space-time super-resolution method and device based on improved deformable convolution correction
CN113034380A
Apparatus and method for image processing, and computer program
JP2003337953A