Video super-resolution method and device based on auxiliary loss feature alignment cycle architecture
By introducing shallow and frequency domain assisted losses in the video super-resolution model, the problem of insufficient constraints on shallow features of images in the prior art is solved, and the recovery and detail recovery of high-resolution videos are achieved.
Patent Information
- Application Number
- CN202411094186.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-09
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2044-08-09
AI Technical Summary
In the prior art, only a single loss function is used for supervision and constraints in the network output part, resulting in insufficient constraints on the shallow feature map of the image, reducing the effect of video super-resolution, and unable to realize the recovery of high-resolution video.
The first auxiliary loss is introduced into the shallow features of the video super-resolution model and the second auxiliary loss is introduced in the frequency domain to construct a new video super-resolution model through which the target video image is super-resolution processed.
It effectively improves the super-resolution effect of video, realizes the restoration of high-resolution video, and improves the ability to recover image details and textures.
Smart Images

Figure CN119168872B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video processing technology, and in particular to a video super-resolution method and device based on an auxiliary loss feature alignment loop architecture. Background Art
[0002] Video super-resolution (VSR) technology aims to generate high-resolution videos from low-resolution videos. This technology has been widely used in many fields, including video surveillance, medical imaging, satellite imagery, video enhancement, and entertainment.
[0003] In related technologies, the network model is trained based on a constructed deformable feature aligned recurrent architecture network and a produced training set. According to the learned model parameters, a 5-frame low-resolution video sequence is used as the input of the network, and the corresponding super-resolution sequence is obtained as the output.
[0004] However, related technologies only use a single loss function for supervision and constraint in the network output part, resulting in insufficient constraints on shallow feature maps, leading to insufficient optimization of shallow image features, reducing the effect of video super-resolution, and unable to achieve high-resolution video restoration, which needs to be solved urgently. Summary of the Invention
[0005] The present application provides a video super-resolution method and device based on an auxiliary loss feature alignment loop architecture to solve the problem that in related technologies, only a single loss function is used for supervision and constraint in the network output part, resulting in insufficient constraints on shallow feature maps, leading to insufficient optimization of shallow image features, reduced video super-resolution effect, and inability to achieve high-resolution video restoration.
[0006] The first aspect of the present application provides a video super-resolution method based on an auxiliary loss feature alignment loop architecture, comprising the following steps: constructing a video super-resolution model that meets preset conditions; introducing a first auxiliary loss into the shallow features of the video super-resolution model, and introducing a second auxiliary loss into the frequency domain of the video super-resolution model to obtain a new video super-resolution model; inputting a target video image into the new video super-resolution model to perform super-resolution processing on the target video image, and obtaining a video super-resolution result based on the auxiliary loss feature alignment loop architecture.
[0007] Optionally, in one embodiment of the present application, the introducing the first auxiliary loss into the shallow features of the video super-resolution model includes: obtaining a fused feature map of the target video image during the second-layer propagation process of the video super-resolution model; upsampling the fused feature map to obtain a processed feature map; performing a residual connection between the processed feature map and the interpolated target resolution image to obtain a predicted super-resolution image; using the predicted super-resolution image and the target video image to obtain the first auxiliary loss, and introducing the first auxiliary loss into the shallow features of the video super-resolution model.
[0008] Optionally, in one embodiment of the present application, the introducing the second auxiliary loss into the frequency domain of the video super-resolution model includes: performing Fourier transform processing on the predicted super-resolution image to obtain a first feature map, and performing the Fourier transform processing on the target video image to obtain a second feature map; using the first feature map and the second feature map to obtain the second auxiliary loss, and introducing the second auxiliary loss into the frequency domain of the video super-resolution model.
[0009] Optionally, in one embodiment of the present application, the Fourier transform formula is:
[0010]
[0011] Among them, F(u,v) represents the signal in the frequency domain (the result after Fourier transform); f(x,y) represents the signal in the spatial domain (the original signal); u and v represent the frequencies in the x-direction and y-direction respectively; j represents the imaginary unit; e -j2π(ux+vy) Represents a complex exponential function, which is used to map spatial domain signals to frequency domain.
[0012] The second aspect of the present application provides a video super-resolution device based on an auxiliary loss feature alignment loop architecture, including: a construction module for constructing a video super-resolution model that meets preset conditions; a determination module for introducing a first auxiliary loss into the shallow features of the video super-resolution model, and introducing a second auxiliary loss into the frequency domain of the video super-resolution model to obtain a new video super-resolution model; a processing module for inputting a target video image into the new video super-resolution model to perform super-resolution processing on the target video image and obtain a video super-resolution result based on the auxiliary loss feature alignment loop architecture.
[0013] Optionally, in one embodiment of the present application, the determination module includes: an acquisition unit for acquiring a fused feature map of the target video image during the second-layer propagation process of the video super-resolution model; a processing unit for upsampling the fused feature map to obtain a processed feature map; a first determination unit for performing a residual connection between the processed feature map and the interpolated target resolution image to obtain a predicted super-resolution image; a first introduction unit for obtaining the first auxiliary loss using the predicted super-resolution image and the target video image, and introducing the first auxiliary loss into the shallow features of the video super-resolution model.
[0014] Optionally, in one embodiment of the present application, the determination module includes: a second determination unit, used to perform Fourier transform processing on the predicted super-resolution image to obtain a first feature map, and perform the Fourier transform processing on the target video image to obtain a second feature map; a second introduction unit, used to use the first feature map and the second feature map to obtain the second auxiliary loss, and introduce the second auxiliary loss into the frequency domain of the video super-resolution model.
[0015] Optionally, in one embodiment of the present application, the Fourier transform formula is:
[0016]
[0017] Among them, F(u,v) represents the signal in the frequency domain (the result after Fourier transform); f(x,y) represents the signal in the spatial domain (the original signal); u and v represent the frequencies in the x-direction and y-direction respectively; j represents the imaginary unit; e -j2π(ux+vy) Represents a complex exponential function, which is used to map spatial domain signals to frequency domain.
[0018] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the video super-resolution method based on the auxiliary loss feature alignment loop architecture as described in the above embodiment.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above-mentioned video super-resolution method based on the auxiliary loss feature alignment loop architecture.
[0020] The fifth aspect of the present application provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned video super-resolution method based on the auxiliary loss feature alignment loop architecture.
[0021] The embodiment of the present application can introduce the first auxiliary loss into the shallow features of the video super-resolution model, and introduce the second auxiliary loss into the frequency domain of the video super-resolution model to obtain a new video super-resolution model. The target video image is input into the new video super-resolution model to perform super-resolution processing on the target video image, thereby obtaining a video super-resolution result based on the auxiliary loss feature alignment loop architecture, effectively improving the effect of video super-resolution and achieving high-resolution video restoration. This solves the problem in the related art that only a single loss function is used for supervision and constraint in the network output part, resulting in insufficient constraints on the shallow feature map, resulting in insufficient optimization of the shallow features of the image, reducing the effect of video super-resolution, and failing to achieve high-resolution video restoration.
[0022] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0024] Figure 1 A flowchart of a video super-resolution method based on an auxiliary loss feature alignment loop architecture provided according to an embodiment of the present application;
[0025] Figure 2 This is a schematic diagram of a loop architecture based on auxiliary loss feature alignment according to a specific embodiment of the present application;
[0026] Figure 3 A schematic diagram of shallow auxiliary loss according to a specific embodiment of the present application;
[0027] Figure 4 This is a schematic diagram of frequency domain auxiliary loss according to a specific embodiment of the present application;
[0028] Figure 5 Schematic diagram of the structure of a video super-resolution device based on an auxiliary loss feature alignment loop architecture provided according to an embodiment of the present application;
[0029] Figure 6 A schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0031] The following describes a video super-resolution method and apparatus based on an auxiliary loss feature alignment recurrent architecture according to an embodiment of the present application with reference to the accompanying drawings. In response to the problem that the related art mentioned in the background art above only uses a single loss function for supervision and constraint at the network output, resulting in insufficient constraints on shallow feature maps, leading to insufficient optimization of shallow image features, reduced video super-resolution effect, and inability to achieve high-resolution video restoration, the present application provides a video super-resolution method based on an auxiliary loss feature alignment recurrent architecture. In this method, a first auxiliary loss can be introduced into the shallow features of a video super-resolution model, and a second auxiliary loss can be introduced into the frequency domain of the video super-resolution model to obtain a new video super-resolution model. The target video image is input into the new video super-resolution model to perform super-resolution processing on the target video image, thereby obtaining a video super-resolution result based on the auxiliary loss feature alignment recurrent architecture, effectively improving the video super-resolution effect and achieving high-resolution video restoration. Thus, the problem that the related art only uses a single loss function for supervision and constraint at the network output, resulting in insufficient constraints on shallow feature maps, leading to insufficient optimization of shallow image features, reduced video super-resolution effect, and inability to achieve high-resolution video restoration, is solved.
[0032] Specifically, Figure 1 A flowchart of a video super-resolution method based on an auxiliary loss feature alignment loop architecture provided in an embodiment of the present application.
[0033] like Figure 1 As shown, the video super-resolution method based on the auxiliary loss feature alignment cycle architecture includes the following steps:
[0034] In step S101, a video super-resolution model that meets preset conditions is constructed.
[0035] It can be understood that the embodiments of the present application can construct a video super-resolution model that meets certain conditions. For example, the overall architecture of the video super-resolution model constructed by the embodiments of the present application can include a feature extraction module, a feature alignment module, a feature fusion module, a feature propagation module and an upsampling module.
[0036] Among them, such as Figure 2 As shown in the figure, the feature extraction module is responsible for extracting useful features from the input low-resolution video frame. Then, the feature map of the low-resolution image is obtained through the residual block, which can capture the local structure and texture information in the video image.
[0037] Feature alignment module: The bidirectional recurrent network has four layers: backward propagation, forward propagation, backward propagation, and forward propagation. Assume that the current layer is the second layer of the propagation process and the current frame is t i frame, corresponding to t i-1 Frame and ti-2 The frames are first aligned by optical flow t i Frame, get the coarsely aligned frame and Then these two frames are passed through the deformable convolution dcnv2 to t i Perform a fine alignment to obtain It can reduce information misalignment caused by object movement and ensure that features between adjacent frames can be correctly matched.
[0038] Feature fusion module: Since it is currently in the second layer, this layer Set as This finely aligned frame With the current frame t i , and the previous layers, that is, this is the first layer Obtain a fused feature through a residual block When reconstructing high-resolution video frames, information from multiple frames is utilized to improve the reconstruction quality.
[0039] Feature propagation module: the feature map obtained Continue to propagate downward and forward, the downward propagated feature map is used as part of the input of the third layer feature fusion module, and the forward propagated feature map is used as the input of the third layer feature fusion module. u+2 , t i+1 Perform optical flow alignment as t i+2 , t i+1 These two frame features align part of the input to the module, which helps to maintain consistency and coherence between video frames.
[0040] Finally, the four-layer structure completes reasoning in the above order. After the features of each layer are fused, they are saved in a dictionary, plus the feature extraction part. Then, these tensors are fused, upsampled in the upsampling module, and a residual connection is made with the interpolated low-resolution image. This allows the model to focus on learning the difference between low resolution and high resolution through residual learning, thereby achieving better reconstruction effects.
[0041] In step S102, the first auxiliary loss is introduced into the shallow features of the video super-resolution model, and the second auxiliary loss is introduced into the frequency domain of the video super-resolution model to obtain a new video super-resolution model.
[0042] In the embodiment of the present application, the first auxiliary loss is a shallow auxiliary loss; the second auxiliary loss is a frequency domain auxiliary loss.
[0043] It can be understood that the embodiment of the present application can introduce the first auxiliary loss in the following steps into the shallow features of the video super-resolution model, and introduce the second auxiliary loss in the following steps into the frequency domain of the video super-resolution model. Two auxiliary losses are introduced to improve the training process of the model to obtain a new video super-resolution model, which can better capture details and texture information during the training process, while also maintaining good temporal consistency, effectively improving the reconstruction quality of the video super-resolution model.
[0044] In one embodiment of the present application, a first auxiliary loss is introduced into the shallow features of a video super-resolution model, including: obtaining a fused feature map of a target video image during the second-layer propagation process of the video super-resolution model; upsampling the fused feature map to obtain a processed feature map; performing a residual connection between the processed feature map and the interpolated target resolution image to obtain a predicted super-resolution image; using the predicted super-resolution image and the target video image to obtain a first auxiliary loss, and introducing the first auxiliary loss into the shallow features of the video super-resolution model.
[0045] In the actual implementation process, Figure 2 As shown, the embodiment of the present application can introduce shallow auxiliary loss in the training process of the video super-resolution model, that is, add a shallow auxiliary loss to the shallow feature map, obtain the fused feature map in the second layer propagation process, and interpolate the low-resolution image, and then pass the final fused feature map through two sub-pixel convolutions (Sub-pixel Convolution) to increase the spatial resolution of the feature map, and then pass it through a convolution layer to keep it consistent with the original super-resolution map size, and obtain the final predicted super-resolution image by performing a residual connection with the interpolated low-resolution image. The original super-resolution map is subjected to two losses, namely the loss between the final predicted image and the high-resolution target image, and the auxiliary loss calculated at the shallow feature level. These two losses are back-propagated at the shallow level to ensure high-quality learning of shallow features.
[0046] Therefore, the embodiment of the present application ensures high-quality learning of shallow features by introducing shallow auxiliary loss into the shallow features of the video super-resolution model, which helps to further capture the basic structure and texture information in the video frame, helps the video super-resolution model to learn rich feature expressions at an early stage, and improves the diversity and robustness of features. Furthermore, in subsequent deep processing, the video super-resolution model can perform more complex combinations and optimizations based on rich shallow features.
[0047] In one embodiment of the present application, a second auxiliary loss is introduced into the frequency domain of the video super-resolution model, including: performing Fourier transform processing on the predicted super-resolution image to obtain a first feature map, and performing Fourier transform processing on the target video image to obtain a second feature map; using the first feature map and the second feature map to obtain a second auxiliary loss, and introducing the second auxiliary loss into the frequency domain of the video super-resolution model.
[0048] In some embodiments, Figure 2 As shown, the embodiment of the present application can introduce frequency domain auxiliary loss into the frequency domain of the video super-resolution model, mainly by performing Fourier transform on the upsampled image and also performing Fourier transform on the original super-resolution image, wherein the Fourier transform formula is as follows:
[0049]
[0050] Among them, F(u,v) represents the signal in the frequency domain (the result after Fourier transform); f(x,y) represents the signal in the spatial domain (the original signal); u and v represent the frequencies in the x-direction and y-direction respectively; j represents the imaginary unit; e -j2π(ux+vy) Represents a complex exponential function, which is used to map spatial domain signals to frequency domain.
[0051] Next, the embodiment of the present application uses the real and imaginary parts of the two Fourier transformed images to calculate L1_loss to obtain the loss in the frequency domain, and assigns the loss to the frequency domain constraint of the video image. This constraint allows the video image to better restore the high-frequency information on the image and increase the image details. It should be noted that the shallow auxiliary loss and the frequency domain auxiliary loss are only used in the training stage, will not affect the speed of image inference, and will not increase the complexity of the model.
[0052] Therefore, the embodiment of the present application can introduce frequency domain auxiliary loss in the frequency domain, so that the video super-resolution model can better capture the high-frequency and low-frequency information in the video frame, thereby enhancing the detail recovery ability. Frequency domain processing helps to eliminate blur and artifacts in the video, improve the clarity and fineness of the picture, and promote the fusion of multi-scale features, so that the video super-resolution model can integrate information at different scales and improve the super-resolution quality of the video frame.
[0053] In step S103 , the target video image is input into a new video super-resolution model to perform super-resolution processing on the target video image to obtain a video super-resolution result based on an auxiliary loss feature alignment loop architecture.
[0054] It can be understood that the embodiments of the present application can input the target video image, such as a low-resolution video image, into a new video super-resolution model to perform super-resolution processing on the low-resolution video image, and obtain a video super-resolution result based on the auxiliary loss feature alignment loop architecture, which effectively improves the model's ability to restore high-resolution video frames. Specifically, shallow features often contain more edge and texture information. By optimizing these features, the model can better capture and restore the details of the image, and improve the resolution and visual quality of the reconstructed image; the frequency domain auxiliary loss can focus on the high-frequency components of the image, and can better restore the details in the image, such as edges and textures, and effectively improve the sharpness and visual effect of the image.
[0055] In addition, the shallow feature loss and frequency domain feature loss proposed in the embodiments of this application can be directly transplanted into other video super-resolution models, are plug-and-play, and will only increase the complexity of the training process, have no effect on the reasoning process, and can play a role in improving model performance.
[0056] For example, Figure 3 As shown in the figure, it is a schematic diagram of shallow auxiliary loss. Here, we will not consider the frequency domain auxiliary loss for the time being. First, after the feature extraction step, the input video image can be represented in the form of a tensor as Ved∈R b×t×c×h×w , the usual loss function is used at the end, and the frames of the first four layers of fine alignment need to be fused before upsampling: by Ved b1 ∈R b×t×c×h×w (Ved_backward_1): represents all finely aligned video image feature maps generated by the first backward propagation; Ved f1 ∈R b×t×c×h×w (Ved_forward_1): represents all finely aligned video image feature maps generated by the first forward propagation; Ved b2 ∈R b×t×c×h×w (Ved_backward_2): represents all finely aligned video image feature maps generated by the second backward propagation; Ved f2 ∈R b×t×c×h×w (Ved_forward_2): indicates that all finely aligned video image feature maps generated by the second forward propagation need to be frame-fused on the finely aligned frames of the previous two layers before upsampling. The normal loss function will fuse Ved, Ved_b1, Ved_f1, Ved_b2, and Ved_f2 through a residual block. However, the problem with this is that the loss function obtained in this way is not effective for shallow features, so shallow auxiliary loss is required. The specific steps are as follows:
[0057] Step S301: During the second layer propagation, an auxiliary loss calculation can be performed, that is, feature fusion of Ved, Ved_b1, and Ved_f1 to obtain feat∈R b×t×c×h×w ;
[0058] Step S302: After two PixelShufflePack blocks, the feature is upsampled by 4 times feat_4∈R b ×t×c×4h×4w , and then pass through a convolution layer to make the channel dimension become 3D;
[0059] Step S303: Perform a residual connection on the low-resolution image after interpolation and upsampling to obtain a prediction pred gt ∈R b×t×3×4h×4w ;
[0060] Step S304: Compare with the original super-resolution graph gt∈R b×t×3×4h×4w Perform an L1_loss (the frequency domain auxiliary loss is not considered here). This loss can make the shallow feature map obtain good results, so that the shallow features can obtain good results.
[0061] For example, Figure 4 As shown, it is a schematic diagram of frequency domain auxiliary loss. The embodiment of the present application uses frequency domain auxiliary loss (fft loss). The embodiment of the present application can calculate the L1 loss back propagation and the frequency domain auxiliary loss for the predicted super-resolution image obtained after upsampling, and can also add frequency domain auxiliary loss to the shallow auxiliary loss. Ordinary L1 loss mainly calculates the difference in pixel values in the time domain (spatial domain). This most intuitive way of calculating loss can have a good restoration effect on the video image as a whole, but it is easy to ignore high-frequency details (such as edges and textures). Therefore, frequency domain auxiliary loss can optimize frequency domain information and can better capture and restore high-frequency details in the image. The combination of the two can comprehensively utilize multi-scale information, which not only ensures the overall consistency of the image, but also retains detailed features, improves the quality of the reconstructed image, and enables the new video super-resolution model to be optimized in both the time domain and the frequency domain, thereby providing higher quality reconstruction results. Among them, the specific steps of frequency domain auxiliary loss are:
[0062] Step S401: Taking the second layer as an example, the predicted super-resolution image pred is obtained after upsampling and residual connection gt ∈R b×t×3×4h×4w , after performing ordinary L1 loss with gt (original super-resolution image), calculating the pixel value difference in the time domain, and after Fourier transform, the predicted super-resolution image time domain is converted into frequency domain, but the image size does not change, and pred gtfft ∈R b×t×3×4h×4w(fft stands for Fourier transform);
[0063] Step S402: Since the tensor obtained in this way is in complex form, the loss cannot be calculated directly. It is necessary to separate the real part and the imaginary part to obtain pred gtreal ∈R b×t×3×4h×4w and pred gtimag ∈R b×t×3×4h×4w , then the two tensors are concat on the channel dimension to get pre_gt_con∈R b×t×6×4h×4w ;
[0064] Step S403: Super-resolution original video image, gt∈R b×t×3×4h×4w After the same processing as above, we get gt_fft∈R b×t×3×4h×4w , gt_real∈R b×t×3×4h×4w ,gt_imag∈R b×t×3×4h×4w , gt_con∈R b×t×6×4h×4w ;
[0065] Step S404: pre_gt_con∈R b×t×6×4h×4w and gt_con∈R b×t×6×4h×4w These two tensors calculate the L1 loss to achieve loss calculation in the frequency domain and further enhance the details of the image.
[0066] Since the loss calculated in this way is several orders of magnitude higher than the L1 loss, the frequency domain loss will mask the L1 loss. Therefore, in order to ensure that both losses have a rough impact on the features during the back propagation process, a coefficient λ is multiplied in front of the frequency domain loss. Therefore, in order to maintain a roughly consistent order of magnitude in the new video super-resolution model, the coefficient used can be λ = 0.05.
[0067] According to the video super-resolution method based on the auxiliary loss feature alignment loop architecture proposed in the embodiment of the present application, the first auxiliary loss can be introduced into the shallow features of the video super-resolution model, and the second auxiliary loss can be introduced into the frequency domain of the video super-resolution model to obtain a new video super-resolution model. The target video image is input into the new video super-resolution model to perform super-resolution processing on the target video image, and a video super-resolution result based on the auxiliary loss feature alignment loop architecture is obtained, which effectively improves the effect of video super-resolution and realizes the restoration of high-resolution video. This solves the problem in the related art that only a single loss function is used for supervision and constraint in the network output part, resulting in insufficient constraints on the shallow feature map, resulting in insufficient optimization of the shallow features of the image, reduced video super-resolution effect, and inability to achieve high-resolution video restoration.
[0068] Next, a video super-resolution device based on an auxiliary loss feature alignment loop architecture proposed in accordance with an embodiment of the present application will be described with reference to the accompanying drawings.
[0069] Figure 5 4 is a block diagram of a video super-resolution device based on an auxiliary loss feature alignment loop architecture according to an embodiment of the present application.
[0070] like Figure 5 As shown, the video super-resolution device 10 based on the auxiliary loss feature alignment cycle architecture includes: a construction module 100, a determination module 200 and a processing module 300.
[0071] Specifically, the construction module 100 is used to construct a video super-resolution model that meets preset conditions.
[0072] The determination module 200 is used to introduce the first auxiliary loss into the shallow features of the video super-resolution model and introduce the second auxiliary loss into the frequency domain of the video super-resolution model to obtain a new video super-resolution model.
[0073] The processing module 300 is used to input the target video image into the new video super-resolution model to perform super-resolution processing on the target video image and obtain a video super-resolution result based on the auxiliary loss feature alignment loop architecture.
[0074] Optionally, in one embodiment of the present application, the determination module 200 includes: an acquisition unit, a processing unit, a first determination unit, and a first introduction unit.
[0075] Among them, the acquisition unit is used to obtain the fusion feature map of the target video image in the second layer propagation process of the video super-resolution model.
[0076] The processing unit is used to upsample the fused feature map to obtain a processed feature map.
[0077] The first determining unit is used to perform a residual connection between the processed feature map and the interpolated target resolution image to obtain a predicted super-resolution image.
[0078] The first introduction unit is used to obtain a first auxiliary loss by using the predicted super-resolution image and the target video image, and introduce the first auxiliary loss into the shallow features of the video super-resolution model.
[0079] Optionally, in one embodiment of the present application, the determination module 200 includes: a second determination unit and a second introduction unit.
[0080] The second determining unit is configured to perform Fourier transform processing on the predicted super-resolution image to obtain a first feature map, and perform Fourier transform processing on the target video image to obtain a second feature map.
[0081] The second introduction unit is used to obtain a second auxiliary loss by using the first feature map and the second feature map, and introduce the second auxiliary loss into the frequency domain of the video super-resolution model.
[0082] Optionally, in one embodiment of the present application, the Fourier transform formula is:
[0083]
[0084] Among them, F(u,v) represents the signal in the frequency domain (the result after Fourier transform); f(x,y) represents the signal in the spatial domain (the original signal); u and v represent the frequencies in the x-direction and y-direction respectively; j represents the imaginary unit; e -j2π(ux+vy) Represents a complex exponential function, which is used to map spatial domain signals to frequency domain.
[0085] It should be noted that the above explanation of the embodiment of the video super-resolution method based on the auxiliary loss feature alignment loop architecture is also applicable to the video super-resolution device based on the auxiliary loss feature alignment loop architecture of this embodiment, and will not be repeated here.
[0086] According to the video super-resolution device based on the auxiliary loss feature alignment loop architecture proposed in the embodiment of the present application, the first auxiliary loss can be introduced into the shallow features of the video super-resolution model, and the second auxiliary loss can be introduced into the frequency domain of the video super-resolution model to obtain a new video super-resolution model. The target video image is input into the new video super-resolution model to perform super-resolution processing on the target video image, and a video super-resolution result based on the auxiliary loss feature alignment loop architecture is obtained, which effectively improves the effect of video super-resolution and realizes the restoration of high-resolution video. This solves the problem in the related art that only a single loss function is used for supervision and constraint in the network output part, resulting in insufficient constraints on the shallow feature map, resulting in insufficient optimization of the shallow features of the image, reduced video super-resolution effect, and inability to achieve high-resolution video restoration.
[0087] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0088] A memory 601 , a processor 602 , and a computer program stored in the memory 601 and executable on the processor 602 .
[0089] When the processor 602 executes the program, the video super-resolution method based on the auxiliary loss feature alignment loop architecture provided in the above embodiment is implemented.
[0090] Furthermore, the electronic device further includes:
[0091] The communication interface 603 is used for communication between the memory 601 and the processor 602 .
[0092] The memory 601 is used to store computer programs that can be run on the processor 602 .
[0093] The memory 601 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0094] If the memory 601, processor 602, and communication interface 603 are implemented independently, the communication interface 603, memory 601, and processor 602 can be interconnected via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0095] Optionally, in a specific implementation, if the memory 601, the processor 602 and the communication interface 603 are integrated on a chip, the memory 601, the processor 602 and the communication interface 603 can communicate with each other through an internal interface.
[0096] The processor 602 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0097] This embodiment also provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the video super-resolution method based on the auxiliary loss feature alignment loop architecture is implemented as described above.
[0098] This embodiment also provides a computer program product, including a computer program. When the computer program is executed, it is used to implement the above-mentioned video super-resolution method based on the auxiliary loss feature alignment cycle architecture.
[0099] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0100] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0101] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0102] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0103] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, it can be implemented using any one or a combination of the following technologies known in the art: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0104] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0105] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0106] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A video super-resolution method based on auxiliary loss feature alignment cycle architecture, characterized in that: The following steps are involved: Constructing a video super-resolution model that meets preset conditions, wherein the video super-resolution model includes a feature extraction module, a feature alignment module, a feature fusion module, a feature propagation module and an upsampling module; The feature extraction module is responsible for extracting features from the input target resolution video frame. Then, the residual block is used to obtain the feature map of the target resolution image to capture the local structure and texture information in the video image. Feature alignment module: The bidirectional recurrent network has four layers: backward propagation, forward propagation, backward propagation, and forward propagation. If the current layer is the second layer of the propagation process, the current frame is t i frame, corresponding to t i-1 Frame and t i-2 The frames are first aligned by optical flow t i Frame, get the coarsely aligned frame and Then these two frames are passed through the deformable convolution dcnv2 to t i Perform a fine alignment to obtain Feature fusion module: If the current layer is the second layer, this layer Set as Align the frames With the current frame t i , and the previous layers, a fused feature is obtained through a residual block Feature propagation module: the feature map obtained Continue to propagate downward and forward, the downward propagated feature map is used as part of the input of the third layer feature fusion module, and the forward propagated feature map is used as the input of the third layer feature fusion module. i+2 , t i+1 Perform optical flow alignment as t i+2 , t i+1 These two frames are part of the input to the feature alignment module; each layer of features is fused and saved in a dictionary; A first auxiliary loss is introduced into the shallow features of the video super-resolution model, and a second auxiliary loss is introduced into the frequency domain of the video super-resolution model to obtain a new video super-resolution model, wherein the introducing the first auxiliary loss into the shallow features of the video super-resolution model includes: obtaining a fusion feature map of the target video image in the second layer propagation process of the video super-resolution model; upsampling the fusion feature map to obtain a processed feature map; performing a residual connection between the processed feature map and the interpolated target resolution image to obtain a predicted super-resolution image; using the predicted super-resolution image and the target video image to obtain the first auxiliary loss, and introducing the first auxiliary loss into the shallow features of the video super-resolution model; introducing the second auxiliary loss into the frequency domain of the video super-resolution model includes: upsampling the predicted super-resolution image; The image is subjected to Fourier transform processing to obtain a first feature map, and the target video image is subjected to the Fourier transform processing to obtain a second feature map; the first feature map and the second feature map are used to obtain the second auxiliary loss, and the second auxiliary loss is introduced into the frequency domain of the video super-resolution model, wherein the processed feature map is residually connected with the interpolated target resolution image to obtain a predicted super-resolution image, including: based on the fused feature map in the second layer propagation process, the target resolution image is interpolated to obtain a final fused feature map, and the final fused feature map is subjected to two sub-pixel convolutions to increase the spatial resolution of the feature map, and then subjected to a convolution layer to keep the size consistent with the original super-resolution map, and the predicted super-resolution image is obtained by performing a residual connection with the interpolated target resolution image; The target video image is input into the new video super-resolution model to perform super-resolution processing on the target video image to obtain a video super-resolution result based on an auxiliary loss feature alignment loop architecture.
2. The method according to claim 1, characterized in that The formula for the Fourier transform is: Among them, F(u,v) represents the signal in the frequency domain (the result after Fourier transform); f(x,y) represents the signal in the spatial domain (the original signal); u and v represent the frequencies in the x-direction and y-direction respectively; j represents the imaginary unit; e -j2π(ux+vy) Represents a complex exponential function, which is used to map spatial domain signals to frequency domain.
3. A video super-resolution device based on an auxiliary loss feature alignment cycle architecture, characterized in that: include: A construction module, configured to construct a video super-resolution model that meets preset conditions, wherein the video super-resolution model includes a feature extraction module, a feature alignment module, a feature fusion module, a feature propagation module, and an upsampling module; The feature extraction module is responsible for extracting features from the input target resolution video frame. Then, the residual block is used to obtain the feature map of the target resolution image to capture the local structure and texture information in the video image. Feature alignment module: The bidirectional recurrent network has four layers: backward propagation, forward propagation, backward propagation, and forward propagation. If the current layer is the second layer of the propagation process, the current frame is t i frame, corresponding to t i-1 Frame and t i-2 The frames are first aligned by optical flow t i Frame, get the coarsely aligned frame and Then these two frames are passed through the deformable convolution dcnv2 to t i Perform a fine alignment to obtain Feature fusion module: If the current layer is the second layer, this layer Set as Align the frames With the current frame t i , and the previous layers, a fused feature is obtained through a residual block Feature propagation module: the feature map obtained Continue to propagate downward and forward, the downward propagated feature map is used as part of the input of the third layer feature fusion module, and the forward propagated feature map is used as the input of the third layer feature fusion module. i+2 , t i+1 Perform optical flow alignment as t i+2 , t i+1 These two frames are part of the input to the feature alignment module; each layer of features is fused and saved in a dictionary; A determination module is provided for introducing a first auxiliary loss into the shallow features of the video super-resolution model and introducing a second auxiliary loss into the frequency domain of the video super-resolution model to obtain a new video super-resolution model, wherein the introducing the first auxiliary loss into the shallow features of the video super-resolution model includes: obtaining a fusion feature map of the target video image in the second layer propagation process of the video super-resolution model; upsampling the fusion feature map to obtain a processed feature map; performing a residual connection between the processed feature map and the interpolated target resolution image to obtain a predicted super-resolution image; using the predicted super-resolution image and the target video image to obtain the first auxiliary loss, and introducing the first auxiliary loss into the shallow features of the video super-resolution model; introducing the second auxiliary loss into the frequency domain of the video super-resolution model includes: The super-resolution image is subjected to Fourier transform processing to obtain a first feature map, and the target video image is subjected to the Fourier transform processing to obtain a second feature map; the first feature map and the second feature map are used to obtain the second auxiliary loss, and the second auxiliary loss is introduced into the frequency domain of the video super-resolution model, wherein the processed feature map is residually connected with the interpolated target resolution image to obtain the predicted super-resolution image, including: based on the fused feature map in the second layer propagation process, the target resolution image is interpolated to obtain a final fused feature map, and the final fused feature map is subjected to two sub-pixel convolutions to increase the spatial resolution of the feature map, and then subjected to a convolution layer to keep the size consistent with the original super-resolution map, and the predicted super-resolution image is obtained by performing a residual connection with the interpolated target resolution image; A processing module is used to input the target video image into the new video super-resolution model to perform super-resolution processing on the target video image to obtain a video super-resolution result based on an auxiliary loss feature alignment loop architecture.
4. The device according to claim 3, characterized in that The formula for the Fourier transform is: Among them, F(u,v) represents the signal in the frequency domain (the result after Fourier transform); f(x,y) represents the signal in the spatial domain (the original signal); u and v represent the frequencies in the x-direction and y-direction respectively; j represents the imaginary unit; r -j2π(ux+vy) Represents a complex exponential function, which is used to map spatial domain signals to frequency domain.
5. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the video super-resolution method based on the auxiliary loss feature alignment loop architecture as described in any one of claims 1-2.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the video super-resolution method based on the auxiliary loss feature alignment cycle architecture as described in any one of claims 1-2.
7. A computer program product comprising a computer program, characterized in that The computer program is executed by a processor to implement the video super-resolution method based on the auxiliary loss feature alignment cycle architecture according to any one of claims 1 to 2.
Citation Information
Patent Citations
Video super-resolution method based on deformable feature alignment loop architecture
CN116309045A