Video super-resolution reconstruction method and system

Through the dual-path architecture of explicit-implicit hybrid alignment and the spatiotemporal Transformer technology, the problem that video super-resolution technology is difficult to balance reconstruction capabilities and robustness in complex motion scenarios is solved, and efficient video super-resolution reconstruction and cross-device adaptation are achieved.

CN120013766AActive Publication Date: 2025-05-16BEIJING UNIV OF TECH

Patent Information

Application Number
CN202510479631.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

Existing video super-resolution technologies are difficult to balance reconstruction capabilities and robustness in complex scenarios where rigid and non-rigid motions coexist, and traditional methods are prone to resolution mismatch, loss of details or redundant computing resources in cross-device scenarios.

Method used

A dual-path architecture with explicit-implicit hybrid alignment is proposed. The motion trajectory field is generated by pre-training the optical flow network and pixel-level motion compensation is completed using a differentiable bilinear sampling algorithm. It combines the spatial-temporal Transformer and multi-scale pyramid fusion strategy to achieve feature alignment and reconstruction.

Benefits of technology

It significantly improves the robustness and accuracy of video super-resolution reconstruction, can balance reconstruction capabilities and robustness in complex motion scenarios, and adapt to the resolution requirements of different display terminals, reducing resolution mismatch and detail loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013766A_ABST
    Figure CN120013766A_ABST
Patent Text Reader

Abstract

The invention discloses a video super-resolution reconstruction method and system, and relates to the technical field of video restoration processing, and the method comprises the steps: collecting a to-be-reconstructed video frame sequence and a reference frame sequence; extracting a dense optical flow field between the reference frame and the target frame; performing geometric alignment on the target frame to the reference frame according to the dense optical flow field to obtain a motion consistency feature; performing feature extraction on the video frame sequence through a parameterized residual scaling module; inputting the extracted features into a bidirectional propagation module to process forward and backward image sequences respectively, extracting time sequence forward features and time sequence reverse features, and fusing the features to obtain detail features; carrying out adaptive fusion on the motion consistency features and the detail features, and sending the fused features into a reconstruction network to obtain deep visual features; pixel-level adaptive reconstruction is realized according to the deep visual features, and a super-resolution video sequence is obtained; according to the method, the feature space consistency is kept, and meanwhile, the feature matching precision in a complex motion scene is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video restoration processing, and in particular to a video super-resolution reconstruction method and system. Background Art

[0002] Video super-resolution technology aims to achieve high-resolution video reconstruction by mining complementary information in video frame sequences. With the rapid popularization of ultra-high-definition display devices, users' requirements for video content quality are increasing, while the resolution specifications of different display terminals vary significantly.

[0003] Traditional video super-resolution technology usually builds independent models for fixed scaling, which makes it difficult to dynamically adapt to diverse display requirements, resulting in resolution mismatch, detail loss, or redundant computing resources in cross-device scenarios. At the same time, the core challenge of traditional video super-resolution methods lies in the efficient alignment and fusion of spatiotemporal features across frames.

[0004] Existing technologies are mainly divided into two paradigms: explicit alignment and implicit alignment. Explicit alignment achieves pixel-level compensation through motion estimation, and is stable in large displacement scenes, but its performance is highly dependent on the accuracy of optical flow estimation, and is prone to compensation deviation in occluded areas or weak motion scenes. Implicit alignment uses neural networks to autonomously learn inter-frame associations. Although it has the advantage of dynamic adaptation, it faces bottlenecks such as insufficient interpretability and limited complex motion modeling capabilities. In short, a single alignment method is difficult to adapt to different types of motion, and cannot balance reconstruction capabilities and robustness in complex scenes where rigid and non-rigid motion coexist. Summary of the invention

[0005] In view of the shortcomings of the existing technology that the single alignment method is difficult to adapt to different types of motion and cannot balance the reconstruction capability and robustness in complex scenes where rigid and non-rigid motions coexist, the present invention proposes a video super-resolution reconstruction method and system, constructs a dual-path architecture of explicit-implicit hybrid alignment, generates a motion trajectory field through a pre-trained optical flow network and uses a differentiable bilinear sampling algorithm to complete pixel-level motion compensation, and realizes feature alignment through the modeling capability of the spatiotemporal Transformer on a bidirectional propagation network, combined with a multi-scale pyramid fusion strategy, thereby greatly improving the problems existing in the existing technology.

[0006] A video super-resolution reconstruction method comprises the following steps: Acquire a target frame image and a center frame image in a video frame sequence to be reconstructed; The dense optical flow field between the target frame image and the center frame image is extracted through the parameter-frozen optical flow prediction model. The target frame image is geometrically aligned with the center frame image according to the dense optical flow field. The aligned frame sequence is shallowly encoded to obtain motion consistency features. The shallow visual features of the video frame sequence are extracted by adaptively adjusting the feature fusion ratio through a parameterized residual scaling module; the extracted shallow visual features are input into a bidirectional propagation module with a spatiotemporal Transformer as a basic unit to obtain temporal forward features and temporal reverse features; wherein 3D position encoding is introduced into the spatiotemporal attention mechanism of the spatiotemporal Transformer; the temporal forward features and the temporal reverse features are fused through a multi-scale feature pyramid to obtain detail features; The motion consistency features and detail features are fused; hierarchical feature extraction is performed on the fused features to obtain deep visual features; pixel-level reconstruction is performed based on the deep visual features to generate the final super-resolution video sequence.

[0007] Furthermore, the method of geometrically aligning the target frame image with the center frame image according to the dense optical flow field, and performing shallow feature encoding on the aligned frame sequence to obtain motion consistency features specifically includes the following steps: According to the dense optical flow field, a differentiable bilinear sampling algorithm is used to geometrically align the target frame image to the center frame image; The aligned frame sequence is shallowly encoded through a lightweight convolutional network, which specifically includes upgrading the aligned frame sequence to a high-dimensional semantic feature space, compressing the upgraded multi-frame features into a single-frame representation vector using a temporal dimension adaptive average pooling operation, and implicitly completing the adaptive fusion of motion information through a dynamic weight allocation mechanism to obtain motion consistency features.

[0008] Furthermore, the shallow visual features of the video frame sequence are extracted by adaptively adjusting the feature fusion ratio through the parameterized residual scaling module, specifically including: adopting a serial cascaded five-layer stacking structure to extract features from local details to global semantics; wherein, the residual connection strength is dynamically adjusted by using learning parameters to adaptively adjust the feature fusion ratio, and the statistical distribution differences between different video frames are eliminated by removing the BatchNorm layer; each level feature is extracted through progressive transmission, and each layer includes a 3×3 convolution layer, a LeakyReLU activation function and a parameterized residual connection; the feature extraction process is expressed as: ; in, represents the learning parameters, and Indicates a convolution layer with a convolution kernel of 3*3. Represented as LeakyReLU activation function, represents the input feature map, Represents the output feature map.

[0009] Furthermore, the extracted shallow visual features are input into a bidirectional propagation module with a spatiotemporal Transformer as a basic unit to obtain temporal forward features and temporal reverse features, which specifically includes the following steps: In the forward processing stage, shallow visual features are input into the forward processing network composed of stacked spatiotemporal Transformers in the original time sequence. The spatiotemporal attention mechanism of 3D position encoding is introduced to aggregate the shallow visual features of multiple frames in the video frame sequence, and the temporal forward features are extracted frame by frame. In the reverse processing stage, the shallow visual features are mirrored along the time dimension and fed into an independent reverse processing network. By introducing the spatiotemporal attention mechanism of 3D position encoding, the shallow visual features of multiple frames in the video frame sequence are aggregated, and the temporal forward features are extracted frame by frame.

[0010] Furthermore, the spatiotemporal attention mechanism of introducing 3D position encoding to aggregate multiple shallow visual features in a video frame sequence specifically includes the following steps: Generate 3D dynamic position encoding through 3D position convolution; The 3D dynamic position encoding and shallow visual features are segmented to obtain local blocks; Generate the query matrix Q with the local block segmented by 3D dynamic position encoding as the core, and generate the key matrix K and value matrix V with the local block segmented by shallow visual features as the core; The similarity matrix is ​​obtained by performing dot product calculation on the query matrix Q and the transpose of the key matrix K; The similarity matrix and the value matrix V are weighted and summed to obtain the aggregated features; The aggregated features are input into the attention mechanism module of Transformer, and the position-driven calculation strategy is used to achieve implicit inter-frame alignment and feature fusion.

[0011] Furthermore, the temporal forward features and the temporal reverse features are fused through a multi-scale feature pyramid to obtain detail features, which specifically includes the following steps: Concatenate the time series forward features and time series reverse features according to the spatial dimension; Using the progressive spatial-temporal downsampling strategy, the concatenated features are subjected to 3D average pooling operations with different downsampling coefficients to construct a feature pyramid. The spatial-temporal dimension of feature pyramid cross-resolution features is aligned through a differentiable trilinear interpolation algorithm; All aligned features are concatenated in the channel dimension and the average is taken to output detail features with detail sensitivity and context perception.

[0012] Furthermore, the motion consistency feature and the detail feature are fused using a dynamic gating fusion mechanism; specifically, the following steps are included: The motion consistency feature and detail feature are concatenated along the channel dimension to obtain the joint feature; The weights of the joint features are estimated through a gating network composed of convolutional layers, and Softmax normalization is performed along the channel dimension to obtain the final fusion features.

[0013] Furthermore, the step of extracting hierarchical features from the fused features to obtain deep visual features specifically includes the following steps: The fused features are extracted hierarchically using multiple stacked residual group modules; each residual group module consists of a residual channel attention block; each residual channel attention block includes two convolutional layers and a channel attention module; the first convolutional layer is used to extract local features, and the spatial information of local features is compressed by global average pooling in the channel attention module. After the second convolutional layer learns the channel weights, the channel attention map of local features is generated by the Sigmoid activation function; Multiply the channel attention map with the extracted local features channel by channel to complete the feature recalibration of the channel dimension; The calibrated features are extracted through the convolution layer to obtain visual features, and the residual of the visual features and the fused features is calculated to obtain deep visual features.

[0014] Furthermore, the pixel-level reconstruction is performed according to the deep visual features to generate the final super-resolution video sequence; specifically, the following steps are included: Generate content-aware weights of deep visual features based on dynamic convolutional kernels; Constructing a normalized coordinate network, generating extended sampling coordinates for each position by superimposing a predefined kernel offset; expanding the extended sampling coordinates to obtain a continuous tensor; the continuous tensor represents the neighborhood coordinates that can be sampled at each output position; According to the continuous tensor, the deep visual features are used to obtain the neighborhood sampling features through the bilinear interpolation algorithm; and the neighborhood sampling features are rearranged to form a feature cube; After aligning the content-aware weighted interpolation to the target resolution, the Einstein summation is performed with the feature cube to complete the neighborhood pixel fusion; The residual calculation is performed between the features after the neighborhood pixel fusion and the central frame image of the video sequence to obtain the reconstructed super-resolution video.

[0015] The present invention also includes a video super-resolution reconstruction system, comprising: An acquisition module, used for acquiring a target frame image and a center frame image in a video frame sequence to be reconstructed; The explicit alignment module is used to extract the dense optical flow field between the target frame image and the center frame image through the parameter-frozen optical flow prediction model, geometrically align the target frame image to the center frame image according to the dense optical flow field, perform shallow feature encoding on the aligned frame sequence, and obtain motion consistency features; An implicit alignment module is used to extract shallow visual features of a video frame sequence by adaptively adjusting the feature fusion ratio through a parameterized residual scaling module; the extracted shallow visual features are input into a bidirectional propagation module with a spatiotemporal Transformer as a basic unit to obtain temporal forward features and temporal reverse features; wherein 3D position encoding is introduced into the spatiotemporal attention mechanism of the spatiotemporal Transformer; the temporal forward features and the temporal reverse features are fused through a multi-scale feature pyramid to obtain detail features; The reconstruction module is used to fuse the motion consistency features and detail features; perform hierarchical feature extraction on the fused features to obtain deep visual features; perform pixel-level reconstruction based on the deep visual features to generate the final super-resolution video sequence.

[0016] The present invention provides a video super-resolution reconstruction method, which has the following beneficial effects: In view of the inherent limitations of a single alignment method, the present invention constructs a dual-path architecture of explicit-implicit hybrid alignment; a dense optical flow field is generated through a parameter-frozen optical flow prediction model, and the target frame image is geometrically aligned to the center frame image, pixel-level motion compensation is completed, and an alignment feature with high fidelity and spatiotemporal consistency is output, which significantly reduces the motion blur and dislocation phenomenon during feature fusion in dynamic scenes; feature alignment is achieved through the modeling capability of the spatiotemporal Transformer on a bidirectional propagation network, combined with a multi-scale pyramid fusion strategy. This bidirectional collaborative processing mechanism effectively breaks through the traditional single-path alignment by jointly modeling historical states and future trends. The proposed method overcomes the limitations of the model's field of view in time series modeling, significantly enhances the system's ability to represent complex spatiotemporal correlations, and introduces a spatiotemporal attention mechanism guided by spatiotemporal position encoding, which effectively solves the attention ambiguity problem caused by the coupling of position and content features in traditional methods, and significantly improves the accuracy of multi-frame alignment. This method effectively solves the limitations of traditional single-path methods in motion blur and artifact suppression through feature fusion of different branches, and significantly improves the feature matching accuracy in complex motion scenes while maintaining the consistency of feature space. This solves the problem that existing technologies cannot balance reconstruction capability and robustness in complex scenes where rigid and non-rigid motion coexist. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is an architecture diagram of a method for super-resolution reconstruction of arbitrary-scale videos based on hybrid alignment in an embodiment of the present invention; Figure 2This is a schematic diagram of the structure of a hybrid alignment module in an embodiment of the present invention; Figure 3 Schematic diagram of the structure of an explicit alignment module in an embodiment of the present invention; Figure 4 This is a diagram of the implicit alignment module architecture in an embodiment of the present invention; Figure 5 A schematic diagram of bidirectional propagation in an embodiment of the present invention; Figure 6 This is a schematic diagram of the spatiotemporal Transformer structure in an embodiment of the present invention; Figure 7 Schematic diagram of the structure of the spatiotemporal attention mechanism in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of an image reconstruction module in an embodiment of the present invention; Fig. 9 Schematic diagram of the structure of the sampling module at any scale in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0019] The present invention proposes a video super-resolution reconstruction method; in view of the inherent limitations of a single alignment method, a dual-path architecture of explicit-implicit hybrid alignment is constructed. The explicit alignment branch generates a motion trajectory field through a pre-trained optical flow network and uses a differentiable bilinear sampling algorithm to complete pixel-level motion compensation, significantly enhancing the model's modeling ability for large displacement motion; the implicit alignment branch designs a spatiotemporal attention mechanism guided by spatiotemporal position encoding, and realizes feature alignment through the modeling ability of the spatiotemporal Transformer on a bidirectional propagation network combined with a multi-scale pyramid fusion strategy. An adaptive gating network is designed to realize feature fusion of different branches. At the same time, a content-aware upsampling module based on a dynamic convolution kernel is designed to achieve upsampling reconstruction of any scale by generating pixel-level adaptive weights and.

[0020] like Figure 1 As shown, the specific steps include: S1. Collect continuous low-resolution video sequences , where the height and width are H×W.

[0021] S2. In the hybrid alignment stage, a dual-path collaborative working mechanism is adopted: the explicit alignment branch generates the motion trajectory field through the pre-trained RAFT optical flow network, and uses differentiable deformation operations to achieve pixel-level motion compensation; the implicit alignment branch constructs a bidirectional temporal propagation network, captures spatiotemporal correlation features through a spatiotemporal attention mechanism based on 3D position encoding, and fuses bidirectional temporal features through a pyramid arbitrary scale fusion strategy. The two-way features are adaptively fused through a gating network, in which the gating unit dynamically allocates the fusion weights of explicit features and implicit features in the form of learnable parameters.

[0022] Temporal alignment in video super-resolution is the core link in solving motion blur and artifact problems. The present invention proposes a hybrid alignment module (Hybrid Alignment Module, HAM), which effectively solves the limitations of traditional single-path methods in motion blur and artifact suppression by constructing a dual-path collaborative architecture of implicit alignment and explicit alignment, and combining it with a dynamic gating fusion mechanism. Specifically, the explicit alignment branch uses deformable convolution to achieve motion compensation, and the implicit alignment branch captures deep feature associations through an attention mechanism, and finally the dynamic gating network realizes the adaptive fusion of different features. This hybrid paradigm significantly improves the feature matching accuracy in complex motion scenes while maintaining the consistency of the feature space, laying a reliable foundation for temporal alignment for subsequent super-resolution reconstruction. The module structure is as follows Figure 2 shown.

[0023] (1) Explicit alignment module The core principle of video super-resolution technology is to use multi-frame image data to reconstruct high-frequency details that are missing in a single frame. The explicit alignment module achieves accurate matching between frames by establishing a motion compensation model based on optical flow estimation. Unlike traditional methods that rely on the limitations of optical flow estimation accuracy, the present invention introduces a RAFT network architecture based on deep learning. Its iterative correction mechanism significantly improves the robustness of motion vectors. RAFT (Recurrent All-Pairs Field Transforms) is an end-to-end optical flow prediction model that is mainly used for pixel-level motion estimation between video frames. Specifically, this module obtains sub-pixel motion trajectories through a pre-trained spatiotemporal feature extractor, and uses a differentiable bilinear sampling algorithm to complete feature field deformation, and finally outputs alignment features with both high fidelity and spatiotemporal consistency. Module details are as follows: Figure 3 shown.

[0024] The present invention uses a pre-trained optical flow estimation model to construct an explicit alignment module, and its core process is divided into two steps: first, the dense optical flow field between the reference frame and the target frame is extracted through the parameter-frozen RAFT-Small model, and a cross-frame pixel-level displacement mapping relationship is established; second, the target frame features are geometrically aligned to the reference frame coordinate system according to the optical flow field using a differentiable spatial transformation to eliminate inter-frame motion offsets. In this process, the optical flow prediction model uses parameter freezing as a priori motion sensor to avoid optimization conflicts introduced by updating optical flow weights during training. This explicit alignment strategy significantly reduces motion blur and dislocation during feature fusion in dynamic scenes through physically interpretable motion modeling, providing a geometrically consistent spatiotemporal representation basis for subsequent tasks.

[0025] The aligned frame sequence is encoded with a lightweight convolutional network to obtain motion consistency features. . In order to eliminate redundant temporal information, an adaptive average pooling operation in the temporal dimension is used to compress multi-frame features into a single-frame compact representation vector, and the adaptive fusion of motion information is implicitly completed through a dynamic weight allocation mechanism. This architecture achieves a trade-off between computational efficiency and representation performance: on the one hand, the gradient transfer overhead is eliminated by freezing the optical flow estimation module, and on the other hand, a lightweight feature extraction architecture is used to reduce the parameter requirement. This step follows the "alignment-fusion" hierarchical processing flow to ensure the interpretability and consistency of spatiotemporal features in a unified physical coordinate system, providing a geometrically robust underlying feature representation for high-order visual tasks.

[0026] (2) Implicit alignment module In the task of video super-resolution reconstruction, the motion blur effect between adjacent frames will significantly weaken the spatial consistency constraint between pixels. Although the traditional explicit motion compensation method achieves inter-frame alignment through the motion estimation module, its computational complexity is high and it lacks robustness to fast-moving scenes. Although the implicit alignment method avoids explicit motion calculation, the one-way propagation model is limited by the temporal modeling capability and is difficult to capture the global spatiotemporal correlation of the video sequence, resulting in structural distortion or artifacts when reconstructing dynamic areas.

[0027] To solve the above problems, a feature fusion framework of any scale based on bidirectional spatiotemporal propagation is proposed, which effectively improves the super-resolution performance of dynamic scenes by constructing a closed-loop feature flow. The architecture consists of three core modules: ① Feature extractor: Robust extraction of video features is achieved through a parameterizable residual scaling strategy, and efficient feature abstraction is performed from shallow texture to deep semantics. The feature extractor consists of five layers of improved residual feature extractors in series, each of which contains a 3×3 convolutional layer, a LeakyReLU activation function, and a parameterized residual connection. In the improved residual feature extractor, learnable parameters are used to αDynamically adjust the residual connection strength so that the network can automatically adjust the feature fusion ratio according to the motion intensity of the video clip. Remove the BatchNorm layer in the ordinary residual extractor to eliminate the statistical distribution differences between different video frames and prevent feature distribution shift. Through progressive feature transfer, complete feature extraction from local details to full sentence semantics.

[0028] ② Bidirectional propagation network: Use the spatiotemporal Transformer to process the forward and backward image sequences respectively, and process the temporal dependencies and reverse context information.

[0029] ③ Feature fusion module: It integrates bidirectional propagation features through multi-scale feature pyramid to achieve effective integration of global information. Figure 4 shown.

[0030] The improved residual feature extractor is used to process the input video sequence. First, a parameterized residual scaling module is constructed using learnable parameters. α Dynamically adjust the residual connection strength so that the network can automatically adjust the feature fusion ratio according to the motion intensity of the video clip. Remove the BatchNorm layer to eliminate the statistical distribution differences between different video frames and prevent feature distribution shift. The main body of the network adopts a five-layer stacked structure, each layer contains a 3×3 convolutional layer, a LeakyReLU activation function, and a parameterized residual connection. The features of each level are progressively transferred to complete the feature extraction from local details to global semantics. The process of the feature extractor can be expressed as: (1) in, represents a learnable scaling factor, and Indicates a convolution layer with a convolution kernel of 3*3. Represented as LeakyReLU activation function. represents the input features, Represents the shallow visual features of the output. Shallow visual features refer to the more basic visual features extracted from the image. These features are usually closer to the input image and contain more pixels and detail information. Shallow visual features mainly include fine-grained information such as color, texture, edge and corner.

[0031] The extracted shallow visual features are then fed into a bidirectional propagation module with the spatiotemporal Transformer as the basic unit to construct bidirectional spatiotemporal features and process temporal dependencies. In the forward processing stage, the video sequence is input into the forward processing network composed of stacked spatiotemporal Transformers in the original time sequence, and the temporal forward features are extracted frame by frame; in the reverse processing stage, the input sequence is mirrored along the time dimension and fed into an independent reverse processing network, and future contextual information is captured through reverse temporal modeling. Finally, feature alignment is achieved through secondary temporal flipping to obtain temporal reverse features. This bidirectional collaborative processing mechanism effectively breaks through the field of vision limitations of traditional unidirectional models in temporal modeling by jointly modeling historical states and future trends, and significantly enhances the system's ability to represent complex spatiotemporal associations. Bidirectional propagation such as Figure 5 shown.

[0032] The forward propagation process is shown in formula (2): (2) in, Indicated by m It consists of cascaded spatiotemporal Transformer blocks. Indicates the positive characteristics of the timing.

[0033] The back propagation process is shown in formula (3): (3) in, Indicated by m A network composed of cascaded spatiotemporal Transformers to model the reverse video sequence. Represents the reverse feature of the time series. The structure of the spatiotemporal Transformer is as follows Figure 6 shown.

[0034] The core design of the spatiotemporal Transformer architecture follows the basic paradigm proposed by Vit, but in order to better capture the spatiotemporal position relationship in the video sequence, an attention mechanism based on 3D position encoding is proposed to optimize the video multi-frame alignment performance. The attention mechanism constructs a position-guided global spatiotemporal interaction module through the structural design of 3D position encoding and high-dimensional visual features. First, the dynamic position encoding is generated through 3D position convolution, and then the 3D position encoding is spliced ​​with the shallow visual features in the channel dimension and input into the attention mechanism module. In the attention mechanism module, the position-driven calculation strategy is used to achieve implicit inter-frame alignment and feature fusion. Specifically, the query matrix (Query) is generated with 3D dynamic position encoding as the core to encode the spatiotemporal position information of pixels in the video sequence: the 3D dynamic position encoding is mapped to the query matrix through the grouped convolution layer to ensure that the position information of different channels is expressed independently and avoid the interference of semantic features. The key matrix (Key) and value matrix (Value) are generated with visual features as the core: the key matrix is ​​generated by grouped convolution of visual features, while the value matrix is ​​generated by ordinary convolution, which retains the ability of cross-channel semantic fusion and ensures that the expression of the value matrix contains richer detailed information. The query matrix Q and the key matrix K are transposed and dot-producted to obtain the similarity matrix, which is then weighted and summed with the value matrix V to obtain the aggregated features. This decoupled design enables the network to actively query the content area to be aligned based on position when calculating attention, effectively solving the attention ambiguity problem caused by the coupling of position and content features in traditional methods, and significantly improving the accuracy of multi-frame alignment.

[0035] By dividing the Q, K, and V feature matrices into 4x4 local blocks, the original pixel-level attention mechanism is transformed into block-level similarity matching, reducing the computational complexity from the square of the number of pixels O(n²) to the square of the number of blocks, achieving a linear increase in computational efficiency. Subsequently, the block sequence is split into multiple attention heads, and each head independently calculates the correlation between blocks. The specific details are as follows Figure 7 shown.

[0036] Finally, the multi-scale pyramid fusion module is used to extract and fuse features in different time and space dimensions to enhance the model's ability to capture multi-scale information and obtain detailed features. Specifically, the temporal forward features and temporal reverse features are spliced ​​according to the spatial dimension, and the spliced ​​features are used to construct a feature pyramid through 3D average pooling operations with different downsampling coefficients according to the progressive spatial-temporal downsampling strategy; the progressive spatial-temporal downsampling strategy is used to construct a feature pyramid through 3D average pooling operations with different downsampling coefficients, forming a multi-level feature expression covering microscopic motion details to macroscopic scene semantics; then the spatial-temporal dimension alignment of cross-resolution features is achieved through a differentiable trilinear interpolation algorithm, effectively eliminating the geometric deviation between multi-scale features; the features of all scales are spliced ​​in the channel dimension and the average is taken, and then the redundant information is compressed through a lightweight convolutional layer, and a fused feature vector with both detail sensitivity and context perception is output.

[0037] (3) Dynamic Gating Fusion The present invention uses a dynamic gating fusion mechanism to achieve adaptive fusion of explicit and implicit alignment features. First, the motion consistency features output by the explicit alignment module are And the detailed features output by the implicit alignment module The concatenation is performed along the channel dimension. The weights are then estimated through a gating network consisting of convolutional layers to obtain a full channel weight map. , = + ; and perform Softmax normalization along the channel dimension to ensure that the sum of explicit-implicit weights is 1. The final fusion feature is shown in formula (4).

[0038] (4).

[0039] In S3 and image reconstruction stages, stacked residual groups are used to implement cross-layer feature reuse and improve the ability to restore high-frequency details.

[0040] Hierarchical feature extraction is performed by stacking multiple residual group modules. A residual link mechanism for shared features is adopted, and the original fused features are weighted after each residual group module to form a cross-level feature reuse link, which effectively alleviates the gradient vanishing problem and enhances the transmission efficiency of low-level features. Each residual group is composed of a residual channel attention block, and each residual attention block is composed of a "convolution layer-channel attention block-convolution layer" in series. The convolution layer is responsible for extracting local features. The channel attention module compresses spatial information through global average pooling, and uses 1×1 convolution to learn channel weights, and finally generates a channel attention map through the Sigmoid activation function. The channel attention map is multiplied channel by channel with the extracted local features to complete the feature recalibration of the channel dimension. This mechanism can dynamically enhance the feature response of important channels and suppress redundant information. The calibrated features are extracted through the convolution layer to obtain visual features, and the visual features are residually calculated with the fused features to obtain deep visual features. The specific details are as follows Figure 8 shown.

[0041] S4. Aiming at the limitation of the current video super-resolution algorithm with fixed magnification, the present invention also constructs an arbitrary scale upsampling module based on a dynamic convolution kernel; realizes nonlinear mapping of the feature map output by the image reconstruction module through cascaded convolution kernels; constructs a normalized coordinate network, and Generate extended sampling coordinates by superimposing predefined kernel offsets; expand the extended sampling coordinates to obtain a continuous tensor; the continuous tensor represents the number of samples required for each output position neighborhood coordinates; according to the continuous tensor, the neighborhood sampling features are obtained through differentiable bilinear interpolation, and the sampling results are rearranged to form a feature cube; after the dynamic convolution kernel weight W is interpolated and aligned to the target resolution, the Einstein summation calculation is performed with the feature cube to complete the neighborhood pixel fusion and obtain the target resolution scale; while maintaining content perception, video super-resolution of different scales is achieved, which significantly broadens the application scenarios of video super-resolution algorithms.

[0042] The arbitrary scale upsampling module generates content-aware weights based on dynamic convolution kernels, realizes pixel-level adaptive reconstruction through geometrically precise neighborhood sampling and Einstein summation, supports stepless scaling and enhances spatial continuity. In order to achieve image reconstruction at any scale, an arbitrary scale upsampling module based on dynamic convolution kernel generation is proposed. The arbitrary scale upsampling module consists of three parts: kernel weight prediction, deformable neighborhood sampling and adaptive fusion. The specific module process is as follows Fig. 9 shown.

[0043] First, the deep visual features output by the image reconstruction module , nonlinear mapping is achieved by cascading convolution kernels. The calculation process is shown in formula (5): (5) in, and is a 3×3 convolutional layer, Represents the RELU activation function. Output tensor Each spatial position correspond The normalized weights of the convolution kernel reflect the local semantic relevance of the input features.

[0044] To achieve any target resolution Adaptation to build a normalized coordinate network , each coordinate position satisfies For each position By superimposing predefined nuclear offsets Generate extended sampling coordinates. The calculation process is shown in formula (6): (6) in, is the scaling factor related to the input resolution. Expand the extended sampling coordinates to obtain a continuous tensor, which represents the number of samples required for each output position. The neighborhood coordinates are calculated as shown in formula (7): (7) The neighborhood sampling features are obtained by differentiable bilinear interpolation, and the sampling results are rearranged to form a feature cube. The calculation process is shown in formula (8) and formula (9): (8) (9) Next, content adaptive fusion is performed. After the dynamic convolution kernel weight W is interpolated and aligned to the target resolution, the Einstein summation calculation is performed with the feature cube to complete the neighborhood pixel fusion and obtain the target resolution scale. The calculation process is shown in formula (10): (10) in, Represents the element-wise multiplication of channel-space positions, and the summation operation is along Dimension.

[0045] Finally, the residual calculation is performed between the features after neighborhood pixel fusion and the central frame image of the video sequence to reconstruct the final super-resolution video.

[0046] This process dynamically generates position-by-position convolution kernel weights by learning the spatial semantic dependencies of input features to achieve high-quality, content-aware image upsampling. In this process, the generation of dynamic convolution kernels and feature sampling operations are jointly optimized through end-to-end gradient propagation, enabling the model to autonomously explore the spatial correlation of input features and adaptively adjust pixel fusion strategies for different semantic regions. Compared with traditional fixed interpolation technology, this module breaks through the inductive bias limitations of traditional fixed interpolation technology, realizes adaptive reconstruction of image content in the frequency domain, and effectively suppresses aliasing distortion of high-frequency components. At the same time, the decoupled design of the weight generation network and the dynamic convolution operation gives the model compatibility with arbitrary target resolutions and supports non-integer upsampling tasks.

[0047] Experimental analysis: Dataset: During the training process, Vimeo-90K is used as the training set of the network. Vimeo-90K is a widely used benchmark dataset in video super-resolution tasks. The training set has a total of 64,612 sets of high-resolution video image sequences, each sequence contains 7 frames of images, and the bicubic downsampling (BI) method is used to generate low-resolution and high-resolution frame pairs that are strictly aligned in time and space.

[0048] When testing network performance, we selected Vid4, a commonly used test dataset for current video super-resolution tasks. Vid4 contains 4 high-definition video clips (such as urban landscapes, human movements, etc.), which are often used to evaluate the algorithm's ability to handle blur, texture details, and motion consistency.

[0049] During the training process, the input image is cropped into 64*64 size image blocks as input. The network optimizer uses the Adam optimizer, where and The decay rates are set to 0.9 and 0.999 respectively, and the initial learning rate is set to , and cooperate with the cosine annealing learning rate scheduling strategy to achieve dynamic attenuation during training.

[0050] The loss function uses L1 norm as error metric, and the batch size is fixed to 4. For the training of different data sets, the model is configured with 100 full training cycles to ensure full convergence. PNSR and SSIM are used as evaluation indicators to evaluate the performance of the proposed model. During the training process, the input sequence is downsampled to an arbitrary scale (×1.0—×4.0) using a dual-scale algorithm to generate LR sequence frames of different resolutions, and a new random scaling factor is used for each batch of data input to the network.

[0051] The loss function uses the Charbonnier loss constraint network, and the calculation formula is shown in formula (11): (11) in, represents the image reconstruction result, represents a high-resolution image, , which is used to control the smoothness of the loss function near zero.

[0052] Experimental comparison: The compared methods are mainly divided into two categories: (1) single-frame image super-resolution methods of arbitrary scale: Meta-SR, LIIF, ArbSR and LTE; (2) video super-resolution methods: EDVR, BasicVSR, IconVSR, BasicVSR++. The "BI" mark added after the compared method indicates that the output result has been further processed so that its output result can adapt to the target resolution.

[0053] Table 1 Experimental results of super-resolution reconstruction under symmetric scaling factors

[0054] The experimental results of super-resolution reconstruction under symmetric scaling factors are shown in Table 1. The experimental results show that the method proposed in the present invention comprehensively surpasses the traditional single-frame arbitrary scale super-resolution method in terms of PNSR and SSIM objective evaluation indicators. The reason is that the single-frame arbitrary scale super-resolution method fails to effectively model the spatiotemporal correlation in the video sequence, resulting in excessive reliance on the local features of the single-frame image during the reconstruction process. This comparison result confirms the key role of spatiotemporal features in the video super-resolution task. The hybrid alignment module proposed in the present invention realizes the collaborative modeling of spatiotemporal features through a bidirectional propagation network, significantly improving the processing capability of video sequences. The experimental results of the video super-resolution algorithm show that the super-resolution results of the method proposed in the present invention under integer multiples of the scaling factor are the same in terms of PNSR and SSIM objective evaluation indicators. The super-resolution results under non-integer multiples of the scaling factor are better than other video super-resolution methods. This result is due to the arbitrary scale upsampling module proposed in the present invention, which realizes the adaptive weight expression based on content features through kernel weight prediction, deformable neighborhood sampling and adaptive fusion, so that the image content can effectively adjust the feature expression under different scaling factors.

[0055] Table 2 Comparison results of video super-resolution reconstruction under asymmetric scaling factors

[0056] The comparison results of video super-resolution reconstruction under asymmetric scaling factors are shown in Table 2. Experimental data show that the method proposed in the present invention exhibits the best super-resolution performance under various scaling factors, demonstrating its strong spatial adaptability and detail reconstruction capabilities; compared with traditional interpolation and single-frame image arbitrary scale super-resolution methods, it maintains excellent results in extreme scaling and medium span tasks, verifying the effectiveness of the arbitrary scale upsampling module and its universality advantage in complex scaling scenarios.

[0057] Ablation experiment: The effectiveness of the method proposed in this invention is verified by ablation experiment. All ablation experiments are verified on the Vid4 dataset. This invention mainly conducts experimental verification on the effectiveness of the hybrid alignment module, the bidirectional propagation mechanism and the adaptive upsampling network. The experimental results of the relevant experiments are shown in Tables 3 to 5, where (√) indicates that the module exists, and (×) indicates that the module does not exist.

[0058] Verification of the effectiveness of the hybrid alignment module: In order to deeply analyze the functions of each module in the hybrid alignment module, this experiment conducts ablation verification on the three key modules in the hybrid alignment module for horizontal comparison. The results of the ablation experiment are tested under different symmetric scaling factors, and the results of the ablation experiment are shown in Table 3. Among them, for the experiment of the dynamic gating network, fixed weight fusion is used for comparison, and the weights of the implicit alignment module and the explicit alignment module are both 0.5.

[0059] Table 3 Ablation experiment of hybrid alignment module

[0060] The results of ablation experiments show that the submodules in the hybrid alignment module have clear complementarity in the task of super-resolution at any scale. The explicit alignment module implements explicit spatial transformation modeling based on optical flow compensation, which can significantly improve the quality of structured feature reconstruction under low-magnification scaling factors. This may be because the information obtained at low resolution is relatively rich, and the optical flow field formed is relatively accurate; the implicit alignment module learns nonlinear mapping relationships through spatiotemporal Transformer, and can effectively restore image quality under high-magnification scaling factors; the dynamic gating network adopts a context-aware weight allocation mechanism, and the evaluation indicators in various scenarios perform well, reflecting the advantages of dynamic fusion of explicit features and implicit features. The three work together to construct a hierarchical feature alignment system. Explicit alignment ensures spatial geometric consistency, implicit alignment enhances local detail representation, and the dynamic gating network uses a scene-adaptive feature fusion strategy to enable the model to maintain good performance under different scaling factors, verifying the effectiveness of the hybrid architecture design.

[0061] Verification of the effectiveness of the bidirectional propagation mechanism: In order to further analyze the effectiveness of the bidirectional propagation mechanism in the implicit alignment module, this experiment will conduct ablation experiments according to different propagation directions for horizontal comparison. The results of the ablation experiment will be tested under different symmetric scaling factors. The results of the ablation experiment are shown in Table 4.

[0062] Table 4 Ablation experiment of bidirectional propagation mechanism

[0063] Experimental results show that in the video super-resolution task, the image quality of the reconstructed image by the two-way propagation mechanism is significantly better than that of the one-way propagation mechanism by fusing the temporal information of future frames and historical frames through forward propagation and backward propagation, and the image quality evaluation under all scaling factors achieves the best performance. Among them, the experimental results of backward propagation are significantly lower than those of forward propagation. This is because videos usually follow temporal causality, that is, historical frames determine future frames, and forward propagation is more in line with the laws of natural motion, while backward propagation requires "reverse time order" modeling, which may introduce noise due to dynamic occlusion of future frames, resulting in a decrease in the reconstructed image effect.

[0064] Verification of the effectiveness of the spatiotemporal attention mechanism: In order to deeply analyze the effectiveness of the spatiotemporal attention mechanism of the spatiotemporal Transformer, the present invention will be compared with the ordinary self-attention. The results of the ablation experiment will be tested under different symmetric scaling factors. The results of the ablation experiment are shown in Table 5.

[0065] Experimental results show that the spatiotemporal Transformer significantly outperforms the traditional self-attention model in video super-resolution tasks by introducing a spatiotemporal attention mechanism with 3D position encoding. Its PSNR and SSIM are significantly improved at all scaling factors, especially at high magnification. 3D encoding enables the model to implicitly align spatiotemporal information across frames, dynamically aggregate multi-frame features to compensate for motion blur and detail loss, thereby enhancing high-frequency detail recovery capabilities and structural coherence, verifying the core value of spatiotemporal modeling for cross-frame feature fusion in video super-resolution tasks.

[0066] Table 5 Ablation experiment of spatiotemporal attention mechanism

[0067] Based on the same inventive concept, the present invention also proposes a video super-resolution reconstruction system, comprising: The acquisition module is used to acquire the target frame image and the center frame image in the video frame sequence to be reconstructed.

[0068] The explicit alignment module is used to extract the dense optical flow field between the target frame image and the center frame image through the parameter-frozen optical flow prediction model, geometrically align the target frame image to the center frame image according to the dense optical flow field, and perform shallow feature encoding on the aligned frame sequence to obtain motion consistency features.

[0069] An implicit alignment module is used to extract shallow visual features of a video frame sequence by adaptively adjusting the feature fusion ratio through a parameterized residual scaling module; the extracted shallow visual features are input into a bidirectional propagation module with a spatiotemporal Transformer as a basic unit to obtain temporal forward features and temporal reverse features; 3D position encoding is introduced into the spatiotemporal attention mechanism of the spatiotemporal Transformer; the temporal forward features and the temporal reverse features are fused through a multi-scale feature pyramid to obtain detail features.

[0070] The reconstruction module is used to fuse the motion consistency features and detail features; perform hierarchical feature extraction on the fused features to obtain deep visual features; perform pixel-level reconstruction based on the deep visual features to generate the final super-resolution video sequence.

[0071] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A video super-resolution reconstruction method, characterized in that: The following steps are involved: Acquire a target frame image and a center frame image in a video frame sequence to be reconstructed; The dense optical flow field between the target frame image and the center frame image is extracted through the parameter-frozen optical flow prediction model. The target frame image is geometrically aligned with the center frame image according to the dense optical flow field. The aligned frame sequence is shallowly encoded to obtain motion consistency features. The shallow visual features of the video frame sequence are extracted by adaptively adjusting the feature fusion ratio through a parameterized residual scaling module; the extracted shallow visual features are input into a bidirectional propagation module with a spatiotemporal Transformer as a basic unit to obtain temporal forward features and temporal reverse features; wherein 3D position encoding is introduced into the spatiotemporal attention mechanism of the spatiotemporal Transformer; the temporal forward features and the temporal reverse features are fused through a multi-scale feature pyramid to obtain detail features; The motion consistency features and detail features are fused; hierarchical feature extraction is performed on the fused features to obtain deep visual features; pixel-level reconstruction is performed based on the deep visual features to generate the final super-resolution video sequence.

2. A video super-resolution reconstruction method according to claim 1, characterized in that: The method geometrically aligns the target frame image with the center frame image according to the dense optical flow field, performs shallow feature encoding on the aligned frame sequence, and obtains motion consistency features; specifically includes the following steps: According to the dense optical flow field, a differentiable bilinear sampling algorithm is used to geometrically align the target frame image to the center frame image; The aligned frame sequence is shallowly encoded through a lightweight convolutional network, which specifically includes upgrading the aligned frame sequence to a high-dimensional semantic feature space, compressing the upgraded multi-frame features into a single-frame representation vector using a temporal dimension adaptive average pooling operation, and implicitly completing the adaptive fusion of motion information through a dynamic weight allocation mechanism to obtain motion consistency features.

3. The video super-resolution reconstruction method according to claim 1, characterized in that: The shallow visual features of the video frame sequence are extracted by adaptively adjusting the feature fusion ratio through the parameterized residual scaling module, specifically including: extracting features from local details to global semantics of the video frame sequence using a serial cascaded five-layer stacking structure; dynamically adjusting the residual connection strength using learning parameters to adaptively adjust the feature fusion ratio, and eliminating the statistical distribution differences between different video frames by removing the BatchNorm layer; extracting features at each level through progressive transmission, and each layer includes a 3×3 convolution layer, a LeakyReLU activation function and a parameterized residual connection; the feature extraction process is expressed as: ; in, represents the learning parameters, and Indicates a convolution layer with a convolution kernel of 3*3. Represented as LeakyReLU activation function, represents the input features, Represents the shallow visual features of the output.

4. A video super-resolution reconstruction method according to claim 3, characterized in that: The extracted shallow visual features are input into a bidirectional propagation module with a spatiotemporal Transformer as a basic unit to obtain temporal forward features and temporal reverse features, which specifically includes the following steps: In the forward processing stage, shallow visual features are input into the forward processing network composed of stacked spatiotemporal Transformers in the original time sequence. The spatiotemporal attention mechanism of 3D position encoding is introduced to aggregate the shallow visual features of multiple frames in the video frame sequence, and the temporal forward features are extracted frame by frame. In the reverse processing stage, the shallow visual features are mirrored along the time dimension and fed into an independent reverse processing network. By introducing the spatiotemporal attention mechanism of 3D position encoding, the shallow visual features of multiple frames in the video frame sequence are aggregated, and the temporal forward features are extracted frame by frame.

5. A video super-resolution reconstruction method according to claim 4, characterized in that: The method of aggregating multiple shallow visual features of frames in a video frame sequence by introducing a spatiotemporal attention mechanism of 3D position encoding specifically includes the following steps: Generate 3D dynamic position encoding through 3D position convolution; The 3D dynamic position encoding and shallow visual features are segmented to obtain local blocks; Generate the query matrix Q with the local block segmented by 3D dynamic position encoding as the core, and generate the key matrix K and value matrix V with the local block segmented by shallow visual features as the core; The similarity matrix is ​​obtained by performing dot product calculation on the query matrix Q and the transpose of the key matrix K; The similarity matrix and the value matrix V are weighted and summed to obtain the aggregated features; The aggregated features are input into the attention mechanism module of Transformer, and the position-driven calculation strategy is used to achieve implicit inter-frame alignment and feature fusion.

6. The video super-resolution reconstruction method according to claim 1, characterized in that: The time series forward features and time series reverse features are fused through a multi-scale feature pyramid to obtain detail features, which specifically includes the following steps: Concatenate the time series forward features and time series reverse features according to the spatial dimension; Using the progressive spatial-temporal downsampling strategy, the concatenated features are subjected to 3D average pooling operations with different downsampling coefficients to construct a feature pyramid. The spatial-temporal dimension of feature pyramid cross-resolution features is aligned through a differentiable trilinear interpolation algorithm; All aligned features are concatenated in the channel dimension and the average is taken to output detail features with detail sensitivity and context perception.

7. The video super-resolution reconstruction method according to claim 1, characterized in that: The motion consistency feature and detail feature are fused using the dynamic gating fusion mechanism; specifically, the following steps are included: The motion consistency feature and detail feature are concatenated along the channel dimension to obtain the joint feature; The weights of the joint features are estimated through a gating network composed of convolutional layers, and Softmax normalization is performed along the channel dimension to obtain the final fusion features.

8. The video super-resolution reconstruction method according to claim 1, characterized in that: The step of extracting hierarchical features from the fused features to obtain deep visual features specifically includes the following steps: The fused features are extracted hierarchically using multiple stacked residual group modules; each residual group module consists of a residual channel attention block; each residual channel attention block includes two convolutional layers and a channel attention module; the first convolutional layer is used to extract local features, and the spatial information of local features is compressed by global average pooling in the channel attention module. After the second convolutional layer learns the channel weights, the channel attention map of local features is generated by the Sigmoid activation function; Multiply the channel attention map with the extracted local features channel by channel to complete the feature recalibration of the channel dimension; The calibrated features are extracted through the convolution layer to obtain visual features, and the residual of the visual features and the fused features is calculated to obtain deep visual features.

9. The video super-resolution reconstruction method according to claim 1, characterized in that: The pixel-level reconstruction is performed according to the deep visual features to generate the final super-resolution video sequence; specifically, the following steps are included: Generate content-aware weights of deep visual features based on dynamic convolutional kernels; Constructing a normalized coordinate network, generating extended sampling coordinates for each position by superimposing a predefined kernel offset; expanding the extended sampling coordinates to obtain a continuous tensor; the continuous tensor represents the neighborhood coordinates that can be sampled at each output position; According to the continuous tensor, the deep visual features are used to obtain the neighborhood sampling features through the bilinear interpolation algorithm; and the neighborhood sampling features are rearranged to form a feature cube; After aligning the content-aware weighted interpolation to the target resolution, the Einstein summation is performed with the feature cube to complete the neighborhood pixel fusion; The residual calculation is performed between the features after the neighborhood pixel fusion and the central frame image of the video sequence to obtain the reconstructed super-resolution video.

10. A video super-resolution reconstruction system, characterized in that: include: An acquisition module, used for acquiring a target frame image and a center frame image in a video frame sequence to be reconstructed; The explicit alignment module is used to extract the dense optical flow field between the target frame image and the center frame image through the parameter-frozen optical flow prediction model, geometrically align the target frame image to the center frame image according to the dense optical flow field, perform shallow feature encoding on the aligned frame sequence, and obtain motion consistency features; An implicit alignment module is used to extract shallow visual features of a video frame sequence by adaptively adjusting the feature fusion ratio through a parameterized residual scaling module; the extracted shallow visual features are input into a bidirectional propagation module with a spatiotemporal Transformer as a basic unit to obtain temporal forward features and temporal reverse features; wherein 3D position encoding is introduced into the spatiotemporal attention mechanism of the spatiotemporal Transformer; the temporal forward features and the temporal reverse features are fused through a multi-scale feature pyramid to obtain detail features; The reconstruction module is used to fuse the motion consistency features and detail features; perform hierarchical feature extraction on the fused features to obtain deep visual features; perform pixel-level reconstruction based on the deep visual features to generate the final super-resolution video sequence.

Citation Information

Patent Citations

  • Video super-resolution processing method and device, and storage medium

    CN112700392A

  • Video super-resolution reconstruction method and system based on multi-scale local self-attention

    CN115082308A

  • Underwater video super-resolution method based on coarse and fine granularity feature alignment

    CN118644393A

  • Image super-resolution method and system, computer equipment and storage medium

    CN119784593A

Cited By

  • Lightweight malicious traffic classification method based on deep learning

    CN120321048A

  • Video understanding method and system for carrying out key frame enhancement based on differential gradient

    CN120656110A

  • Video understanding method and system based on key frame enhancement by differential gradient

    CN120656110B

  • Gallium nitride radio frequency device defect detection method and system based on deep learning

    CN120707528A

  • Artificial intelligence prediction method for frame offset of main and standby audio and video streams

    CN120894728A