Video prediction method based on multi-scale stream fusion

Through the multi-scale stream fusion method, adaptive selection of processing scale and combination of improved residual structure and feature fusion module, the challenges of complex motion and high-resolution video in traditional video prediction are solved, and high-quality future frame reconstruction is achieved.

CN120672792APending Publication Date: 2025-09-19JILIN UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510512898.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional video prediction methods have difficulty handling complex, multi-scale motions, resulting in a decrease in the quality of predicted frames. In addition, the inconsistency of motion scale in high-resolution videos leads to incorrect relative positions of objects. Existing methods have high computational costs and lose high-frequency detail information.

Method used

A multi-scale stream fusion method is adopted to adaptively select the processing scale through a scale selector. Combined with an improved residual structure and a multi-scale feature fusion module, the structural information is retained to achieve efficient motion estimation and future frame reconstruction.

Benefits of technology

It improves the ability to represent large-scale motion in real scenes, reduces artifact flickering, maintains the correct relative positions of objects in future frames, and generates natural and realistic future frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672792A_ABST
    Figure CN120672792A_ABST
Patent Text Reader

Abstract

A video prediction method based on multi-scale stream fusion belongs to the technical field of computer vision, and comprises the following steps: applying a scale selector to an input continuous video frame to obtain an appropriate processing scale factor; respectively extracting inter-frame motion features by using an improved residual structure under a selected scale path and an original scale path; fully interacting motion features of two scales by means of a multi-scale feature fusion network; estimating optical flow and shielding conditions between each current frame and a future frame; reversely distorting the current frame by using the optical flow, and generating a prediction frame in combination with the shielding condition; an existing estimation result is continuously refined in an iteration mode; and combining the reconstruction loss and the perception loss to optimize the method model. According to the method, the scale selector and the multi-scale feature fusion network are provided, so that multi-scale moving objects can be processed, a prediction algorithm is enabled to adapt to a high-resolution video, and the application universality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a video prediction method based on multi-scale stream fusion. Background Art

[0002] Video prediction is a technique for inferring future frames from a limited number of historical frames. Its core is to model motion in natural scenes. This technology has broad application prospects in areas such as autonomous driving, anomaly detection, and robot navigation.

[0003] Traditional methods use recurrent neural networks to model motion between video frames. This approach performs well in scenes with fixed backgrounds and small-scale motion, but it can easily lead to a significant decrease in the quality of predicted frames when faced with complex real-world scenes. This is because the model over-relies on local, small-scale spatiotemporal features and has difficulty covering complex real-world motion, resulting in flickering artifacts or image structure distortion. For some time, some improvement schemes have introduced complex state transition modules to enhance the ability to model long-term dependencies, but such designs significantly increase computational costs, and high-frequency detail information is easily lost during processing.

[0004] In practical applications, varying video resolutions further complicate video prediction. Compared to low-resolution video frames, high-resolution frames not only contain richer detail but also involve inconsistent motion scales. To address this challenge, existing technologies attempt to generalize resolution by determining iterative scaling factors in a fixed sequence based on routing vectors. However, these methods lack flexibility and cannot dynamically select the optimal processing scale based on the input content. Furthermore, existing methods are prone to erroneous key spatial information during downscaling, resulting in incorrect relative positioning of objects when reconstructing the predicted frame.

[0005] Therefore, video prediction methods need to be multi-scale adaptable and able to process video data at appropriate scale factors to address the challenge of inconsistent motion scales in high-resolution videos. They must also be able to efficiently estimate motion features and preserve the structural information of video frames to achieve high-quality reconstruction of future frames. Summary of the Invention

[0006] The purpose of the present invention is to provide a video prediction method based on multi-scale stream fusion, so as to solve the complex, multi-scale motion problems that are difficult to handle with traditional video prediction methods. Through the proposed scale selector, the present invention can adaptively select the appropriate scale factor according to the input to meet the challenge of inconsistent motion scale. The present invention processes the video frame at both the original scale and the selected scale, supplementing the context and structural information of the image. In addition, the multi-scale feature fusion module can retain more key structural information and realize the prediction of future frames under the premise of efficiently guiding the interaction of multi-scale features. The present invention adopts a combination of multiple modules and training methods to solve the problems raised in the above-mentioned background technology and realize natural and realistic video prediction.

[0007] The present invention proposes a video prediction method based on multi-scale stream fusion, which includes the following steps:

[0008] 1) Construct a scale selector to generate a weight vector based on the input continuous video frames and select the processing scale. Specifically, the two-branch selection problem with scale factors of 1 / 2 and 1 / 4 is regarded as a binary classification task. A network consisting of a 1×1 convolutional layer, 3 linear layers and a sigmoid output layer is used as the scale selector architecture. The probability of selecting a scaling factor of 1 / 2 is estimated. 1 / 2 , and generate branch weights w based on Bernoulli sampling 1 / 2 and w 1 / 4 ; The training and inference process of the scale selector includes the following steps:

[0009] 1.1 During the training process, the variable w is introduced 1 / 2 and w 1 / 4 , extract motion features h from the input video frame by scaling the branches to 1 / 2 and 1 / 4 respectively 1 / 2 and h 1 / 4 , and obtain the large-scale motion feature h′ by weighted addition, that is: h′=w 1 / 2 *h 1 / 2 +w 1 / 4 *h 1 / 4 ; Among them, weighted addition refers to multiplying two features point by point according to their respective weights and then adding them together to obtain the fused feature representation;

[0010] 1.2 Use straight-through estimator for gradient transfer; Specifically, the straight-through estimator is a technique that uses the gradient of continuous variables to approximate the gradient of discrete operations, which can achieve end-to-end gradient transfer; In this step, by setting w 1 / 2 The gradient is approximately p 1 / 2 To ensure the differentiability of the sampling operation and achieve effective training of sampling branch selection, the mathematical expression of the backpropagation gradient is:

[0011]

[0012] Where: O is the objective function;

[0013] 1.3 In the reasoning process, directly according to w 1 / 2 The value of activates the only branch, eliminating the need for steps 1.1 and 1.2:

[0014]

[0015] 2) Constructing an improved residual structure, dividing the input into three sub-components, and introducing a hierarchical structure and cross-connection operations to extract motion features in parallel at both the original scale and the processing scale selected in step 1). The hierarchical structure refers to the stepwise feature extraction in a multi-level manner, enabling the model to model motion information from coarse to fine scales. The cross-connection operation establishes skip connections between different layers, facilitating the transfer of low-level features to higher levels and enhancing the ability to integrate contextual information.

[0016] 3) By combining the attention of local features and global features, a multi-scale feature fusion network is constructed to align the motion features extracted in parallel in step 2) while preserving the structural information, including the following steps;

[0017] 3.1 Pyramid pooling is used to calculate global feature attention. Specifically, the pyramid pooling operation first adds the motion features extracted in parallel in step 2) element by element. Then, it passes through three adaptive average pooling layers with output sizes of 1, 2, and 4 in parallel. The three pooling results are then flattened and concatenated, and finally input into a fully connected layer to generate a global attention map. Element-by-element addition refers to the direct addition of the pixel values ​​at each corresponding position in the two feature maps.

[0018] 3.2 Use lightweight point-by-point convolution (1×1 convolution) to complete channel interaction of spatial positions and calculate the attention of local features;

[0019] 3.3 Based on the calculation results of steps 3.1 and 3.2, apply the sigmoid activation function to the motion features extracted in parallel in step 2) to obtain a fused feature map;

[0020] 4) Input the fused feature map obtained in step 3.3 into the transposed convolution layer and decode it to generate the optical flow and occlusion map between the input frames and future frames;

[0021] 5) Using the optical flow and occlusion map generated in step 4) to predict future frames, including the following steps:

[0022] 5.1 Apply pixel-level inverse warping to each input frame to generate a preliminary predicted frame. Pixel-level inverse warping involves finding the corresponding pixel value in the original image based on the position offset information of each pixel in the optical flow map to generate a new image position, thereby achieving frame distortion mapping.

[0023] 5.2 Use the occlusion map to perform weighted fusion on the preliminary predicted frames to generate the final future frame. Weighted fusion refers to assigning different weights to multiple predicted frames based on their corresponding occlusion probabilities and then combining them into a unified frame.

[0024] 6) Using the future frame predicted in step 5) as the new input, repeat steps 1) to 5) to iteratively refine the optical flow and occlusion generated in step 4) by element-by-element addition, and update the result of step 5);

[0025] 7) Using reconstruction loss and perceptual loss to jointly optimize the model, including the following steps:

[0026] 7.1 Calculating the reconstruction loss L R :

[0027] Calculate the predicted frame generated at each iteration With the real frame I t The L1 loss on the Laplacian pyramid is denoted as d and is given an exponentially increasing weight to minimize the quality difference between the predicted image and the real image, the reconstruction loss L R The mathematical expression is:

[0028]

[0029] 7.2 Calculating the Perceptual Loss L P :

[0030] Extract the final estimated future frame using the VGG19 network and real frame I t features, and calculate the L1 distance between high-level representations, the perceptual loss L P The mathematical expression is:

[0031]

[0032] Where: φ l Represented as the lth layer of the VGG19 network; h l 、w l and c l They represent the height, width and number of channels of the lth layer respectively; N is the number of iterations;

[0033] 7.3 Reconstruction loss L calculated according to step 7.1 R and the perceptual loss L calculated in step 7.2 P, calculate the overall loss L total , joint optimization method model, overall loss L total The mathematical expression is:

[0034] L total =L R +0.05*L P (5).

[0035] Compared with existing video prediction methods, the present invention can adaptively select the processing scale based on the input, improve the ability to characterize large-scale motion in real scenes, and further expand the application scenarios of the present invention. In addition, the improved residual structure can efficiently estimate motion, reduce artifact flickering in predicted future frames, and significantly reduce distortion of human senses. At the same time, the parallel processing of multiple scales and the fusion operation of preserving structural features can keep the relative positions of objects in future frames correct and maintain structural consistency as the prediction progresses. In summary, the present invention can adapt to the challenges of high-resolution video prediction and can predict more natural and realistic future frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Flowchart of a video prediction method based on multi-scale stream fusion;

[0037] Figure 2 This is a structural diagram of the scale selector;

[0038] Figure 3 Schematic diagram of the improved residual structure;

[0039] Figure 4 Schematic diagram of the multi-scale feature fusion module;

[0040] Figure 5 A schematic diagram of the overall structure. DETAILED DESCRIPTION

[0041] The present invention will be described below with reference to the accompanying drawings.

[0042] Figure 1 This is a flowchart of the present invention. This invention implements processing scale selection by constructing a scale selector, achieves efficient inter-frame motion estimation by improving the residual structure, and implements multi-scale feature fusion that preserves structural information through a multi-scale feature fusion module. Furthermore, an iterative refinement approach makes the estimated optical flow more accurate, facilitating the generation of realistic future frames. Figure 5 It is a schematic diagram of the overall structure of the present invention, and the present invention includes the following steps:

[0043] 1) Build Figure 2The scale selector shown in Figure 1 generates a weight vector based on the input continuous video frames to select the processing scale. Specifically, the two-branch selection problem with scale factors of 1 / 2 and 1 / 4 is regarded as a binary classification task. A network consisting of a 1×1 convolutional layer, 3 linear layers and a sigmoid output layer is used as the scale selector architecture. The probability p of selecting a scaling factor of 1 / 2 is estimated. 1 / 2 , and generate branch weights w based on Bernoulli sampling 1 / 2 and w 1 / 4 ; The training and inference process of the scale selector includes the following steps:

[0044] 1.1 During the training process, the variable w is introduced 1 / 2 and w 1 / 4 , extract motion features h from the input video frame by scaling the branches to 1 / 2 and 1 / 4 respectively 1 / 2 and h 1 / 4 , and obtain the large-scale motion feature h′ by weighted addition, that is: h′=w 1 / 2 *h 1 / 2 +w 1 / 4 *h 1 / 4 ; Among them, weighted addition refers to multiplying two features point by point according to their respective weights and then adding them together to obtain the fused feature representation;

[0045] 1.2 Use straight-through estimator for gradient transfer; Specifically, the straight-through estimator is a technique that uses the gradient of continuous variables to approximate the gradient of discrete operations, which can achieve end-to-end gradient transfer; In this step, by setting w 1 / 2 The gradient is approximately p 1 / 2 To ensure the differentiability of the sampling operation and to achieve effective training of sampling branch selection, the mathematical expression of the backpropagation gradient is:

[0046]

[0047] Where: O is the objective function;

[0048] 1.3 In the reasoning process, directly according to w 1 / 2 The value of activates the only branch, eliminating the need for steps 1.1 and 1.2:

[0049]

[0050] 2) If Figure 3As shown in the figure, the residual structure is improved, the input is divided into three sub-parts, and a hierarchical structure and cross-connection operation are introduced to extract motion features in parallel at the original scale and the processing scale selected in step 1). The hierarchical structure refers to the stepwise feature extraction in a multi-level manner, which enables the model to model motion information from coarse to fine. The cross-connection operation refers to the establishment of jump connections between different layers, which facilitates the transmission of low-level features to high-level features and enhances the ability to integrate contextual information.

[0051] 3) By combining the attention of local features and global features, we can construct Figure 4 The multi-scale feature fusion network shown in FIG2 aligns and fuses the motion features extracted in parallel in step 2 while preserving the structural information, including the following steps:

[0052] 3.1 Pyramid pooling is used to calculate global feature attention. Specifically, the pyramid pooling operation first adds the motion features extracted in parallel in step 2) element by element. Then, it passes through three adaptive average pooling layers with output sizes of 1, 2, and 4 in parallel. The three pooling results are then flattened and concatenated, and finally input into a fully connected layer to generate a global attention map. Element-by-element addition refers to the direct addition of the pixel values ​​at each corresponding position in the two feature maps.

[0053] 3.2 Use lightweight point-by-point convolution (1×1 convolution) to complete channel interaction of spatial positions and calculate the attention of local features;

[0054] 3.3 Based on the calculation results of steps 3.1 and 3.2, apply the sigmoid activation function to the motion features extracted in parallel in step 2) to obtain a fused feature map;

[0055] 4) Input the fused feature map obtained in step 3.3 into the transposed convolution layer and decode it to generate the optical flow and occlusion map between the input frames and future frames;

[0056] 5) Using the optical flow and occlusion map generated in step 4) to predict future frames, including the following steps:

[0057] 5.1 Apply pixel-level inverse warping to each input frame to generate a preliminary predicted frame. Pixel-level inverse warping involves finding the corresponding pixel value in the original image based on the position offset information of each pixel in the optical flow map to generate a new image position, thereby achieving frame distortion mapping.

[0058] 5.2 Use the occlusion map to perform weighted fusion on the preliminary predicted frames to generate the final future frame. Weighted fusion refers to assigning different weights to multiple predicted frames based on their corresponding occlusion probabilities and then combining them into a unified frame.

[0059] 6) Using the future frame predicted in step 5) as the new input, repeat steps 1) to 5) to iteratively refine the optical flow and occlusion generated in step 4) by element-by-element addition, and update the result of step 5);

[0060] 7) Using reconstruction loss and perceptual loss to jointly optimize the model, including the following steps:

[0061] 7.1 Calculating the reconstruction loss L R :

[0062] Calculate the predicted frame generated at each iteration With the real frame I t The L1 loss on the Laplacian pyramid is denoted as d and is given an exponentially increasing weight to minimize the quality difference between the predicted image and the real image, the reconstruction loss L R The mathematical expression is:

[0063]

[0064] 7.2 Calculating the Perceptual Loss L P :

[0065] Extract the final estimated future frame using the VGG19 network and real frame I t features, and calculate the L1 distance between high-level representations, the perceptual loss L P The mathematical expression is:

[0066]

[0067] Where: φ l Represented as the lth layer of the VGG19 network; h l 、w l and c l They represent the height, width and number of channels of the lth layer respectively; N is the number of iterations;

[0068] 7.3 Reconstruction loss L calculated according to step 7.1 R and the perceptual loss L calculated in step 7.2 P , calculate the overall loss L total , joint optimization method model, overall loss L total The mathematical expression is:

[0069] L total =L R +0.05*L P (5).

Claims

1. A video prediction method based on multi-scale stream fusion, characterized in that: The following steps are involved: 1) Construct a scale selector to generate a weight vector based on the input continuous video frames and select the processing scale; specifically, the two-branch selection problem with scale factors of 1 / 2 and 1 / 4 is regarded as a binary classification task, and a network consisting of a 1×1 convolutional layer, 3 linear layers and a sigmoid output layer is used as the scale selector architecture. The probability p of selecting a scaling factor of 1 / 2 is estimated. 1 / 2 , and generate branch weights w based on Bernoulli sampling 1 / 2 and w 1 / 4 ; The training and inference process of the scale selector includes the following steps: 1.1 During the training process, the variable w is introduced 1 / 2 and w 1 / 4 , extract motion features h from the input video frame by scaling the branches to 1 / 2 and 1 / 4 respectively 1 / 2 and h 1 / 4 , and obtain the large-scale motion feature h′ by weighted addition, that is: h′=w 1 / 2 *h 1 / 2 +w 1 / 4 *h 1 / 4 ; Among them, weighted addition refers to multiplying two features point by point according to their respective weights and then adding them together to obtain the fused feature representation; 1.2 Use the straight-through estimator for gradient transfer, let w 1 / 2 The gradient is approximately p 1 / 2 The gradient of is used to realize the differentiability of the sampling operation. The mathematical expression of the backpropagation gradient is: Where: O is the objective function; 1.3 In the reasoning process, directly according to w 1 / 2 The value of activates the only branch, eliminating the need for steps 1.1 and 1.2: 2) Constructing an improved residual structure, dividing the input into three sub-components, and introducing a hierarchical structure and cross-connection operations to extract motion features in parallel at both the original scale and the processing scale selected in step 1). The hierarchical structure refers to the stepwise feature extraction in a multi-level manner, enabling the model to model motion information from coarse to fine scales. The cross-connection operation establishes skip connections between different layers, facilitating the transfer of low-level features to higher levels and enhancing the ability to integrate contextual information. 3) By combining the attention of local features and global features, a multi-scale feature fusion network is constructed to align the motion features extracted in parallel in step 2) while preserving the structural information, including the following steps; 3.1 Pyramid pooling is used to calculate global feature attention. Specifically, the pyramid pooling operation first adds the motion features extracted in parallel in step 2) element by element. Then, it passes through three adaptive average pooling layers with output sizes of 1, 2, and 4 in parallel. The three pooling results are then flattened and concatenated, and finally input into a fully connected layer to generate a global attention map. Element-by-element addition refers to the direct addition of the pixel values ​​at each corresponding position in the two feature maps. 3.2 Use lightweight point-by-point convolution (1×1 convolution) to complete channel interaction of spatial positions and calculate the attention of local features; 3.3 Based on the calculation results of steps 3.1 and 3.2, apply the sigmoid activation function to the motion features extracted in parallel in step 2) to obtain a fused feature map; 4) Input the fused feature map obtained in step 3.3 into the transposed convolution layer and decode it to generate the optical flow and occlusion map between the input frames and future frames; 5) Using the optical flow and occlusion map generated in step 4) to predict future frames, including the following steps: 5.1 Apply pixel-level inverse warping to each input frame to generate a preliminary predicted frame. Pixel-level inverse warping involves finding the corresponding pixel value in the original image based on the position offset information of each pixel in the optical flow map to generate a new image position, thereby achieving frame distortion mapping. 5.2 Use the occlusion map to perform weighted fusion on the preliminary predicted frames to generate the final future frame. Weighted fusion refers to assigning different weights to multiple predicted frames based on their corresponding occlusion probabilities and then combining them into a unified frame. 6) Using the future frame predicted in step 5) as the new input, repeat steps 1) to 5) to iteratively refine the optical flow and occlusion generated in step 4) by element-by-element addition, and update the result of step 5); 7) Using reconstruction loss and perceptual loss to jointly optimize the model, including the following steps: 7.1 Calculating the reconstruction loss L R : Calculate the predicted frame generated at each iteration With the real frame I t The L1 loss on the Laplacian pyramid is denoted as d and is given an exponentially increasing weight to minimize the quality difference between the predicted image and the real image, the reconstruction loss L R The mathematical expression is: 7.2 Calculating the Perceptual Loss L P : Extract the final estimated future frame using the VGG19 network and real frame I t features, and calculate the L1 distance between high-level representations, the perceptual loss L P The mathematical expression is: in: φ l Represented as the lth layer of the VGG19 network; h l 、w l and c l They represent the height, width and number of channels of the lth layer respectively; N is the number of iterations; 7.3 Reconstruction loss L calculated according to step 7.1 R and the perceptual loss L calculated in step 7.2 P , calculate the overall loss L total , joint optimization method model, overall loss L total The mathematical expression is: L total =L R +0.05*L P (5)