A video prediction method based on deterministic and probabilistic decoupling modeling
By using decoupling modeling method of decoupling and probabilistic modeling in video prediction, and using spectrum perception enhancers and motion diffusers to model appearance and motion, the problem of inaccurate appearance motion prediction in the prior art is solved, and high-quality video prediction is achieved.
Patent Information
- Application Number
- CN202411084075.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-08-08
AI Technical Summary
Existing video prediction technologies are difficult to achieve accurate appearance motion prediction in the real world, and are affected by complex spatial correlations, motion trends, and multi-objective interactions.
The video prediction method based on decoupling modeling of decoupling and probabilistic decoupling modeling is adopted to decouple appearance and motion through decoupling and probabilistic paths. The deterministic path uses a spectrum-aware enhancer to retain and enhance appearance details, while the probabilistic paths model motion flow through a motion diffuser.
Accurate modeling and prediction of appearance and motion is achieved, and the generated predicted video sequence quality is significantly improved, which is better than the current state-of-the-art methods.
Smart Images

Figure CN119052530B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video prediction under multi-scenario requirements, and in particular to a video prediction method based on deterministic and probabilistic decoupling modeling. Background Art
[0002] Video prediction is a self-supervised learning task that mines latent structures in unlabeled spatiotemporal data and models temporal evolution by predicting future video frames from given information. It has important applications in important fields such as autonomous driving, climate modeling, traffic flow, and world simulation. However, in the real world, complex spatial correlations, motion trends, and multi-target interactions are usually exhibited, so the quality of the output sequence of video prediction is affected and limited, and accurate appearance motion prediction cannot be achieved. Video sequences usually exhibit stable low-level appearance and dynamic high-level motion. The prediction performance can be improved by designing models based on video sequence attributes, which is not fully explored by existing methods.
[0003] Therefore, it is an urgent problem for those skilled in the art to propose a video prediction method based on deterministic and probabilistic decoupling modeling to solve the difficulties existing in the prior art. Summary of the invention
[0004] In view of this, the present invention provides a video prediction method based on deterministic and probabilistic decoupling modeling, which can decouple appearance and motion through deterministic paths and probabilistic paths to achieve accurate prediction tasks.
[0005] In order to achieve the above object, the present invention adopts the following technical solution:
[0006] A video prediction method based on deterministic and probabilistic decoupling modeling includes the following steps:
[0007] Step 1: Video sequence acquisition: Collect video data, sample video frames at specific time intervals, and adjust the size of video frames to form a T-frame real video sequence.
[0008] Step 2, real feature extraction: input the real video sequence collected in step 1 into the encoder composed of two-dimensional convolution, and independently extract the real image features of each frame of the real video image in the latent space;
[0009] Step 3, deterministic path: input the real image features extracted in step 2 into the spectrum-aware enhancer, retain and enhance the appearance details in the frequency domain, and obtain the enhanced appearance features of each frame of the real video image;
[0010] Step 4, probabilistic path: input the real image features extracted in step 2 into the motion diffuser based on the spatiotemporal diffusion state space model, model the proposed motion flow distribution, and obtain the motion flow estimation;
[0011] Step 5, appearance motion mapping: Based on the motion flow estimation obtained in step 4, the enhanced appearance features of the multi-frame real video sequence output in step 3 are integrated to obtain the multi-frame predicted image features of the predicted video sequence;
[0012] Step 6: Generate prediction sequence: Input the multi-frame prediction image features obtained in step 5 into the decoder composed of two-dimensional convolution, generate each frame of prediction video image independently, and further obtain T' frame prediction video sequence
[0013] In the above method, optionally, the spectrum sensing enhancer in step 3 includes: a plurality of spectrum sensing modules;
[0014] Among them, a re-parameterized deep dilated convolution for token mixing is introduced, the feature representation is improved by integrating parallel small dilated convolutions, and channel mixing is performed through the energy-based frequency channel mixing method.
[0015] In the above method, optionally, the energy frequency channel mixing method includes:
[0016] 1) Calculate the mean and variance of the real image features extracted in step 2:
[0017]
[0018] in, is the mean, is the variance, N = H × W is the number of tokens, H and W are the number of tokens in the length and width dimensions respectively; x i is the token feature;
[0019] 2) Based on the mean and variance calculated in 1), calculate the target token energy value;
[0020] The energy value is determined by minimizing the following process:
[0021]
[0022] Among them, e i,j The target token t i,j energy value, i∈H, j∈W, δ is a hyperparameter; the energy feature is composed of all e i,j Composition; after that, re-adjust the energy value:
[0023]
[0024] Among them, LeakyReLU is a specific activation function;
[0025] 3) Process the energy feature as a global vector C is the number of channels, c is a constant; then the global vector is transferred to Fourier space:
[0026]
[0027] An attention-based operation is introduced to enhance the appearance features.
[0028] In the above method, optionally, the content of the motion diffuser in step 4 includes:
[0029] Appearance-independent motion flow is established by computing token similarity across frames {F t→t+1} t∈{1,2,...,T-1} ;
[0030] The motion diffuser simulates the past motion flow MDiff (F t→t+1 ), to iteratively estimate the future motion flow It contains two key parts: motion flow establishment and motion flow estimation.
[0031] In the above method, optionally, the motion flow establishment content is as follows:
[0032] For real image features C is the number of channels, N = H × W is the number of tokens, H and W are the number of tokens in length and width dimensions respectively; for the real image features X of two adjacent frames t and X t+1 , calculate the dot product similarity to establish motion flow:
[0033]
[0034] in, and They are the i-th token features of the t-th frame and the j-th token features of the t+1-th frame respectively; different from the optical flow algorithm that establishes a one-to-one pixel correspondence, the motion flow captures the many-to-many token relationship between flow frames, reflecting the influence of the i-th token on the j-th token in different frames.
[0035] In the above method, optionally, the motion flow estimation comprises the following steps:
[0036] 1) Spatiotemporal state space model: The spatiotemporal state space model is constructed from the spatiotemporal state space model sequence and represented by a linear ordinary differential equation;
[0037] 2) Discretization method: The continuous parameters are converted into discrete parameters through global convolution to obtain the discretized motion flow sequence;
[0038] 3) Motion flow estimation: The motion diffuser is used to perform spatiotemporal selective scanning on the discretized motion flow sequence, including aggregation after forward and reverse scanning of the flip sequence; the estimated motion flow is expressed as:
[0039]
[0040] Among them, the motion flow from the Tth frame to the T+t'th frame is directly estimated, and the estimated motion flow is shared across scales through convolution projection; the diffusion process gradually adds noise to the motion flow; let β t represents the noise variance ratio at time t, α t =1-β t , and input the context motion flow as condition c and the denoised motion flow as condition z; motion diffuser ∈ φ (z t ; The training loss of t) is:
[0041]
[0042] In the above method, optionally, the appearance motion mapping in step 5 is specifically as follows:
[0043] Appearance motion mapping converts motion flow Appearance features X applied to T-frame real-world video sequences t Get the image features of the predicted video sequence at frame T+t' It is expressed as:
[0044]
[0045] All previous appearance features are integrated via motion flow at multiple scales; the decoder then produces predicted video sequence frames by transforming the features from latent space to pixel space.
[0046] It can be seen from the above technical solution that, compared with the prior art, the video prediction method based on deterministic and probabilistic decoupling modeling of the present invention has the following beneficial effects:
[0047] 1) We use decoupled deterministic and probabilistic modeling to construct two different latent space paths; the deterministic path contains a spectrum-aware enhancer to preserve and amplify appearance details in the frequency domain, ensuring the consistency of appearance while capturing complex long-term motion dynamics; the probabilistic path defines motion flows, explicitly describes complex many-to-many motion patterns between tokens, and models their distribution using a motion diffuser;
[0048] 2) The present invention significantly outperforms the current state-of-the-art methods and generates superior predicted video sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0050] Figure 1 A flow chart of a video prediction method based on deterministic and probabilistic decoupling modeling provided by the present invention;
[0051] Figure 2 The overall framework diagram of a video prediction method based on deterministic and probabilistic decoupling modeling provided by the present invention;
[0052] Figure 3 A structural diagram of a spectrum sensing module provided in an embodiment of the present invention;
[0053] Figure 4 A schematic diagram of appearance motion mapping provided by an embodiment of the present invention;
[0054] Figure 5 A schematic diagram of a moving diffuser provided by an embodiment of the present invention;
[0055] Figure 6 A visual comparison chart of UCF Sports provided by an embodiment of the present invention;
[0056] Figure 7 A KITTI & Caltech visualization comparison diagram provided for an embodiment of the present invention;
[0057] Figure 8 A TaxiBJ visualization comparison diagram provided by an embodiment of the present invention;
[0058] Fig. 9 A visual comparison chart of WeatherBench provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0060] In this application, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of more restrictions, the elements defined by the sentence "comprise one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0061] The present invention can be used in many general or special computing device environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multi-processor devices, distributed computing environments including any of the above devices or devices, etc.
[0062] Reference Figure 1-2 As shown, the present invention discloses a video prediction method based on deterministic and probabilistic decoupling modeling, comprising the following steps:
[0063] Step 1: Video sequence acquisition: Collect video data, sample video frames at specific time intervals, and adjust the size of video frames to form a T-frame real video sequence.
[0064] Step 2, real feature extraction: input the real video sequence collected in step 1 into the encoder composed of two-dimensional convolution, and independently extract the real image features of each frame of the real video image in the latent space;
[0065] Step 3, deterministic path: input the real image features extracted in step 2 into the spectrum-aware enhancer, retain and enhance the appearance details in the frequency domain, and obtain the enhanced appearance features of each frame of the real video image;
[0066] Step 4, probabilistic path: input the real image features extracted in step 2 into the motion diffuser based on the spatiotemporal diffusion state space model, model the proposed motion flow distribution, and obtain the motion flow estimation;
[0067] Step 5, appearance motion mapping: Based on the motion flow estimation obtained in step 4, the enhanced appearance features of the multi-frame real video sequence output in step 3 are integrated to obtain the multi-frame predicted image features of the predicted video sequence;
[0068] Step 6: Generate prediction sequence: Input the multi-frame prediction image features obtained in step 5 into the decoder composed of two-dimensional convolution, generate each frame of prediction video image independently, and further obtain T' frame prediction video sequence
[0069] Further, the spectrum sensing enhancer in step 3 includes: a plurality of spectrum sensing modules;
[0070] Among them, a re-parameterized deep dilated convolution for token mixing is introduced, the feature representation is improved by integrating parallel small dilated convolutions, and channel mixing is performed through the energy-based frequency channel mixing method.
[0071] Specifically, the spectrum sensing enhancer proposed in the present invention is composed of multiple spectrum sensing modules to improve frequency domain details; among them, a re-parameterized deep dilated convolution for token mixing is introduced, which can improve feature representation by integrating parallel small dilated convolution machines without incurring any inference cost; for channel mixing, an energy-based frequency channel mixing method is proposed.
[0072] Furthermore, the frequency channel mixing method of energy includes:
[0073] 1) Calculate the mean and variance of the real image features extracted in step 2:
[0074]
[0075] in, is the mean, is the variance, N = H × W is the number of tokens, H and W are the number of tokens in the length and width dimensions respectively; x i is the token feature;
[0076] 2) Based on the mean and variance calculated in 1), calculate the target token energy value;
[0077] The energy value is determined by minimizing the following process:
[0078]
[0079] Among them, e i,j The target token t i,j energy value, i∈H, j∈W, δ is a hyperparameter; the energy feature is composed of all e i,j Composition; after that, re-adjust the energy value:
[0080]
[0081] Among them, LeakyReLU is a specific activation function;
[0082] 3) Process the energy feature as a global vector C is the number of channels, c is a constant; then the global vector is transferred to Fourier space:
[0083]
[0084] An attention-based operation is introduced to enhance the appearance features.
[0085] Specifically, the spectrum sensing module structure is as follows: Figure 3 Shown
[0086] in, The amplitude component and phase angle component represents different information; therefore, the present invention introduces an attention-based operation to enhance:
[0087]
[0088] in and is the filter of amplitude and phase components, ⊙ is the Hadamard product; the Fourier feature is converted to the original space by inverse Fourier transform Finally transform to align the input.
[0089] Furthermore, the content of the motion diffuser in step 4 includes:
[0090] Appearance-independent motion flow is established by computing token similarity across frames {F t→t+1} t∈{1,2,...,T-1} ;
[0091] The motion diffuser simulates the past motion flow MDiff (F t→t+1 ), to iteratively estimate the future motion flow It contains two key parts, motion flow establishment and motion flow estimation, such as Figure 5 shown.
[0092] Furthermore, the motion flow is established as follows:
[0093] For real image features C is the number of channels, N = H × W is the number of tokens, H and W are the number of tokens in length and width dimensions respectively; for the real image features X of two adjacent frames t and X t+1 , calculate the dot product similarity to establish motion flow:
[0094]
[0095] in, and They are the i-th token features of the t-th frame and the j-th token features of the t+1-th frame respectively; different from the optical flow algorithm that establishes a one-to-one pixel correspondence, the motion flow captures the many-to-many token relationship between flow frames, reflecting the influence of the i-th token on the j-th token in different frames.
[0096] Furthermore, motion flow estimation includes the following steps:
[0097] 1) Spatiotemporal state space model: The spatiotemporal state space model is constructed from the spatiotemporal state space model sequence and represented by a linear ordinary differential equation;
[0098] 2) Discretization method: The continuous parameters are converted into discrete parameters through global convolution to obtain the discretized motion flow sequence;
[0099] 3) Motion flow estimation: The motion diffuser is used to perform spatiotemporal selective scanning on the discretized motion flow sequence, including aggregation after forward and reverse scanning of the flip sequence; the estimated motion flow is expressed as:
[0100]
[0101] Among them, the motion flow from the Tth frame to the T+t'th frame is directly estimated, and the estimated motion flow is shared across scales through convolution projection; the diffusion process gradually adds noise to the motion flow; let β t represents the noise variance ratio at time t, α t =1-β t , and input the context motion flow as condition c and the denoised motion flow as condition z; motion diffuser ∈ φ (z t ; The training loss of t) is:
[0102]
[0103] Specifically, in order to effectively capture long-term motion patterns, the present invention explores a new spatiotemporal state space model; the model is a system constructed by a sequence of spatiotemporal state space models, which is composed of a mapping one-dimensional function or sequence is composed; it can be expressed as a linear ordinary differential equation:
[0104] h'(t)=Ah(t)+Bx(t),y(t)=Ch(t)
[0105] in, Hidden State M is the matrix dimension; its discrete version includes a time scale parameter Δ, which converts continuous parameters A and B into discrete parameters and A common discretization method is the zero-order hold:
[0106]
[0107] y t =Ch t
[0108] The above process can be simplified to calculate a global convolution:
[0109]
[0110] Where L is the length of the space-time sequence, is a structured convolution kernel; given the motion flow of the past T frames Where N = C; these sequences are input into the motion diffuser for spatiotemporal selective scanning; specifically, after forward scanning and reverse scanning of the flipped sequences, the sequences are aggregated; and the estimated motion flow is obtained.
[0111] Among them, the motion flow from the Tth frame to the T+t'th frame is directly estimated to reduce error accumulation; the estimated motion flow is shared across scales through convolution projection to reduce multi-scale iterations; the diffusion process gradually adds noise to the motion flow; let β t represents the noise variance ratio at time t, α t =1-β t ; Take the context motion flow as condition c and the denoised motion flow as condition z, which are used as input; and get the training loss of the motion diffuser.
[0112] Furthermore, the appearance motion mapping in step 5 is specifically as follows:
[0113] Appearance motion mapping converts motion flow Appearance features X applied to T-frame real-world video sequences t Get the image features of the predicted video sequence at frame T+t' It is expressed as:
[0114]
[0115] All previous appearance features are integrated via motion flow at multiple scales; the decoder then produces predicted video sequence frames by transforming the features from latent space to pixel space.
[0116] Specifically, Figure 4 As shown, instead of only mapping the optical flow of the current frame, the present invention integrates all previous appearance features through motion flow; this process is performed at each scale, and then the decoder produces a predicted video sequence frame by converting the features from the latent space to the pixel space.
[0117] In a specific embodiment:
[0118] The present invention is based on PyTorch on NVIDIA A100 GPU, with 16 sequences in a single batch, trained using Adam optimizer and One Cycle scheduler; 5e -2 The weight decay, and from {1e -2 ,5e -3 ,1e -3}The learning rate is chosen to maintain stability; MSE loss is used to supervise the training and random depth is used for regularization.
[0119] The above steps are a description of the multi-target tracking method of the present invention. In order to further demonstrate the advancement and effectiveness of the present invention, experimental verification was carried out on the public datasets UCF Sports, KITTI&Caltech, TaxiBJ and WeatherBench.
[0120] UCF Sports contains 150 videos from different sports scenes, depicting 10 different human motion patterns; Table 1 shows the PSNR and LPIPS indicators. Compared with other methods, the proposed method has a significant performance improvement. The proposed design effectively solves related challenges by separating appearance and motion, such as Figure 6 As shown, the potential for real-world applications and scalability for high-resolution spatiotemporal data are demonstrated.
[0121] Table 1 UCF Sports test set verification results
[0122]
[0123] Generalization ability is crucial for real driving scenarios. The KITTI & Caltech datasets evaluate the generalization ability of different datasets, and predict the next frame based on the 10 frames observed previously. Table 2 shows the comparison results between the present invention and mainstream methods. The present invention achieves state-of-the-art performance in all indicators, verifying its effectiveness in modeling spatiotemporal driving data. Qualitative visualization Figure 7 As shown, it can better predict lane lines and accurately locate small entities in the distance.
[0124] Table 2 KITTI&Caltech test set verification results
[0125]
[0126]
[0127] TaxiBJ contains GPS trajectory data of Beijing taxis at intervals of 30 minutes. The model predicts 4 future frames based on 4 observed frames. The complex road network dependencies and nonlinear time dynamics have brought unprecedented challenges to traffic prediction methods. The quantitative results are shown in Table 3, and the qualitative visualization results are shown in Figure 8 In comparison, the present invention is always better than other methods and shows the smallest error.
[0128] Table 3. TaxiBJ test set verification results
[0129] method MSE↓ MAE↓ SSIM↑ PSNR↑ ConvLSTM 0.485 17.7 0.978 37.38 PredRNN 0.464 17.1 0.971 38.52 PredRNN++ 0.448 16.9 0.977 38.71 E3D-LSTM 0.432 16.9 0.979 38.75 PhyDNet 0.419 16.2 0.982 39.18 PredRNNv2 0.383 15.6 0.983 39.38 SimVP 0.414 16.2 0.982 39.17 DMVFN 3.395 45.5 0.832 31.14 TAU 0.344 15.6 0.983 39.50 iTrendRNN 0.431 17.3 0.977 38.71 Method of the present invention 0.302 14.9 0.984 39.78
[0130] WeatherBench contains climate data from 1979 to 2018. Temperature predictions are evaluated at 5.625 resolution, using 2010-2015 for training, 2016 for validation, and 2017-2018 for testing. Table 4 gives a quantitative comparison with the state-of-the-art methods, and the visualization results are shown in Fig. 9 This demonstrates the powerful capability of the present invention in capturing weather patterns for global forecasting tasks.
[0131] Table 4. WeatherBench test set verification results
[0132]
[0133]
[0134] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0135] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented by electronic hardware, computer software, or a combination of both.
[0136] In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0137] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A video prediction method based on deterministic and probabilistic decoupling modeling, characterized in that: The following steps are involved: Step 1: Video sequence acquisition: Collect video data, sample video frames at specific time intervals, and adjust the size of video frames to form a T-frame real video sequence. Step 2, real feature extraction: input the real video sequence collected in step 1 into the encoder composed of two-dimensional convolution, and independently extract the real image features of each frame of the real video image in the latent space; Step 3, deterministic path: input the real image features extracted in step 2 into the spectrum-aware enhancer, retain and enhance the appearance details in the frequency domain, and obtain the enhanced appearance features of each frame of the real video image; Step 4, probabilistic path: input the real image features extracted in step 2 into the motion diffuser based on the spatiotemporal diffusion state space model, model the proposed motion flow distribution, and obtain the motion flow estimation; Step 5, appearance motion mapping: Based on the motion flow estimation obtained in step 4, the enhanced appearance features of the multi-frame real video sequence output in step 3 are integrated to obtain the multi-frame predicted image features of the predicted video sequence; Step 6: Generate prediction sequence: Input the multi-frame prediction image features obtained in step 5 into the decoder composed of two-dimensional convolution, generate each frame of prediction video image independently, and further obtain T' frame prediction video sequence 2. The video prediction method based on deterministic and probabilistic decoupling modeling according to claim 1, characterized in that: The spectrum sensing enhancer in step 3 includes: a plurality of spectrum sensing modules; Among them, a re-parameterized deep dilated convolution for token mixing is introduced, the feature representation is optimized by integrating parallel dilated convolutions, and channel mixing is performed through the energy-based frequency channel mixing method.
3. The video prediction method based on deterministic and probabilistic decoupling modeling according to claim 2 is characterized in that: The content of the motion diffuser in step 4 includes: Appearance-independent motion flow is established by computing token similarity across frames {F t→t+1 } t∈{1,2,...,T-1} ; The motion diffuser simulates the past motion flow MDiff (F t→t+1 ), to iteratively estimate the future motion flow T→T+t' is from the Tth frame to the T+t'th frame; it contains two key parts: motion flow establishment and motion flow estimation.
4. The video prediction method based on deterministic and probabilistic decoupling modeling according to claim 3 is characterized in that: The motion flow is established as follows: For real image features C is the number of channels, N = H × W is the number of tokens, H and W are the number of tokens in length and width dimensions respectively; for the real image features X of two adjacent frames t and X t+1 , calculate the dot product similarity to establish motion flow: in, and They are the i-th token features of the t-th frame and the j-th token features of the t+1-th frame respectively; different from the optical flow algorithm that establishes a one-to-one pixel correspondence, the motion flow captures the many-to-many token relationship between flow frames, reflecting the influence of the i-th token on the j-th token in different frames.
5. The video prediction method based on deterministic and probabilistic decoupling modeling according to claim 3 is characterized in that: Motion flow estimation includes the following steps: 1) Spatiotemporal state space model: The spatiotemporal state space model is constructed from the spatiotemporal state space model sequence and represented by a linear ordinary differential equation; 2) Discretization method: The continuous parameters are converted into discrete parameters through global convolution to obtain the discretized motion flow sequence; 3) Motion flow estimation: The motion diffuser is used to perform spatiotemporal selective scanning on the discretized motion flow sequence, including aggregation after forward and reverse scanning of the flip sequence; the estimated motion flow is expressed as: Among them, the motion flow from the Tth frame to the T+t'th frame is directly estimated, and the estimated motion flow is shared across scales through convolutional projection.
6. The video prediction method based on deterministic and probabilistic decoupling modeling according to claim 3 is characterized in that: The appearance motion mapping in step 5 is as follows: Appearance motion mapping converts motion flow Appearance features X applied to T-frame real-world video sequences t Get the image features of the predicted video sequence at frame T+t' It is expressed as: Integrate all previous appearance features through motion flow at multiple scales; The decoder then produces predicted video sequence frames by transforming features from the latent space to the pixel space.
Citation Information
Patent Citations
Decoupling propagation and cascade optimization optical flow estimation method based on confidence guidance
CN117237656A
Bidirectional enhanced confrontation video prediction method
CN117315054A