A bidirectional enhanced adversarial video prediction method

CN117315054BActive Publication Date: 2026-09-25TONGJI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311124502.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2026-09-25
Estimated Expiration
2043-09-01

AI Technical Summary

Technical Problem

然而,GANs很容易在基于条件时面临模式坍塌,随机隐变量容易被模型忽略,从而难以根据观测视频帧生成真实的未来帧

Benefits of technology

[0033](1)本发明构建的双向增强随机对抗视频预测框架,通过两组变分自编码器-生成对抗网络,分别用于顺序预测和逆序预测,充分利用相邻帧之间2种不同方向的转换关系,充分挖掘像素运动趋势变化与长时间预测的不确定性,使得逆向预测能为正向预测提供额外有用信息,提高了视频预测结果的准确性,可适应时空变化的各种预测场景,生成真实且多样的轨迹。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315054B_ABST
    Figure CN117315054B_ABST
Patent Text Reader

Abstract

The application relates to a bidirectional enhanced adversarial video prediction method, which comprises the following steps: constructing a bidirectional enhanced random adversarial video prediction framework, including two groups of variational autoencoder-generative adversarial networks, which are respectively used for sequential prediction and reverse sequence prediction; wherein an input video frame sequence sequentially passes through the encoder of the sequential prediction, the decoder and the encoder of the reverse sequence prediction, and the decoder of the sequential prediction for cyclic reconstruction; and the trained bidirectional enhanced random adversarial video prediction framework is used for video prediction. Compared with the prior art, the application can predict and generate a higher-quality video frame sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video prediction technology, and in particular to a bidirectional enhanced adversarial video prediction method. Background Technology

[0002] Video prediction has high application value in tasks such as predicting human behavior, traffic flow, and typhoon rainfall. Based on existing video frame sequences, video prediction generates a sequence of future video frames containing the original pixels, which can present information about the future actions of moving objects.

[0003] The movements of objects in a video can be complex, diverse, and constantly changing over time. While predicting the next few frames can yield relatively accurate predictions in the short term, the probability space changes drastically after just a few frames due to spatial information. This makes the future prediction in video inherently multimodal. How to exploit this uncertainty in long-term prediction remains a challenge.

[0004] Deterministic Video Prediction: Recurrent Neural Networks (RNNs) are widely used for deterministic video prediction to model temporal dependencies in videos. ConvLSTM extends FC-LSTM by incorporating convolutional structures in both input-to-state and state-to-state transitions, thus better capturing spatiotemporal relationships. PredRNN enables the memory state to be updated along the vertical direction of stacked RNN layers and the horizontal direction of all RNN states; updates are performed through an ST-LSTM unit that can simultaneously capture and memorize spatiotemporal representation features. A major challenge in video prediction is the highly dynamic and stochastic nature of the motion of moving objects. RNN-based deterministic video prediction methods achieve relatively ideal results in terms of accuracy. However, deterministic predicted video outputs average all possible futures, resulting in deterministic predictions that are highly fitted to the dataset. Methods using deterministic models and loss functions cannot handle the inherent uncertainty of object motion, such as mean squared error (MSE), which averages possible futures, leading to fuzzy predictions.

[0005] Stochastic video prediction: Modeling the uncertainty of future motion of a moving object, these methods are mainly based on VAEs or GANs. VAEs perform predictions by training a latent variable model. VoxelFlow is an architecture combining traditional optical flow prediction and autoencoders, which encodes the movement of pixels in existing video frames into precise 3D pixel feature representations. SV2P can generate different possible futures based on each sample of the latent variables. Although VAE-based methods can model the distribution of different possible futures, the predicted distribution is still completely factored down to the pixel level, which often produces fuzzy predictions.

[0006] GANs can achieve realistic generation in the process of the generator and discriminator competing against each other, and have recently been widely used for video prediction. Vondrick et al. used a spatiotemporal convolutional architecture to extract the foreground from the background and used GANs for unconditional video generation. MocoGAN uses a sequence of random variables to generate predicted video frames, where each random variable contains entity and action information. However, GANs are prone to mode collapse when conditional, and random latent variables are easily ignored by the model, making it difficult to generate realistic future frames based on observed video frames. In addition, existing methods have few effective attempts at modeling the common latent space based on video sequences.

[0007] Therefore, there is an urgent need to design a video prediction method that can predictively generate higher quality video frame sequences. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a bidirectional enhanced adversarial video prediction method that can predict and generate higher quality video frame sequences.

[0009] The objective of this invention can be achieved through the following technical solutions:

[0010] This invention provides a bidirectional enhanced adversarial video prediction method, which includes the following steps:

[0011] A bidirectional enhanced stochastic adversarial video prediction framework is constructed, comprising two sets of variational autoencoder-generative adversarial networks, used for sequential prediction and reverse prediction respectively; wherein, the input video frame sequence is cyclically reconstructed by sequential prediction encoder, reverse prediction decoder and encoder, sequential prediction decoder;

[0012] A trained bidirectional augmented stochastic adversarial video prediction framework is used for video prediction.

[0013] Preferably, each set of variational autoencoder-generative adversarial networks includes an encoder E for variational inference, a generator G for generating sequential / reverse video predictions, and a discriminator D for distinguishing whether the reconstructed and direction-transformed frame sequences are real; where k=1 represents sequential prediction and k=2 represents reverse prediction, that is, sequential prediction uses E1, G1 and D1, and reverse prediction uses E2, G2 and D2.

[0014] Preferably, there is a potential space for bidirectional data sharing between the two sets of variational autoencoders-generative adversarial networks, specifically:

[0015] For sequential video frame sequence x (1) and reverse video frame sequence x (2)Define a shared sequence of latent variables z that can cover distributions in two directions and can be generated from distributions in either direction;

[0016] Encoders E1 and E2 share the weights of the last c1 layer, and generators G1 and G2 share the weights of the beginning c2 layer, constructing a bidirectional data sharing potential space, where c1 and c2 are set values.

[0017] Preferably, for a potential space for bidirectional data sharing, using The prior Gaussian distribution associated with the latent space is represented by the latent variables obtained, which are used for regular prediction, self-reconstruction, and cyclic reconstruction.

[0018] Preferably, the self-reconstruction and cyclic reconstruction process in the sequential prediction case specifically includes:

[0019] Self-reconstruction: Real Adjacency Frames Compared to the previous moment The frames are grouped and encoded into latent variables by the E1 encoder. Generators with implicit variables To generate the next time frame conditionally Achieve self-reconstruction;

[0020] Cyclic Reconstruction: Sequential Prediction The process is repeated cyclically through E1, G2, E2, G1.

[0021] Preferably, the cyclic reconstruction specifically refers to: the truth value at the current time. The truth value of the previous moment Encoded as latent variables by encoder E1 Latent variables and Input to generator G2, generate The latent variables are obtained after E2 encoding by the encoder. Latent variables and The data is input into generator G1 to generate the reconstructed frame for the current time step. in, Loss reconstruction through loop Conduct training;

[0022] Will and The input is sent to the discriminator D1 for differentiation. and The input is fed into the discriminator D2 for differentiation.

[0023] Preferably, for the generator, the method includes using a prior distribution for latent variable sampling in the adversarial generative network and using an encoder to encode adjacent frames to approximate the posterior distribution for latent variable sampling in the variational autoencoder; the two distributions are connected to the adversarial generative network and the variational autoencoder by minimizing the KL divergence loss.

[0024] Preferably, the objective function expression for optimizing the bidirectional enhanced stochastic adversarial video prediction framework is:

[0025]

[0026] In the formula, Let be the objective function of the variational autoencoder; The objective function for generating the GAN (Generative Adversarial Network); Let this be the objective function corresponding to the cycle consistency constraint;

[0027] The training process is carried out using an alternating training method. The specific training process is as follows: First, fix E1, E2, G1 and G2, and update D1 and D2; then, fix D1 and D2, and update E1, E2, G1 and G2.

[0028] Preferably, the objective function of the variational autoencoder includes KL divergence. To standardize approximate posterior For the prior distribution p(z) t-1 Approximate terms of )

[0029] The objective function corresponding to the cyclic consistency constraint includes a KL term for penalizing latent codes that deviate from the prior distribution, and a term for ensuring that the sequential prediction restores the original input after two direction transformations.

[0030] Preferably, the objective function of the generative adversarial network (GAN) includes a regular adversarial loss term, a loss term for self-reconstruction, and a loss term for cyclic reconstruction.

[0031] Among them, the loss term used for cyclic reconstruction uses a distribution. The latent code sampling represents the approximate posterior distribution in the opposite direction to the prediction in direction k.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] (1) The bidirectional enhanced random adversarial video prediction framework constructed in this invention uses two sets of variational autoencoders-generative adversarial networks for sequential prediction and reverse prediction, respectively. It makes full use of the two different directions of transformation relationship between adjacent frames, fully explores the uncertainty of pixel motion trend changes and long-term prediction, so that reverse prediction can provide additional useful information for forward prediction, improve the accuracy of video prediction results, adapt to various prediction scenarios with spatiotemporal changes, and generate realistic and diverse trajectories.

[0034] (2) Based on the assumption of bidirectional data sharing latent space and the advantages of cyclic consistency in image transformation, cyclic consistency is used to further promote the unification and sharing of the bidirectional frame sequence latent space, thereby enhancing sequential prediction and providing more information for capturing long-term temporal dependencies in sequential prediction; by establishing the joint distribution of the two video directional domains through weight sharing and cyclic consistency, the commonality and consistency of the latent joint distribution are ensured, so that inverse prediction can enhance sequential prediction; the uniformity of the common latent space is guaranteed by cyclic consistency loss, and the assistance of inverse prediction enables sequential prediction to more accurately grasp the spatiotemporal changes of actions. Attached Figure Description

[0035] Figure 1 This is a schematic diagram of the bidirectional enhanced random adversarial video prediction framework of the present invention;

[0036] Figure 2 This is a schematic diagram of the VAE-GANs structure for bidirectional prediction according to the present invention;

[0037] Figure 3 This is a schematic diagram of the direction conversion process;

[0038] Figure 4 The results of the metric changes of the four models on the KTH dataset over time, frame by frame;

[0039] Figure 5 This is a qualitative result of the Moving-MNIST dataset used in the examples;

[0040] Figure 6 This is a qualitative result of the KTH dataset in the example. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0042] Example

[0043] This embodiment provides a bidirectional enhanced adversarial video prediction method. The method constructs a bidirectional enhanced stochastic adversarial video prediction framework BEAVP, which includes two sets of variational autoencoder-generative adversarial networks (VAE-GANs) for random sequential prediction and reverse prediction, respectively. The input video frame sequence is cyclically reconstructed by passing through a sequential prediction encoder, a reverse prediction decoder, and an encoder and a sequential prediction decoder.

[0044] The method of this embodiment will now be described in detail.

[0045] The bidirectional augmented stochastic adversarial video prediction framework fully utilizes the transformation relationships between two different directions between adjacent frames, enabling inverse prediction to provide additional useful information for forward prediction. The two different video frame sequences are treated as two distinct video domains. Sequential prediction is trained using the original video dataset, while inverse video sequence prediction is trained using the original dataset after reversing its order.

[0046] The bidirectional augmented random adversarial video prediction framework is based on VAE-GANs, such as... Figure 1 As shown, it includes six sub-networks: two video sequence encoders E1 and E2, frame generators G1 and G2 for two domains, and two adversarial discriminators D1 and D2. Sequential prediction uses E1, G1, and D1, while reverse prediction uses E2, G2, and D2.

[0047] For a more detailed breakdown, see [link to original text]. Figure 2 and 3 The superscript (k) in the following text represents the input, output and intermediate variables for the two directions of prediction, where k=1 indicates sequential prediction and k=2 indicates reverse prediction.

[0048] For ease of recording during the derivation, the initial video frame sequence (2 video frames) is recorded as using the first frame.

[0049] During training, taking sequential prediction as an example, the generator is expected to learn a deterministic mapping of the initial video frames. and a sequence of latent variables Mapping to a predicted future frame sequence The latent variable sequence encompasses various factors that influence future predictions and the interaction between two predictions; where T is the sequence length.

[0050] Assuming a potential space for bidirectional data sharing, using The prior Gaussian distribution associated with the latent space is represented by the latent variables obtained, which are used as conditions for regular prediction, self-reconstruction, and cyclic reconstruction.

[0051] During testing, the prior Gaussian distribution p(z) is used.t Obtain the sequence of hidden variables and input it into the generator to direct the generator's generation.

[0052] The framework consists of a pair of VAE-GANs. Each VAE-GAN contains an encoder for variational inference, a generator for generating video predictions in a certain order, and a discriminator for distinguishing whether the reconstructed and orientation-transformed frame sequences are real. and This represents the reconstructed video frame. This represents the video frame after the direction has been changed.

[0053] Encoders E1 and E2 share the weights of the last c1 layer (represented by dashed lines), and generators G1 and G2 share the weights of the beginning c2 layer, constructing a bidirectional data sharing potential space, where c1 and c2 are set values.

[0054] The directional shift of weight sharing and cycle consistency promotes the sharing and unification of the common potential space for the two directions of enhanced prediction.

[0055] The generator takes a frame from a previous time step and a latent variable to generate the frame for the current time step. The previous time step frame is denoted as... This indicates that the frame may be the ground truth in the initial stage. Or the previous prediction result In the standalone GANs module, latent variables are sampled using a prior distribution; while in the VAE module, the latent variables are sampled using an encoder that encodes adjacent frames to approximate the posterior distribution. These two distributions bridge the gap between the standalone GANs module and the VAE module by minimizing the KL divergence loss.

[0056] VAE-GANs: A pair of VAE-GANs is used for sequential and reverse prediction respectively. For prediction in a single direction, this part of the structure is similar to SAPV. VAE-GANs consist of E k G k and D k The composition is as follows: k=1 indicates sequential prediction, and k=2 indicates reverse prediction.

[0057] For sequential prediction, the generator G1 needs to be derived from the prior distribution in the separate GANs components. Sampling; also requires conditional probability distributions in a separate VAE section. Modeling is performed where the distribution uses a Laplace distribution with fixed variance, and the mean is given by... Given.

[0058] Data Likelihood Through the posterior probability density function of the latent variables It is difficult to maximize directly and requires approximate estimation through variational inference; therefore, a... Parameterized distribution To approximate it. via encoder Modeling. An encoder composed of deep neural networks receives adjacent sequential frames, containing various ambiguous information from frame order transitions. An approximate distribution will... Sampling yields latent variables used for generation

[0059] During training, real adjacent frames Compared to the previous moment The frame group is encoded into latent variables by the E1 encoder. The generator uses this latent variable as a condition to generate the next time frame. This is the self-reconstruction process that connects VAEs and GANs.

[0060] Discriminator D1 distinguishes between real frame samples and those generated by generator G1, outputting true and false respectively. In practice, three binary classifiers are used to differentiate between samples generated by a regular generator, self-reconstructed samples, and cyclically reconstructed samples. Reverse prediction is based on another set of VAE-GANs, and the process is almost identical to the sequential prediction process.

[0061] Weight sharing: An assumption is made about the potential space for bidirectional data sharing, that is, for sequential and reverse-order video frame sequences x. (1) and x (2) There exists a shared latent variable sequence z that can cover distributions in both directions and can be generated from distributions in either direction. Previous work has demonstrated the role of weight sharing in promoting the unification and sharing of this latent space. By sharing the weights of the last few layers of the encoder and the first few layers of the generator, the unification and sharing of the latent space of bidirectional frame sequences can be facilitated.

[0062] Cyclic Consistency: Both backward and forward prediction can capture different transformation relationships between neighboring frames of the same object, making bidirectional data sharing latent space naturally adaptable. Based on the assumption of bidirectional data sharing latent space and the advantages of cyclic consistency in image translation, using cyclic consistency can further promote the unification and sharing of bidirectional frame sequence latent space, thereby enhancing forward prediction and providing more information for capturing long-term temporal dependencies.

[0063] Define the transformation function F 1→2 Sequential video frame sequences can be mapped to reverse order:

[0064]

[0065] Similarly, there are:

[0066]

[0067] A cyclic consistency constraint for a single directional transition has been established:

[0068]

[0069] Figure 4 This explains the process of direction conversion.

[0070] Sequential prediction The reconstruction is achieved by sequentially passing through E1, G2, E2, and G1.

[0071] use This represents the loss from cyclic reconstruction.

[0072] Learning: Framework optimization is performed from three loss components: VAE, GAN, and cycle consistency.

[0073] The optimization objective for a single frame is to maximize the log-likelihood, which can be transformed into optimizing the ELBO (Elastic Low Beginning) lower bound of the log-likelihood:

[0074]

[0075] Based on the assumptions in VAE Obtain the parameter as The Laplace distribution of (location parameter) and β (scale parameter). Since this is the output of the generator network, the simplified form of the first term in the above equation can be obtained as follows:

[0076]

[0077] Where C is a constant. Based on the above derivation, the lower bound of log-likelihood ELBO is extended to time series, yielding the objective function of the VAE part:

[0078]

[0079] In the first item, The encoder, i.e., the distribution of the inference network, is described in this paper by... Model it; the second term is the KL divergence loss. in Denotes the KL divergence, used to normalize the approximate posterior. For the prior distribution p(z) t-1 The approximation of ) is given by λ1 and λ2, which control the weights of the two terms above.

[0080] The objective function for the GAN part used for cyclic reconstruction is as follows:

[0081]

[0082] Among them, the latent code is from the distribution The sampling is performed, and the distribution is the approximate posterior distribution in the opposite direction to the prediction in direction k.

[0083] The discriminator in the cyclic reconstruction learns to make the frame sequences after all two directional transformations as close as possible to the corresponding target ground sequences. For the GAN objective function, a parameter related to... Similar and Items, but using from distribution respectively and common potential distribution The sampled latent codes are used as both the conventional adversarial loss and the loss for self-reconstruction. The effects of these three losses are controlled by λ3, λ4, and λ5, respectively.

[0084] loss function

[0085]

[0086] In the formula, Using distribution respectively The sampled latent codes are used as conventional adversarial loss, for self-reconstruction, and for cyclic reconstruction, respectively.

[0087] Cyclic consistency constraints:

[0088]

[0089] In the formula, the two KL terms are used to penalize latent codes that deviate from the prior distribution; the last term is used to ensure that the sequential prediction restores the original input as much as possible after two directional transformations; λ6 and λ7 are used to control the weights of the two different categories of target terms.

[0090] Framework objective function:

[0091]

[0092] Constructing a shared latent space using an alternating training method:

[0093] First, fix E1, E2, G1, and G2, and update D1 and D2;

[0094] Then, fix D1 and D2, and update E1, E2, G1 and G2.

[0095] Experimental setup

[0096] Datasets: The method was tested on two widely used datasets for video prediction. The MovingMNIST handwritten digit movement dataset contains various variations of all handwritten Arabic numerals, including rotation, displacement, scaling, and shrinking. The KTH human motion dataset encompasses six common human actions performed by different people in four different scenes: walking, jogging, running, boxing, clapping, and waving. The frame sequence length was set to 20, and the prediction of the next 10 frames was conditioned on the first 10 frames. Sequential prediction used the original dataset, while reverse prediction reversed the original dataset to obtain the video frame sequence in the opposite direction. Sequences predicted in both directions were randomly sampled during training and testing according to the set sequence length.

[0097] Evaluation indicators:

[0098] For quantitative evaluation, three commonly used frame-by-frame evaluation metrics were used: PSNR, SSIM, and LPIPS. LPIPS differs from the other two in that it is a perceptual metric more closely related to human judgment. For the first two metrics, higher values ​​correspond to better results, while for LPIPS, the opposite is true.

[0099] Implementation details:

[0100] The model was trained using the Adam optimizer with an initial learning rate of 0.0002 and an exponential decay rate of 0.25. Based on the metrics calculated on the validation set, hyperparameters λ1 to λ7 were set to 1000, 0.1, 100, 10, 10, 0.1, and 1000, respectively. Each mini-batch contained a sequence of frames in both sequential and reverse order.

[0101] Compared with existing methods:

[0102] The method of this invention was compared with several state-of-the-art video prediction methods based on image generation, including: temporal variants of SV2P, SVG-LP, SAVP and VAE-based variants, and SLAMP. The method of this invention achieved excellent results on four metrics of the KTH dataset, as shown in Table 1.

[0103] The method of this invention achieves state-of-the-art results in SSIM, PSNR, and LPIPS metrics compared to other state-of-the-art methods. Compared to SAPV, which is also based on VAE-GANs, the method of this invention improves performance by over 20% in LPIPS. Figure 4 BEAVP exhibits a more gradual performance degradation in frame-by-frame prediction over time steps compared to other methods, maintaining a degree of superior prediction capability over longer periods. Figure 5 and Figure 6The quantization results for the Moving-MNIST and KTH datasets are presented separately, with SV2P using a time-variant variant. Figure 6 In the analysis, all three methods can achieve intuitive semantic modeling on the first 10 frames of the image, clearly indicating the character's movements. In frames t=26 and later, the method of this invention still maintains its ability to clearly model, capturing the long-term temporal dependence of the character's movements.

[0104] Table 1. Average evaluation results of different indicators on the KTH dataset.

[0105] SV2P time-variant 0.79 26.11 0.21 SVG-LP 0.77 24.03 0.14 SAVP-VAE 0.77 25.98 0.13 SAVP 0.68 23.90 0.12 SLAMP 0.80 24.85 0.10 BEAVP (This invention) 0.82 26.87 0.09

[0106] Figure 4 This section quantifies the frame-by-frame metrics changes of four models on the KTH dataset. All models are conditioned on the initial 10 frames and then predict the next 30 frames. The bold vertical line at frame 20 indicates that the model itself was trained to predict the next 10 frames. The mean PSNR, SSIM, and LPIPS for all test videos are shaded with 95% confidence intervals. Higher PSNR and SSIM indicate better results, while lower LPIPS indicate even better results.

[0107] Figure 6 The qualitative results from the KTH dataset showcase three examples: predicting human high-fives, jogging, and running. Even in unidirectional video prediction, pixel movement is not synchronous; generally, human movements exhibit antisynchronicity. For example, when a person walks forward, one arm tends to move forward while the other moves backward. The KTH dataset contains abundant video data of this recurring antisynchronicity in human movements. The model can fully utilize the regular human movement information from the antisynchronous prediction to enhance sequential prediction and achieve better prediction performance.

[0108] ablation experiment

[0109] The effectiveness of each component was validated through ablation experiments on the KTH dataset by removing key components from the architecture, as shown in Table 2. Since reverse prediction uses pedestrian motion data from the opposite direction, the bidirectional prediction architecture is considered a form of data augmentation. However, using only this component does not fully extract the additional information from reverse prediction. But by adding weight sharing, the LPIPS decreased from 0.20 to 0.15. Similarly, a structure with cycle consistency, allowing reverse predicted trajectories to participate in directional translation, reduced LPIPS from 0.20 to 0.13. Ultimately, by utilizing both weight sharing and cycle consistency, the advantages of both are combined, resulting in improved performance across all metrics.

[0110] Table 2 Ablation Experiment

[0111] 24.50 0.75 0.21 √ 24.79 0.75 0.20 √ √ 26.05 0.77 0.15 √ √ 25.70 0.78 0.13 √ √ √ 26.56 0.82 0.10

[0112] In summary, the bidirectional enhanced stochastic video prediction framework proposed in this invention assumes bidirectional sharing of the latent space and establishes a bridge between forward and backward prediction through weight sharing and cycle consistency, yielding convincing results in both quantitative and qualitative evaluation.

[0113] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A bidirectional enhanced adversarial video prediction method, characterized in that, The method includes the following steps: A bidirectional enhanced stochastic adversarial video prediction framework is constructed, comprising two sets of variational autoencoder-generative adversarial networks, used for sequential prediction and reverse prediction respectively; wherein, the input video frame sequence is cyclically reconstructed by sequential prediction encoder, reverse prediction decoder and encoder, sequential prediction decoder; A pre-trained bidirectional augmented stochastic adversarial video prediction framework is used for video prediction. Each set of variational autoencoder-generative adversarial networks includes an encoder. A generator used for variational inference Used to generate sequential / reverse video predictions, and a discriminator. Used to determine whether the reconstructed and orientation-transformed frame sequence is authentic; among which, Indicates sequential prediction, This indicates reverse prediction, i.e., sequential prediction uses... , and Reverse prediction uses , and ; There is a potential space for bidirectional data sharing between the two sets of variational autoencoder-generative adversarial networks, specifically: For sequential video frame sequences and reverse video frame sequence Define a shared sequence of latent variables. It can cover distributions in two directions and can be generated from distributions in either direction. ; encoder and at last Layer weight sharing, generator and start Layer weight sharing constructs a potential space for bidirectional data sharing, in which... Set value; For the potential space of bidirectional data sharing, use The prior Gaussian distribution associated with the latent space is represented, and the resulting latent variables are used for regular prediction, self-reconstruction, and cyclic reconstruction. The self-reconstruction and cyclic reconstruction processes in the sequential prediction case are specifically as follows: Self-reconstruction: Real Adjacency Frames Compared to the previous moment Frame groups are formed and encoded. Encoding as latent variables The generator uses implicit variables To generate the next time frame conditionally To achieve self-reconstruction; Cyclic Reconstruction: Sequential Prediction Passing through in sequence Perform cyclic reconstruction; The cyclic reconstruction specifically refers to: the true value at the current moment. The truth value of the previous moment via encoder Encoding as latent variables , will hidden variables and Input to generator , generated via encoder Hidden variables are obtained after encoding , will hidden variables and Input together into the generator Generate the reconstructed frame at the current moment. ;in, Loss reconstruction through loop Conduct training; Will and Input to discriminator To distinguish, and Input to discriminator To distinguish them.

2. The bidirectional enhanced adversarial video prediction method according to claim 1, characterized in that, For the generator, latent variable sampling is performed using a prior distribution in the adversarial generative network, and latent variable sampling is performed using an encoder-encoded adjacent frame approximation of the posterior distribution in the variational autoencoder; the two distributions are connected by minimizing the KL divergence loss between the adversarial generative network and the variational autoencoder.

3. The bidirectional enhanced adversarial video prediction method according to claim 2, characterized in that, The objective function expression for optimizing the bidirectional enhanced stochastic adversarial video prediction framework is: In the formula, Let be the objective function of the variational autoencoder; The objective function for generating the GAN (Generative Adversarial Network); Let this be the objective function corresponding to the cycle consistency constraint; The training method uses alternating training, and the specific training process is as follows: First, fix... and ,renew and Then, fix and and update and .

4. The bidirectional enhanced adversarial video prediction method according to claim 3, characterized in that, The objective function of the variational autoencoder includes KL divergence. To standardize approximate posterior For the prior distribution Approximate terms; The objective function corresponding to the cyclic consistency constraint includes a KL term for penalizing latent codes that deviate from the prior distribution, and a term for ensuring that the sequential prediction restores the original input after two direction transformations.

5. The bidirectional enhanced adversarial video prediction method according to claim 3, characterized in that, The objective function of the Generative Adversarial Network (GAN) includes a regular adversarial loss term, a loss term for self-reconstruction, and a loss term for cyclic reconstruction. Among them, the loss term used for cyclic reconstruction uses a distribution. Latent code sampling, the representation direction is The approximate posterior distribution in the opposite direction of the prediction.

Citation Information

Patent Citations

  • Video anomaly detection model training method, video anomaly detection method and device

    CN113435432A