Video generation method based on generative adversarial network
Through the depth autoencoder explicitly extracting and integrating inter-frame motion characteristics, the problem of insufficient inter-frame dynamic relationship capture in multi-source real scene videos is solved, and high-quality video generation effect is achieved.
Patent Information
- Application Number
- CN202510316006.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-25
AI Technical Summary
When existing video generation models deal with multi-source real scene videos, it is difficult to capture the dynamic relationship between continuous frames, resulting in problems such as incoherence, motion distortion, and loss of details between generated videos.
The video frame is encoded and decoded using a depth autoencoder, explicitly extracts inter-frame motion features, and integrates them into an ordered time series, and splices them with noise vectors and constant tensors as generator inputs. The generator is trained through improved loss functions to ensure the timing consistency of features and dynamic modeling capabilities.
The coherence and dynamic feature capture capabilities of generated videos have been significantly improved. The FVD16 and FVD128 indicators have been increased by about 36.0% and 47.3% respectively, enhancing the perception of subtle movement changes, and the generated video is more natural and realistic.
Smart Images

Figure CN120378706A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image generation, and particularly to a video generation method based on a generative adversarial network. Background Art
[0002] With the rapid development of the digital content industry, video generation technology has become the core driving force in fields such as film and television special effects, game animation, and virtual reality. The generation of high-resolution dynamic videos (1024x1024 resolution) can significantly enhance the visual effect and provide key support for immersive interactive experiences. Multi-source real-scene video data is real-world video data collected from multiple device sources, with diverse content, dynamic patterns, and background environments, which can cover a wide range of real-world scenarios and provide support for video generation. Although current models based on generative adversarial networks such as StyleGAN-V have made progress in static image synthesis, they still face challenges in video sequence generation. Existing models have limited ability to extract motion features between consecutive frames, making it difficult to capture subtle changes in dynamic relationships, and there are significant limitations in processing multi-source real-scene video data, resulting in problems such as frame-to-frame incoherence, action distortion, and detail loss in the generated videos. These problems severely restrict the authenticity and application scope of the generated videos. Therefore, designing an improved method that can accurately model temporal dynamic features and enhance video coherence and detail expressiveness is crucial for promoting the practical implementation of video generation technology.
[0003] Most existing video generation methods have some problems. For example, the document with the application number "CN201911012257.8" discloses a "video generation method, device, and equipment", which generates a synthetic video that matches the video scene attribute information and contains the specified video elements by obtaining the source video, the specified video elements, and the video attribute information, and using an adversarial neural network model. Its disadvantage is that although this method can match the specified video attribute information when generating a video, it is difficult to accurately grasp the dynamic correlation between frames when dealing with complex scenarios, resulting in poor motion smoothness of the generated video. Another document with the application number "CN202210359767.8" discloses a "method, device, and storage medium for video reconstruction based on VAE-GAN". This method obtains a video sequence and preprocesses it to obtain single-frame images and video attributes, and then inputs them into a pre-trained VAE-GAN model for video reconstruction. Its disadvantage is that although this method improves video coherence and clarity through a dual-channel decoder, due to its coarse-to-fine generation method, it still cannot well grasp the dynamic change details between frames when dealing with dynamic changes, thus affecting the natural transition of the generated video during motion. The common problem of the above solutions is that when facing complex scenarios and dynamic changes, they all have insufficient capture of the dynamic relationship between frames, thereby affecting the motion coherence and detail expressiveness of the generated videos. Summary of the Invention
[0004] The present invention proposes a video generation method based on a generative adversarial network to address the deficiencies of existing video generation models in motion feature extraction, particularly the poor performance in capturing the dynamic relationships between consecutive frames, resulting in unnatural motion and missing details in the generated videos.
[0005] To achieve the above objective, the technical solution of the present invention is as follows: A video generation method based on a generative adversarial network, comprising the following steps:
[0006] Step 1: Use a sequence of original video frames of continuous multi-source real scenes as input data to provide temporal dynamic information, preprocess the original video frames, and construct an input data set;
[0007] Step 2: The deep autoencoder encodes and decodes the original video frames in the preprocessed input data set during the pre-training stage to explicitly extract the inter-frame motion features;
[0008] Step 3: Integrate the extracted motion features into an ordered temporal feature sequence;
[0009] Step 4: Concatenate the integrated motion features with other inputs of the generator to form an optimized generator input. The generator and the discriminator together form a video generation model, and the discriminator is used to evaluate the authenticity of the generated video;
[0010] Step 5: Update the parameters of the video generation model, continuously train, and determine whether the iteration number is reached and convergence occurs. If so, the training of the generation model is completed; if not, return to Step 1 to continue training;
[0011] Further, the deep autoencoder constructed in Step 2 above includes an encoder and a decoder:
[0012] Among them, the encoder maps the high-dimensional video frame sequence to a low-dimensional motion feature space through a multi-layer convolutional network to extract the motion features of the inter-frame dynamic changes;
[0013] The decoder decodes the motion features back to the video frame sequence through deconvolution to ensure the temporal consistency and effectiveness of the extracted motion features.
[0014] Further, during the pre-training of the deep autoencoder in Step 2 above, the optimization of the reconstruction loss function is expressed as
[0015]
[0016] where X is the input video frame sequence, X′ is the output reconstructed by the decoder, and L rec is the reconstruction loss, representing the mean square error between the input data and the reconstructed data.
[0017] Further, in the integration process of Step 3 above, first, the motion features extracted by the deep autoencoder are arranged in the order of video frame time to generate a corresponding feature sequence
[0018] m = {m1, m2,..., m n} (5)
[0019] Then, for the motion feature m t attach the timestamp information t of the corresponding frame to form a motion feature pair with timestamp (t, m t )
[0020] {(1, m1), (2, m2),..., (n, m n )} (6)
[0021] Further, other inputs in Step 4 above refer to the noise vector and the constant tensor, which are concatenated in the channel dimension and expressed as
[0022]
[0023] where Concat represents the concatenation operation in the channel dimension, is the motion feature m at dimension select the motion feature of the current time step t0 according to the dynamic demand of the time step t, Z is the noise vector, and C is the constant tensor.
[0024] Compared with the existing technologies, the beneficial effects of the present invention are:
[0025] (1) Aiming at the problems that may exist in the existing styleGAN-V video generation technology when processing multi-source real-scene video frames, such as insufficient extraction of motion features, resulting in discontinuous video frames, distorted actions, and lost details in the generated video, the present invention proposes a method of encoding and decoding the preprocessed video frames using a deep autoencoder. The advantage is that this method can explicitly extract the inter-frame motion features, and through the collaborative work of the encoder and the decoder, ensure that the extracted features are both effective and temporally consistent. Compared with other methods, the performance of the present invention is significantly improved when processing complex dynamic scenes and multi-source data. Especially in the high-resolution video generation task, it can better capture and model the dynamic changes between frames, thus generating higher-quality video content.
[0026] (2) In view of the problems in the existing video generation technology that the integration of motion features lacks orderliness, resulting in poor inter-frame information flow, insufficient dynamic modeling ability, and easy occurrence of inter-frame incoherence, etc., the present invention integrates the extracted motion features into an ordered time series. This method arranges the motion features strictly in the time order of the original video frames and attaches the time stamp information of the corresponding frames, strongly binding the features to the time dimension, strengthening the inter-frame information flow and dynamic modeling ability. Compared with other methods, the present invention has been significantly improved in integrating motion features into a time series, can ensure that the features at each time step are completely aligned with the time points of the original frames, avoid the loss of dynamic information caused by feature misalignment, retain details well, and at the same time provide a reliable time reference for subsequent dynamic video generation, thereby improving the dynamic coherence and authenticity of the generated video, and the generated motion video shows naturalness.
[0027] (3) In view of the problems in the existing video generation technology that the input of the generator is single and it is difficult to effectively integrate motion features and randomness control, the present invention splices the motion features with other inputs of the generator (such as noise vectors and constant tensors). By combining the explicit motion features extracted by the deep autoencoder, the noise vectors that control randomness and diversity, and the constant tensors that provide a stable starting point, the generator can accurately control the dynamic changes between frames when generating videos, while maintaining sample diversity and stability, realizing the effective integration of various input information. Compared with other methods, the quality of the videos generated by the present invention has been significantly improved, and they are richer and more realistic in terms of texture, light and shadow, and local details. This method significantly improves the coherence and dynamic feature capture ability of the generated videos, and improves by about 36.0% and 47.3% respectively compared with StyleGAN-V in terms of the FVD16 and FVD128 metrics, enhances the perception ability of subtle motion changes, and performs excellently in terms of dynamic features and picture continuity. Brief Description of the Drawings
[0028] Figure 1 It is the overall flowchart of the present invention. Detailed Embodiments
[0029] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the drawings and embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, and are only used to illustrate the present invention, but not to limit the scope of the present invention.
[0030] See Figure 1, the basic idea of the present invention is as follows: First, a continuous video frame sequence is used as input data to provide temporal dynamic information, which serves as the basis for motion feature extraction. Subsequently, the input video frames are encoded and decoded by a deep autoencoder to explicitly extract the inter-frame motion features. Then, the extracted motion features are integrated into an ordered time series to enhance the flow of inter-frame information and the dynamic modeling ability. Next, these refined motion features are concatenated with the constant tensor and noise vector input to the generator to form the optimized generator input. Finally, the optimized input is fed into the generator, the improved loss function is calculated, the video generation model parameters are updated, and continuous training is carried out until the model is completed.
[0031] Based on the above basic idea, the present invention provides a video generation method based on an adversarial generative network, and the specific steps are as follows:
[0032] Step 1. Use the original video frame sequence of continuous multi-source real scenes as input data to provide temporal dynamic information, which serves as the basis for motion feature extraction. The specific process includes:
[0033] Step 1.1 Collect the original video data of multi-source real scenes, extract the continuous video frame sequence, ensure coverage of different dynamic change patterns, and label the temporal correlation. The method for extracting the video sequence is as follows:
[0034] X t =ExtractFrame(V,t) (1)
[0035] where X t is the t-th frame, V is the video data, and ExtractFrame is the function for extracting frames.
[0036] Step 1.2 Preprocess the original video frames of the extracted multi-source real scenes, including inter-frame alignment, denoising, resolution unification, and timestamp standardization, to construct a temporally coherent high-resolution input data set, retaining dynamic details and removing redundant interference.
[0037] Step 2. The deep autoencoder encodes and decodes the preprocessed video frames of the multi-source real scenes in the input data set to explicitly extract the inter-frame motion features. The specific process includes:
[0038] Step 2.1 Construct a deep autoencoder: The deep autoencoder consists of two parts, an encoder f enc and a decoder f dec . Its main objective is to compress the high-dimensional processed video frame data into a low-dimensional motion feature space and reconstruct the input data through decoding to ensure the effectiveness and temporal consistency of the extracted features. The encoding and decoding process is learned through a neural network to minimize the difference between the input data and the reconstructed data.
[0039] In the encoder f enc the input is a sequence of video frames X = {x1, x2,..., xt} of continuous multi-source real scenes preprocessed in Step 1, where x t represents the video frame at time step t.
[0040] The input frames are encoded through a multi-layer convolutional network, mapping the spatio-temporal features of the video frames to a low-dimensional motion feature space:
[0041]
[0042] where t is the number of time steps and c is the number of channels of the motion features. The core task of the encoder f enc is to extract motion features m with explicit temporal relationships from the input high-dimensional frame sequence, and these features are used to reflect the dynamic changes between video frames.
[0043] In the decoder f dec the input is the motion feature m extracted by the encoder, and the output is the reconstructed frame sequence X′, which is used to compare with the original input frame X to ensure that the extracted motion feature m can completely express the information of the original frame sequence:
[0044] X′ = f dec (m), X′ ≈ X (3)
[0045] The decoder f dec performs spatial reduction on the motion feature m and decodes it back to the form of a video frame sequence through multi-layer transposed convolution.
[0046] Step 2.2 pre-trains the deep autoencoder: The training objective of the deep autoencoder is to ensure that the motion features m extracted by the encoder can completely retain the inter-frame dynamic relationships, while the decoder can successfully reconstruct the input frames. To achieve this goal, the following reconstruction loss function is mainly optimized during the training process:
[0047]
[0048] where X is the input video frame sequence, X′ is the output reconstructed by the decoder, and L rec is the reconstruction loss, representing the mean square error between the input data and the reconstructed data. By minimizing L rec , the deep autoencoder can optimize the parameters of the encoder and the decoder, so that the encoder f enc can extract effective low-dimensional motion features m. The decoder f dec can verify whether the extracted motion features contain the complete dynamic information between frames.
[0049] Step 2.3 The pre-trained deep autoencoder encodes and decodes the original video frames in the input dataset to explicitly extract the inter-frame motion features:
[0050] First, the encoder compresses the high-dimensional video frame sequence X into the motion feature space m, explicitly modeling the dynamic correlation between frames to extract key low-dimensional motion features. Second, the decoder decodes and reconstructs the motion feature m to verify the effectiveness and integrity of the extracted features, ensuring that the motion features can accurately express the temporal information of the video frame sequence. Finally, the motion feature m extracted by the trained encoder is further processed as the input of the generator, and the generator uses the motion feature to generate new video frames, providing high-quality feature support for the generation process.
[0051] The input of the deep autoencoder is a continuous motion sequence, and the motion features are extracted by the encoder. The encoder outputs the compressed motion features, and then the decoder attempts to reconstruct the input sequence. By continuously updating the parameters and optimizing the reconstruction error, it is ensured that the extracted motion features have temporal consistency and dynamic accuracy.
[0052] Step 3. Integrate the extracted motion features into an ordered time feature sequence to strengthen the flow of inter-frame information and the dynamic modeling ability;
[0053] Step 3.1 First, arrange the discrete motion features extracted by the deep autoencoder in Step 2 in strict chronological order of the original video frames to generate a motion feature sequence corresponding one-to-one with the input frames. The specific implementation is as follows:
[0054] X t represents the t-th video frame, and m t represents the motion feature corresponding to the t-th frame. By traversing the time axis T = {1, 2,..., n} of the input video frames, the motion features corresponding to each frame are sequentially extracted and stored in an ordered feature list in the order of time steps. This process can be expressed as:
[0055] m = {m1, m2,..., m n} (5)
[0056] where m is the motion feature sequence arranged in chronological order. This process ensures that the feature m t at each time step is exactly aligned with the time point of the original frame, avoiding the loss of dynamic information caused by feature misalignment.
[0057] Step 3.2 Attach the timestamp information t of the corresponding frame to each motion feature m t to form a motion feature pair with timestamp (t, m t ), ensuring a strong binding between the feature and the time dimension. The entire feature sequence can be expressed as:
[0058] {(1, m1), (2, m2),..., (n, m n )}(6)
[0059] This timestamp binding mechanism can not only accurately record the time attributes of features, but also provide a reliable time reference for subsequent dynamic modeling, ensuring the continuity and consistency of dynamic changes between frames.
[0060] Step 4: Concatenate the motion features with other inputs of the generator, and use the optimized features as the input content to the generator:
[0061] Step 4.1 The motion features m extracted by the trained deep autoencoder t need to be combined with other inputs of the generator: the noise vector Z and the constant tensor C to form the final generator input. The noise vector Z is usually a random vector used to control the randomness and sample diversity of video generation. Adjusting the value of Z can generate output videos with different styles or contents:
[0062]
[0063] where represents the standard normal distribution. The noise vector usually affects the low-level features of the generator, such as texture and local details. The constant tensor C is the "basis" for the generator to generate video frames, providing a stable starting point for feature modeling of subsequent networks. At the same time, the constant tensor also contains the core information of the static scene in the video.
[0064] Step 4.2 The way of combining features with the generator is similar to the original model. The dimension of the extracted motion feature m will select the feature of the current time step t0 according to the dynamic requirements of the time step t m0 is adjusted to the tensor shape after expansion to make it consistent with the shape of the constant tensor C. Finally, and the noise vector Z and the constant tensor C are concatenated in the channel dimension to form the complete input G of the generator. The specific formula is:
[0065]
[0066] where Concat represents the concatenation operation in the channel dimension. Through this concatenation, the generator can accurately control the dynamic changes between frames when generating videos, generating a video sequence with high dynamic consistency. The improved motion feature extraction method uses a deep autoencoder. Through the design of the encoder and decoder, it can not only extract the explicit motion features of the video frame sequence, but also optimize the accuracy and temporal consistency of feature extraction through training.
[0067] Step 5: Input the optimized motion features into the generator in the video generation model, calculate the updated loss function in the comparison between the generator and the real data, update the parameters of the video generation model, continuously perform iterative training of the model, and determine whether the iteration times are reached and convergence occurs. If so, the generation model is trained successfully; if not, return to Step 1 to continue training.
[0068] FVD is an index for measuring the difference between the generated video and the real video. Therefore, the present invention uses and FVD128 as indexes to evaluate the quality of the generated video:
[0069]
[0070] where and are the mean values of the features of the i-th frame of the real video and the generated video, respectively.
[0071] and are the feature covariances of the i-th frame and the j-th frame of the real video and the generated video, respectively.
[0072] The lower the values of FVD16 and FVD128, the smaller the difference between the generated video and the real video, that is, the better the quality of the generated video. The difference between FVD16 and FVD128 is that FVD16 uses a 16-frame video, while FVD128 uses a 128-frame video for calculation.
[0073] As shown in Table 1, when comparing the FVD16 and FVD128 indexes of the present invention with those of the original network StyleGAN-V, the improved algorithm of the present invention has a significant reduction in the FVD16 and FVD128 indexes, indicating that the algorithm proposed by the present invention has made remarkable progress.
[0074] Table 1
[0075]
[0076] The above description is an illustration of the specific implementation of the present invention, rather than a limitation of the present invention. Those skilled in the relevant technical field can also make various equivalent technical solutions without departing from the scope of the present invention. Therefore, all equivalent technical solutions should be included in the protection scope of the present invention.
Claims
1. A video generation method based on a generative adversarial network, characterized in that: It includes the following steps: Step 1: Use the original video frame sequence of continuous multi-source real scenes as input data to provide temporal dynamic information, preprocess the original video frames, and construct an input data set; Step 2: In the pre-training stage, the deep autoencoder encodes and decodes the original video frames in the preprocessed input data set to explicitly extract the inter-frame motion features; Step 3: Integrate the extracted motion features into an ordered time feature sequence; Step 4: Concatenate the integrated motion features with other inputs of the generator to form an optimized generator input. The generator and the discriminator together form a video generation model, and the discriminator is used to evaluate the authenticity of the generated video; Step 5: Update the parameters of the video generation model, continuously train, and determine whether the iteration times are reached and convergence occurs. If so, the generation model is trained; if not, return to Step 1 to continue training.
2. The video generation method based on the adversarial generative network according to claim 1, wherein: The deep autoencoder constructed in Step 2 includes an encoder and a decoder: Among them, the encoder maps the high-dimensional video frame sequence to a low-dimensional motion feature space through a multi-layer convolutional network to extract the motion features of the inter-frame dynamic changes; The decoder decodes the motion features back to the video frame sequence through deconvolution to ensure that the extracted motion features have temporal consistency and effectiveness.
3. The video generation method based on the adversarial generative network according to claim 2, wherein: In Step 2, during the pre-training of the deep autoencoder, the optimization of the reconstruction loss function is expressed as Where X is the input video frame sequence, X′ is the output reconstructed by the decoder, and L rec is the reconstruction loss, representing the mean square error between the input data and the reconstructed data.
4. A video generation method based on a generative adversarial network according to claim 3, characterized in that: The integration process in Step 3 is that first, the motion features extracted by the deep autoencoder are arranged in the time order of the video frames to generate a corresponding feature sequence m = {m1, m2,..., m n} (5) Then, for the motion feature m t Attach the timestamp information t of the corresponding frame to form a motion feature pair with timestamp (t, m t ) {(1, m1), (2, m2),..., (n, m n )}(6).
5. A video generation method based on a generative adversarial network according to claim 4, characterized in that: The other inputs in Step 4 refer to the noise vector and the constant tensor, and the concatenation in the channel dimension is expressed as G = Concat(m t0 , Z, C) (8) Among them, Concat represents the concatenation operation in the channel dimension. is the motion feature m at dimension Select the motion feature of the current time step t0 according to the dynamic requirements of the time step t. Z is the noise vector, and C is the constant tensor.
Citation Information
Patent Citations
Video generation methods, apparatus and equipment
CN110753264B
Video reconstruction method and device based on VAE-GAN and storage medium
CN114708459A
Cited By
Children story video generation method and system based on AI
CN120812370A