Few-sample video generation method based on space-time adaptive adjustment and signal-to-noise ratio optimization
By introducing space-time adaptive adjustment and signal-to-noise ratio optimization methods in text-generated video tasks, the STAN-SNR framework is designed, and the dependence problem on large amounts of data and high resource consumption in the prior art is solved, and the effect of generating high-quality videos under the condition of few samples is achieved.
Patent Information
- Application Number
- CN202510229563.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art requires a large amount of text in text-generating video tasks - videos consume data and high resource, which limits its wide application, especially under the conditions of few samples.
A small sample video generation method based on spatiotemporal adaptive adjustment and signal-to-noise ratio optimization is proposed. By designing the STAN-SNR framework, a pre-trained text-to-image generation model is used, and combined with spatiotemporal feature regulation, feature rolling enhancement and dynamic signal-to-noise ratio weighting strategies, the diffusion process is optimized to generate high-quality video content.
Videos consistent with the training center's motion mode were successfully generated under a small number of video samples, with strong generalization capabilities, reducing training costs and resource consumption, and achieving flexibility and quality improvement in video generation.
Smart Images

Figure CN120075495A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video generation, and in particular, to a few-shot video generation method based on spatio-temporal adaptive adjustment and signal-to-noise ratio optimization. Background Art
[0002] Recently, image generation technology has rapidly emerged with the wave of AIGC. The generation models mainly include methods based on generative adversarial networks (GANs), variational autoencoders (VAEs), autoregressive models, and diffusion models. Among them, diffusion models have made remarkable progress in the task of generating images from text prompts, especially in the field of text-to-image (T2I). However, despite the great success in the T2I field, the text-to-video (T2V) task still faces many challenges, especially in how to reduce the dependence on a large amount of text videos for data and reduce the resource consumption during training, which remains an urgent problem to be solved.
[0003] Several recent works have been dedicated to solving this problem by directly training diffusion-based T2V models using large-scale text-video pairs. However, most of them need to train the model from scratch, and such methods require a large amount of labeled data. Due to the complex video data, the training cost is extremely high, and many researchers cannot afford this burden, which limits the wide application of such methods. In addition, there is another type of method that combines the powerful generative ability of diffusion models with the structural guidance of video templates to manipulate video content. Although this method can generate high-quality detailed content and has strong semantic generalization ability, it needs to comply with the constraints of the template, and when generating complex videos, the computational overhead and resource requirements for model training are relatively high.
[0004] Therefore, although the text-to-image technology based on diffusion models has made remarkable progress, it still faces great challenges in extending to the field of video generation, especially under few-shot conditions.
[0005] In order to solve the problem that existing methods usually rely on a large amount of text-video pair data or consume a large amount of training resources, which limits their wide application, the present invention proposes a few-shot video generation method based on spatio-temporal adaptive adjustment and signal-to-noise ratio optimization. By optimizing the pre-trained text-to-image generation model, it can learn and generate high-quality video content with a small number of video samples. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a few-shot video generation method based on spatio-temporal adaptive regulation and signal-to-noise ratio optimization. By designing a brand-new few-shot T2V task framework STAN-SNR, the goal is to optimize the pre-trained T2I diffusion model so that it can learn common motion patterns from a small amount of video data. The pre-trained T2I diffusion model already has strong semantic understanding ability under text prompt guidance. Therefore, only a small amount of data is needed to help the model master the correspondence between prompts and motion patterns, and then generate diverse video content. When tuning the T2I model into a T2V model through few-shot learning, three main problems need to be solved: (1) Since video data is scarce and the annotation cost is high, this may lead to weak generalization ability, and the generated new videos are easily limited within the range of training samples. (2) The T2I model usually focuses on capturing the spatial dimension, while the T2V model also needs to handle the dynamic changes in the time dimension simultaneously. (3) The key difference between video and image lies in the time dimension. The T2V model must ensure the temporal continuity and consistency of the generated video frames, while the T2I model only needs to generate a single static image. Few-shot learning may result in a lack of smooth transition between different frames and discontinuous actions.
[0007] In response to the challenges of the current T2V generation task, the present invention proposes a few-shot T2V generation framework called STAN-SNR. The present invention evaluated STAN-SNR in 8 motion scenarios. The results show that after a small amount of tuning on 8 to 16 videos on a single GPU, STAN-SNR can successfully generate videos consistent with the motion patterns in the training set and has strong generalization ability.
[0008] The few-shot video generation method based on spatio-temporal adaptive regulation and signal-to-noise ratio optimization provided by the present invention includes the following steps:
[0009] S1. Input text prompts and a small number of video samples, and generate the first high-resolution image through a pre-trained text-to-image generation model (based on the StableDiffusion architecture);
[0010] S2. Construct a spatio-temporal feature regulation module. By combining the depth convolution and pointwise convolution structures, efficient downsampling and upsampling are achieved in the time dimension, while reducing the computational amount. And by setting up a squeeze-and-excitation mechanism to enhance the feature expression ability and ensure that important features are fully utilized; in addition, through an improved LayerNorm operation, only the channel dimension is normalized to ensure the stability of numerical calculations;
[0011] S3. By designing a feature rolling enhancement module, perform various dynamic processes on the spatio-temporal latent features to enhance the feature expression ability and diversity, and use them to improve the quality and stability of the generated videos.
[0012] S4. Optimize the diffusion process by adopting a dynamic signal-to-noise ratio (SNR) weighting strategy, specifically as follows:
[0013] a. First, calculate the SNR for each time step;
[0014] b. Adopt a cosine annealing strategy to calculate the smoothing control weight for the time step;
[0015] c. Apply Softmax normalization to ensure that the sum of the weighting coefficients is 1 and apply it to the weighted mean squared error (MSE) loss;
[0016] S5. Reconstruct the denoised features into video frames through a decoder to generate a complete video sequence.
[0017] Preferably, the spatio-temporal feature regulation module includes a depth three-dimensional convolutional block and a squeeze-and-excitation mechanism for efficiently extracting video features and reducing the computational amount.
[0018] Preferably, the spatio-temporal feature regulation module can be formalized as:
[0019] ST(X) = SE(W up2 (W up1 ((3D-Conv(W down2 (W down1
[0020] (X)))))))+X(4).
[0021] Preferably, the feature rolling enhancement module enhances the expression ability and diversity of features by reshaping the input features, applying 3D convolution, rolling offset operation, random flipping, and noise perturbation.
[0022] Preferably, the formula for calculating the SNR for each time step in step S4 is:
[0023]
[0024] where: α t is the cumulative product factor of the diffusion process, σ t represents the variance of the noise, calculated as 1 - α t ; The Max function is used to ensure that the value of SNR is not lower than 1e-6.
[0025] Preferably, the formula for calculating the smoothing control weight for the time step in step S4 is:
[0026]
[0027] where: γ min is the minimum value of γ, γ maxis the maximum value of γ; t is the current time step; T is the maximum number of time steps; this strategy gradually adjusts the weights throughout the training process to prevent unstable training caused by mutations.
[0028] Compared with the related technologies, the few-shot video generation method based on spatio-temporal adaptive regulation and SNR optimization provided by the present invention has the following beneficial effects:
[0029] The present invention proposes a new few-shot T2V generation framework to achieve a balance between training cost and generation flexibility;
[0030] The present invention introduces spatio-temporal feature regulation, significantly enhancing the extraction efficiency and reducing the computational amount;
[0031] The present invention designs a feature rolling enhancement and dynamic SNR weighting strategy, significantly improving the generation efficiency, diversity, and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is the STAN-SNR framework diagram in the present invention;
[0033] Figure 2 is the schematic diagram of the spatio-temporal feature adjustment module in the present invention;
[0034] Figure 3 is the schematic diagram of the rolling offset operation of the feature rolling enhancement module in the present invention;
[0035] Figure 4 is the result diagram of a horse running in the present invention;
[0036] Figure 5 is the comparison diagram between the STAN-SNR model and the baseline model in the present invention;
[0037] Figure 6 is the scene diagram of fireworks over the ocean in the present invention;
[0038] Figure 7 is the scene diagram of a horse running in the ablation experiment in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0039] The method of the present invention will be further described below with reference to the drawings and embodiments.
[0040] 1. Method:
[0041] First, I will elaborate on the basic knowledge (preliminary knowledge) of the diffusion model in the present invention.
[0042] 1.1 Preliminary Knowledge
[0043] In this section, the present invention introduces the preliminary knowledge based on the diffusion model. Given data x 0∈X, a Markov chain can be defined as:
[0044]
[0045] where t = 1, ..., T, and T is the total number of steps. β t is the coefficient that controls the noise intensity at step t. Iterative noise addition can be simplified to:
[0046]
[0047] where The diffusion model learns the distribution of the dataset X by minimizing the training objective, which can be written as:
[0048]
[0049] ∈ θ (·) represents the noise prediction function of the diffusion model, and c is the condition similar to the text prompt. After training, the diffusion model can generate data from noise by reversing the noise addition process.
[0050] 1.2 Process
[0051] Existing T2V methods either require a large amount of text-video pair data or extremely high computing power. To make video generation consume less resources while ensuring continuity in the time dimension, the present invention proposes a new few-shot learning framework STAN-SNR. The STAN-SNR of the present invention is as Figure 1 shown. This model is built on the pre-trained T2I model StableDiffusion v1-4. In this experiment, 8 to 16 videos are used for each motion. It is assumed that there is a video set N = {Ni|i ∈ [1, n]} containing n videos and a text prompt for describing the motion as training data. The present invention first obtains the corresponding latent features Xi through a pre-trained encoder. Then, the latent features are input into the forward diffusion process, and noise is added to the latent features in an increasing manner. In the backward diffusion process, a U-Net network is used to predict the added noise. The present invention extends the Resnet block to a 3D block to adapt to the input in the time dimension. In addition, the present invention introduces a lightweight spatio-temporal feature regulation method and a feature rolling enhancement mechanism. Finally, the present invention uses a pre-trained decoder to reconstruct the video from the denoised features. A dynamic signal-to-noise ratio (SNR) weighting strategy is used to calculate the loss, ensuring a smooth transition of noise during the generation process, thereby significantly improving the quality and stability of the generated video and making the generated video more visually coherent.
[0052] Figure 1The STAN - SNR framework in the present invention. The framework of the present invention learns motion patterns based on 8 to 16 videos and text prompts describing the videos. This method effectively reduces the use of a large number of text - video data pairs and reduces the computing power. The present invention evenly extracts 16 frames, generates latent features through an encoder, adds noise to subsequent frames through a pre - trained diffusion model, and predicts the latent noise through the U - Net network designed by the present invention. Finally, the decoder uses a dynamic signal - to - noise ratio weighting strategy to accelerate convergence, thereby restoring the image.
[0053] 1.3 Model
[0054] In this section, the present invention will describe in detail the proposed spatio - temporal feature regulation, feature rolling enhancement, and dynamic signal - to - noise ratio weighting strategies.
[0055] Spatio - temporal feature regulation:
[0056] In video generation and spatio - temporal data modeling tasks, accurately capturing temporal dynamics is crucial for model performance. To alleviate the computational burden of traditional spatio - temporal convolutions and attention modules when dealing with high - dimensional inputs, the present invention introduces a spatio - temporal feature regulation module based on depth three - dimensional convolutions, as Figure 2 shown. This module consists of a depth three - dimensional convolution block. By combining depth convolution with a point - wise convolution structure, it achieves efficient down - sampling and up - sampling in the temporal dimension while reducing the computational amount. The squeeze - and - excitation mechanism embedded in the module further enhances the feature expression ability, ensuring that important features are fully utilized. In addition, through an improved LayerNorm operation, normalization is only performed on the channel dimension, thus ensuring the stability of numerical calculations. The spatio - temporal feature regulation module can be formalized as
[0057] ST(X) = SE(W up2 (W up1 ((3D - Conv(W down2 (W down1
[0058] (X)))))))+X(4)
[0059] By applying 3D convolution on low - dimensional inputs, the method of the present invention effectively simplifies the complexity of temporal modeling, thereby significantly improving the memory utilization efficiency during training. The specific results are shown in Table 1.
[0060] Table 1. Comparison of model sizes with several publicly available methods.
[0061]
[0062] Feature rolling enhancement:
[0063] Traditional text-to-video (T2V) generation frameworks typically rely on spatio-temporal modeling within a single frame through self-attention mechanisms. Although this method can effectively capture local information within each frame, it performs inadequately in terms of temporal dependencies between frames. This shortcoming causes the model to struggle to accurately capture the dynamic correlations between frames when generating videos, thereby affecting the overall coherence and generation quality of the video. Although some improved methods for spatio-temporal joint attention have been proposed in recent years, aiming to enhance the model's temporal modeling ability by integrating multi-frame information, these methods are often accompanied by huge computational complexities. Specifically, for a video containing L frames and N tokens, the computational complexity of its global spatio-temporal attention mechanism is as high as O(L 2 N 2 ). This extremely high computational requirement significantly increases the time and resource consumption for training and inference, making these methods difficult to apply to video tasks that require efficient generation.
[0064] To address the above problems, the present invention designs a feature rolling enhancement module. This module effectively enhances the expression ability and diversity of features, and significantly improves the quality and stability of the generated video through various means such as reshaping the input features, 3D convolution processing, rolling offset operations, random flipping, and noise perturbation. Since time (T) and space (H, W) are interrelated in video data processing, in order to enhance the flexibility of the model for input features, the present invention re-partitions the one-dimensional sequence length (Seq_len) of the input features into two dimensions of height and width, making it convenient for subsequent processing in the spatial and temporal dimensions. Then, the module introduces 3D convolution to extract features in both the spatial and temporal dimensions, thereby enhancing the coherence and logic of the generation results. To further enrich the expression ability of the temporal dimension, the module designs a rolling offset operation (as shown in Figure 3 . Step one represents the operation of the first column sub-region shifting downward by an offset of 4, and step two represents the operation of the second column sub-region shifting downward by an offset of 3). This operation processes the features in blocks, divides the spatial dimension into 4 sub-regions, and each sub-region corresponds to a specific spatial position (as shown in Figure 3One column represents a sub-region). In the time dimension, by using the torch.roll function to apply different rolling offsets (set to 4, 3, 2, 1) to each sub-region respectively, the simulation of time dynamic changes is realized, so that the time features of different sub-regions show diverse changes locally, effectively enhancing the dynamic changes in the time dimension. After the rolling offset operation, the module introduces random horizontal flipping and vertical flipping, thereby increasing the data diversity and reducing the overfitting to specific spatial patterns. In addition, the module also adds noise perturbations to further enhance the robustness of the data. Finally, to ensure that the output features match the processing requirements of the model, the module restores the enhanced features to the original shape, so that the generated features can be seamlessly connected to the subsequent processing steps.
[0065] The formula is shown as follows:
[0066] out(:,s i :,j::4,i::4,:) = roll(x(:,:,j::4,i::4,:),s i ,dim = 1)[:,s i :,:,:]
[0067] Assume the input tensor is x, the two-dimensional feature map is a 4×4 matrix, and the output tensor is out. Among them, s i is the rolling offset shifts[i] = [4, 3, 2, 1], j ∈ {0, 1, 2, 3} represents four sub-blocks, and rolling is performed in the spatial dimension to enhance the data diversity. Set a random number to determine whether to perform horizontal flipping or vertical flipping. Finally, to increase the anti-interference ability of the model, a random noise perturbation (set to 0.05) is added.
[0068] Dynamic signal-to-noise ratio weighting:
[0069] In the training of diffusion models, the volatility of the signal-to-noise ratio (SNR) and the impact of time steps on the loss often lead to instability in the training process. Traditional methods may over-rely on certain specific time steps, or have too high weights in the high-noise stage, and cannot effectively balance the contributions of each time step, thus affecting the convergence speed and accuracy of the model. To solve this problem, the present invention proposes a dynamic signal-to-noise ratio weighting strategy, the core idea of which is similar to an automatic volume adjustment system - when the signal is weak (high noise), appropriately increase the weight of this time step to enable the model to learn key features more stably; when the signal is strong (low noise), reduce the weight to avoid the model over-relying on simple patterns, thereby improving the overall training effect.
[0070] Specifically, the present invention first calculates the SNR of each time step:
[0071]
[0072] Where: α t is the cumulative product factor of the diffusion process. σ t represents the variance of the noise, calculated as 1 - α t . The Max function is used to ensure that the value of SNR is not lower than 1e-6.
[0073] Based on the SNR calculation, the present invention adjusts the weights for each time step to ensure a smooth transition of weight distribution throughout the training process. To avoid drastic fluctuations of the weights over time steps, the present invention further adopts a cosine decay strategy to calculate the smooth regulation weights for time steps:
[0074]
[0075] Where: γ min is the minimum value of γ, γ max is the maximum value of γ. t is the current time step. T is the maximum number of time steps. This strategy enables the weights to be gradually adjusted throughout the training process, preventing unstable training caused by mutations.
[0076] The final time step weights are jointly controlled by SNR and cosine decay:
[0077] w t = min(max(SNR(t), 10 -6 ), γ t )
[0078] While ensuring that SNR is not lower than 10 -6 , it does not exceed the dynamically regulated γ t . In addition, to keep the weights stable at all time steps, the present invention adopts Softmax normalization:
[0079]
[0080] To ensure that the sum of the weighting coefficients is 1 and apply it to the weighted mean squared error (MSE) loss:
[0081]
[0082] Where N represents the number of samples. is the predicted value of the model. y i is the target value. weight factor is the proportionality factor used to adjust the influence of the weights on the loss.
[0083] The method of the present invention will be further described by means of experiments below.
[0084] 2. Experiments:
[0085] 2.1 Experimental implementation
[0086] The experiments of the present invention are based on the SD-v1.4 and DALL-E models. During training, the resolution sampled from the input video by the present invention is set to 320×512 and 16 uniform frames are sampled. The model is fine-tuned for 25,000 steps with a batch size of 1 and a learning rate set to 3.0×10 -5 , and only the parameters in the spatio-temporal feature regulation module are adjusted, greatly reducing the computational amount. During inference, the first frame is generated using the Transformer-based DALL-E model to accurately match the text description and ensure high-quality image generation. Subsequent frames are then inferred using the lighter SD-v1.4 model to optimize costs while ensuring generation efficiency. All experiments are carried out on a single A6000 GPU, consuming 12GB vRAM during the training phase and 7GB vRAM during the inference phase.
[0087] 2.2 Evaluation
[0088] Dataset: To evaluate the method of the present invention, the present invention uses 100 videos collected from the DAVIS dataset, covering 8 types of actions, including helicopter (rigid motion), waterfall (fluid motion), rain and fireworks (particle motion), horse running (animal motion), bird flying (multi-body motion), smile change (emotional expression), and playing the guitar (human motion). Six (the results of horse running are as Figure 4 shown) prompts are designed for each action, and an evaluation set containing 48 videos is constructed.
[0089] Baseline models: As shown in Table 2, the present invention selects four publicly available methods for comparison with the method of the present invention: 1) Plug-and-Play: an advanced image editing model capable of independently editing videos frame by frame. 2) AnimateDiff with large-scale pre-training. 3) Tune-A-Video, a single-shot-based video editing method. 4) SimDA: a simple diffusion adapter for efficient video generation. These methods are evaluated in a variety of mainstream environments, demonstrating the advantages of few-shot learning.
[0090] Table 2. Quantitative comparison with text-to-video methods
[0091]
[0092] Evaluation on the WebVid dataset: As shown in Table 3, the method of the present invention was evaluated on 48 generated videos and the baseline model after 12,300 iterations. As shown in the table, the STAN-SNR model of the present invention is superior to the baseline model in many aspects. Specifically, it has a lower computational cost, with the number of parameters reduced by 16%, the inference speed increased by 11%, and the memory usage reduced by 4.8%. In addition, the videos generated by the model of the present invention have higher quality, as reflected by a 12.2% reduction in FVD, a 6.6% increase in alignment, and a 10.2% increase in inter-frame consistency. These results indicate that, compared with the baseline, the method of the present invention not only improves the video quality but also improves the efficiency.
[0093] Table 3. Quantitative comparison with the baseline model
[0094]
[0095] Quantitative results: The present invention evaluated the texture alignment, frame consistency, and generation diversity of STAN-SNR. As shown in Table 2, the model of the present invention achieved the highest alignment (31.5874) and consistency (97.9974), ensuring the accuracy and stability of video generation. It also obtained the lowest diversity score (72.1256), reducing unnecessary variations and improving stability. These results highlight the excellent consistency, text fidelity, and smooth frame transitions of STAN-SNR.
[0096] Text alignment:
[0097] The alignment between the video and the text is represented by calculating the CLIP score for each frame and taking its average. Assuming that each frame in the video is f i , and the text prompt is t, the CLIP model returns the alignment score CLIP(f i , t) for each frame with respect to the text. Then the text alignment of the video is:
[0098]
[0099] where N is the total number of frames in the video, and f i represents the i-th frame of the video.
[0100] Frame consistency:
[0101] Frame consistency is measured by calculating the average cosine similarity of the CLIP image embeddings between video frame pairs. Assuming that the CLIP image embeddings of each frame f i and f j in the video are E(f i ) and E(f j ), then the frame consistency is:
[0102]
[0103] Among them, · represents the dot product of vectors, and ∥·∥ represents the norm of vectors.
[0104] Generative diversity:
[0105] Generative diversity is measured by the cosine distance between CLIP image embeddings of videos. First, the average of the CLIP embeddings of all frames of each video is taken to represent the entire video. Suppose the average embeddings of videos A and B are and respectively. Then the cosine distance of the video pair is:
[0106]
[0107] where M is the total number of videos. The larger the cosine distance, the greater the difference between the videos, that is, the higher the generative diversity.
[0108] Fréchet inception distance:
[0109] The Fréchet inception distance (FID) is a widely used metric for evaluating the quality of generated images in a generative model. It quantifies the difference between generated images and real images by calculating the distance between their feature distributions. It uses a pre-trained Inception network to extract high-dimensional features and maps the images to the feature space, thus focusing on the perceptual similarity rather than just the pixel similarity. The lower the FID value, the more similar the generated images are to the real images, indicating better image quality and diversity. The calculation formula of FID can be expressed as:
[0110]
[0111] where μ r and μ g are the means of the feature vectors extracted from real images and generated images respectively. ∑ r and ∑ g are the covariance matrices of the feature vectors of real images and generated images respectively. ||μ r - μ g || 2 is the Euclidean distance between the mean vectors, representing the difference in the global feature distributions of generated images and real images. Tr represents the trace operation of the matrix, that is, the sum of the elements on the diagonal of the matrix, representing the difference between the covariance matrices. is the square root of the matrix product of the covariance matrices, used to measure the overlap of the feature distributions of generated images and real images.
[0112] To compare the effectiveness of the dynamic SNR weighting strategy of the present invention, the present invention calculates the inference video generated every 100 iterations. First, the present invention extracts the video into 16-frame pictures. Since the first frame of the present invention is generated using the DALL-E model, the present invention calculates the FID metric between the subsequent 15 frames and the first frame and takes the average of the 15 results as the FID value for this iteration. As Figure 5 shown, the present invention compares the model of the present invention with the baseline model and finds that the model of the present invention starts to converge at 12,300 iterations, while the baseline model converges at 30,000 iterations, and the convergence speed is 2.44 times faster.
[0113] From Figure 5 it can be seen that after adopting the dynamic signal-to-noise ratio weighting strategy, the convergence speed of the present invention is 2.44 times faster than that of the baseline model, achieving satisfactory results.
[0114] Depthwise Separable 3D Convolution:
[0115] In a standard 3D convolution, each output channel is a weighted sum of all input channels. Specifically, each input channel is convolved with a convolution kernel of size K T ×K H ×K W and the results are summed over all input channels to finally generate C out output channels. Although this method can effectively extract spatial and channel features, the computational cost is huge, especially when dealing with high-dimensional data, it is easy to cause waste of computing resources. To reduce the computational cost, the present invention introduces depthwise separable convolution. It splits the standard 3D convolution into two steps:
[0116] (1) Depthwise Convolution: Each input channel is separately convolved with a K T ×K H ×K W 3D convolution to extract spatial features.
[0117] (2) Pointwise Convolution: A 1×1×1 convolution kernel is used to fuse the channels, thereby restoring the number of output channels.
[0118] From the perspective of computational cost:
[0119] (1) Computational cost of standard 3D convolution:
[0120] FLOPs Standard =T×H×W×C in ×C out ×K T ×K H ×K W
[0121] (2) Computational cost of depthwise separable 3D convolution:
[0122] FLOPs Depth = T × H × W × C in × K T × K H × K W
[0123] FLOPs Point = T × H × W × C in × C out
[0124] FLOPs Separable = FLOPs Depth + FLOPs Point
[0125] (3) Calculation of speedup ratio:
[0126]
[0127] For example, when C in = 256, C out = 256, and K T = K H = K W = 3, the computational complexity of depthwise separable convolution is reduced by about 8.4 times. If C out = 512, it is reduced by about 16.5 times.
[0128] Qualitative results
[0129] In Figure 6 shows the scene of "Fireworks over the Ocean". The method of the present invention is visually compared with four baseline models. The lack of sufficient data in Tune-A-Video leads to model overfitting. The SimDA model significantly improves its performance and robustness in various tasks through representation learning, handling diversity, and adaptive learning strategies. However, due to the small dataset, there are deficiencies in detail processing and clarity, the overall picture is relatively blurred, lacking a sense of hierarchy and detailed texture performance, resulting in a less eye-catching visual effect. AnimateDiff learns the motion layer on a large-scale dataset and embeds it into a personalized T2I model to generate videos with a specific style and high visual quality. Although the effect of the fireworks and the color of the lights are vivid, the overall detail processing and clarity of the image are insufficient, resulting in blurred outlines of the background and buildings. The actual output of Plug-and-Play fails to present the effect of the fireworks, the image content is relatively single, lacking the expected visual impact and richness. In contrast, the STAN-SNR of the present invention achieves good frame consistency and generates videos with reasonable motion patterns, benefiting from the spatio-temporal feature regulation, feature rolling enhancement, and dynamic SNR weighting strategies proposed by the present invention.
[0130] 2.3 Ablation Experiments
[0131] An ablation experiment was conducted for the present invention to evaluate the importance of spatio-temporal feature regulation, feature rolling enhancement, and dynamic SNR weighting strategy. Each design was ablated individually to analyze its impact. As Figure 7 (b) shows, compared with the complete model, the horse's movements appear stiff, the details are handled roughly, the background atmosphere is monotonous, and there is a lack of coordination with the horse's movements. Figure 7 (c) Although there is also a certain degree of smoothness in the action performance, the handling of details and background is relatively inferior, lacking a sense of hierarchy and vividness, and the overall visual effect appears relatively plain. Figure 7 (d) shows that although the movements still maintain a certain degree of vividness, the details are handled roughly, lacking a sense of hierarchy and visual impact, and the overall effect appears relatively monotonous. These results demonstrate the important contributions of each key module to the final complete model. The quantitative results of the ablation experiment are shown in Table 4. The alignment (31.59) and consistency (98.00) of the complete model are the best. It is proved that optimizing and combining these modules can improve the accuracy and stability of text-video. The data are from 48 videos generated by the present invention.
[0132] Table 4. Results of the ablation experiment. S-T represents spatio-temporal feature regulation, Scr represents feature rolling enhancement, and SNR represents dynamic SNR weighting strategy.
[0133]
[0134] 3. Conclusion
[0135] This paper proposes a video diffusion model STAN-SNR based on few-shot learning for the generation and editing of text-guided videos. This method learns motion patterns from a small amount of video data and achieves a good balance between the training burden and generation flexibility. In the method of the present invention, first, a high-quality first-frame image is generated using a T2I model, and then subsequent frames are predicted, effectively avoiding the problem of overfitting of data content in the few-shot case. At the same time, the innovative design of the network structure and inference strategy significantly enhances the effect of T2V generation. Experimental results show that this paper achieves the optimal performance in various evaluation metrics and has good generalization ability. The work of this paper is based on few-shot research and paves the way for a broader future research field of T2V.
[0136] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural or equivalent process transformation made using the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for generating a few-sample video based on spatiotemporal adaptive adjustment and signal-to-noise ratio optimization, characterized in that: The following steps are involved: S1, input text prompts and a small number of video samples, and generate the first frame of high-resolution image through the pre-trained text-to-image generation model; S2. Construct a spatiotemporal feature control module, which combines deep convolution with point-by-point convolution structure to achieve efficient downsampling and upsampling in the time dimension, while reducing the amount of calculation, and enhances the feature expression ability by setting a squeeze-excitation mechanism to ensure that important features are fully utilized; in addition, through the improved LayerNorm operation, only the channel dimension is normalized to ensure the stability of numerical calculation; S3. By designing a feature rolling enhancement module, a variety of dynamic processing is performed on the temporal and spatial potential features to enhance the expressiveness and diversity of the features, and to improve the quality and stability of the generated video; S4, using dynamic signal-to-noise ratio weighting strategy to optimize the diffusion process, specifically: a. First calculate the SNR at each time step; b. Use the cosine decay strategy to calculate the smoothing control weight of the time step; c. Use Softmax normalization to ensure that the sum of weighted coefficients is 1, and apply it to the weighted mean square error loss; S5. Reconstruct the denoised features into video frames through the decoder to generate a complete video sequence.
2. The method for generating a few-sample video based on spatiotemporal adaptive adjustment and signal-to-noise ratio optimization according to claim 1, characterized in that: The spatiotemporal feature control module includes a deep three-dimensional convolution block and a squeeze-excitation mechanism, which is used to efficiently extract video features and reduce the amount of calculation.
3. The method for generating a small number of samples of video based on spatiotemporal adaptive adjustment and signal-to-noise ratio optimization according to claim 1, characterized in that: The spatiotemporal feature control module can be formalized as: ST(X)=SE(W up2 (W up1 ((3D-Conv(W down2 (W down1 (X)))))))+X(4)。 4. The method for generating a few-sample video based on spatiotemporal adaptive adjustment and signal-to-noise ratio optimization according to claim 1, characterized in that: The feature rolling enhancement module enhances the expressiveness and diversity of features by reshaping input features, applying 3D convolution, rolling offset operations, random flipping and noise perturbations.
5. The method for generating a few-sample video based on spatiotemporal adaptive adjustment and signal-to-noise ratio optimization according to claim 1, characterized in that: The SNR calculation formula for each time step in step S4 is: Where: α t is the cumulative multiplication factor of the diffusion process, σ t Represents the variance of the noise, calculated as 1-α t ; The Max function is used to ensure that the SNR value is not less than 1e-6.
6. The method for generating a few-sample video based on spatiotemporal adaptive adjustment and signal-to-noise ratio optimization according to claim 1, characterized in that: The calculation formula of the smoothing control weight of the time step in step S4 is: Where: γ min is the minimum γ value, γ max is the maximum γ value; t is the current time step; T is the maximum number of time steps.