Consistency video generation method and device based on wavelet transform and frequency decomposition
Through the wavelet transformation and frequency decomposition methods, combined with cross-modal attention and dynamic alignment mechanism, the problems of medium and high-frequency details drifting and dynamic scene identity distortion are solved in video generation, achieving high-quality and stable video generation.
Patent Information
- Application Number
- CN202510388116.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-31
AI Technical Summary
Existing video generation methods have bottlenecks in cross-modal feature fusion, dynamic identity maintenance and long sequence stability, especially in the problems of high-frequency detail drift, dynamic scene identity distortion and poor long sequence generation stability.
The wavelet transformation and frequency decomposition methods are adopted to extract low-frequency contour features and high-frequency texture features through multi-scale Haar wavelet decomposition, and text embedding is obtained in combination with CLIP text encoder, and the fused features are reconstructed by band adaptive convolution module and inverse wavelet, and identity distribution alignment and timing smooth constraints of latent variable space are performed through the dual-channel cross attention module. Finally, high-quality video is generated through Hamilton-Jacobi optimization.
It significantly improves the quality and stability of the generated video, realizes identity representation fusion across time and space dimensions, solves the problems of high-frequency details drift and identity distortion in dynamic scenes, and improves the stability of long sequence generation.
Smart Images

Figure CN120264095A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more particularly to a method and apparatus for generating consistent videos based on wavelet transform and frequency decomposition. Background Art
[0002] In recent years, significant progress has been made in video synthesis technology based on generative artificial intelligence. Among them, Identity-Preserving Text-to-Video (IPT2V) has attracted much attention due to its application value in fields such as virtual character driving and personalized content creation. This technology needs to simultaneously meet two core requirements: text-driven dynamic scene generation and identity feature preservation of reference images. However, existing methods still have significant bottlenecks in cross-modal feature fusion, dynamic identity preservation, and long-sequence stability.
[0003] In the prior art, models based on the U-Net architecture inject identity information through a text-image cross-attention mechanism, but in practical applications, problems such as high-frequency detail drift and temporal action interference are exposed. For example, high-frequency identity markers (such as facial textures, pupil details) are easily affected by motion deformation, resulting in insufficient frame-to-frame consistency in the generated images; low-frequency contour features (such as facial proportions, poses) are difficult to decouple dynamic scenes and static identity information due to the high coupling degree of shallow features in the model. Although recent research has attempted to introduce frequency decomposition strategies to optimize feature expressions, it has not fully utilized the multi-scale frequency band analysis ability of wavelet transform, resulting in low cross-frequency band information fusion efficiency and vulnerable to noise interference in complex motion scenes.
[0004] Regarding the temporal modeling problem, video generation methods based on diffusion models (such as Stable Video Diffusion) achieve motion control by adding spatio-temporal coupling layers, but their distribution alignment mechanism fails to effectively coordinate the statistical property differences between identity features and diffusion latent variables. In addition, existing methods generally rely on post-processing tools to repair generation defects, resulting in video domain mismatches, increased computational overhead, and inability to achieve end-to-end optimization.
[0005] Summarizing the bottlenecks of the prior art, the core problems can be summarized as follows: 1) Insufficient multi-scale frequency domain decoupling, and the adaptive fusion of cross-level frequency band features has not been achieved through wavelet transform; 2) Lack of dynamic distribution alignment, and the mean-variance statistical properties of identity embedding and diffusion latent variables have not been explicitly aligned, resulting in identity distortion under the interference of the temporal layer; 3) Poor long-sequence generation stability, lack of optical flow constraints and sliding window fusion mechanisms, and it is difficult to balance motion coherence and identity consistency. To address the above problems, a new technical framework integrating frequency domain decomposition, dynamic distribution alignment, and end-to-end optimization is urgently needed. Summary of the Invention
[0006] In view of the above technical problems in the related art, the present invention proposes a consistent video generation method and device based on wavelet transform and frequency decomposition.
[0007] In a first aspect, the present invention provides a consistent video generation method based on wavelet transform and frequency decomposition, including the following steps:
[0008] S1. Based on the band decoupling of multi-scale Haar wavelet decomposition, extract the low-frequency contour features and high-frequency texture features of the reference image, obtain the text embedding of the text prompt through the CLIP text encoder, and perform cross-modal attention on the low-frequency contour features, high-frequency texture features, and text embedding to obtain multi-scale wavelet features F wavelet ;
[0009] S2. Fusion of multi-scale wavelet features F wavelet of multi-level features through a band adaptive convolution module and inverse wavelet reconstruction, and retain cross-scale details to output fused image features F fused ;
[0010] S3. The facial embedding E arc of the reference image obtained by the ArcFace network and the global context feature E clip of the reference image obtained by the CLIP-ViT encoder are fused by Q-Former to generate a refined facial embedding E final ;
[0011] S4. Input the diffusion latent variable Z t , fused image features F fused and identity embedding E final into the dual-path cross-attention module for identity distribution alignment and temporal smoothing constraint in the latent variable space to obtain the aligned latent variable Z aligned ;
[0012] S5. Inject the aligned latent variable Z aligned and text embedding E text into the spatio-temporal DiT architecture to generate a preliminary temporal latent variable Z time , and generate an optimized latent variable Z t from the diffusion latent variable Z through Hamilton-Jacobi optimization. Then, generate a high-quality video V t-1 from the preliminary temporal latent variable Z time and the optimized latent variable Z t-1 through a multi-frame autoregressive generator output .
[0013] Specifically, step S1 specifically includes the following steps:
[0014] S11. The reference image Input the multi-level Haar wavelet transform, and obtain the low-frequency component after recursively decomposing L levels and the high-frequency component H is the height of the image, and W is the width of the image;
[0015] S12. Band feature recombination. Align and splice the low-frequency components of each layer through bicubic interpolation, and fuse them to generate the global low-frequency feature F low , and sum the high-frequency components of each layer hierarchically weighted to generate the local high-frequency feature F high ;
[0016] S13. Input the text prompt T into the CLIP text encoder to extract the text embedding E text , and interact with the wavelet features through cross-modal attention, and finally output the multi-scale wavelet feature F wavelet =(F low ,F high ); where, F low is the low-frequency feature; F high is the high-frequency feature.
[0017] Specifically, step S2 specifically includes the following steps:
[0018] S21. Inject the multi-scale wavelet feature F wavelet =(F low ,F high ) into the band adaptive convolution module for hierarchical reverse reconstruction and inverse wavelet transform reconstruction to obtain the reconstructed feature R (i) ;
[0019] S22. Cross-level feature aggregation. Aggregate the reconstructed feature R (i) through skip connections and hierarchical weighting, and output the fused image feature F fused :
[0020] Specifically, step S3 specifically includes the following steps:
[0021] S31. Input the reference image I ref into the ArcFace network, and generate the face embedding E arc through multi-scale feature fusion and global average pooling;
[0022] S32. Input the reference image I ref into the CLIP-ViT encoder, and output the global context feature E clip after block embedding, position encoding and multiple layers of Transformer;
[0023] S33. The face embedding E arc and the global context feature E clipInput to the Q-Former module, and through the learnable Query matrix projection and cross-attention calculation of the Q-Former module, fuse and generate the refined facial embedding E final 。
[0024] Specifically, step S4 specifically includes the following steps:
[0025] S41. Inject the diffusion latent variable Z t into the dual-path cross-attention module, and interact with the fused image feature F fused respectively to obtain the image feature diffusion variable Z img , and interact with the identity embedding E final to obtain the identity embedding diffusion variable Z face ; the dual-path cross-attention module includes image-path cross-attention and identity-path cross-attention;
[0026] S42. Input the image feature diffusion variable Z img and the identity embedding diffusion variable Z face into the dynamic distribution alignment module. First, calculate the image feature mean μ img and the identity embedding mean μ face as well as the image feature standard deviation σ img and the identity embedding standard deviation σ face through the statistic calculation formula, and obtain the projection distribution by the distribution projection formula, fuse the residuals and add the gating weight γ, and output the aligned latent variable Z aligned ;
[0027] S43. Add a temporal smoothing constraint to the aligned latent variable Z aligned : Among them, the temporal smoothing constraint formula is as follows:
[0028]
[0029] Among them, L align is the loss function of the temporal smoothing constraint, which is used to measure the difference between the aligned latent variable and its average pooling result at different time steps; is the representation of the aligned latent variable Z aligned at the t-th time step; AvgPool() is the average pooling operation function; T is the total denoising steps of the diffusion model.
[0030] Specifically, step S5 specifically includes the following steps:
[0031] S51. Inject the aligned latent variable Z aligned and the text embedding E text into the spatio-temporal DiT architecture, and generate the preliminary temporal latent variable Z through shallow low-frequency feature fusion and deep high-frequency cross-attentiontime ; The shallow - layer low - frequency feature fusion is performed by bilinear upsampling followed by convolution for fusion;
[0032] S52. Input the diffusion latent variable Z t into the Hamilton - Jacobi optimization module, and perform reverse optimization based on the identity similarity loss to adjust the distribution of the diffusion latent variable Z t , and output the optimized latent variable Z t-1 ;
[0033] S53. Inject the optimized latent variable Z t-1 into the multi - frame autoregressive generator, and through the optical flow constraint loss L flow and sliding window weighted fusion, finally output a high - quality video with consistent identity
[0034] Specifically, step S11 specifically includes the following steps:
[0035] S111. Initialize wavelet transform, input the reference image and the text prompt T, and use the predefined Haar wavelet filter bank {f LL , f LH , f HL , f HH}, where:
[0036]
[0037] Perform the i - th level two - dimensional discrete wavelet transform (2D - DWT) on I ref :
[0038]
[0039] where and the output resolution of each level is reduced to 1 / 2 of the previous level; the value range of i is from 1 to L; stride = 2 means the output resolution of each level is reduced to 1 / stride of the previous level; is the image of the previous level of the i - th level two - dimensional discrete wavelet transform;
[0040] S112. Multilevel recursive decomposition, recursively perform L - level wavelet decomposition on the low - frequency component :
[0041]
[0042] Output the multi - scale feature group where is the low - frequency component, is the high - frequency component, and L is a positive integer greater than or equal to 3.
[0043] Specifically, step S12 specifically includes the following steps:
[0044] S121. Global low-frequency feature F low Align and splice the low-frequency components of each layer through bicubic interpolation to obtain:
[0045]
[0046] where Concat() is a fusion function; Upsample() is an upsampling function;
[0047] S122. Local high-frequency feature F high Perform hierarchical weighted fusion through the high-frequency components of each layer:
[0048]
[0049] where is the attenuation weight, is the channel splicing.
[0050] Specifically, step S13 specifically includes the following steps:
[0051] S131. Text condition injection. Input the text prompt T into the CLIP text encoder. After Tokenize and positional encoding, generate the text embedding
[0052] E text = CLIP text (T) = Transformer 12L (Embed(T))
[0053] CLIP text (T) means inputting the text prompt T into the text branch of the CLIP model for processing to obtain the preliminary text features; Transformer 12L (Embed(T)) means converting the text prompt T into an embedding vector through Embed(T), and then inputting this embedding vector into a 12-layer Transformer model for Tokenize and positional encoding, and finally obtaining the text embedding of the text prompt T where N represents the number of video frames generated in a single generation in the multi-frame autoregressive generator, which is used to control the step size of the sliding window; d is the feature dimension of the diffusion latent variable, which is used to characterize the spatial information of the compressed video frames;
[0054] S132. Inject E text into the wavelet features through cross-modal attention to output the multi-scale wavelet feature F wavelet = (F low ,Fhigh )
[0055] F wavelet = CrossAttn(Q = F low ‖F high , K = V = E text )
[0056] where CrossAttn is a cross-modal attention function, || represents concatenation along the channel dimension, the number of attention heads is 8, and the scaling factor is Q is the query vector, which is formed by concatenating different-scale features of the image; K is the key vector; V is the value vector; F low is the low-frequency feature; F high is the high-frequency feature.
[0057] In a second aspect, the present invention provides a consistent video generation device based on wavelet transform and frequency decomposition. Based on the consistent video generation method based on wavelet transform and frequency decomposition described in the above first aspect, it includes the following units:
[0058] Wavelet feature extraction unit, based on the frequency band decoupling of multi-scale Haar wavelet decomposition, extracts the low-frequency contour feature and high-frequency texture feature of the reference image, obtains the text embedding of the text prompt through the CLIP text encoder, and performs cross-modal attention on the low-frequency contour feature, high-frequency texture feature, and text embedding to obtain the multi-scale wavelet feature F wavelet ;
[0059] Cross-scale fusion unit, used to fuse the multi-level features of the multi-scale wavelet feature F wavelet through band adaptive convolution and inverse wavelet reconstruction and retain the cross-scale details to output the fused image feature F fused ;
[0060] Facial embedding generation unit, used to fuse the facial embedding E arc of the reference image obtained by the ArcFace network and the global context feature E clip of the reference image obtained by the CLIP-ViT encoder through Q-Former to generate a refined facial embedding E final ;
[0061] Alignment latent variable generation unit, used to input the diffusion latent variable Z t , the fused image feature F fused and the identity embedding E final into the dual-path cross-attention module for identity distribution alignment and temporal smoothing constraint in the latent variable space to obtain the aligned latent variable Z aligned ;
[0062] Consistent video generation unit, used to use the aligned latent variable Z alignedWith text embedding E text Inject into the spatio-temporal DiT architecture to generate a preliminary temporal latent variable Z time , and the diffusion latent variable Z t Generate an optimized latent variable Z through Hamilton-Jacobi optimization t-1 , then the preliminary temporal latent variable Z time And the optimized latent variable Z t-1 Generate a high-quality video V through a multi-frame autoregressive generator output .
[0063] Through the wavelet multi-level frequency band decoupling-fusion framework and the dynamic cross-modal attention mechanism, combined with the dynamic alignment mechanism of the distribution-aware adapter and the HJB optimization strategy of hierarchical diffusion, the present invention realizes the fusion of identity representations across spatio-temporal dimensions while ensuring parameter efficiency, solves the problems of high-frequency detail drift, dynamic scene identity distortion, and poor long-sequence generation stability in traditional video generation methods, significantly improves the quality and stability of the generated video, and provides more advanced technical support for video processing and related applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0065] Figure 1 Schematic diagram of a consistency video generation method based on wavelet transform and frequency decomposition provided by an embodiment of the present invention;
[0066] Figure 2 Schematic diagram of a consistency video generation device based on wavelet transform and frequency decomposition provided by an embodiment of the present invention;
[0067] Figure 3 Schematic diagram of a consistency video generation device based on wavelet transform and frequency decomposition provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The present invention can be explained in detail through the following embodiments. The purpose of providing the present invention is to protect all technical improvements within the scope of the present invention. In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0069] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0070] Embodiment 1
[0071] Reference Figure 1 , this embodiment provides a method for generating consistent videos based on wavelet transform and frequency decomposition, including the following steps:
[0072] S1. Based on the frequency band decoupling of multi-scale Haar wavelet decomposition, extract the low-frequency contour features and high-frequency texture features of the reference image, and obtain the text embedding of the text prompt through the CLIP text encoder, and perform cross-modal attention on the low-frequency contour features, high-frequency texture features, and text embedding to obtain multi-scale wavelet features F wavelet ;
[0073] Step S1 specifically includes the following steps:
[0074] S11. Input the reference image into a multi-level Haar wavelet transform, and recursively decompose it into L levels to obtain the low-frequency component and the high-frequency component where H is the image height and W is the image width;
[0075] Two-dimensional discrete wavelet transform (2D-DWT) is a method that extends one-dimensional discrete wavelet transform to two dimensions for processing two-dimensional signals such as images. Its core idea is to decompose an image into sub-bands (low-frequency and high-frequency components) of different scales through filtering and downsampling, thereby achieving multi-resolution analysis. 2D-DWT is a general image decomposition method that can use different wavelet basis functions (such as Haar, Daubechies, Symlets, etc.), and the effects produced when using different wavelet basis functions are different.
[0076] In this embodiment, the multi-scale Haar wavelet decomposition is a two-dimensional discrete wavelet transform using the Haar wavelet basis function.
[0077] The core of the Haar wavelet transform is the binary hierarchical decomposition of signals. For one-dimensional signals, the low-frequency and high-frequency components are separated by calculating the average and difference of adjacent data; for two-dimensional signals (such as images), four sub-bands are generated through two one-dimensional transforms in the row and column directions:
[0078] Low-frequency component (LL): preserves the overall contour of the signal or the average information of the image;
[0079] Horizontal high-frequency component (HL): captures changes in the vertical direction (such as horizontal edges);
[0080] Vertical high-frequency component (LH): captures changes in the horizontal direction (such as vertical edges);
[0081] Diagonal high-frequency component (HH): captures diagonal edges or texture details;
[0082] In the Haar wavelet transform, multiple wavelet transforms can be performed on a signal or an image. Each transform decomposes the signal or image into low-frequency and high-frequency components. Similarly, in 2D-DWT, multi-level decomposition can also be performed on an image, and each decomposition decomposes the image into four sub-bands.
[0083] Specifically, it includes the following steps:
[0084] S111. Initialize the wavelet transform, input the reference image and the text prompt T, and adopt the predefined Haar wavelet filter bank {f LL , f LH , f HL , f HH}, where:
[0085]
[0086] Perform the i-th level two-dimensional discrete wavelet transform (2D-DWT) on I ref :
[0087]
[0088] where and the output resolution of each level is reduced to 1 / 2 of the previous level; the value range of i is from 1 to L; stride = 2 means that the output resolution of each level is reduced to 1 / stride of the previous level; is the image of the previous level of the i-th level two-dimensional discrete wavelet transform;
[0089] S112. Perform multi-level recursive decomposition, and recursively perform L-level wavelet decomposition on the low-frequency component :
[0090]
[0091] Output multi-scale feature group wherein is the low-frequency component, is the high-frequency component, L is the wavelet decomposition level, and L is a positive integer greater than or equal to 3;
[0092] S12. Band feature recombination. Align and splice the low-frequency components of each layer through bicubic interpolation, and fuse them to generate the global low-frequency feature F low , that is, the low-frequency contour feature (contour information); sum the high-frequency components of each layer hierarchically with weights to generate the local high-frequency feature F high , that is, the high-frequency texture feature (texture details);
[0093] S121. Global low-frequency feature F low Align and splice the low-frequency components of each layer through bicubic interpolation to obtain:
[0094]
[0095] wherein, Concat() is the fusion function; Upsample() is the upsampling function;
[0096] S122. Local high-frequency feature F high is hierarchically weighted and fused through the high-frequency components of each layer:
[0097]
[0098] wherein is the attenuation weight, is the channel splicing.
[0099] S13. Input the text prompt T into the CLIP text encoder to extract the text embedding E text , and interact with the wavelet features through cross-modal attention, and finally output the multi-scale wavelet feature F wavelet =(F low ,F high ); wherein, F low is the low-frequency feature; F high is the high-frequency feature.
[0100] S131. Text condition injection. Input the text prompt T into the CLIP text encoder, and generate the text embedding after Tokenize and positional encoding whose expression formula is:
[0101] E text =CLIP text (T)=Transformer 12L (Embed(T))
[0102] CLIPtext (T) represents inputting the text prompt T into the text branch of the CLIP model for processing to obtain preliminary text features; Transformer 12L (Embed(T)) represents converting the text prompt T into an embedding vector through Embed(T), and then inputting this embedding vector into a 12-layer Transformer model for Tokenize and position encoding to finally obtain the text embedding of the text prompt T where N is the number of video frames generated in a single time by the multi-frame autoregressive generator in the subsequent steps, used to control the step size of the sliding window; d is the feature dimension of the diffusion latent variable (LatentVariable) in the subsequent steps, used to characterize the spatial information of the compressed video frames
[0103] Tokenize is a key step in natural language processing (NLP) to split text into meaningful units (such as words, punctuation, etc.), and its core role is to provide structured data support for subsequent tasks
[0104] S132. Inject the wavelet features E through cross-modal attention text to output multi-scale wavelet features F wavelet =(F low , F high ):
[0105] F wavelet =CrossAttn(Q = F low ‖F high , K = V = E text )
[0106] where CrossAttn is the cross-modal attention function, || represents concatenation along the channel dimension, the number of attention heads is 8, and the scaling factor is Q is the query vector, composed of the concatenated features of different scales of the image; K is the key vector; V is the value vector; F low is the low-frequency feature; F high is the high-frequency feature
[0107] S2. Fuse the multi-level features of the multi-scale wavelet feature F wavelet through the band adaptive convolution module and inverse wavelet reconstruction and retain the cross-scale details to output the fused image feature F fused ;
[0108] Specifically, it includes the following steps
[0109] S21. The multi-scale wavelet feature F wavelet =(F low , F high)The injection frequency band adaptive convolution module performs hierarchical reverse reconstruction and inverse wavelet transform reconstruction to obtain the reconstructed feature R (i) ;
[0110] S211. Construct the frequency band adaptive convolution kernel of the frequency band adaptive convolution module; the frequency band adaptive convolution kernel includes a low-frequency channel convolution kernel and a high-frequency channel convolution kernel; the total number of levels of the hierarchical structure of the frequency band adaptive convolution module is 3 levels;
[0111] The low-frequency channel convolution kernel adopts a decomposed sparse convolution structure:
[0112]
[0113] The high-frequency channel convolution kernel adopts a dense small kernel structure:
[0114]
[0115] Among them, the low-frequency channel uses 51×5 and 5×51 kernels; the high-frequency channel uses 3×3 dense kernels; C low is the number of low-frequency channels, and C high is the number of high-frequency channels;
[0116] S212. Hierarchical reverse reconstruction: Through the frequency band adaptive convolution kernel, recursively perform inverse wavelet transform and convolution operations on the multi-scale wavelet feature to generate the reconstructed feature R (i) , and the reconstruction formula for the i-th level is:
[0117]
[0118] Among them, is the reconstructed low-frequency feature of the i-th level; is the reconstructed high-frequency feature of the i-th level; IWT(·) is the inverse wavelet transform, and IWT(·) uses the Haar wavelet inverse transform kernel Conv(·) is the convolution function;
[0119] It can be understood that the multi-scale wavelet feature F wavelet =(F low ,F high ) is equivalent to , and both F low and F high contain L-level features, where L is the wavelet decomposition level; and are equivalent; the value range of i is from 1 to L;
[0120] In the hierarchical reverse reconstruction in step S212, the number of levels is the same as the decomposition level L in step S11 (usually L = 3), that is, the reverse reconstruction is recursively executed L times and then ends. For example, if the original image is decomposed by wavelet at 3 levels (L = 3), 3-level reverse operations are also required during reconstruction. At each level, the resolution is gradually restored through the band adaptive convolution module, and finally the fused image features are output.
[0121] S22. Cross-level feature aggregation, aggregating the reconstructed features R (i) through skip connections and hierarchical weighting to output the fused image features F fused :
[0122]
[0123] where is the hierarchical attenuation weight; Upsample(·, k) represents bicubic interpolation upsampling by k times.
[0124] S3. Fuse the facial embedding E of the reference image obtained by the ArcFace network arc and the global context features E of the reference image obtained by the CLIP-ViT encoder clip through Q-Former to generate the refined facial embedding E final ;
[0125] In this step, the ArcFace biometrics and the CLIP global context are combined, and a global content-aware facial encoder is constructed through Q-Former to achieve identity-scene dynamic interaction; specifically, it includes the following steps:
[0126] S31. Input the reference image I ref into the ArcFace network, and generate the facial embedding E arc through multi-scale feature fusion and global average pooling;
[0127] Specifically, the ArcFace network is a ResNet-50 backbone network;
[0128] Specifically, the ArcFace network is used to extract facial features: input the reference image extract multi-scale features through the ResNet-50 backbone network
[0129]
[0130] The multi-scale features are equivalent to stride 2 k , then fuse the multi-scale features and generate the facial embedding E arc :
[0131]
[0132] Among them, are the backbone network parameters; Concat() is the fusion function; GAP(·) is the global average pooling; FC(·) is the fully connected layer with a dimension of 512;
[0133] The classification loss function L used in the multi-scale feature fusion and global average pooling stage of the ArcFace network arc adopts the additive angular margin loss:
[0134]
[0135] where e is the base of the natural logarithm; log is the natural logarithm function; θ y is the angle of the true sample; m is the additive angular margin; y represents the true class label; j represents the possible class indices of the traversal index; θ j is the angle of the j-th sample; s is the scaling factor used to adjust the influence of the angle, making the distance between samples of similar classes smaller and the distance between different classes larger;
[0136] The classification loss function L arc (additive angular margin loss) is used in step S31 to optimize the face embedding generation of the ArcFace network. By expanding the inter-class angular margin of different identities and reducing the intra-class differences, the discriminability of the identity features is enhanced. In step S31, this loss function acts on the multi-scale feature fusion and global average pooling stage of the ArcFace network, directly driving the network to generate highly discriminative face embeddings E arc ;
[0137] S32. Input the reference image I ref into the CLIP-ViT encoder. After patch embedding, positional encoding, and multiple layers of Transformer, output the global context feature E clip ;
[0138] Both the CLIP text encoder and the CLIP-ViT encoder are part of the CLIP model. The former processes text to generate semantic embeddings, and the latter processes images to generate visual embeddings.
[0139] The CLIP-ViT encoder is an image encoder that combines the CLIP (Contrastive Language-Image Pre-training) encoder and the Vision Transformer (ViT). It is a key component in the CLIP model and is used to encode image content into high-dimensional feature vectors. These feature vectors can be compared with text feature vectors in different tasks to achieve the matching between images and texts. In the CLIP model, the image encoder is usually based on the ViT architecture. The CLIP-ViT encoder can use ViT as its backbone network to extract image features.
[0140] The CLIP-ViT encoder first divides the input image into image patches, and then performs linear embedding and positional encoding on these patches to preserve spatial information. Next, these embedded image patches are fed into a multi-layer Transformer network. The Transformer network learns the relationships between the image patches through the self-attention mechanism and outputs the global context features of the image.
[0141] Specifically, it includes: inputting the reference image I ref to the CLIP encoder to extract global features, generating global embeddings through the ViT structure, that is, the global context feature E clip , with a dimension of 768:
[0142]
[0143] The specific operation process of VIT:
[0144] Image patching:
[0145]
[0146] Transformer encoding:
[0147] E clip = MultiHeadAttn(LN(P pos ))
[0148] where P is the image patch; PatchEmbed() is the image patching function, MultiHeadAttn() is the multi-head attention function, LN() is the layer normalization function, and P pos is the patch position encoding of the image patch, and in ViT, position information is assigned to each image patch to preserve the spatial relationship.
[0149] S33. Embed the face E arc and the global context feature Eclip Input to the Q-Former module, and through the learnable Query matrix projection and cross-attention calculation of the Q-Former module, a refined facial embedding E is fused and generated. final ;
[0150] The Q-Former module is a cross-modal interaction component used to fuse the text / image features of CLIP and external identity features (such as the facial embedding of ArcFace), and achieve multimodal alignment through learnable Queries.
[0151] The dual encoders of CLIP provide the basic semantic-visual representation, and the Q-Former introduces a dynamic attention mechanism on this basis to enhance the fine-grained fusion of identity-related features.
[0152] Specifically, it includes: performing cross-modal interaction through the Q-Former module, inputting the facial embedding and the global feature projecting the facial embedding E through the learnable Query matrix W Q to project the facial embedding Q, and projecting the global context feature E arc to the key feature projection K and value feature projection V through the Key-Value projection, and then obtaining the attention facial embedding E through the cross-attention calculation of the facial embedding projection Q, key feature projection K, and value feature projection V clip , and finally obtaining the refined facial embedding E through feature fusion face ; final ;
[0153] The formula for the projection of the learnable Query matrix W Q is shown as follows:
[0154]
[0155] The Key-Value projection formula is as follows:
[0156]
[0157] The cross-attention calculation formula is as follows:
[0158]
[0159] The feature fusion formula is as follows:
[0160] E final = LayerNorm(E arc + FC(Flatten(E face )))
[0161] Among them, M is the number of heads in the multi-head attention (Multi-Head Attention), and its value is 32; D q is the dimension of the query vector (Query), and its value is 256; W K is the dimension of the weight matrix of the key (Key); W V is the dimension of the weight matrix of the value (Value); D k is the final dimension of the key vector (Key), and its value is 256; D v is the final dimension of the value vector (Value), and its value is 256; H' is the spatial height of the image; W' is the spatial width of the image; E face is the attention face embedding; LayerNorm() is the layer normalization (Layer Normalization) function; Flatten() is the flattening operation used to expand a multi-dimensional tensor into a one-dimensional sequence; E final is the enhanced face embedding (Enhanced Face Embedding) with a dimension of 512.
[0162] S4. Input the diffusion latent variable Z t , the fused image feature F fused and the identity embedding E final into the dual-path cross-attention module to perform identity distribution alignment and temporal smoothing constraint in the latent variable space to obtain the aligned latent variable Z aligned ;
[0163] Input Z aligned into the temporal constraint module, and optimize it through the sliding average optical flow loss, and finally output a temporally consistent latent variable sequence;
[0164] Specifically, it includes the following steps:
[0165] S41. Inject the diffusion latent variable Z t into the dual-path cross-attention module, and interact with the fused image feature F fused respectively to obtain the image feature diffusion variable Z img , and interact with the identity embedding E final to obtain the identity embedding diffusion variable Z face ; The dual-path cross-attention module includes image-path cross-attention and identity-path cross-attention;
[0166] The dual-path cross-attention module is an independent cross-modal fusion component in this method, used to align the diffusion latent variable and the identity feature (as described in step S41), while the CLIP model is used to provide the basic features of text / image encoding (text encoder and CLIP-ViT encoder).
[0167] The diffusion latent variable Z tInitialized as random Gaussian noise; the diffusion latent variable Z t has a feature dimension of d.
[0168] Specifically, the input diffusion latent variable image feature identity embedding are fed into the dual-path cross-attention module, and the image feature diffusion variable Z img is obtained through the image-path cross-attention formula. The image-path cross-attention formula is as follows:
[0169]
[0170] where CrossAttn is the cross-attention function; is the query projection weight matrix in the image path; is the key projection weight matrix in the image path; D h is the feature dimension of a single attention head (HeadDimension); is the value projection weight matrix in the image path; T is the total number of denoising steps of the diffusion model;
[0171] The identity embedding diffusion variable Z face is obtained through the identity-path cross-attention formula. The identity-path cross-attention formula is as follows:
[0172]
[0173] where Expand() is the tensor dimension expansion operation; is the query projection weight matrix in the identity path; is the key projection weight matrix in the identity path; is the value projection weight matrix in the identity path; T is the total number of denoising steps of the diffusion model;
[0174] The image feature diffusion variable Z img is the output result of the image-path cross-attention module, representing the diffusion latent variable guided by the fused image features;
[0175] The identity embedding diffusion variable Z face is the latent variable generated through the identity-path cross-attention, representing the distribution correction result guided by the identity embedding. Function: Inject facial identity features (such as pupils, contours) into the latent variable space to ensure the consistency of the character identity in the generated video.
[0176] S42. Feed the image feature diffusion variable Z img and the identity embedding diffusion variable Z faceInput to the dynamic distribution alignment module, first calculate the mean μ of the image features through the statistic calculation formula img and the mean μ of the identity embedding face and the standard deviation σ of the image features img and the standard deviation σ of the identity embedding face , and obtain the projection distribution through the distribution projection formula Fuse the residuals and add the gating weight γ to output the aligned latent variable Z aligned ;
[0177] The dynamic distribution alignment module includes statistic calculation, distribution projection calculation and residual fusion calculation;
[0178] The statistic calculation formula is as follows:
[0179]
[0180] Among them, Mean() is the mean function; Std() is the standard deviation function; Dim is the dimension; the 4 formulas belong to statistical formulas.
[0181] The distribution projection formula is as follows:
[0182]
[0183] The residual fusion formula is as follows:
[0184]
[0185] Among them, LayerNorm is the layer normalization function, which is used to reduce the feature distribution shift through normalization, making the alignment of the latent variables of the identity and image paths more efficient and reliable; γ is the gating weight.
[0186] S43. Enhance the temporal robustness. For the aligned latent variable Z aligned Add a temporal smoothing constraint: Among them, the temporal smoothing constraint formula is as follows:
[0187]
[0188] Among them, L align is the loss function of the temporal smoothing constraint, which is used to measure the difference between the aligned latent variable and its average pooling result at different time steps; is the representation of the aligned latent variable Z aligned at the t-th time step; AvgPool() is the average pooling operation function; T is the total number of denoising steps of the diffusion model.
[0189] S5. Combine the aligned latent variable Z aligned with the text embedding E textInject the spatio-temporal DiT architecture to generate the initial temporal latent variable Z time , and the diffusion latent variable Z t Generate the optimized latent variable Z through Hamilton-Jacobi optimization t-1 , and then the initial temporal latent variable Z time And the optimized latent variable Z t-1 Generate a high-quality video V through a multi-frame autoregressive generator output .
[0190] The diffusion latent variable Z t Is an intermediate state of the diffusion model, representing the latent space features during the denoising process. The diffusion model realizes video generation by adjusting the latent variable distribution, and the two are in a "framework-core variable" relationship.
[0191] The diffusion model is a generative model that generates data by gradually adding noise (forward diffusion) and then denoising backward. In the embodiment, the diffusion model refers to a video generation framework based on latent variable optimization, that is, step S5, which converts random Gaussian noise into video latent variables through multi-step denoising.
[0192] Inject the aligned latent variable Z aligned And the text embedding E text Into the spatio-temporal DiT architecture, and generate the initial temporal latent variable through shallow low-frequency feature fusion (bilinear upsampling + convolution) and deep high-frequency cross-attention; Inject the diffusion latent variable Z t Into the Hamilton-Jacobi optimization module, and perform reverse optimization based on the identity similarity loss L id , adjust the diffusion latent variable Z t Distribution, and output the optimized Z t-1 ; Inject the optimized latent variable sequence into the multi-frame autoregressive generator, and finally output a high-quality video with consistent identity through the optical flow constraint loss L flow And sliding window weighted fusion
[0193] S51. Inject the aligned latent variable Z aligned And the text embedding E text Into the spatio-temporal DiT architecture, and generate the initial temporal latent variable Z time Through shallow low-frequency feature fusion and deep high-frequency cross-attention; The shallow low-frequency feature fusion adopts bilinear upsampling followed by convolution for fusion;
[0194] The spatio-temporal DiT architecture is an improved structure that combines a diffusion model and a Transformer. Its core is to introduce a spatio-temporal separation attention mechanism (spatial 3D convolution + temporal self-attention) on the basis of the standard DiT to synchronously model the spatial details within video frames and the motion continuity between frames. This architecture belongs to the customized design of the method in this embodiment, and its basic configuration includes 6 stacked Transformer blocks. Each layer integrates a spatial local perception and a global temporal interaction module, supporting multi-scale latent variable optimization;
[0195] In step S51, the spatio-temporal DiT architecture is used to iteratively denoise and generate videos.
[0196] Specifically: Align the latent variables Text embedding Input them into the spatio-temporal DiT architecture;
[0197] Shallow low-frequency injection. The first 3 DiT blocks inject multi-scale wavelet low-frequency features The fusion formula is:
[0198]
[0199] Among them, Upsample() is an upsampling function; Conv() is a convolution function; l is the layer index of the spatio-temporal DiT architecture, representing the l-th Transformer block being processed currently; DiTBlock() is a spatio-temporal separated Transformer module, composed of spatial 3D convolution and temporal self-attention, supporting local spatial perception and global motion modeling; is the input latent variable of the l-th layer, the intermediate feature after being processed by the previous layer (l - 1); is the output projection weight matrix of the l-th layer, used to map the output features of the DiTBlock to the input dimension of the next layer;
[0200] Deep high-frequency injection. The 4th - 6th DiT blocks inject high-frequency features Adopt the cross-attention mechanism CrossAttn:
[0201]
[0202] In this embodiment and have the same meaning; and have the same meaning.
[0203] The cross-attention mechanism CrossAttn adopts a spatio-temporal attention separation mechanism. Each DiT block contains two kinds of attention executed in sequence:
[0204] Spatial attention:
[0205]
[0206] Temporal attention:
[0207]
[0208] where Q space is the spatial query matrix; is the spatial key matrix; V space is the spatial value matrix; Softmax() is the normalization function; Q time is the temporal query matrix; is the temporal key matrix; V time is the temporal value matrix; H' is the height of the feature map after bilinear upsampling; W' is the width of the feature map after bilinear upsampling, scaled synchronously with H'; D h is the diffusion latent variable Z t 's feature dimension, representing the length of each latent variable vector; T is the total number of denoising steps of the diffusion model, controlling the number of iterations from noise to clear video;
[0209] S52. Input the diffusion latent variable Z t into the Hamilton-Jacobi optimization module, and perform backpropagation optimization based on the identity similarity loss to adjust the distribution of the diffusion latent variable Z t and output the optimized latent variable Z t-1 .
[0210] The Hamilton-Jacobi (HJB) optimization module is a diffusion model training framework based on optimal control theory. Its role is to adjust the gradient update path of the diffusion latent variable Z t through dynamic programming strategies to ensure the stability and distribution alignment of the video generation process. The Hamilton-Jacobi (HJB) optimization module is the core technology (not part of CLIP) in the diffusion model training of the method in this embodiment. It is independent of the encoding function of CLIP and only acts on the latent variable optimization stage (step S52) to achieve high-quality video generation by minimizing the loss function.
[0211] For Hamilton-Jacobi optimization, in the denoising steps t ∈ {T,..., 1}, apply the face optimization constraint to the latent variable Z t to obtain the optimized latent variable Z t-1 , and the expression is shown as follows:
[0212]
[0213] where EDM_Solver is an optimization solver based on the Hamilton-Jacobi equation, using the backpropagation gradient Dynamically adjust the latent variable distribution and gradually transform Z t towards the direction of identity consistency, and finally output Z t-1 ; L id () is the identity similarity loss function; represents the identity similarity loss function L id with respect to the diffusion latent variable Z t partial derivative (i.e., gradient); η is the learning rate;
[0214] S53. Inject the optimized latent variable Z t-1 into the multi-frame autoregressive generator, and through the optical flow constraint loss L flow and sliding window weighted fusion, finally output a high-quality video with consistent identity
[0215] The diffusion latent variable Z t is initialized as random Gaussian noise and undergoes multi-step iterative denoising through the spatio-temporal DiT architecture to gradually update Z t to approximate the real video distribution.
[0216] Z time is the latent variable after the time dimension expansion in the spatio-temporal DiT. Z t is the main body for optimizing the diffusion model, and Z time is its spatio-temporal extended form. The two cooperate to achieve the temporal and spatial consistency of video generation.
[0217] The multi-frame autoregressive generator is the final module for video temporal generation in the generation method of this embodiment. Its function is to generate a video sequence frame by frame in an autoregressive manner based on the optimized diffusion latent variable Z t-1 , specifically including:
[0218] The multi-frame autoregressive generator gradually generates the latent variable sequence injected with the optimized latent variable Z t-1 in a sliding window manner. When generating a new window each time, weighted average fusion is performed on the overlapping region with the previous window to generate the final latent variable The final latent variable is mapped to the pixel space through the decoder module (VAE decoder) in the diffusion model framework and transformed into a visual video frame; through the optical flow constraint loss L flow force the motion consistency between adjacent frames and reduce jitter; finally, the frames generated by all windows are spliced in chronological order, and the video frames are output as a complete video through temporal smoothing, that is, a high-quality video with consistent identity The temporal smoothing process is temporal average pooling;
[0219] The optical flow constraint loss L flow is calculated as follows:
[0220]
[0221] Among them, the Flow() function is used to calculate the optical flow between adjacent frames, output a pixel-level motion vector field, and characterize the motion direction and amplitude of objects between frames (such as the micro-expression changes of the face between adjacent frames);
[0222] The Warp() function performs spatio-temporal warping on the latent variables based on the optical flow field, maps the latent variables of the current frame to the coordinate system of the next frame according to the optical flow, forces the alignment of the feature spaces of adjacent frames, and reduces temporal jitter;
[0223] It can be understood that in the formula, Z i , Z i+1 is the process latent variable used to represent the generation of the optimized latent variable Z t-1 ;
[0224] The final latent variable is gradually generated by a sliding window Specifically, it includes: generating with a window step of 4 frames and generating the final latent variable by weighted averaging in the overlapping area
[0225]
[0226] Among them, Z new is the new latent variable sequence generated by the current sliding window, representing the latest generated 4-frame latent features;
[0227] Z prev is the latent variable sequence generated by the previous window, which is used for weighted fusion in the overlapping area with the current window;
[0228] is the final latent variable from time window t to t + 3, which is the optimized result after optical flow constraint and weighted averaging;
[0229] α is the fusion weight parameter, with a value range of (0, 1), and preferably takes a value of 0.7 in this embodiment;
[0230] Finally, the frames generated by all windows are spliced in chronological order, and a complete video is output through temporal smoothing processing (such as temporal average pooling);
[0231] The finally generated latent variable is used to characterize the sequence of latent space features of video frames, and this latent variable needs to be mapped to the pixel space through the decoder module (VAE decoder) in the diffusion model framework and converted into a visual video frame.
[0232] Specifically, all the After the sequences are spliced in chronological order, the input decoder performs feature reconstruction to generate the original pixel-level video frames. Subsequently, temporal smoothing processing (such as temporal average pooling) is performed to further eliminate inter-frame jitter, and finally, high-quality videos are output.
[0233] The number of video frames generated at one time in the multi-frame autoregressive generator is N;
[0234] Temporal Average Pooling is an operation commonly used in the processing of time series data, especially in the fields of video processing and audio processing. Its basic idea is to reduce the dimension of data in the time dimension by calculating the average value of features within a certain time window to obtain a representative feature vector, which is prior art and will not be elaborated here.
[0235] In this embodiment, through the wavelet multi-level frequency band decoupling-fusion framework and the dynamic cross-modal attention mechanism, combined with the dynamic alignment mechanism of the distribution-aware adapter and the HJB optimization strategy of hierarchical diffusion, the identity representation fusion across spatio-temporal dimensions is achieved while ensuring parameter efficiency, solving the problems of high-frequency detail drift, dynamic scene identity distortion, and poor long-sequence generation stability in traditional video generation methods, significantly improving the quality and stability of the generated videos, and providing more advanced technical support for video processing and related applications.
[0236] Embodiment 2
[0237] Reference Figure 2 , this embodiment provides a consistent video generation device based on wavelet transform and frequency decomposition, including the following units:
[0238] Wavelet feature extraction unit, based on the frequency band decoupling of multi-scale Haar wavelet decomposition, extracts the low-frequency contour features and high-frequency texture features of the reference image, obtains the text embedding of the text prompt through the CLIP text encoder, and performs cross-modal attention on the low-frequency contour features, high-frequency texture features, and text embedding to obtain multi-scale wavelet features F wavelet ;
[0239] Cross-scale fusion unit, used to fuse the multi-level features of multi-scale wavelet features F wavelet through band adaptive convolution and inverse wavelet reconstruction and retain cross-scale details to output fused image features F fused ;
[0240] Facial embedding generation unit, used to fuse the facial embedding E arc of the reference image obtained by the ArcFace network and the global context feature E clip of the reference image obtained by the CLIP-ViT encoder through Q-Former to generate a refined facial embedding E final ;
[0241] An alignment latent variable generation unit for aligning the diffusion latent variable Z t , the fused image feature F fused and the identity embedding E final are input into a dual-path cross-attention module for aligning the identity distribution in the latent variable space and performing temporal smoothing constraints to obtain the aligned latent variable Z aligned ;
[0242] A consistency video generation unit for using the aligned latent variable Z aligned and the text embedding E text to inject into a spatio-temporal DiT architecture to generate a preliminary temporal latent variable Z time , and using the diffusion latent variable Z t to generate an optimized latent variable Z through Hamilton-Jacobi optimization t-1 , then using the preliminary temporal latent variable Z time and the optimized latent variable Z t-1 to generate a high-quality video V through a multi-frame autoregressive generator output .
[0243] Embodiment 3
[0244] Reference Figure 3 , Figure 3 FIG. is a schematic structural diagram of a consistency video generation device based on wavelet transform and frequency decomposition according to this embodiment. The consistency video generation device 20 based on wavelet transform and frequency decomposition in this embodiment includes a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, the steps in the above method embodiment are implemented. Alternatively, when the processor 21 executes the computer program, the functions of each module / unit in the above device embodiments are implemented.
[0245] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the consistency video generation device 20 based on wavelet transform and frequency decomposition. For example, the computer program can be divided into the respective modules in Embodiment 2. For the specific functions of each module, please refer to the working process of the device described in the above embodiment, and details are not described herein again.
[0246] The consistency video generation device 20 based on wavelet transform and frequency decomposition may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art can understand that the schematic diagram is only an example of the consistency video generation device 20 based on wavelet transform and frequency decomposition, and does not constitute a limitation on the consistency video generation device 20 based on wavelet transform and frequency decomposition. It may include more or fewer components than shown, or combine certain components, or different components. For example, the consistency video generation device 20 based on wavelet transform and frequency decomposition may also include input / output devices, network access devices, buses, etc.
[0247] The processor 21 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 21 is the control center of the consistency video generation device 20 based on wavelet transform and frequency decomposition, and connects various parts of the entire consistency video generation device 20 based on wavelet transform and frequency decomposition through various interfaces and lines.
[0248] The memory 22 can be used to store the computer programs and / or modules. The processor 21 realizes various functions of the consistency video generation device 20 based on wavelet transform and frequency decomposition by running or executing the computer programs and / or modules stored in the memory 22, and by calling the data stored in the memory 22. The memory 22 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.), etc. In addition, the memory 22 can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0249] Among them, if the modules / units integrated in the consistency video generation device 20 based on wavelet transform and frequency decomposition are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 21, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0250] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0251] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate for implementing in the process Figure 1 one process multiple processes and / or blocksFigure 1 a device with functions specified in one or more boxes
[0252] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in one or more processes and / or boxes Figure 1 one or more processes Figure 1 a device with functions specified in one or more boxes
[0253] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or boxes Figure 1 one or more processes Figure 1 a device with functions specified in one or more boxes
[0254] The parts not detailed in the present invention are prior art. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and are intended to encompass all changes falling within the meaning and scope of the equivalent elements within the present invention.
Claims
1. A method for generating consistent videos based on wavelet transform and frequency decomposition, characterized in that, Including the following steps: S1. Based on the frequency band decoupling of multi-scale Haar wavelet decomposition, extract the low-frequency contour features and high-frequency texture features of the reference image, obtain the text embedding of the text prompt through the CLIP text encoder, and perform cross-modal attention on the low-frequency contour features, high-frequency texture features, and text embedding to obtain the multi-scale wavelet feature F wavelet ; S2. Multiscale wavelet features F are fused through a band adaptive convolution module and inverse wavelet reconstruction, and multi-level features are retained while cross-scale details are output to obtain the fused image feature F wavelet ; fused ; S3. Embed the face of the reference image obtained by the ArcFace network, E arc with the global context feature E of the reference image obtained by the CLIP-ViT encoder clip to generate a refined face embedding E through fusion by Q-Former final ; S4. Input the diffusion latent variable Z t , fuse the image feature F fused and the identity embedding E final into the dual-path cross-attention module for identity distribution alignment and temporal smoothing constraint in the latent variable space to obtain the aligned latent variable Z aligned ; S5. Inject the aligned latent variable Z aligned and the text embedding E text into the spatio-temporal DiT architecture to generate a preliminary temporal latent variable Z time , and generate an optimized latent variable Z t for the diffusion latent variable Z through Hamilton-Jacobi optimization t-1 . Then, generate a high-quality video V time from the preliminary temporal latent variable Z t-1 and the optimized latent variable Z output through a multi-frame autoregressive generator 2. The consistency video generation method based on wavelet transform and frequency decomposition according to claim 1, wherein Step S1 specifically includes the following steps: S11. Input the reference image into a multi-level Haar wavelet transform, and recursively decompose it by L levels to obtain the low-frequency component and the high-frequency component where H is the image height and W is the image width; S12. Band feature recombination: Align and splice the low-frequency components of each layer through bicubic interpolation, and fuse them to generate the global low-frequency feature F low , sum the high-frequency components of each layer hierarchically with weights to generate the local high-frequency feature F high ; S13. Input the text prompt T into the CLIP text encoder to extract the text embedding E text , and through cross-modal attention and wavelet feature interaction, finally output the multi-scale wavelet feature F wavelet =(F low ,F high ); where F low is the low-frequency feature; F high is the high-frequency feature.
3. The consistency video generation method based on wavelet transform and frequency decomposition according to claim 2, wherein Step S2 specifically includes the following steps: S21. Inject the multi-scale wavelet feature F wavelet =(F low , F high ) into the frequency-band adaptive convolution module for hierarchical reverse reconstruction and inverse wavelet transform reconstruction to obtain the reconstructed feature R (i) ; S22. Cross-level feature aggregation, aggregating the reconstructed feature R through skip connections and hierarchical weighting (i) to output the fused image feature F fused .
4. The consistency video generation method based on wavelet transform and frequency decomposition according to claim 2, wherein Step S3 specifically includes the following steps: S31. Input the reference image I ref into the ArcFace network, and generate a face embedding E through multi-scale feature fusion and global average pooling arc ; S32. Input the reference image I ref into the CLIP-ViT encoder. After patch embedding, positional encoding, and multiple layers of Transformer, output the global context feature E clip ; S33. Embed the face into E arc and the global context feature E clip Input them into the Q-Former module. Through the learnable query matrix projection and cross-attention calculation of the Q-Former module, a refined face embedding E is fused and generated final .
5. The consistency video generation method based on wavelet transform and frequency decomposition according to claim 1, wherein Step S4 specifically includes the following steps: S41. Inject the diffusion latent variable Z t into the dual-path cross-attention module, and interact with the fused image feature F fused respectively to obtain the image feature diffusion variable Z img , and interact with the identity embedding E final to obtain the identity embedding diffusion variable Z face ; the dual-path cross-attention module includes image-path cross-attention and identity-path cross-attention; S42. Input the image feature diffusion variable Z img and the identity embedding diffusion variable Z face into the dynamic distribution alignment module. First, calculate the image feature mean μ img and the identity embedding mean μ face as well as the image feature standard deviation σ img and the identity embedding standard deviation σ face through the statistic calculation formula, and obtain the projection distribution by the distribution projection formula. Then, fuse the residuals and add the gating weight γ to output the aligned latent variable Z aligned ; S43. To align the latent variable Z aligned Add a time-domain smoothing constraint: where the formula for the time-domain smoothing constraint is as follows: Among them, L align is the loss function of the temporal smoothing constraint, which is used to measure the difference between the aligned latent variable and its average pooling result at different time steps; is the representation of the aligned latent variable Z aligned at the t-th time step; AvgPool() is the average pooling operation function; T is the total number of denoising steps of the diffusion model.
6. The consistency video generation method based on wavelet transform and frequency decomposition according to claim 1, characterized in that Step S5 specifically includes the following steps: S51. Align the latent variable Z aligned with the text embedding E text and inject them into the spatio-temporal DiT architecture to generate a preliminary temporal latent variable Z through shallow low-frequency feature fusion and deep high-frequency cross-attention time ; the shallow low-frequency feature fusion is performed by bilinear upsampling followed by convolution for fusion; S52. Input the diffusion latent variable Z t into the Hamilton-Jacobi optimization module, and perform reverse optimization based on the identity similarity loss to adjust the distribution of the diffusion latent variable Z t and output the optimized latent variable Z t-1 ; S53. Inject the optimized latent variable Z t-1 into the multi-frame autoregressive generator, and through the optical flow constraint loss L flow and sliding window weighted fusion, finally output a high-quality video with consistent identity 7. The consistency video generation method based on wavelet transform and frequency decomposition according to claim 2, characterized in that Step S11 specifically includes the following steps: S111. Initialize wavelet transform, input reference image and text prompt T, using a predefined Haar wavelet filter bank {f LL , f LH , f HL , f HH}, where: For I ref perform the i-th level two-dimensional discrete wavelet transform (2D-DWT): Among them and the output resolution of each level is reduced to 1 / 2 of the previous level; the value range of i is from 1 to L; stride = 2 means that the output resolution of each level is reduced to 1 / stride of the previous level; is the image of the previous level of the 2D discrete wavelet transform at the i-th level; S112. Multistage recursive decomposition is performed on the low-frequency components Recursively perform L-level wavelet decomposition: Output multi-scale feature groups Among them is the low-frequency component, is the high-frequency component, and L is a positive integer greater than or equal to 3.
8. The consistency video generation method based on wavelet transform and frequency decomposition according to claim 2, characterized in that Step S12 specifically includes the following steps: S121, Global low-frequency feature F low Align and splice the low-frequency components of each layer through bicubic interpolation Obtain: Among them, Concat() is a fusion function; Upsample() is an upsampling function; S122, Local high-frequency feature F high By hierarchically weighted fusion of high-frequency components of each layer: Among them is the attenuation weight, is the channel splicing.
9. The consistency video generation method based on wavelet transform and frequency decomposition according to claim 1, wherein Step S13 specifically includes the following steps: S131. Text condition injection: Input the text prompt T into the CLIP text encoder. After tokenization and positional encoding, generate a text embedding E text = CLIP text (T) = Transformer 12L (Embed(T)) CLIP text (T) means that the text prompt T is input into the text branch of the CLIP model for processing to obtain preliminary text features; Transformer 12L (Embed(T)) means that the text prompt T is converted into an embedding vector through Embed(T), and then this embedding vector is input into a 12-layer Transformer model for Tokenize and positional encoding to finally obtain the text embedding of the text prompt T where N represents the number of video frames generated in a single time in the multi-frame autoregressive generator, which is used to control the step size of the sliding window; d is the feature dimension of the diffusion latent variable, which is used to characterize the spatial information of the compressed video frames S132. Inject E into the wavelet features through cross-modal attention to output multi-scale wavelet features F text =(F wavelet , F low , F high ): F wavelet = CrossAttn(Q = F low ‖F high , K = V = E text ) Among them, CrossAttn is a cross-modal attention function, || represents concatenation along the channel dimension, the number of attention heads is 8, and the scaling factor is Q is a query vector, which is formed by concatenating different scale features of the image; K is a key vector; V is a value vector; F low is a low-frequency feature; F high is a high-frequency feature.
10. A consistent video generation device based on wavelet transform and frequency decomposition, characterized in that, Including the following units: Wavelet feature extraction unit, based on the frequency band decoupling of multi-scale Haar wavelet decomposition, extracts the low-frequency contour features and high-frequency texture features of the reference image, obtains the text embedding of the text prompt through the CLIP text encoder, and performs cross-modal attention on the low-frequency contour features, high-frequency texture features, and text embedding to obtain multi-scale wavelet features F wavelet ; Cross-scale fusion unit, which is used to fuse multi-scale wavelet features F through band adaptive convolution and inverse wavelet reconstruction wavelet of multi-level features and retain cross-scale details to output fused image features F fused ; A face embedding generation unit for generating a face embedding E of a reference image obtained by an ArcFace network arc and a global context feature E of the reference image obtained by a CLIP-ViT encoder clip are fused through a Q-Former to generate a refined face embedding E final ; Alignment latent variable generation unit, used to align the diffusion latent variable Z t , fuse the image feature F fused and the identity embedding E final Input them into the dual-path cross-attention module to perform identity distribution alignment and temporal smoothing constraint in the latent variable space to obtain the aligned latent variable Z aligned ; A consistency video generation unit for injecting the aligned latent variable Z aligned and the text embedding E text into the spatio-temporal DiT architecture to generate a preliminary temporal latent variable Z time and generating an optimized latent variable Z by optimizing the diffusion latent variable Z t through Hamilton-Jacobi. Then, the preliminary temporal latent variable Z t-1 and the optimized latent variable Z time are used to generate a high-quality video V t-1 through a multi-frame autoregressive generator output .
Citation Information
Patent Citations
Image feature extraction method and system based on multi-modal model, and electronic equipment
CN118506374A
Video generation method and device, electronic equipment and readable storage medium
CN119052529A
Controllable video generation method and system based on multi-modal fusion
CN119091362A
Consistency identity picture generation method based on diffusion model
CN119478090A
Image processing method and apparatus, device, medium and program product
WO2025015824A1
Cited By
Multi-modal diffusion-based long video role scene decoupling generation method and system
CN120583276A
Small sample visual anomaly detection method and device for power line
CN120599235A
A small sample visual anomaly detection method and device for power lines
CN120599235B
Video generation control method and system based on multi-mode lens trend
CN122248236A