Dance video generation method based on multi-mode music driving and frequency domain-space double-flow decomposition

Through the multimodal music driving and frequency domain-space dual-stream decomposition method, the problems of motion lag, visual details loss and instability in the scenes of occlusion in dance generation in the prior art are solved, and high-fidelity, synchronous and stable dance video generation are achieved.

CN120238708AActive Publication Date: 2025-07-01湖南马栏山视频先进技术研究院有限公司

Patent Information

Application Number
CN202510387352.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing music-driven dance generation methods have problems such as movement lag or missed shooting, loss of visual details, and instability in generation in occlusion scenes.

Method used

The dance video generation method based on multimodal music driving and frequency domain-space dual-stream decomposition is adopted to achieve space-time alignment of music-visual features by gated cross-modal attention, and the joint motion trajectory is predicted using a segmented Transformer decoder, and the spatial posture and frequency domain features are separated through graph convolution networks and Butterworth filter groups, and a parallel diffusion framework is constructed for optimization. Finally, high-fidelity dance video is generated through Laplace pyramid reconstruction and subpixel convolution.

Benefits of technology

The millisecond synchronization of the action and music beats is achieved, which improves the fidelity of visual details and maintains the stability and rationality of generation in the occluded scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238708A_ABST
    Figure CN120238708A_ABST
Patent Text Reader

Abstract

The invention provides a dance video generation method based on multi-mode music driving and frequency domain-space double-flow decomposition, and the method comprises the steps: extracting multi-granularity music features through a composite encoder (Librosa + Jukebox), employing a beat gating attention mechanism, enabling the key actions such as dancing hand raising, kicking and the like to be strictly aligned with a music re-beat point, and enabling the synchronization error to be reduced to 118 ms through the verification of a test data set; for the problem of visual detail loss, a frequency domain-space double-flow decomposition architecture is provided, a Butterworth filter bank is used to decouple a reference image into a low-frequency energy diagram and a high-frequency residual error, and a double-flow diffusion mechanism is used to optimize a global attitude and local details respectively; a joint confidence prediction module is introduced for the generation stability in a shielding scene, and the motion trail of an abnormal joint point is dynamically corrected through a time domain sliding window weighted fusion strategy, so that a reasonable action conforming to ergonomics can still be generated under a 50% limb shielding rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and artificial intelligence-generated content. Specifically, it relates to a dance video generation method based on multi-modal music driving and frequency domain-spatial double-stream decomposition. Background Art

[0002] Existing music-driven dance generation methods have three defects: First, the single-modal feature alignment strategy is difficult to capture the complex temporal correlations between music beats and dance movements, resulting in lagging or off-beat movements. For example, in existing CNN-LSTM-based models, the action synchronization error exceeds 200 ms under strong-rhythm music. Second, the single-stream generation framework cannot effectively separate spatial pose movements and frequency domain detail features. Especially when the limbs move rapidly, edge blurring and texture distortion are likely to occur. Test data shows that the PSNR value of traditional methods in the clothing fold area is only 28.4 dB. Third, the global optimization strategy lacks dynamic evaluation of local joint confidence. When there are occlusions in the input image (such as an arm being covered by an object), the generated pose sequence is prone to limb breakage or anti-joint abnormalities.

[0003] The background description provided in this article is for the purpose of presenting the context of the present disclosure generally. Unless otherwise indicated herein, the materials described in this section are not prior art to the claims of this application and should not be admitted as prior art by including them in this section. Summary of the Invention

[0004] In view of the above technical problems in the related art, the present invention proposes a dance video generation method based on multi-modal music driving and frequency domain-spatial double-stream decomposition, including:

[0005] S1. Obtain the beat sequence and music style features of the music; and use gated cross-modal attention to achieve spatio-temporal alignment of music-visual features and output multi-modal alignment features;

[0006] S2. Use a part-based Transformer decoder to predict the joint movement trajectory, and optimize the frequency domain consistency between high-frequency actions and music beats through wavelet decomposition to obtain the global pose sequence; the part-based Transformer is specifically a four-head attention mechanism that processes different body parts respectively;

[0007] S3. Process the global pose sequence based on a graph convolutional network to generate a body region mask, and use a Butterworth filter bank to process the non-body region in the reference image to separate the low-frequency energy map and the high-frequency residual features;

[0008] S4. Construct a parallel diffusion framework for spatial and frequency domain streams, and optimize the original video sequence by combining the optical flow consistency constraint to maintain temporal smoothness to obtain an optimized video sequence, where the spatial stream is guided by AdaIN pose and the frequency domain stream is a wavelet modulation network;

[0009] S5. Reconstruct cross-scale features through the Laplacian pyramid, enhance high-frequency details using sub-pixel convolution, and generate a high-fidelity dance video.

[0010] Further, the step S1 specifically includes:

[0011] S11. Perform dual feature encoding on the music waveform signal, and use the Librosa toolkit to extract the binary beat sequence Extract music style features using the Transformer encoder of the Jukebox pre-trained model;

[0012] S12. Extract the underlying texture features and high-level semantic features of the reference images in the video, and splice the temporal attention results with the low-level texture features upsampled by bilinear interpolation through the gated cross-modal attention mechanism to output multi-modal alignment features; among them, a beat-gated cross-modal attention mechanism is constructed based on the beat features Dynamically adjust the music-visual feature fusion weight; σ represents the Sigmoid activation function.

[0013] Further, in the step S12, the splicing of the temporal attention result with the underlying texture feature through the gated cross-modal attention mechanism to output multi-modal alignment features is specifically: through the learnable parameter matrix Project the music style feature into a query vector Map the visual high-level semantic feature to a key And a value

[0014] Based on the beat feature Generate a dynamic gating factor Modulate the attention weight of each frame through γ t Finally, splice the temporal attention result with the underlying feature to output multi-modal alignment features.

[0015] Further, the step S2 specifically includes:

[0016] S21. Perform spatio-temporal window partitioning on the multi-modal alignment features to obtain overlapping window sequences {W k = F align [8k - L + 1:8k]}, and each window is compressed to 256 channels in dimension through a 1×1 convolution to obtain ​Among them, T represents the total number of video frames, and H / 4 and W / 4 are the spatial dimensions after downsampling;

[0017] S22. Use a part-based Transformer to predict joint movement trajectories The part-based Transformer specifically uses a four-head attention mechanism to process different body parts; among them, the query vector is generated by the position encoding feature and the projection matrix ; the key-value pairs K i , V i are mapped from the context feature through ; the four-head attention results are fused through a fully connected layer and then the pose prediction within the window is output

[0018] S23. The discrete wavelet transform decomposes the pose sequence of the sequence of the window into low-frequency components and high-frequency components and calculates the frequency-domain consistency loss L between the high-frequency components after the short-time Fourier transform and the music beat feature freq , and performs weighted averaging on the prediction results of all windows containing the t-th frame to obtain the global pose sequence

[0019] Further, step S3 specifically includes the following steps:

[0020] S31. Model the human body topological relationship through a graph convolutional network for the pose to generate a body region mask

[0021] S32. Weight the reference image through the body region mask for to obtain the low-frequency component For the non-body region it is processed by a fourth-order Butterworth high-pass filter bank to separate the high-frequency components in the horizontal LH t , vertical HL t and diagonal HH t three directions, and the low-frequency component LL t generates a low-frequency energy map through 3×3 convolution and layer normalization The high-frequency components are compressed to 32 dimensions through channel concatenation Cat(LH t , LH t , HH t ) and 1×1 convolution, and then through instance normalization, finally generate a high-frequency residual map

[0022] Further, step S4 specifically includes the following steps:

[0023] S41. Process the initial video sequence based on AdaIN to obtain the first optimized video sequence;

[0024] S42. Use the wavelet modulation network to process the frequency domain features

[0025] to obtain...

[0026] Further, the pose projection network of AdaIN adopts a 3-layer MLP structure.

[0027] Further, the depthwise separable convolution kernel size of the wavelet modulation network is 3×3, and the dilation rate is 2.

[0028] Further, step S5 specifically includes the following steps:

[0029] S51. Based on the optimized video sequence construct a four-level Laplacian pyramid;

[0030] S52. Map the features of each level of the Laplacian pyramid to a unified dimension through cross-scale attention fusion to obtain projection features Construct a query vector key vector and value vector Calculate the cross-scale correlation weight through the multi-head attention mechanism synthetic feature

[0031] S53. Calculate the Laplacian residual between the original frame and the upsampled low-frequency component Restore the detail information through the improved sub-pixel convolutional network Decov(·).

[0032] Further, for the improved sub-pixel convolutional network, first expand the number of channels to 3×r through a 1×1 convolutional layer 2 = 12, and the weight matrix adopts He initialization and performs operations Subsequently, apply the PixelShuffle operation to increase the spatial resolution of the feature map to 2H×2W and reduce the channel dimension to 3 to form intermediate features To further enhance the detail expression ability, a three-layer convolutional structure is used for processing: the first layer of 3×3 convolution extracts local texture features; the second layer of 5×5 dilated convolution expands the receptive field; the third layer of 1×1 convolution generates a detail weight map Finally, it is fused with the residual features through skip connections to calculate the detail enhancement result where ⊙ represents element-wise multiplication, and after bicubic interpolation downsampling to the original resolution, it is synthesized with the cross-scale fusion features in proportion and output Truncated to the range of [0, 255].

[0033] The present invention extracts multi-granularity music features through a composite encoder (Librosa + Jukebox), and uses a beat gating attention mechanism to strictly align key dance actions such as raising hands and kicking legs with the music beat points. After verification by the test dataset, the synchronization error is reduced to 118 ms. Aiming at the problem of visual detail loss, a frequency-domain-spatial dual-stream decomposition architecture is proposed. The reference image is decoupled into a low-frequency energy map (encoding the overall body movement trend) and a high-frequency residual (retaining clothing texture and light and shadow details) using a Butterworth filter bank. The dual-stream diffusion mechanism optimizes the global pose and local details respectively. Aiming at the generation stability in occlusion scenarios, a joint confidence prediction module is introduced, and the motion trajectory of abnormal joint points is dynamically corrected through a time-domain sliding window weighted fusion strategy, so that reasonable actions that conform to ergonomics can still be generated under a 50% limb occlusion rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0035] Figure 1 It is a schematic diagram of a dance video generation method based on multi-modal music driving and frequency-domain-spatial dual-stream decomposition provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present invention.

[0037] Embodiment 1

[0038] Reference Figure 1 , this embodiment discloses a dance video generation method based on multi-modal music driving and frequency-domain-spatial dual-stream decomposition, including:

[0039] S1. Obtain the beat sequence and music style features of the music; and use gated cross-modal attention to achieve spatio-temporal alignment of music-visual features and output multi-modal alignment features;

[0040] In this embodiment, music features are extracted through Librosa beat detection and Jukebox style encoding; specifically, the following steps are included:

[0041] S11. Perform dual feature encoding on the music waveform signal, and use the Librosa toolkit to extract the binary beat sequence; use the Transformer encoder of the Jukebox pre-trained model to extract the music style features;

[0042] Specifically, in this embodiment, the original audio can be decoded to obtain time series data to form the music waveform signal.

[0043] For the original music waveform signal perform dual feature encoding, where T m represents the number of audio sampling points, and F = 128 is the Mel spectrogram dimension. Extract the binary beat sequence through the Librosa toolkit Each element of which represents whether the corresponding video frame contains a beat point. At the same time, use the Transformer encoder of the Jukebox pre-trained model to extract the music style features of the original music waveform signal And resample the music style features through cubic spline interpolation to the video frame rate to obtain aligned music style features

[0044] S12. Extract the low-level texture features and high-level semantic features of the reference image in the video, splice the temporal attention result with the low-level texture features through the gated cross-modal attention mechanism, and output multi-modal alignment features; among them, a beat gated cross-modal attention mechanism is constructed based on the beat features Dynamically adjust the music-visual feature fusion weight; where σ represents the Sigmoid activation function;

[0045] Specifically, the splicing of the temporal attention result with the low-level texture features through the gated cross-modal attention mechanism and outputting multi-modal alignment features is specifically: project the music style features into query vectors through a learnable parameter matrix map the visual high-level semantic features to keys and values and

[0046] Based on the beat features generate a dynamic gating factor modulate the attention weight of each frame through γ t to Finally, the temporal attention result is concatenated with the low-level features upsampled by bilinear interpolation to output the multimodal alignment feature. in Used to generate dynamic weight factors for the beat-gated cross-modal attention mechanism to modulate the fusion strength of music-visual features.

[0047] Combined with the reference image, a 3×3 convolutional network is used to extract low-level texture features. and high-level semantic features encoded by residual blocks Constructing a beat-gated cross-modal attention mechanism Dynamically adjust the music-visual feature fusion weights, where is a learnable gating parameter.

[0048] For the reference image The low-level texture features are extracted through a 3×3 convolutional layer The high-level semantic features are then further encoded through a residual block consisting of two 3×3 convolutional layers and skip connections.

[0049] The cross-modal alignment phase uses a learnable parameter matrix Project the music style features into a query vector Visual high-level feature maps as keys Sum

[0050] Based on beat features Generating dynamic gating factors By γ t Modulating the attention weights of each frame Finally, the temporal attention result is concatenated with the low-level features upsampled by bilinear interpolation to output the multimodal alignment feature. In this process, d = 256 is the feature projection dimension, σ represents the Sigmoid activation function, and the residual block operation is defined as F out =Conv(ReLU(Conv(F in )))+F in , where spatial downsampling is achieved through convolution operations with a stride of 2.

[0051] S2, using the part-by-part Transformer decoder to predict the joint motion trajectory, and optimizing the frequency domain consistency of high-frequency movements and music beats through wavelet decomposition to obtain the global posture sequence; the part-by-part Transformer is specifically a four-head attention mechanism that processes different body parts respectively;

[0052] Step S2 specifically includes the following steps:

[0053] S21. Align multimodal features Perform spatio-temporal window partitioning to obtain overlapping window sequences {W k = F align [8k - L + 1:8k]}, where each window is compressed to 256 channels in the channel dimension through 1×1 convolution to obtain where T represents the total number of video frames, and H / 4 and W / 4 are the spatial dimensions after downsampling;

[0054] First, perform spatio-temporal window partitioning on the multi-modal alignment features , where T represents the total number of video frames, and H / 4 and W / 4 are the spatial dimensions after downsampling. The temporal features are segmented into overlapping window sequences {W k = F align [8k - L + 1:8k]} of length L = 16 through sliding window slicing operation, and each window is compressed to 256 channels in the channel dimension through 1×1 convolution to obtain

[0055] S22. Use the part-based Transformer to predict joint movement trajectories The part-based Transformer specifically uses a four-head attention mechanism to process different body parts; among them, the query vector is generated by the position encoding feature and the projection matrix , and the key-value pairs K i , V i are mapped from the context feature through . The four-head attention results are fused through a fully connected layer and then the pose prediction within the window is output

[0056] In the improved Transformer decoding stage, a four-head attention mechanism is used to process different body parts respectively, where the query vector is generated by the position encoding feature and the projection matrix , and the key-value pairs K i , V i are mapped from the context feature through . The four-head attention results are fused through a fully connected layer and then the pose prediction within the window is output where the three-dimensional coordinates of 25 joint points include the x, y positions and the confidence c.

[0057] S23. Discrete wavelet transform decomposes the pose sequence of the window sequence into low-frequency components and high-frequency components and calculate the frequency-domain consistency loss L between the high-frequency components after short-time Fourier transform and the music beat features of freq Perform weighted average on the prediction results of all windows containing the t-th frame to obtain the global pose sequence where C k ∈[0,1] L×25 is the joint confidence, and the global pose sequence includes the three-dimensional coordinates and confidence of the joints.

[0058] Decompose the pose sequence into low-frequency components and high-frequency components by discrete wavelet transform, and calculate the frequency-domain consistency loss L between the high-frequency components after short-time Fourier transform and the music beat features of freq The frequency-domain consistency loss calculates the L1 norm difference between the action spectrum and the music spectrum in the beat region: specifically, the amplitude spectra of the high-frequency components of the action and the amplitude spectra of the music beat features are subtracted point by point in the masked area, and the weights in the uncovered area are set to zero. The loss function is expressed as where M beat is the beat mask, and S motion and S music are the STFT results of the action and the music respectively. This loss function forces the frequency-domain energy distribution of high-frequency limb movements (such as hand jitter and foot frequency) to align with the music signal at the corresponding moments of the music beat.

[0059] where STFT uses a frame length of 512 points and a Hamming window, and the beat region mask M mask is generated by binarization with a threshold τ = 0.7 for The beat region mask M mask is mainly used to constrain the synchronization optimization of high-frequency action components and music beats in the frequency domain. Its core function is to mark the time-frequency regions corresponding to music beats through binarization, focus on the spectral alignment of these regions during the calculation of the frequency-domain consistency loss, and force the frequency-domain energy distribution of high-dynamic actions such as hand jitter and foot rhythm to strictly match the music beat features. The confidence prediction module processes the pose sequence through a one-dimensional convolutional network with a kernel size of 3, and outputs the joint confidence C k ∈[0,1] L×25 Finally, perform weighted average on the prediction results of all windows containing the t-th frame to obtain the global pose sequence

[0060] In the window slicing strategy of this embodiment, the multiple relationship between L = 16 and the step size of 8 ensures the action coherence. The Haar wavelet basis function adopts the standard orthogonal form and Implement frequency domain decomposition.

[0061] S3. Process the global pose sequence based on the graph convolutional network to generate a body region mask, and use a Butterworth filter bank to process the non-body region in the reference image to separate the low-frequency energy map and the high-frequency residual features;

[0062] Step S3 specifically includes the following steps:

[0063] S31. Model the human body topological relationship through the graph convolutional network for the said pose Generate a body region mask

[0064] Model the human body topological relationship through the graph convolutional network (GCN) to generate a body region mask. Based on the SMPL human anatomical structure, pre-define the adjacency matrix A ∈ {0, 1} 25×25 , if there is a physical connection between joints i and j (such as shoulder-elbow), then A ij = 1, capture the joint dependency relationship through the normalized adjacency matrix (where D is the degree matrix, D ii represents the number of adjacent nodes of joint i). The first layer of graph convolutional operation multiplies the pose coordinates P t with the weight matrix and outputs a 64-dimensional feature after ReLU activation The second layer of operation generates a spatial attention score through the weight and obtains the body region mask after Sigmoid activation and bilinear interpolation upsampling Precisely segment the main body action region.

[0065] S32. Weight the reference image through the said body region mask For Obtain the low-frequency component Non-body region Then, it is processed by a fourth-order Butterworth high-pass filter bank to separate the horizontal LH t , vertical HL t and diagonal HH t high-frequency components in three directions. The low-frequency component LL t generates a low-frequency energy map through 3×3 convolution and layer normalization The high-frequency components are compressed to 32 dimensions through channel concatenation Cat(LH t , HL t , HH t ) and 1×1 convolution, and then through instance normalization, finally generate a high-frequency residual map

[0066] ​Subsequently, it enters the wavelet domain decomposition stage: the masked weighted reference image Generates the low-frequency component through 3×3 average pooling (AvgPool) with a step size of 2 Captures the overall motion energy; non-body regions Then passes through a fourth-order Butterworth high-pass filter bank (transfer function ) for processing to separate the horizontal LH t , vertical HL t and diagonal HH t high-frequency components in three directions, with the cut-off frequency set to 0.1 times the Nyquist frequency to match the spectral characteristics of dance movements. The low-frequency component LL t Generates the low-frequency energy map through 3×3 convolution (Xavier initialization) and layer normalization (LayerNorm) Stabilizes the global feature distribution; the high-frequency components are compressed to 32 dimensions through channel concatenation Cat(LH t , HL t , HH t ) and 1×1 convolution, and then the instance normalization (InstanceNorm) is used to retain the local contrast. Finally, the high-frequency residual map is generated according to the fusion weight α = 0.3

[0067] S4. Construct a parallel diffusion framework for the spatial flow and frequency domain flow, and optimize the original video sequence by combining the optical flow consistency constraint to maintain temporal smoothness to obtain the optimized video sequence, where the spatial flow is guided by AdaIN pose and the frequency domain flow is a wavelet modulation network;

[0068] Step S4 specifically includes the following steps:

[0069] S41. Process the initial video sequence based on AdaIN to obtain the first optimized video sequence;

[0070] Based on the initial video sequence (where T is the total number of frames, H = 1024 and W = 768 are the frame resolutions) and frequency domain features The spatial flow processing adopts a dynamic adaptive instance normalization mechanism, and its mathematical expression is V t (k) =(1 - α k )·V t (k-1) +α k ·AdaIN(V t (k-1) , P t ), where k ∈ {1,..., K} represents the number of diffusion iterations (default K = 50), α k= 0.1 + 0.8(k - 1) / (K - 1) to achieve parametric linear growth, and the AdaIN operation is defined as μ p , σ p = MLP(P t ), where μ v , σ v are the statistics of the input video frame. When processing the initial video sequence based on AdaIN in this embodiment, the input of its pose projection network is the concatenated feature of the frequency-domain feature and the pose sequence, and the affine parameters are generated through a 3-layer MLP to dynamically modulate the feature distribution of the video frame, ensuring that the global motion is aligned with the music beat.

[0071] S42. Use the wavelet modulation network to process the frequency-domain feature to obtain the cross-modal modulated low-frequency energy map and the time-frequency enhanced high-frequency residual feature, and the modulated low-frequency energy map and high-frequency residual feature are respectively input to the spatial stream;

[0072] The frequency-domain stream detail injection is implemented through the wavelet modulation network where represents the depthwise separable convolution operation, and SpatialDropout randomly masks the spatial region with a probability p = 0.2 during the training phase.

[0073] S43. Through the PWC-Ne optical flow estimation network and differentiable bilinear interpolation, realize the pose-driven optical flow fusion optimization to obtain the optimized video sequence; among them, PWC-net calculates the motion field of adjacent frames in the video sequence, and differentiable bilinear interpolation realizes the pose-driven optical flow to calculate the pose sequence.

[0074] The temporal consistency constraint module uses the pre-trained PWC-Net optical flow estimation network φ(·) to calculate the motion field of adjacent frames and realizes the pose-driven optical flow through differentiable bilinear interpolation Finally, construct the optical flow consistency loss where the TV regularization term coefficient λ = 0.05 is used to smooth abnormal motion. The video frame after the two-stream fusion is output after 5 iterations of optimization

[0075] In this embodiment, the pose projection network of AdaIN adopts a 3-layer MLP structure (256-128-64 nodes) to extract pose parameters; the depthwise separable convolution kernel size of the wavelet modulation network is 3×3, and the dilation rate is set to 2 to expand the receptive field; the pose difference during the calculation of the optical flow loss is smoothed through quaternion interpolation, and the interpolation weight coefficient β = 0.7 is obtained by end-to-end learning. The description of the source of each parameter is as follows: V t (k) represents the intermediate video frame of the kth diffusion iteration, α kThe linear growth strategy ensures the preservation of image fidelity in the initial stage; Affine parameters generated for pose conditions; Generated by mapping joint displacements to the image plane, with its normalized coordinate transformation matrix Initialized with camera parameters and fine-tuned through backpropagation.

[0076] S5. Reconstruct cross-scale features through the Laplacian pyramid, enhance high-frequency details using sub-pixel convolution, and generate high-fidelity dance videos.

[0077] Specifically, the high-frequency residual features And the enhanced high-frequency features output by the wavelet modulation network After cross-scale attention fusion, they are input into the Laplacian pyramid, where And the upsampled low-frequency component LL t Are concatenated to generate multi-scale features, Then, a detail weight map is generated through a sub-pixel convolution network, and the texture details of the Laplacian residuals of the original frame are enhanced channel by channel, and finally a high-fidelity video sequence is synthesized.

[0078] Step S5 specifically includes the following steps:

[0079] S51. Based on the optimized video sequence Construct a four-level Laplacian pyramid;

[0080] Based on the optimized video sequence Construct a four-level Laplacian pyramid, where the l-th level pyramid Is generated by downsampling through bicubic interpolation, satisfying V t l+1 = PyrDown(V t l ) and the weight of the downsampling kernel matrix Is in the two-dimensional separable form of [1, 4, 6, 4, 1] / 16.

[0081] S52. Through cross-scale attention fusion, map the features of each level of the Laplacian pyramid to a unified dimension to obtain projection features Construct a query vector A key vector And a value vector Calculate the cross-scale correlation weights through the multi-head attention mechanism Synthesize features

[0082] Through cross-scale attention fusion, map the features of each level through a 3×3 convolution to a unified dimension d = 256 to obtain projection features Furthermore, construct a query vector A key vector Sum vector where the projection matrix is a learnable parameter. Calculate the cross-scale correlation weights through the multi-head attention mechanism Finally, synthesize the feature where is the hierarchical importance coefficient, initialized to [0.4, 0.3, 0.2, 0.1] and optimized through backpropagation;

[0083] S53. Calculate the Laplacian residual R between the original frame and the upsampled low-frequency component t = V t (5) - PyrUp(V t 1 ), and restore the detailed information through the improved sub-pixel convolutional network Deconv(·).

[0084] In the high-frequency detail compensation stage, calculate the Laplacian residual R between the original frame and the upsampled low-frequency component t = V t (5) - PyrUp(V t 1 ), and restore the detailed information through the improved sub-pixel convolutional network Deconv(·), and its core operation is D t = Conv3×3(PixelShuffle(R t W sub ))), where is the sub-pixel convolution kernel (enlargement factor r = 2), and the final output frame is truncated to the range of [0, 255] through numerical truncation.

[0085] The low-frequency component of this embodiment is the low-frequency component obtained in step S32

[0086] When constructing the Laplacian pyramid in this embodiment, a five-layer Gaussian kernel with a standard deviation of σ = 1.6 is adopted, the cross-scale attention adopts a 4-head parallel computing mechanism, and the channel number configuration of the sub-pixel convolutional network is a three-layer structure of [32, 64, 32]; the symbol PyrUp(·) represents bilinear interpolation upsampling, and its interpolation coefficient matrix is The source explanations of each entity mathematical symbol are as follows: V t l is generated by downsampling the initial video frame l times, and K down The weight distribution conforms to the characteristics of the second derivative of the Gaussian; β l,m reflects the dependence strength of the l-th layer on the m-th layer, and is parameterized and modeled through multi-head attention; R t characterizes the detailed difference between the original image and the low-frequency reconstruction, and its calculation process uses zero-padding to process the boundary; Wsub Channel dimension of 3×r 2 Ensure that the number of output channels matches the input after PixelShuffle.

[0087] The technical process of constructing a four - level Laplacian pyramid for the video sequence in step S5 is as follows:

[0088] Using the initial video frame as the 0 - th level V of the pyramid t 0 , perform a progressive downsampling operation using a two - dimensional separable convolution kernel. Specifically, the downsampling kernel matrix K down is composed of the horizontal kernel k h = [1, 4, 6, 4, 1] / 16 and the vertical kernel . The resolution is halved through a convolution operation with a stride of 2. For the generation of the l - th level pyramid, first perform reflection padding (padding = 2) on V t l-1 to handle boundary effects, and then perform one - dimensional convolution along the horizontal direction to calculate the intermediate feature F t l,inter = Conv1D(V t l-1 , k h ), and then perform convolution of the same kernel on the intermediate feature along the vertical direction to obtain the downsampling result V t l = Conv1D(F t l,inter , k v ). This process is repeated three times to generate a four - level pyramid with gradually decreasing resolutions: the 1 - st level the 2 - nd level the 3 - rd level where the equivalent standard deviation σ of the five - layer Gaussian kernel is 1.6, which is approximated by discretizing the kernel coefficients to approximate the continuous Gaussian function generated. After the construction of each level of the pyramid, it is upsampled back to the original resolution through bicubic interpolation and the Laplacian residual is calculated where PyrUp(·) uses a coefficient matrix for bilinear interpolation, and finally a complete pyramid representation containing the low - frequency main structure (V t 3 ) and multi - level high - frequency details is formed.

[0089] The specific implementation process of high - frequency detail compensation in step S5 is as follows:

[0090] First, the optimized original video frame and the upsampled low - frequency component PyrUp(V t1 ) Perform per-pixel difference operation to generate a Laplacian residual matrix Among them, the upsampling operation is performed using a bilinear interpolation kernel, and the interpolation coefficient matrix is defined as And handle the boundary effect through reflection padding. The residual features are input into the improved sub-pixel convolution network. First, the number of channels is expanded to 3×r 2 = 12 (upscaling factor r = 2), and the weight matrix Adopt He initialization and perform the operation Subsequently, apply the PixelShuffle operation to increase the spatial resolution of the feature map to 2H×2W and reduce the channel dimension to 3 to form intermediate features To further enhance the detail expression ability, a three-layer convolution structure is adopted for processing: the first layer is a 3×3 convolution (input / output channels 32, ReLU activation) to extract local texture features; the second layer is a 5×5 dilated convolution (dilation rate 2, number of channels 64) to expand the receptive field; the third layer is a 1×1 convolution (number of channels 32, Sigmoid activation) to generate a detail weight map Finally, fuse with the residual features through skip connection to calculate the detail enhancement result D t = D w ⊙F t shuffle +(1 - D w )⊙Conv3×3(F t shuffle ), where ⊙ represents element-wise multiplication, and after bicubic interpolation downsampling to the original resolution, it is synthesized proportionally with the cross-scale fusion features And output Truncated to the range of [0, 255].

[0091] In this embodiment, by integrating music multi-modal analysis and image frequency domain decomposition technology, an end-to-end generation framework is constructed: extract the music beat sequence based on the Librosa tool Combine with the 128-dimensional style features encoded by the Jukebox model Innovatively design a beat gating attention mechanism Dynamically adjust the fusion weight of visual features and music features, so that strong beat points correspond to large body movements. In the image processing link, generate a human body region mask through a graph convolutional network Use a Butterworth filter bank to decompose the image into a low-frequency energy map And high-frequency residuals Among them, the low-frequency component encodes the overall body movement energy, and the high-frequency component retains the clothing texture and light and shadow details. The two drive spatial pose and detail enhancement respectively through a two-stream diffusion mechanism. This technology can be widely applied to scenarios such as virtual idol real-time driving, personalized digital entertainment creation, and film and television special effects preview.

[0092] Furthermore, the beat-gated cross-modal attention mechanism in this embodiment dynamically adjusts the visual-music feature fusion weight through music beat features to achieve millisecond-level action-beat synchronization;

[0093] The frequency-domain-spatial two-stream decomposition architecture uses a Butterworth filter bank to separate the low-frequency energy map and high-frequency residual features, and optimizes the pose movement and detailed texture respectively;

[0094] The local mask generation technology based on graph convolution accurately segments the body area through a graph convolutional network with human body topology constraints, improving the generation robustness in occlusion scenarios;

[0095] The weighted fusion strategy driven by joint confidence combines a time-domain sliding window and confidence evaluation to dynamically correct abnormal joint movement trajectories;

[0096] The two-stream diffusion collaborative optimization framework has the spatial stream (AdaIN pose guidance) and the frequency-domain stream (wavelet modulation network) iterating in parallel, taking into account global rationality and local fidelity;

[0097] The cross-scale attention fusion mechanism realizes the efficient fusion and detail enhancement of multi-scale features through Laplacian pyramid reconstruction and sub-pixel convolution.

[0098] It should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative work.

[0099] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A dance video generation method based on multimodal music drive and frequency domain-space dual stream decomposition, characterized by: include: S1. Obtain the beat sequence and music style characteristics of the music; And use gated cross-modal attention to achieve spatiotemporal alignment of music and visual features to output multimodal aligned features; S2, using the part-by-part Transformer decoder to predict the joint motion trajectory, and optimizing the frequency domain consistency of high-frequency movements and music beats through wavelet decomposition to obtain the global posture sequence; the part-by-part Transformer is specifically a four-head attention mechanism that processes different body parts respectively; S3, based on the graph convolutional network, the global posture sequence is processed to generate the body area mask, and the Butterworth filter group is used to process the non-body area in the reference image to separate the low-frequency energy map and high-frequency residual features; S4, build a parallel diffusion framework of spatial stream and frequency domain stream, combine the optical flow consistency constraint to maintain temporal smoothness to optimize the original video sequence to obtain the optimized video sequence, where the spatial stream is guided by AdaIN posture and the frequency domain stream is a wavelet modulation network; S5. Reconstruct cross-scale features through Laplacian pyramid and use sub-pixel convolution to enhance high-frequency details to generate high-fidelity dance videos.

2. The method according to claim 1, characterized in that: The step S1 specifically includes: S11. Perform dual feature encoding on the music waveform signal and extract the binary beat sequence using the Librosa toolkit The Transformer encoder of the Jukebox pre-trained model is used to extract music style features; S12. Extract the underlying texture features and high-level semantic features of the reference image in the video, concatenate the temporal attention results with the low-level texture features upsampled by bilinear interpolation through the gated cross-modal attention mechanism, and output multi-modal alignment features; among them, a beat gated cross-modal attention mechanism is constructed based on the beat features. Dynamically adjust the music-visual feature fusion weight; σ represents the Sigmoid activation function.

3. The method according to claim 2, characterized in that In step S12, the temporal attention result is concatenated with the underlying texture feature through the gated cross-modal attention mechanism to output the multi-modal alignment feature. Specifically, the learnable parameter matrix Project the music style features into a query vector Visual high-level semantic feature mapping as key Sum Based on beat features Generating dynamic gating factors By γ t Modulating the attention weights of each frame Finally, the temporal attention results are concatenated with the underlying features to output multimodal alignment features.

4. The method according to claim 3, characterized in that: The step S2 specifically includes: S21. Align multimodal features Perform time-space window division to obtain A sequence of overlapping windows {W k =F align [8k-L+1:8k]}, each window is compressed to 256 channel dimensions by 1×1 convolution. Where T represents the total number of video frames, H / 4 and W / 4 are the spatial dimensions after downsampling; S22. Predict joint motion trajectory using Transformer with different parts The Transformer for each part is specifically a four-head attention mechanism that processes different body parts respectively; the query vector Features encoded by position With the projection matrix Generate a key-value pair K i ,V i By context features through The four attention results are fused by the fully connected layer and the posture prediction in the output window is obtained. S23, discrete wavelet transform decomposes the posture sequence of the window sequence into low-frequency components and high frequency components And calculate the high-frequency component after short-time Fourier transform and the music beat characteristics The frequency domain consistency loss L freq , weighted average of all window prediction results including the tth frame Get the global pose sequence 5. The method according to claim 4, characterized in that: Step S3 specifically includes the following steps: S31, modeling the topological relationship of the human body through a graph convolutional network to the posture Generating body region masks S32: Weighting the reference image by the body region mask right Get low frequency components Non-body areas Then it is processed by a fourth-order Butterworth high-pass filter bank to separate the horizontal LH t , vertical HL t With diagonal HH t High frequency components in three directions, low frequency components LL t The low-frequency energy map is generated by 3×3 convolution and layer normalization. The high frequency component is spliced ​​through the channel Cat (LH t , L.H. t , HH t ) and 1×1 convolution to 32 dimensions, and then normalized by instance to generate a high-frequency residual map 6. The method according to claim 5, characterized in that: Step S4 The specific steps include: S41, processing the initial video sequence based on AdaIN to obtain a first optimized video sequence; S42, using wavelet modulation network to analyze frequency domain features Processing to obtain... S43: Optimize the timing consistency of the first optimized video sequence by optical flow to obtain a timing optimized video sequence.

7. The method according to claim 6, characterized in that: AdaIN's posture projection network adopts a 3-layer MLP structure.

8. The method according to claim 7, characterized in that: The depthwise separable convolution kernel size of the wavelet modulation network is 3×3, and the dilation rate is 2.

9. The method according to claim 8, characterized in that: Step S5 specifically includes the following steps: S51, based on the optimized video sequence Construct a four-level Laplace pyramid; S52, mapping each level of the Laplacian pyramid feature to a unified dimension through cross-scale attention fusion to obtain projection features Constructing query vector Key Vector Sum value vector Calculate cross-scale association weights through multi-head attention mechanism Synthetic Features S53, calculate the Laplace residual of the original frame and the upsampled low-frequency component The detail information is restored by an improved sub-pixel convolutional network Deconv(·).

10. The method according to claim 9, characterized in that: The improved sub-pixel convolutional network first expands the number of channels to 3×r through a 1×1 convolutional layer. 2 =12, weight matrix Initialize with He and perform calculations Then the PixelShuffle operation is applied to increase the spatial resolution of the feature map to 2H×2W, and the channel dimension is reduced to 3 to form the intermediate feature To further enhance the ability to express details, a three-layer convolution structure is used: the first layer of 3×3 convolution extracts local texture features; The second layer of 5×5 hole convolution expands the receptive field; the third layer of 1×1 convolution generates a detail weight map Finally, the skip connection is used to fuse the residual features to calculate the detail enhancement results. Where ⊙ represents channel-by-channel multiplication, bicubic interpolation downsampling to the original resolution and cross-scale fusion features Proportional synthesis, output Truncated to the range [0,255].

Citation Information

Patent Citations

  • Music dance posture generation method based on multi-feature fusion strategy

    CN114998984A

  • Audio driving action synthesis method and device

    CN115604529A

  • Multi-modal feature fusion-based music dance posture generation method and device, and storage medium

    CN117316129A

  • Harmony-aware human motion synthesis with music

    US20230005201A1

  • Audio-visual fusion with cross-modal attention for video action recognition

    WO2021184026A1

Cited By

  • Music performance posture real-time driving method and system based on time-space cooperation

    CN120411317A

  • Intelligent dancing garment generation method and system based on multi-modal action analysis

    CN121302465A

  • A method and system for generating an intelligent dance costume based on multi-modal motion analysis

    CN121302465B

  • Dance video generation method based on motion focusing attention and decoupling control

    CN121531199A

  • Compliant control method and system based on semantic anchoring and frequency domain pulse residual error

    CN121733585A