Method for generating dance video based on music and lyrics
By integrating the time synchronization mechanism of lyric semantics and music beats and the U-Net and Reference Net networks, the problem of lyrics and music alignment in dance video generation is solved, improving the generation effect, and achieving high-quality dance video generation.
Patent Information
- Application Number
- CN202510622376.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-15
AI Technical Summary
The existing dance generation technology relies on music rhythm generation movements, lacks richness and expressiveness, the problem of lyric semantics and time alignment of music cannot be effectively solved, and the high-quality multimodal data sets are lacking, which limits the model training effect.
The time synchronization mechanism that combines lyric semantic analysis and music beat analysis is adopted, combined with U-Net and Reference Net network structures, through self-attention, cross-attention and timing attention mechanisms, dance movements are generated aligned with lyric keywords and music rhythms, and noise masks and initial images are introduced to enhance the generation effect.
The generated dance video movements are aligned with the lyrics content in timing, with higher quality and expressiveness, conforming to user aesthetic standards, and their diversity and innovation have been improved.
Smart Images

Figure CN120499462A_ABST
Abstract
Description
Technical Field
[0001] The present invention involves the intersection of computer vision, artificial intelligence, and multimedia technology. It is a multimodal dance video generation method driven by music melody under the constraints of lyrics semantics. It can be applied to virtual character animation, game development, digital entertainment, digital art and other fields. Background Art
[0002] In today's digital entertainment era, people are increasingly interested in the creation and dissemination of dance videos. Traditional dance video production requires professional dancers, choreographers, and complex filming and post-production processes, resulting in high costs and low efficiency. With the rapid development of artificial intelligence (AI), the use of computer vision and deep learning techniques to automatically generate dance videos has become a research hotspot. Existing dance generation techniques mostly rely on musical rhythm or melody to generate movements, resulting in dance videos that lack richness and expressiveness. Temporal alignment of lyrics and music remains a challenge, resulting in a mismatch between movements and lyrics. The lack of high-quality multimodal datasets (synchronized data containing dance movements, music, and lyrics) also limits model training effectiveness. Research that can fully integrate the semantics and emotions of lyrics and music to create vivid and smooth dance content is of practical significance.
[0003] The U-Net framework effectively integrates data from different modalities. Within the U-Net model, self-attention, cross-attention, and temporal attention mechanisms work together to comprehensively process the input Gaussian noise, fused feature vectors, and static guidance images. The self-attention mechanism focuses on dependencies within the action sequence, the cross-attention mechanism focuses on the associations between features from different modalities, and the temporal attention mechanism ensures temporal coherence of the action sequence. However, research on applying these techniques to generate dance videos based on lyrics and music is still in its exploratory stages and faces numerous technical challenges. Summary of the Invention
[0004] The present invention provides a feasible method for generating dance videos based on music and lyrics. It adopts a time synchronization mechanism that integrates lyrics semantic analysis and music beat analysis to ensure the alignment of dance movements with lyrics keywords and music rhythm. It introduces advanced network structures such as U-Net and ReferenceNet to further improve the quality and expressiveness of generated dance movement sequences. The U-Net network structure is used to fine-tune and optimize the generated dance movements, using its powerful feature fusion and context perception capabilities to ensure the continuity and naturalness of the movements. On the other hand, Reference Net is used to introduce reference information, and by comparing and drawing on predefined dance movement templates or examples, the generated dance movements are more in line with expectations and artistic requirements.
[0005] Furthermore, techniques such as noise masking and initial images are incorporated to further enhance dance video generation. Noise masking helps the model better handle uncertainty and randomness, increasing the diversity and innovation of generated movements. The introduction of initial images provides visual guidance and constraints for dance video generation, ensuring that the generated videos better meet user expectations and aesthetic standards.
[0006] The general process of the present invention is mainly composed of the following steps:
[0007] Step 1: Multi-module feature encoding: Focusing on the multimodal data processing of songs and dances, data collection, video framing, audio track separation, and lyrics analysis are carried out in sequence to build a multimodal dataset. In addition, posture, audio, and lyrics features are extracted to complete multimodal feature encoding.
[0008] Step 2: Multi-module feature fusion: The extracted multi-modal feature vectors are fused into one feature vector, the importance of different modal features is dynamically calculated through the attention mechanism, and weighted fusion is performed based on this.
[0009] Step 3: Feature Processing and Action Generation: The fused features are dimensionally resized, temporally aligned, and renormalized to ensure the proper functioning of the motion decoder and the quality of the generated action sequences. The U-Net model input data includes the pose sequence output by the motion decoder, the guidance image processed by the Reference Net, and Gaussian noise. Through the synergistic effects of self-attention, cross-attention, and temporal attention mechanisms, a continuous action sequence is generated.
[0010] Step 4: Dance Movement Output: Rendering is performed to convert each pose in the U-Net action sequence into a visual sequence of frames. The discrete video frames are synthesized into a coherent video file. Using video editing techniques, special effects, transitions, and other elements are added to produce a harmonious dance video. Furthermore, the specific steps involved in Step 1 are as follows:
[0011] Step 1.1 Data Collection: Use web crawling technology to collect dance video footage from multiple sources, including popular video sharing platforms (such as YouTube and Bilibili) and open-source dance datasets. Convert the collected footage to .MP4 format to build an initial dataset containing dance movements, music, and lyrics.
[0012] Step 1.2 Video framing: Use EasyMocap to obtain high-fidelity body estimation from the video at a rate of 60FPS and output the pose sequence in .csv format.
[0013] Step 1.3: Separate the audio tracks: Use Adobe Audition to perform noise reduction on the video and then save the audio as a .WAV file. Use the open-source audio source separation tool Spleeter to extract the vocals from the mixed audio and save the tracks as separate .wav files.
[0014] Step 1.4 Lyrics parsing: Perform speech recognition (ASR) on the vocal track based on the Whisper pre-trained model, generate a text transcription with timestamps, and output it in .SRT format.
[0015] Step 1.5: Multimodal dataset construction: Synchronize pose (60FPS), audio (librosa beat tracking), and lyrics (.SRT semantic segments) using a global time base, combined with dynamic time warping to achieve cross-modal temporal alignment.
[0016] Step 1.6 Pose Feature Extraction: The pose sequence is encoded into a low-dimensional latent space using the POSE VAE (Variational Autoencoder), obtaining distribution parameters that represent the pose features. Using the reparameterization technique, the latent variable z is sampled from this Gaussian distribution defined by μ and σ. This latent variable z is the direct pose feature. The latent variable z can be expressed as:
[0017] Z=μ+ε⊙σ
[0018] Where μ represents the mean, ε is a random variable sampled from a standard normal distribution, ⊙ represents element-wise multiplication, and σ represents the standard deviation. In this way, the specific posture feature z is obtained from the distribution parameters.
[0019] Step 1.7 Audio Feature Extraction: Use the Librosa library to extract basic audio features, such as Mel-frequency cepstral coefficients (MFCC), constant Q-value chromaticity maps, timing maps, and the intensity of the original audio. Jukebox uses multi-scale VQ-VAE (Vector Quantized Variational Autoencoder) to further encode the extracted audio features into discrete vector representations, while capturing the dynamic changes of the audio through an autoregressive generation mechanism. The generated audio vector is input into the Transformer module, and the temporal dependency between audio features is captured through the self-attention mechanism, thereby optimizing the audio representation. The implementation of calculating the attention weight in the Transformer module is shown in the following formula:
[0020]
[0021] Among them, Q, K, and V represent query, key, and value vectors respectively. k is the dimension of the key vector.
[0022] Step 1.8 Lyrics (text) feature extraction: Use the jieba library to perform semantic segmentation on the lyrics text. Each word or character obtained by the segmentation forms a corresponding embedding vector. The embedded vector sequence is sent to the pre-trained CLIP model. After the multi-layer Transformer module operation, the hidden representation h is obtained, and the feature vector expressing the key content such as the lyrics semantics and emotions is extracted.
[0023] Furthermore, the specific steps involved in step 2 are as follows:
[0024] Step 2.1 uses the additive attention mechanism in the feature-level fusion method to fuse the posture, audio, and lyrics features. The implementation of the weighted fusion feature vector is shown in the following formula:
[0025]
[0026] Among them, f fusion Represents the final fusion feature vector, a i is the attention weight of the i-th modal feature, satisfying f i Usually represents the i-th eigenvector.
[0027] Furthermore, the specific steps of step 3 are as follows:
[0028] Step 3.1 Feature processing: In this step, we need to perform dimension adjustment, temporal alignment, and renormalization on the fused features to ensure that the action decoder can work properly.
[0029] Step 3.2 Action Generation: Use Reference Net to process the guidance image and extract key features, which are used to maintain consistency in style or details in subsequent generation tasks. Gaussian noise is used as the input condition of U-Net, introducing controllable randomness to enhance the diversity of the generated results. The extracted key features, Gaussian noise, and pose sequence are input into the U-Net model. The self-attention mechanism, cross-attention mechanism, and temporal attention mechanism work together to achieve dance movement generation. The Gaussian noise probability density function formula is as follows:
[0030] G(x, y) = f(x, y) + n(x, y)
[0031] Wherein, f(x, y) represents the original signal, and n(x, y) is a function that obeys Gaussian distribution. Further, the specific steps of step 4 are as follows:
[0032] Step 4.1 Rendering: Use the 3D modeling software Blender to build a 3D model of the dancer and perform skeletal rigging and skinning. Map the dance movement sequence data output by U-Net to the skeletal system of the 3D model and export the 3D model as an FBX file. Import this FBX file into the Unity game engine, configure the camera perspective, lighting environment, and rendering parameters, and use Unity's rendering engine to generate a frame-by-frame image sequence.
[0033] Step 4.2 Video editing stage: Use professional video editing software such as Adobe Premiere Pro to combine discrete video frames into a complete video file, and apply editing techniques such as special effects and transitions to ultimately create a harmonious and unified dance video work. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 What is reflected is the flow chart of the present invention.
[0035] Figure 2 It reflects the multimodal model diagram. DETAILED DESCRIPTION
[0036] The present invention provides a method for generating a dance video based on music and lyrics. The present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0037] Stage 1 Unimodal feature encoding training: The original audio is trained using the music feature encoder (Jukebox), the lyrics are trained using the text feature encoder (CLIP), and the pose sequence is trained using the pose feature encoder (Pose VAE).
[0038] Stage 2: Cross-modal feature fusion training: The extracted feature vectors are fused into one feature vector, the importance of different modal features is dynamically calculated through the attention model, and weighted fusion is performed based on this.
[0039] Stage 3: Diffusion model-driven action generation: Based on the fusion features, the U-Net diffusion model is used to gradually denoise and generate diverse dance action sequences, and then fine-tuned with the reference network (Reference Net) to improve the consistency of the action style and the guidance image.
[0040] In summary, the present invention discloses a method for generating a dance video based on music and lyrics, which can successfully generate a dance video that combines lyrics and music features.
[0041] The protection scope of the present invention is not limited to the above-mentioned embodiments. Any equivalent replacement or improved technical solutions implemented according to the principles of the present invention should be included in the protection scope of the present invention. The contents not covered by the present invention can be achieved by using existing technologies.
Claims
1. This invention provides a method for generating dance videos based on music and lyrics. The technical solutions adopted are as follows: Step 1: Multi-module feature encoding: Focusing on the multimodal data processing of songs and dances, data collection, video framing, audio track separation, and lyrics analysis are carried out in sequence to build a multimodal dataset. In addition, posture, audio, and lyrics features are extracted to complete multimodal feature encoding. Step 2: Multi-module feature fusion: The extracted multi-modal feature vectors are fused into one feature vector, the importance of different modal features is dynamically calculated through the attention mechanism, and weighted fusion is performed based on this. Step 3: Feature Processing and Action Generation: The fused features are dimensionally resized, temporally aligned, and renormalized to ensure the proper functioning of the motion decoder and the quality of the generated action sequences. The U-Net model input data includes the pose sequence output by the motion decoder, the guidance image processed by the Reference Net, and Gaussian noise. Through the synergistic effects of self-attention, cross-attention, and temporal attention mechanisms, a continuous action sequence is generated. Step 4: Dance Movement Output: Rendering is performed to convert each pose in the U-Net action sequence into a visual sequence of frames. The discrete video frames are synthesized into a coherent video file. Using video editing techniques, special effects, transitions, and other elements are added to produce a harmonious dance video.
2. The method for generating a dance video based on music and lyrics according to claim 1, wherein the specific steps of step 1 are as follows: Step 1.1 Data Collection: Use web crawling technology to collect dance video footage from multiple sources, including popular video sharing platforms (such as YouTube and Bilibili) and open-source dance datasets. Convert the collected footage to .MP4 format to build an initial dataset containing dance movements, music, and lyrics. Step 1.2: Video framing: Use EasyMocap to obtain high-fidelity body estimation from the video at a rate of 60FPS and output the pose sequence in .csv format. Step 1.3: Separate the audio tracks: Use Adobe Audition to perform noise reduction on the video and then save the audio as a .WAV file. Use the open-source audio source separation tool Spleeter to extract the vocals from the mixed audio and save the tracks as separate .wav files. Step 1.4 Lyrics parsing: Perform speech recognition (ASR) on the vocal track based on the Whisper pre-trained model, generate a text transcription with timestamps, and output it in .SRT format. Step 1.5 Multimodal dataset construction: Synchronize pose (60FPS), audio (1ibrosa beat tracking) and lyrics (.SRT semantic segments) with a global time base, and combine it with dynamic time warping to achieve cross-modal temporal alignment. Step 1.6 Pose Feature Extraction: The pose sequence is encoded into a low-dimensional latent space using the POSE VAE (Variational Autoencoder), obtaining distribution parameters that represent the pose features. Using the reparameterization technique, the latent variable z is sampled from this Gaussian distribution defined by μ and σ. This latent variable z is the direct pose feature. The latent variable z can be expressed as: Z=μ+ε⊙σ in, μ represents the mean, ε is a random variable sampled from a standard normal distribution, ⊙ represents element-wise multiplication, and σ represents the standard deviation. In this way, the specific posture feature z is obtained from the distribution parameters. Step 1.7 Audio Feature Extraction: Use the Librosa library to extract basic audio features, such as Mel-frequency cepstral coefficients (MFCC), constant Q-value chromaticity maps, timing maps, and the intensity of the original audio. Jukebox uses multi-scale VQ-VAE (Vector Quantized Variational Autoencoder) to further encode the extracted audio features into discrete vector representations, while capturing the dynamic changes of the audio through an autoregressive generation mechanism. The generated audio vector is input into the Transformer module, and the temporal dependency between audio features is captured through the self-attention mechanism, thereby optimizing the audio representation. The implementation of calculating the attention weight in the Transformer module is shown in the following formula: Among them, Q, K, and V represent query, key, and value vectors respectively. k is the dimension of the key vector. Step 1.8 Lyrics (text) feature extraction: Use the jieba library to perform semantic segmentation on the lyrics text. Each word or character obtained by the segmentation forms a corresponding embedding vector. The embedded vector sequence is sent to the pre-trained CLIP model. After the multi-layer Transformer module operation, the hidden representation h is obtained, and the feature vector expressing the key content such as the lyrics semantics and emotions is extracted.
3. The method for generating a dance video based on music and lyrics according to claim 1, wherein the specific steps of step 2 are as follows: Step 2.1 uses the additive attention mechanism in the feature-level fusion method to fuse the posture, audio, and lyrics features. The implementation of the weighted fusion feature vector is shown in the following formula: in, f fusion Represents the final fusion feature vector, a i is the attention weight of the i-th modal feature, satisfying f i Usually represents the i-th eigenvector.
4. The method for generating a dance video based on music and lyrics according to claim 1, wherein the specific steps of step 3 are as follows: Step 3.1 Feature processing: In this step, we need to perform dimension adjustment, temporal alignment, and renormalization on the fused features to ensure that the action decoder can work properly. Step 3.2 Action Generation: Use Reference Net to process the guidance image and extract key features, which are used to maintain consistency in style or details in subsequent generation tasks. Gaussian noise is used as the input condition of U-Net, introducing controllable randomness to enhance the diversity of the generated results. The extracted key features, Gaussian noise, and pose sequence are input into the U-Net model. The self-attention mechanism, cross-attention mechanism, and temporal attention mechanism work together to achieve dance movement generation. The Gaussian noise probability density function formula is as follows: G(x, y) = f(x, y) + n(x, y) in, f(x, y) represents the original signal, and n(x, y) is a function that obeys Gaussian distribution.
5. The method for generating a dance video based on music and lyrics according to claim 1, wherein the specific steps of step 4 are as follows: Step 4.1 Rendering: Use the 3D modeling software Blender to build a 3D model of the dancer and perform skeletal rigging and skinning. Map the dance movement sequence data output by U-Net to the skeletal system of the 3D model and export the 3D model as an FBX file. Import this FBX file into the Unity game engine, configure the camera perspective, lighting environment, and rendering parameters, and use Unity's rendering engine to generate a frame-by-frame image sequence. Step 4.2 Video editing stage: Use professional video editing software such as Adobe Premiere Pro to combine discrete video frames into a complete video file, and apply editing techniques such as special effects and transitions to ultimately create a harmonious and unified dance video work.