3D dance generation method and device, equipment and storage medium

CN116309891BActive Publication Date: 2026-09-29SHENZHEN YUANXIANG INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310092026.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-29
Publication Date
2026-09-29
Estimated Expiration
2043-01-29

AI Technical Summary

Technical Problem

[0005]本申请提供了一种3D舞蹈生成方法、装置、设备和存储介质,旨在解决目前自动生成的舞蹈与音乐不匹配,舞蹈自动生成质量差的技术问题

Benefits of technology

[0040]本申请实施例提供了一种3D舞蹈生成方法、装置、设备和存储介质,所述方法通过对给定音乐进行提取操作,获得所述给定音乐的能量、音乐特征和梅尔谱,根据所述给定音乐的能量、音乐特征和梅尔谱,生成第一目标向量,从而加强每个流派与其对应音乐之间的相关性;将舞蹈片段的骨骼关节位置输入矢量量化自动编码器,生成初始上半身姿态编码和初始下半身姿态编码,根据所述初始上半身姿态编码、所述初始下半身姿态编码和所述第一目标向量,生成第二目标向量,将所述第二目标向量输入生成式的预训练模型,预测目标上半身姿态编码和目标下半身姿态编码,提高舞蹈生成框架的整体流派一致性;将所述目标上半身姿态编码和所述目标下半身姿态编码输入矢量量化自动解码器,获得未来3D舞蹈,基于流派一致性,提高了舞蹈生成质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116309891B_ABST
    Figure CN116309891B_ABST
Patent Text Reader

Abstract

The application provides a 3D dance generation method, device and equipment and a storage medium. The method extracts a given music to obtain the energy, music features and mel spectrum of the given music, generates a first target vector according to the energy, music features and mel spectrum, and strengthens the correlation between each genre and the corresponding music. The bone joint positions of a dance segment are input into a vector quantization autoencoder to generate an initial upper body posture code and an initial lower body posture code. According to the initial upper body posture code, the initial lower body posture code and the first target vector, a second target vector is generated. The second target vector is input into a pre-trained generative model to predict a target upper body posture code and a target lower body posture code, thereby improving the overall genre consistency of the dance generation framework. The target upper body posture code and the target lower body posture code are input into a vector quantization auto-decoder to obtain a future 3D dance, thereby improving the dance generation quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a 3D dance generation method, apparatus, device and computer-readable storage medium. Background Technology

[0002] Existing dance generation methods mainly fall into two categories: concatenation retrieval-based and autoregressive generation, both aiming to align dance movements with the music's beat. Autoregressive generation directly feeds the music and dance into a single network to regressively generate a sequence of dance movements. However, the accumulation of errors during the autoregression process can lead to the generation of non-dance movements. Concatenation retrieval divides the dance into fixed-length units and then reassembles these units according to the music's rhythm. While this method ensures the quality of the generated dance, it is not suitable for music with different beats. Another approach, using VQ-VAE and GPT, addresses the problems of the autoregressive and concatenation retrieval methods, achieving high-quality, beat-accurate dance generation. However, none of these methods consider dance genres during the generation process.

[0003] Without considering dance genres, generated dances might incorporate movements from multiple genres within a single piece of music, leading to a mismatch with the music—for example, generating ballet moves within a hip-hop track. Recent work has begun to consider genre information. One approach uses a one-dimensional vector to represent the genre and embeds it into a transformer to generate dances with genre consistency. Another uses a mapping network to transform latent codes into multi-genre style codes while simultaneously generating dance movements. While both methods can generate dances with specific genres, they require manual genre determination during the generation process. Furthermore, the specific correlation between dance type and its background music, allowing choreographers to determine the dance type based on the music, is not considered in the above methods.

[0004] Therefore, how to generate dances that adapt to the background music and improve the quality of automatic dance generation are technical problems that urgently need to be solved. Summary of the Invention

[0005] This application provides a 3D dance generation method, apparatus, device, and storage medium, aiming to solve the technical problems of mismatch between automatically generated dances and music, and poor quality of automatically generated dances.

[0006] In a first aspect, embodiments of this application provide a 3D dance generation method, the method comprising:

[0007] Extraction operations are performed on a given piece of music to obtain its energy, musical features, and Mel spectrum.

[0008] Based on the energy, musical features, and Mel spectrum of the given music, a first target vector is generated;

[0009] The skeletal joint positions of the dance segment are input into a vector quantization autoencoder to generate initial upper body posture codes and initial lower body posture codes;

[0010] A second target vector is generated based on the initial upper body posture code, the initial lower body posture code, and the first target vector;

[0011] The second target vector is input into the pre-trained generative model to predict the target's upper body pose encoding and lower body pose encoding;

[0012] The target's upper body posture code and lower body posture code are input into a vector quantization automatic decoder to obtain the future 3D dance.

[0013] Preferably, the step of generating the first target vector based on the energy, musical features, and Mel spectrum of the given music includes:

[0014] The energy and musical features of the given music are embedded into the first learnable vector and the second learnable vector, respectively, to obtain the corresponding first embedded learnable vector and the second embedded learnable vector.

[0015] The Mel score of the given music is fed into the genre token network to generate a genre embedding vector.

[0016] The first embedded learnable vector and the second embedded learnable vector are connected in the time dimension, and the connected vector is added to the genre embedding vector to form the first target vector.

[0017] Preferably, the step of generating a second target vector based on the initial upper body posture code, the initial lower body posture code, and the first target vector includes:

[0018] The skeletal joint positions of the dance segment are input into a vector quantization autoencoder to generate initial upper body posture codes and initial lower body posture codes;

[0019] The initial upper body posture code and the initial lower body posture code are embedded into the third learnable vector and the fourth learnable vector, respectively, to obtain the corresponding third embedded learnable vector and the fourth embedded learnable vector;

[0020] The first target vector, the third embedded learnable vector, and the fourth embedded learnable vector are connected in the time dimension. A positional embedding is added to the connected vector to obtain the second target vector.

[0021] Preferably, the step of inputting the second target vector into a pre-trained generative model to predict the target's upper body pose encoding and lower body pose encoding includes:

[0022] The second target vector is input into the generative pre-trained model, which outputs the probability of upper body pose encoding and the probability of lower body pose encoding.

[0023] Based on the probability of the upper body posture code and the probability of the lower body posture code, predict the target's upper body posture code and the target's lower body posture code.

[0024] Preferably, the step of feeding the Mel spectrum of the given music into the genre token network to generate a genre embedding vector includes:

[0025] The Mel spectrum of the given music is fed into the genre token network, and after passing through the reference encoder, genre token layer and genre embedding in the genre token network, a genre embedding vector is generated.

[0026] Preferably, before the step of feeding the Mel spectrum of the given music into the genre token network to generate the genre embedding vector, the method further includes:

[0027] Collect dance background music data and tag the dance genres corresponding to the dance background music data;

[0028] Based on the dance background music data and the corresponding dance genre tags, a genre token network is trained, and the genre token network is frozen when the training reaches a preset number of iterations.

[0029] Preferably, the step of inputting the target upper body pose code and the target lower body pose code into a vector quantization automatic decoder to obtain the future 3D dance includes:

[0030] The target's upper body pose encoding and lower body pose encoding are input into a vector quantization automatic decoder, and the loss is calculated according to the dance generation framework loss function to obtain the future 3D dance.

[0031] Secondly, embodiments of this application provide a 3D dance generation device, the device comprising:

[0032] The extraction module is used to extract data from a given piece of music to obtain its energy, musical features, and Mel spectrum.

[0033] The generation module is used to generate a first target vector based on the energy, musical features, and Mel spectrum of the given music;

[0034] The generation module is also used to input the skeletal joint positions of the dance segment into a vector quantization autoencoder to generate initial upper body posture codes and initial lower body posture codes.

[0035] The generation module is further configured to generate a second target vector based on the initial upper body posture code, the initial lower body posture code, and the first target vector;

[0036] The prediction module is used to input the second target vector into the pre-trained model of the generative formula to predict the target's upper body pose encoding and lower body pose encoding;

[0037] The decoding module is used to input the target's upper body posture encoding and lower body posture encoding into a vector quantization automatic decoder to obtain the future 3D dance.

[0038] Thirdly, embodiments of this application provide a 3D dance generation device, which includes a processor, a memory, and a 3D dance generation program stored in the memory and executable by the processor, wherein when the 3D dance generation program is executed by the processor, it implements the steps of the 3D dance generation method as described above.

[0039] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the 3D dance generation method described above.

[0040] This application provides a 3D dance generation method, apparatus, device, and storage medium. The method extracts energy, musical features, and Mel spectrum from given music. Based on the energy, musical features, and Mel spectrum, a first target vector is generated to enhance the correlation between each genre and its corresponding music. The skeletal joint positions of the dance segment are input into a vector quantization autoencoder to generate initial upper body posture codes and initial lower body posture codes. Based on the initial upper body posture codes, the initial lower body posture codes, and the first target vector, a second target vector is generated. The second target vector is input into a generative pre-trained model to predict target upper body posture codes and target lower body posture codes, improving the overall genre consistency of the dance generation framework. The target upper body posture codes and target lower body posture codes are input into a vector quantization autodecoder to obtain the future 3D dance. Based on genre consistency, the quality of dance generation is improved.

[0041] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit the disclosure of the embodiments of this application. Attached Figure Description

[0042] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This is a schematic diagram of the hardware structure of the 3D dance generation device provided in the embodiments of this application;

[0044] Figure 2 This is a flowchart illustrating the first embodiment of the 3D dance generation method provided in this application.

[0045] Figure 3 This is a schematic diagram of the dance generation framework structure in an embodiment of the 3D dance generation method provided in this application.

[0046] Figure 4 A schematic diagram of the genre token network structure in an embodiment of the 3D dance generation method provided in this application.

[0047] Figure 5 This is a T-SNE visualization of genre embedding in the 3D dance generation method embodiment provided in this application;

[0048] Figure 6 This is a visualization of genre consistency in the 3D dance generation method embodiments provided in this application.

[0049] Figure 7 This is a schematic diagram of the functional modules of the 3D dance generation device provided in the embodiments of this application. Detailed Implementation

[0050] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0051] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0052] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0053] Reference Figure 1 , Figure 1 This is a schematic diagram of the hardware structure of the 3D dance generation device involved in an embodiment of the present invention. In this embodiment, the 3D dance generation device may include a processor 1001 (e.g., CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to realize communication between these components; the user interface 1003 may include a display screen or an input unit such as a keyboard; the network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface); the memory 1005 may be a high-speed RAM memory or a stable non-volatile memory, such as a disk storage device, and the memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.

[0054] Those skilled in the art will understand that Figure 1 The hardware structure shown does not constitute a limitation on the 3D dance generation device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0055] Reference Figure 1 , Figure 1 The memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, and a 3D dance generation program.

[0056] exist Figure 1 In this embodiment, the network communication module is mainly used to connect to the server and communicate with the server for data; while the processor 1001 can call the 3D dance generation program stored in the memory 1005 and execute the 3D dance generation method provided in this embodiment of the invention.

[0057] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the 3D dance generation method of the present invention.

[0058] like Figure 2 As shown, the 3D dance generation method includes steps S100 to S600.

[0059] S100. Perform an extraction operation on the given music to obtain the energy, musical characteristics, and Mel spectrum of the given music.

[0060] It is understood that the execution entity in this embodiment is the 3D dance generation device, which is typically an electronic device such as a personal computer or server; this embodiment does not limit this. Since there is a correlation between the speed of dance movements and the energy of the music, energy features are further considered to improve the motion quality of the generated dance. First, the energy and music features of the given music are extracted, and then the energy and music features of the given music are embedded into a first learnable vector Z. e The second learnable vector Z m Simultaneously, the Mel spectrum of the given music segment is extracted, and then the extracted Mel spectrum is fed into a genre token network to generate a genre embedding vector Z. g .

[0061] S200. Generate a first target vector based on the energy, musical characteristics, and Mel spectrum of the given music.

[0062] It should be noted that, as Figure 3 Using a given piece of music as input, the energy and musical features of the given music are first extracted, and then the energy and musical features are embedded into a first learnable vector Z. e The second learnable vector Z m In the process, the corresponding first embedded learnable vector Z is obtained. e′ The second embedded learnable vector Z m′ Simultaneously, the Mel spectrum of the given music fragment is extracted and fed into a Genre Token Network to generate a genre embedding vector Z. g Then connect Z in the time dimension. e′ and Z m′ and with Z g The two vectors are added together to form the first target vector m.

[0063] In this embodiment, step S100 specifically includes:

[0064] The energy and musical features of the given music are embedded into a first learnable vector and a second learnable vector, respectively, to obtain the corresponding first embedded learnable vector and second embedded learnable vector; the Mel spectrum of the given music is sent to a genre token network to generate a genre embedding vector; the first embedded learnable vector and the second embedded learnable vector are connected in the time dimension, and the connected vector is added to the genre embedding vector to form a first target vector.

[0065] Furthermore, in this embodiment, the step of feeding the Mel spectrum of the given music into the genre token network to generate a genre embedding vector includes:

[0066] The Mel spectrum of the given music is fed into the genre token network, and after passing through the reference encoder, genre token layer and genre embedding in the genre token network, a genre embedding vector is generated.

[0067] It should be noted that the Genre Token Network infers genre information by learning the connection between music and genre. Given a music clip as input, the Genre Token Network can infer the genre and use it as a condition for dance generation, ensuring that each generated dance pose conforms to the constraints of that genre's dance movements.

[0068] This embodiment proposes supervised training of genres to achieve a correlation between genres and music. For example... Figure 4 As shown, the architecture of the Genre Token Network consists of three modules: Reference Encoder, Genre Token Layer, and Genre Embedding.

[0069] The reference encoder is used to compress the audio signal into a vector of a set length. In this embodiment, the Mel-spectrogram of the given music segment is fed to the reference encoder and compressed into a learnable reference embedding.

[0070] The Genre Token Layer comprises a set of Genre Token Embeddings and an Attention Module, which uses a reference embedding as the query vector. The Attention Module learns a similarity metric between the reference embedding and each token in a set of randomly initialized embeddings. This set of embeddings, also known as Genre Tokens, is shared across all music clips. The output of the Genre Token Layer is the probability that the given music belongs to each genre. To improve the robustness of the Genre Token Network, a soft embedding method is used to represent genres; that is, tokens are weighted by probability to form embeddings.

[0071] To enhance the correlation between music and genre, the number of tokens is set to match the number of genres. Simultaneously, the genre label is transformed into a one-dimensional embedding and introduced into the genre token layer as the target for token weights. Therefore, the genre token network is optimized through supervised training, and the cross-entropy loss between genre labels and genre token weights is as follows:

[0072]

[0073] Where g t and represents the genre tag vector and genre token weight vector of the t-th time segment in the given music, respectively; T represents the total number of segments in the given music; and CE represents the cross-entropy loss.

[0074] S300: Input the skeletal joint positions of the dance segment into the vector quantization autoencoder to generate the initial upper body posture code and the initial lower body posture code.

[0075] It should be understood that, for dance, the skeletal joint positions of the dance segment are first input into a Vector Quantized-Variational AutoEncoder (VQ-VAE encoder) to generate the initial upper body posture code and the initial lower body posture code corresponding to the dance segment.

[0076] S400. Generate a second target vector based on the initial upper body posture code, the initial lower body posture code, and the first target vector.

[0077] In a specific implementation, the initial upper body posture encoding and the initial lower body posture encoding are then embedded into the third learnable vector and the fourth learnable vector, respectively, to obtain the corresponding third embedded learnable vector u and the fourth embedded learnable vector l. After obtaining the first target vector m, the third embedded learnable vector u, and the fourth embedded learnable vector l, they are concatenated in the time dimension, and a position embedding is added to generate the second target vector. In this embodiment, step S400 includes: inputting the skeletal joint positions of the dance segment into a vector quantization autoencoder to generate the initial upper body posture encoding and the initial lower body posture encoding; embedding the initial upper body posture encoding and the initial lower body posture encoding into the third learnable vector and the fourth learnable vector, respectively, to obtain the corresponding third embedded learnable vector and the fourth embedded learnable vector; concatenating the first target vector, the third embedded learnable vector, and the fourth embedded learnable vector in the time dimension, and adding a position embedding to the concatenated vector to obtain the second target vector.

[0078] S500: Input the second target vector into the pre-trained model of the generative formula to predict the target's upper body posture code and the target's lower body posture code.

[0079] It should be understood that the second target vector is input into the generative pre-training (GPT). Finally, the output of the GPT is the probability of upper and lower body pose encodings, and the upper and lower body pose encodings are predicted and input into the VQ-VAE decoder to obtain the future dance. A teacher-forcing method is used on the genre token network to improve the overall genre consistency of the dance generation framework. In this embodiment, step S500 includes: inputting the second target vector into the generative pre-training model, outputting the probability of upper body pose encoding and the probability of lower body pose encoding; predicting the target upper body pose encoding and the target lower body pose encoding based on the probability of the upper body pose encoding and the probability of the lower body pose encoding.

[0080] S600: Input the target's upper body posture code and the target's lower body posture code into a vector quantization automatic decoder to obtain the future 3D dance.

[0081] Understandably, the target's upper body pose encoding and lower body pose encoding are input into a Vector Quantization Autodecoder (VQ-VAE decoder) to obtain the future 3D dance. A teacher-forcing approach is used on the genre token network to improve the overall genre consistency of the dance generation framework.

[0082] Furthermore, in this embodiment, the step of inputting the target upper body pose code and the target lower body pose code into a vector quantization automatic decoder to obtain the future 3D dance includes:

[0083] The target's upper body pose encoding and lower body pose encoding are input into a vector quantization automatic decoder, and the loss is calculated based on the dance generation framework loss function to obtain the future 3D dance.

[0084] It should be noted that GPT is optimized through supervised training, and the cross-entropy loss between the predicted action probability 'a' and the ground truth pose code 'p' is as follows:

[0085]

[0086] Where T′ represents the total number of segments in the dance segment, CE represents the cross-entropy loss, t is the t-th time segment of the dance segment, u is the third embedded learnable vector, and l is the fourth embedded learnable vector.

[0087] Based on this, the loss function of the dance generation framework can be calculated as follows:

[0088]

[0089] In this embodiment, by extracting the energy, musical features, and Mel spectrum of a given piece of music, a first target vector is generated based on the energy, musical features, and Mel spectrum, thereby strengthening the correlation between each genre and its corresponding music. The skeletal joint positions of the dance segment are input into a vector quantization autoencoder to generate initial upper body posture codes and initial lower body posture codes. Based on the initial upper body posture codes, the initial lower body posture codes, and the first target vector, a second target vector is generated. The second target vector is input into a pre-trained generative model to predict target upper body posture codes and target lower body posture codes, improving the overall genre consistency of the dance generation framework. The target upper body posture codes and target lower body posture codes are input into a vector quantization autodecoder to obtain the future 3D dance. Based on genre consistency, the quality of dance generation is improved.

[0090] Continue to refer to Figure 2 Based on the first embodiment of the above method, a second embodiment of the 3D dance generation method of the present invention is proposed.

[0091] In this embodiment, before the step of feeding the Mel spectrum of the given music into the genre token network to generate the genre embedding vector, the method further includes:

[0092] Collect dance background music data and tag the dance genres corresponding to the dance background music data;

[0093] Based on the dance background music data and the corresponding dance genre tags, a genre token network is trained, and the genre token network is frozen when the training reaches a preset number of iterations.

[0094] It should be understood that this embodiment proposes a pre-training and fine-tuning strategy for a genre token network to improve its generalization ability. Due to the insufficient music data in existing dance datasets, genre token networks struggle to accurately infer genres from music. Therefore, to strengthen the correlation between each genre and its corresponding music, this embodiment pre-collects a large amount of dance background music labeled with genre tags to pre-train the genre token network. Subsequently, within the dance generation framework, we use a dataset aligned with the dance music to fine-tune the genre token network so that it can more effectively infer music genres, further enhancing genre consistency.

[0095] Although the Genre Token Network can infer genres, it still suffers from poor generalization ability, making it difficult to establish a correlation between genres and music, and also making it difficult to ensure genre consistency in generated dances. This is because the cumulative length of all ten genres is less than one hour, resulting in insufficient music training data in the dance dataset. Therefore, this embodiment enhances the generalization ability of the Genre Token Network and improves the genre consistency of generated dances by pre-training it. First, a large amount of dance background music data is collected from the Internet, labeled with corresponding dance genre tags, and then this data is used to train the Genre Token Network. Afterwards, the parameters of the pre-trained Genre Token Network are loaded for fine-tuning, while the GPT in the dance generation framework is trained. To prevent the Genre Token Network from overfitting during fine-tuning, this embodiment freezes the Genre Token Network after a certain number of iterations.

[0096] It should be noted that the dance generation framework and pre-training strategy proposed in this embodiment show significant improvements in evaluation metrics and visualization effects. The quality of the future 3D dance generated in this embodiment is better than that obtained by methods in the prior art. For each method, this embodiment generates 20 dance segments on the AIST++ test set and cuts the generated dances into 20-second lengths.

[0097] The evaluation method is as follows:

[0098] (1) Subjective evaluation

[0099] Subjective evaluation was conducted to further assess the visual performance of the dances generated by the proposed dance generation framework. The test was conducted by 24 participants. Participants were asked to rate the consistency of dance quality and genre, scoring the dances on a scale of 1-5 in 1-point increments. As shown in Table 1, the last two columns report the mean opinion score (MOS) for dance quality and genre consistency. The proposed model outperforms all baseline models, demonstrating that the genre token network can establish a correlation between music and genre. Furthermore, conditioned on the inferred genre, the proposed dance generation framework can generate higher quality and more genre-consistent dances.

[0100]

[0101] Table 1: Evaluation Results of Different Schemes

[0102] (2) Objective evaluation

[0103] Objective evaluation primarily assesses the quality and diversity of the bio-dance, as well as its alignment with musical beats. Specifically, for quality, this embodiment calculates the Fréchet distance (FID) for both dynamics and geometry; the lower the FID, the closer the generated dance is to the ground truth. Similarly, this embodiment calculates the dynamic and geometric diversity (DIV) for the generated dance movements; the higher the DIV, the more diverse the generated dance movements. For beat alignment, this embodiment calculates the beat alignment score (BAS) between the musical beat and the movement beat; the higher the BAS, the more rehearsed the generated dance.

[0104] As shown in Table 1, the framework proposed in this scheme is superior to other schemes in all aspects. This indicates that by taking into account genre and energy, this scheme can generate higher quality and more diverse dances, and improve the alignment between dance movements and musical beats.

[0105] (3) Ablation test

[0106] As shown in the last four rows of Table 1, when there is no Without teacher-forcing, the generated dances will differ significantly from the ground truth and exhibit less diversity. The expressiveness of the dance is reduced if the correlation between the speed of the dance movements and the energy of the music is not considered. While the framework can generate dances similar to the ground truth without pre-training and fine-tuning strategies, its genre consistency is limited due to the poor generalization ability of the genre token network.

[0107] (4) Visualization of genre embedding

[0108] like Figure 5 As shown, this embodiment uses the AIST++ test set to verify the genre token network in the dance generation framework. This demonstrates that different genre embeddings can be well separated from each other, proving that the genre token network can correctly infer genres from music.

[0109] (5) Visualization of school of thought consistency

[0110] To further evaluate the genre consistency of the generated dances, this embodiment visualizes the dance results generated by the proposed framework and existing techniques. We randomly select a 20-second segment from the "LO" (locking) genre dances generated by each framework and sample the results at a frequency of 1 FPS. Figure 6 As shown, for a given music clip, the dance performances generated by existing technologies are displayed in the second column, exhibiting various dance types with low matching accuracy to the music melody. However, the framework proposed in this embodiment can infer the genre and generate dances that match the music melody and are consistent with the "LO" genre, such as... Figure 6As shown in the first column.

[0111] In this embodiment, by collecting dance background music data and labeling the dance genre tags corresponding to the dance background music data, a genre token network is trained based on the dance background music data and the corresponding dance genre tags, and the genre token network is frozen when the training reaches a preset number of iterations, so as to enhance the generalization ability of the genre token network and improve the genre consistency of the generated dance.

[0112] In addition, as Figure 2 The specific implementation of the method shown in this application provides a 3D dance generation device.

[0113] Please see Figure 7 , Figure 7 This is a functional module diagram of a 3D dance generation device according to an embodiment of this application. The device includes:

[0114] Extraction module 10 is used to perform extraction operations on a given piece of music to obtain the energy, musical features, and Mel spectrum of the given music;

[0115] The generation module 20 is used to generate a first target vector based on the energy, musical features, and Mel spectrum of the given music;

[0116] The generation module 20 is also used to input the skeletal joint positions of the dance segment into a vector quantization autoencoder to generate an initial upper body posture code and an initial lower body posture code.

[0117] The generation module 20 is further configured to generate a second target vector based on the initial upper body posture code, the initial lower body posture code, and the first target vector;

[0118] Prediction module 30 is used to input the second target vector into the pre-trained model of the generative formula to predict the target's upper body pose code and the target's lower body pose code;

[0119] The decoding module 40 is used to input the target's upper body posture encoding and the target's lower body posture encoding into a vector quantization automatic decoder to obtain the future 3D dance.

[0120] Other embodiments or specific implementations of the 3D dance generation device of the present invention can be referred to the above-described method embodiments, and will not be repeated here.

[0121] In addition, embodiments of the present invention also provide a computer-readable storage medium.

[0122] The present invention provides a computer-readable storage medium storing a 3D dance generation program, wherein when the 3D dance generation program is executed by a processor, it implements the steps of the 3D dance generation method as described above.

[0123] The method implemented when the 3D dance generation program is executed can be referred to in various embodiments of the 3D dance generation method of the present invention, and will not be repeated here.

[0124] It should be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application.

[0125] It should also be understood that the term “and / or” as used in this application and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0126] It should be noted that the descriptions involving "first," "second," etc., in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.

[0127] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating 3D dance, characterized in that, The method includes: Extraction operations are performed on a given piece of music to obtain its energy, musical features, and Mel spectrum. Based on the energy, musical features, and Mel spectrum of the given music, a first target vector is generated; The skeletal joint positions of the dance segment are input into a vector quantization autoencoder to generate initial upper body posture codes and initial lower body posture codes; A second target vector is generated based on the initial upper body posture code, the initial lower body posture code, and the first target vector; The second target vector is input into the pre-trained generative model to predict the target's upper body pose encoding and lower body pose encoding; The target's upper body posture code and lower body posture code are input into a vector quantization automatic decoder to obtain the future 3D dance. The step of generating the first target vector based on the energy, musical features, and Mel spectrum of the given music includes: The energy and musical features of the given music are embedded into the first learnable vector and the second learnable vector, respectively, to obtain the corresponding first embedded learnable vector and the second embedded learnable vector. The Mel score of the given music is fed into the genre token network to generate a genre embedding vector. The first embedded learnable vector and the second embedded learnable vector are connected in the time dimension, and the connected vector is added to the genre embedding vector to form the first target vector; The step of feeding the Mel spectrum of the given music into the genre token network to generate a genre embedding vector includes: The Mel spectrum of the given music is fed into the genre token network, and after passing through the reference encoder, genre token layer and genre embedding in the genre token network, a genre embedding vector is generated.

2. The 3D dance generation method according to claim 1, characterized in that, The step of generating a second target vector based on the initial upper body posture code, the initial lower body posture code, and the first target vector includes: The skeletal joint positions of the dance segment are input into a vector quantization autoencoder to generate initial upper body posture codes and initial lower body posture codes; The initial upper body posture code and the initial lower body posture code are embedded into the third learnable vector and the fourth learnable vector, respectively, to obtain the corresponding third embedded learnable vector and the fourth embedded learnable vector; The first target vector, the third embedded learnable vector, and the fourth embedded learnable vector are connected in the time dimension. A positional embedding is added to the connected vector to obtain the second target vector.

3. The 3D dance generation method according to claim 1, characterized in that, The step of inputting the second target vector into the pre-trained generative model to predict the target's upper body pose encoding and lower body pose encoding includes: The second target vector is input into the generative pre-trained model, which outputs the probability of upper body pose encoding and the probability of lower body pose encoding. Based on the probability of the upper body posture code and the probability of the lower body posture code, predict the target's upper body posture code and the target's lower body posture code.

4. The 3D dance generation method according to claim 2, characterized in that, Before the step of feeding the Mel spectrum of the given music into the genre token network to generate the genre embedding vector, the method further includes: Collect dance background music data and tag the dance genres corresponding to the dance background music data; Based on the dance background music data and the corresponding dance genre tags, a genre token network is trained, and the genre token network is frozen when the training reaches a preset number of iterations.

5. The 3D dance generation method according to any one of claims 1-4, characterized in that, The step of inputting the target's upper body pose encoding and lower body pose encoding into a vector quantization automatic decoder to obtain the future 3D dance includes: The target's upper body pose encoding and lower body pose encoding are input into a vector quantization automatic decoder, and the loss is calculated according to the dance generation framework loss function to obtain the future 3D dance.

6. A 3D dance generation device, characterized in that, The 3D dance generation device includes: The extraction module is used to extract data from a given piece of music to obtain its energy, musical features, and Mel spectrum. The generation module is used to generate a first target vector based on the energy, musical features, and Mel spectrum of the given music; The generation module is also used to input the skeletal joint positions of the dance segment into a vector quantization autoencoder to generate initial upper body posture codes and initial lower body posture codes. The generation module is further configured to generate a second target vector based on the initial upper body posture code, the initial lower body posture code, and the first target vector; The prediction module is used to input the second target vector into the pre-trained model of the generative formula to predict the target's upper body pose encoding and lower body pose encoding; The decoding module is used to input the target's upper body posture encoding and lower body posture encoding into a vector quantization automatic decoder to obtain the future 3D dance. The generation module is further configured to: The energy and musical features of the given music are embedded into the first learnable vector and the second learnable vector, respectively, to obtain the corresponding first embedded learnable vector and the second embedded learnable vector. The Mel score of the given music is fed into the genre token network to generate a genre embedding vector. The first embedded learnable vector and the second embedded learnable vector are connected in the time dimension, and the connected vector is added to the genre embedding vector to form the first target vector; The generation module is further configured to: The Mel spectrum of the given music is fed into the genre token network, and after passing through the reference encoder, genre token layer and genre embedding in the genre token network, a genre embedding vector is generated.

7. A 3D dance generation device, characterized in that, The 3D dance generation device includes a processor, a memory, and a 3D dance generation program stored in the memory and executable by the processor, wherein when the 3D dance generation program is executed by the processor, it implements the steps of the 3D dance generation method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the 3D dance generation method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Dance motion prediction model training method, dance synthesis method, equipment and product

    CN115375806A

  • Model training method and apparatus, action posture generation method and apparatus, and device and medium

    WO2022227208A1