Training method, generation method and device for controllable music-driven three-dimensional dance movement generation model
By designing an initial generation model including the backbone network and the control network, and using music and two-dimensional key point data to control the generation of three-dimensional dance moves, the problem of difficult generation results in the prior art is solved, and controllable three-dimensional dance moves are achieved.
Patent Information
- Application Number
- CN202411486370.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-10-23
AI Technical Summary
The existing music-driven dance generation scheme lacks the ability to control the generation results and it is difficult to control the specific movement trajectory of the generated human joints.
Using the training set containing music-3D dance movement data and two-dimensional key point data, the initial generation model is designed including the backbone network and the control network. The backbone network takes music and three-dimensional dance movement data as inputs, and the control network takes music, three-dimensional dance movement data and two-dimensional key point data as inputs. Through the stitching of motion features and intermediate features extracted by the backbone network, the generated predicted action data is controlled.
The controllability of the three-dimensional dance movement generation results is achieved, and three-dimensional dance movement data that matches the music and conforms to the user's designed movement trajectory.
Smart Images

Figure CN119516051B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of music-driven dance movement choreography methods, and particularly to a controllable music-driven three-dimensional dance movement generation model training method, generation method and device. Background Art
[0002] The music-driven dance movement generation model aims to automatically generate high-quality and diverse dance movements given a piece of music. It can ultimately be presented in the form of a two-dimensional video, or three-dimensional motion data can be generated and rendered through tools such as post-modeling software to obtain the desired visual effects.
[0003] Such artificial intelligence generation models not only have broad application prospects in film and television and video games, but can also be used as tools to assist choreographers or animators to improve their production efficiency. For example, a large number of character models and character movements may be required in game production, and task movements in special scenarios may be required in the shooting of film and television works.
[0004] In view of the current technology, there are generally two solutions for computer character movement production. One is to hire professionals and motion capture technology teams to collect motion data using motion capture technology, and then through post-computer processing, finally apply it to the actual scenario; the other is to hire professional animators to frame-by-frame produce the movements of the character model using relevant computer software. However, no matter which solution is adopted, it requires a large amount of manpower, material resources and financial resources. Therefore, if artificial intelligence technology can be used to automatically generate, it will save a lot of time and cost and improve production efficiency.
[0005] The diffusion model is a generative neural network model and is widely used in dance movement generation due to its powerful generation ability. Among them, the classic work of the diffusion model in three-dimensional dance movement generation comes from the EDGE model proposed by Tseng et al. It designed a model for predicting and generating dance movement sequences, which can predict and generate a three-dimensional dance movement data of the same duration according to the input music features. However, this music-driven dance movement generation scheme cannot control the generated three-dimensional dance movements. Since random samples are randomly sampled from noise during the generation process of the diffusion model, the final generated results are random, and it is difficult to control the generation of the same three-dimensional dance movements. Thus, the existing music-driven dance movement generation scheme lacks the ability to control the generation results and is difficult to control the specific movement trajectories of human joints. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a controllable music-driven three-dimensional dance movement generation model training method, generation method and device to eliminate or improve one or more defects existing in the prior art.
[0007] On the one hand, the present invention provides a method for training a controllable music-driven three-dimensional dance movement generation model, the method comprising the following steps:
[0008] Construct a first training sample set, the first training sample set comprising a plurality of samples, each sample comprising music and its corresponding three-dimensional dance movement data; construct a second training sample set, the second training sample set comprising a plurality of samples, each sample comprising the music, the three-dimensional dance movement data corresponding to the music, and two-dimensional key point data of the human joints corresponding to the three-dimensional dance movement data;
[0009] Construct an initial generation model, the initial generation model comprising a backbone network and a control network; the backbone network comprises a plurality of linear layers and Transformer decoder layers, the control network comprises a plurality of linear layers, Transformer decoder layers and zero linear layers; the backbone network takes the music and its corresponding three-dimensional dance movement data as input, extracts the three-dimensional dance movement data features, and outputs predicted movement data; the control network takes the music, the three-dimensional dance movement data and its corresponding two-dimensional key point data as input, extracts the features of the two-dimensional key point data, and outputs movement features; input the movement features into the backbone network and splice them with the intermediate features extracted by the backbone network to control the predicted movement data;
[0010] Train the initial generation model using a diffusion model, and add noise to the three-dimensional dance movement data in the first training sample set and the second training sample set to obtain noisy movement data; train the backbone network using the first training sample set, construct the loss between the predicted movement data and the three-dimensional dance movement data, and update the parameters of the backbone network; fix the structure and parameters of the backbone network that have been trained, and apply the structure and parameters to the control network, train the initial generation model using the second training sample set, construct the loss between the predicted movement data controlled by the control network and the three-dimensional dance movement data, and update the parameters of the control network, and finally train a three-dimensional dance movement generation model composed of the backbone network and the control network.
[0011] In some embodiments of the present invention, inputting the movement features into the backbone network and splicing them with the intermediate features extracted by the backbone network comprises:
[0012] Output the movement features extracted by the control network to the backbone network via the zero linear layer;
[0013] Add the output of the zero linear layer to the output of the corresponding Transformer decoder layer in the backbone network; wherein, the zero linear layer refers to a linear layer whose parameters are initialized to zero.
[0014] In some embodiments of the present invention, the backbone network is a U-shaped structure, and the number of Transformer decoder layers on the left and right arms of the U-shaped structure is the same. There is one Transformer decoder layer in the bottom layer. The method further includes:
[0015] Input the music into the Transformer decoder layers on the left and right arms of the U-shaped structure, and input the noise action data into the Transformer decoder layer of the left arm via the linear layer of the left arm;
[0016] In each layer, the output of the Transformer decoder layer of the left arm is concatenated with the output of the Transformer decoder layer of the right arm in the next layer according to the feature dimension, and after adjusting the feature dimension size via the linear layer of the current layer, it is input into the Transformer decoder layer of the right arm of the current layer; finally, the predicted action data is output by the linear layer of the right arm.
[0017] In some embodiments of the present invention, constructing the loss between the predicted action data and the three-dimensional dance action data includes:
[0018] Extract the predicted attribute information and the real attribute information from the predicted action data and the three-dimensional dance action data respectively; the predicted attribute information includes the predicted coordinates of human joints, predicted speed, and predicted acceleration; the real attribute information includes the real coordinates of human joints, real speed, and real acceleration;
[0019] Calculate the root mean square error of each attribute, and perform a weighted sum of the root mean square errors of each attribute to obtain the final loss function value.
[0020] In some embodiments of the present invention, constructing the loss between the predicted action data generated under the control of the control network and the three-dimensional dance action data includes:
[0021] Extract the predicted attribute information and the real attribute information from the predicted action data generated under the control of the control network and the three-dimensional dance action data respectively; the predicted attribute information includes the predicted coordinates of human joints, predicted speed, and predicted acceleration; the real attribute information includes the real coordinates of human joints, real speed, and real acceleration;
[0022] Calculate the root mean square error of each attribute; and perform a weighted sum of the root mean square errors of each attribute to obtain the final loss function value.
[0023] On the other hand, the present invention provides a controllable music-driven three-dimensional dance movement generation method, and the method includes:
[0024] Obtain the music for generating the three-dimensional dance movement and the two-dimensional key point data for controlling the generated movement;
[0025] Input the music and the two-dimensional key point data into a three-dimensional dance movement generation model trained by any one of the controllable music-driven three-dimensional dance movement generation model training methods mentioned above to control the generation of three-dimensional dance movement data.
[0026] In some embodiments of the present invention, inputting the music and the two-dimensional key point data into a three-dimensional dance movement generation model trained by any one of the controllable music-driven three-dimensional dance movement generation model training methods mentioned above to control the generation of three-dimensional dance movement data includes:
[0027] Input the music, the two-dimensional key point data, and the random noise movement data generated based on the Gaussian distribution into the three-dimensional dance movement generation model, and obtain the first predicted movement data after denoising the random noise movement data;
[0028] Use a diffusion model to add noise to the first predicted movement data to obtain the first noise movement data;
[0029] Input the music, the two-dimensional key point data, and the first noise movement data into the three-dimensional dance movement generation model, and obtain the second predicted movement data after denoising the first noise movement data;
[0030] Repeat the above steps a preset number of times to obtain the final predicted movement data.
[0031] In some embodiments of the present invention, the controllable music-driven three-dimensional dance movement generation method further includes:
[0032] Convert the three-dimensional dance movement data into a three-dimensional modeling data format for rendering and displaying in three-dimensional modeling software.
[0033] On the other hand, the present invention also provides a controllable music-driven three-dimensional dance movement generation device, including a processor, a memory, and a computer program / instructions stored on the memory. The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the device implements the steps of any one of the methods mentioned above.
[0034] On the other hand, the present invention also provides a computer-readable storage medium, on which computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of any one of the methods mentioned above are implemented.
[0035] The present invention provides a method, a generation method and a device for training a controllable music-driven three-dimensional dance motion generation model, including: constructing a training set containing music-three-dimensional dance motion data and two-dimensional key point data; constructing an initial generation model, including a backbone network and a control network; the backbone network takes music and three-dimensional dance motion data as inputs and outputs predicted motion data; the control network takes music, three-dimensional dance motion data and two-dimensional key point data as inputs and outputs motion features; splicing the motion features with the intermediate features extracted by the backbone network to control the predicted motion data; training the initial generation model in two stages using the training set, and finally obtaining a three-dimensional dance motion generation model. The three-dimensional dance motion generation model obtained by training in the present invention can be used to generate three-dimensional dance motion data matching the music and has the ability of controllable generation.
[0036] Further, the backbone network is designed as a "U" - shaped structure. Based on the connection method of the left and right arms of the "U" - shaped structure, not only can the sequence front - back time information be fully utilized, but also it is compatible with the design of the control network, thereby realizing the controllable generation effect.
[0037] Further, based on the connection method between the control network and the backbone network, and the method that the control network copies the neural network structure and parameters corresponding to the left arm in the backbone network and combines with a zero linear layer, the two - dimensional key point data can be converted into deviation values for adjusting the three - dimensional dance motion.
[0038] The additional advantages, objectives, and features of the present invention will be partially elaborated in the following description, and will become partially obvious to those of ordinary skill in the art after studying the following text, or can be learned according to the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the specification and the drawings.
[0039] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present invention are not limited to the above - specific descriptions, and the above - mentioned and other objectives that the present invention can achieve will be more clearly understood according to the following detailed description. Brief Description of the Drawings
[0040] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:
[0041] Figure 1 It is a schematic diagram of the steps of the method for training a controllable music - driven three - dimensional dance motion generation model in an embodiment of the present invention.
[0042] Figure 2 It is a schematic diagram of the structure of the initial generation model (three - dimensional dance motion generation model) in an embodiment of the present invention.
[0043] Figure 3 Schematic diagram of the connection mode between the control network and the backbone network in an embodiment of the present invention.
[0044] Figure 4 Schematic diagram of the backbone network structure in an embodiment of the present invention.
[0045] Figure 5 Schematic diagram of the principle of the backbone network training process in an embodiment of the present invention.
[0046] Figure 6 Schematic diagram of the steps of a controllable music-driven three-dimensional dance motion generation method in an embodiment of the present invention.
[0047] Figure 7 Schematic diagram of the principle of a controllable music-driven three-dimensional dance motion generation method in an embodiment of the present invention. Detailed implementation manners
[0048] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the implementation manners and the drawings. Herein, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0049] Herein, it also needs to be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, and other details less related to the present invention are omitted.
[0050] It should be emphasized that the term "including / containing" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0051] Herein, it also needs to be noted that if not otherwise specified, the term "connection" in this article can not only refer to direct connection, but also represent indirect connection with an intermediate.
[0052] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0053] To solve the problem that the existing music-driven dance generation scheme lacks the ability to control the generation result and it is difficult to control the specific movement trajectories of the generated human joints, the present invention provides a training method for a controllable music-driven three-dimensional dance motion generation model, as Figure 1 shown, the method includes the following steps S101 to S103:
[0054] Step S101: Construct a first training sample set, which contains multiple samples, and each sample contains music and its corresponding three-dimensional dance motion data; construct a second training sample set, which contains multiple samples, and each sample contains music, the three-dimensional dance motion data corresponding to the music, and the two-dimensional key point data of the human joints corresponding to the three-dimensional dance motion data.
[0055] Step S102: Construct an initial generation model. Among them, the initial generation model includes a backbone network and a control network. The backbone network includes multiple linear layers and Transformer decoder layers. The control network includes multiple linear layers, Transformer decoder layers, and zero linear layers. The backbone network takes music and its corresponding three-dimensional dance motion data as inputs, extracts the three-dimensional dance motion data features, and outputs predicted motion data; the control network takes music, three-dimensional dance motion data, and its corresponding two-dimensional key point data as inputs, extracts the features of the two-dimensional key point data, and outputs motion features; input the motion features into the backbone network and splice them with the intermediate features extracted by the backbone network to control the predicted motion data.
[0056] Step S103: Use a diffusion model to train the initial generation model, and add noise to the three-dimensional dance motion data in the first training sample set and the second training sample set to obtain noisy motion data. Use the first training sample set to train the backbone network, construct the loss between the predicted motion data and the three-dimensional dance motion data, and update the parameters of the backbone network; fix the structure and parameters of the backbone network that have been trained, and apply this structure and parameters to the control network. Use the second training sample set to train the initial generation model, construct the loss between the predicted motion data generated by the control network and the three-dimensional dance motion data, and update the parameters of the control network. Finally, train to obtain a three-dimensional dance motion generation model composed of the backbone network and the control network.
[0057] In step S101, first construct a first training sample set and a second training sample set for subsequent model training. The first training sample set is a large-scale sample, and each sample contains music and the three-dimensional dance motion data corresponding to the music, with the two corresponding one by one. The second training sample set is a small-scale sample, and each sample contains music, the three-dimensional dance motion data corresponding to the music, and the two-dimensional key point data of the human joints corresponding to the three-dimensional dance motion data, with the three corresponding one by one. Among them, the two-dimensional key point data represents human motion data, and the key points correspond to some joints of the human body, from which the motion trajectory of the human body can be described.
[0058] In some embodiments, a two-dimensional key point in the COCO format disclosed by Microsoft is used to represent human actions, and this format can also be compatible with existing human pose estimation methods. Among them, human pose estimation refers to a task of automatically identifying and extracting human action postures from data such as images and videos. Thus, two-dimensional key point data can be extracted by means of picture extraction, video extraction, manual editing, etc.
[0059] In the present invention, the denoising diffusion principle of the diffusion model is used to train the model in a supervised learning manner, so that the model has the ability to predict actual action data. Furthermore, in the denoising generation process of the diffusion model, the noise samples are gradually denoised to restore the predicted actual action data. Therefore, in step S101, noise is added to the three-dimensional dance action data by using the noise addition process of the diffusion model to obtain noise action data.
[0060] In the present invention, for the real music in the sample, a preset music feature extraction tool, model, etc. are used to extract the features of the real music, such as volume, duration, pitch, spectrum, etc.
[0061] In step S102, considering that in the prior art, three-dimensional dance action data matching the music is generated by inputting music data. Since random sampling of samples from noise is adopted in the generation process of the diffusion model, the final generated result has randomness, and it is very difficult to control the generation of the same three-dimensional dance action. On the other hand, in the actual application scenarios in the fields of film and television, games, etc., users often hope to control the generated three-dimensional dance actions to run according to the motion trajectories they design. However, the prior art is difficult to provide such a function to accept inputs for controlling the motion trajectory in the form of user texts, images, etc. to achieve the control that the generated result conforms to the user design. Therefore, the present invention designs a control network for the above technical problems.
[0062] Construct an initial generation model, as Figure 2 shown, the initial generation model includes a backbone network and a control network. The backbone network includes multiple linear layers and Transformer decoder layers, and the control network includes multiple linear layers, Transformer decoder layers and zero linear layers.
[0063] The backbone network takes music and its corresponding three-dimensional dance action data as inputs, extracts the three-dimensional dance action data features, and outputs predicted action data. Therefore, the backbone network learns the association between the two modalities of music and three-dimensional dance action data from a large-scale music-three-dimensional dance action dataset, so as to have rich prior knowledge for three-dimensional dance action generation.
[0064] The control network takes music, 3D dance movement data, and their corresponding 2D key point data as input, extracts features from the 2D key point data, and converts them into deviation values of the corresponding movement trajectories of the 3D dance movements, that is, the movement features described above. Then, the movement features are input into the backbone network and concatenated with the intermediate features extracted by the backbone network, so as to control the generated predicted action data to be similar to the movement trajectories of the 2D key points, achieving the effect of controllable predicted action data.
[0065] In Figure 2 , the dashed arrows and ellipses in the backbone network and the control network indicate that the repeated model structures are omitted. The black arrows indicate that the output of the upper layer structure is input into the lower layer structure, the gray arrows indicate that they are input into the corresponding structure as conditions, and the specific connection method of the red arrows can be as Figure 3 shown, indicating that the output of the zero linear layer of the control network is added to the output of the Transformer decoder layer in the corresponding backbone network. Among them, the zero linear layer refers to a linear layer whose parameters are initialized to zero, mainly used to convert the extracted features into deviation values of the corresponding movement trajectories of the 3D dance movements.
[0066] In some embodiments, in order for the backbone network to have stronger sequence modeling capabilities and be adapted to the design of the control network, as Figure 4 shown, its structure as a whole presents a "U" - shaped structure. The number of Transformer decoder layers on the left and right arms of this "U" - shaped structure is the same, and there is also one Transformer decoder layer at the bottom layer. In order to make full use of the information of the sequence before and after in time, as Figure 4 shown, the output of the left - hand decoder layer is concatenated and input into the corresponding decoder layer on the right. Specifically: the music is input into the Transformer decoder layers on the left and right arms of the "U" - shaped structure, and the noisy action data is input into the Transformer decoder layer on the left arm via the linear layer on the left arm; in each layer, the output of the Transformer decoder layer on the left arm is concatenated with the output of the Transformer decoder layer on the right arm of the next layer according to the feature dimension, and after adjusting the feature dimension size via the linear layer of the current layer, it is input into the Transformer decoder layer on the right arm of the current layer; finally, the predicted action data is output by the linear layer on the right arm. Figure 4 In, the black arrows indicate that the output of the upper layer structure is input into the lower layer structure, the gray arrows indicate that they are input into the corresponding structure as conditions, and the dashed arrows indicate that the repeated model structures are omitted.
[0067] In step S103, a diffusion model is used to train the initial generation model. Noise is added to the three-dimensional dance motion data of the first training sample set and the second training sample set to obtain noisy motion data. In the present invention, the training process is divided into two stages. In the first stage, only the backbone network is trained, and in the second stage, mainly the control network is trained.
[0068] In the first stage, the first training sample set, i.e., the large-scale music-three-dimensional dance motion dataset, is used to perform supervised training on the backbone network. As Figure 5 shown, a diffusion model is used to add noise to the three-dimensional dance motion data to obtain noisy motion data, which is input into the backbone network together with the music features. Finally, the predicted motion data is output, and a loss function between the predicted motion data and the real motion data (three-dimensional dance motion data) is constructed to update the parameters of the backbone network.
[0069] In some embodiments, the root mean square error loss between the predicted motion data and the three-dimensional dance motion data is constructed, including: extracting the predicted attribute information and the real attribute information from the predicted motion data and the three-dimensional dance motion data respectively. Among them, the predicted attribute information includes the predicted coordinates of human joints, the predicted speed, and the predicted acceleration; the real attribute information includes the real coordinates of human joints, the real speed, and the real acceleration. Calculate the root mean square error of each attribute, and perform weighted summation on the root mean square errors of each attribute to obtain the final loss function value.
[0070] In the second stage, the parameters of the trained backbone network are loaded and the parameters of this part of the structure are fixed so that they will not be updated during the training process of the second stage. After loading the parameters of the backbone network, the control network will copy the structure and parameters in the corresponding backbone network. The parameters of this part of the structure are not fixed and will be updated as the data is input during the training process of the second stage. In Figure 2 it, corresponding annotations are also made. Figure 2 The structure marked with a snowflake pattern in it indicates that its parameters are fixed and will not be updated during the training process. The structure without this mark will be updated during the training process. The "Transformer decoder layer (trainable copy)" in the control network means that this structure is the same as the corresponding "Transformer decoder layer" structure in the backbone network, and the parameters are also the same as the corresponding decoder layer. "Trainable" means that the parameters of this structure are not fixed and can be updated during the training process.
[0071] Use the second training sample set, i.e., the music-three-dimensional dance movement data and its corresponding two-dimensional key point data set, to train the entire initial generation model (mainly the control network). The diffusion model is used to add noise to the three-dimensional dance movement data to obtain noisy movement data, which is input into the initial generation model together with the music features and two-dimensional key point data. Finally, the predicted movement data is output, and a loss function between the predicted movement data controlled by the control network and the three-dimensional dance movement data is constructed to update the parameters of the control network.
[0072] In some embodiments, constructing the root mean square error loss between the predicted movement data controlled by the control network and the three-dimensional dance movement data includes: extracting predicted attribute information and real attribute information from the predicted movement data controlled by the control network and the three-dimensional dance movement data respectively. Among them, the predicted attribute information includes the predicted coordinates of human joints, predicted speed, and predicted acceleration; the real attribute information includes the real coordinates of human joints, real speed, and real acceleration. Calculate the root mean square error of each attribute, and perform weighted summation on the root mean square error of each attribute to obtain the final loss function value.
[0073] Corresponding to the above controllable music-driven three-dimensional dance movement generation model training method, the present invention also provides a controllable music-driven three-dimensional dance movement generation method, as Figure 6 shown, this method includes the following steps S201-S202:
[0074] Step S201: Obtain the music of the three-dimensional dance movement to be generated and the two-dimensional key point data for controlling the generated movement.
[0075] Step S202: As Figure 7 shown, input the music and two-dimensional key point data into the three-dimensional dance movement generation model trained based on the above-mentioned controllable music-driven three-dimensional dance movement generation model training method to control the generation of three-dimensional dance movement data.
[0076] In some embodiments, the quality of the generated three-dimensional dance movement data is improved by using the method of denoising multiple times, including:
[0077] Input the music, two-dimensional key point data, and randomly generated noisy movement data based on the Gaussian distribution into the three-dimensional dance movement generation model, and obtain the first predicted movement data after denoising the randomly generated noisy movement data.
[0078] Use the diffusion model to add noise to the first predicted movement data to obtain the first noisy movement data.
[0079] Input the music, two-dimensional key point data, and the first noisy movement data into the three-dimensional dance movement generation model, and obtain the second predicted movement data after denoising the first noisy movement data.
[0080] Repeat the above steps a preset number of times to obtain the final predicted action data, that is, the finally generated three-dimensional dance action data.
[0081] In some embodiments, the generated three-dimensional dance action data is converted into a three-dimensional modeling data format, such as FBX, OBJ, GLTF, etc., for rendering and displaying in three-dimensional modeling software.
[0082] Corresponding to the above method, the present invention also provides a controllable music-driven three-dimensional dance action generation device, which includes a computer device. The computer device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the method described above.
[0083] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the foregoing edge computing server deployment method. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the art.
[0084] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to execute in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to execute the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0085] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0086] In the present invention, features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0087] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A controllable music-driven three-dimensional dance movement generation model training method, characterized in that: The method comprises the following steps: Constructing a first training sample set, wherein the first training sample set includes a plurality of samples, each sample includes music and its corresponding three-dimensional dance movement data; constructing a second training sample set, wherein the second training sample set includes a plurality of samples, each sample includes the music, the three-dimensional dance movement data corresponding to the music, and the two-dimensional key point data of the human joints corresponding to the three-dimensional dance movement data; Constructing an initial generation model, the initial generation model includes a backbone network and a control network; the backbone network includes multiple linear layers and Transformer decoder layers, and the control network includes multiple linear layers, Transformer decoder layers and zero linear layers; the backbone network takes the music and its corresponding three-dimensional dance motion data as input, extracts the features of the three-dimensional dance motion data, and outputs predicted motion data; the control network takes the music, the three-dimensional dance motion data and its corresponding two-dimensional key point data as input, extracts the features of the two-dimensional key point data, and outputs motion features; the motion features are input into the backbone network, and spliced with the intermediate features extracted by the backbone network to control the predicted motion data; The diffusion model is used to train the initial generation model, and noise is added to the three-dimensional dance motion data of the first training sample set and the second training sample set to obtain noise motion data; the first training sample set is used to train the backbone network, the loss of the predicted motion data and the three-dimensional dance motion data is constructed, and the parameters of the backbone network are updated; the structure and parameters of the trained backbone network are fixed, and the structure and parameters are applied to the control network, the second training sample set is used to train the initial generation model, the loss of the predicted motion data and the three-dimensional dance motion data generated by the control network is constructed, the parameters of the control network are updated, and finally the three-dimensional dance motion generation model composed of the backbone network and the control network is trained.
2. The controllable music-driven three-dimensional dance movement generation model training method according to claim 1, characterized in that: Inputting the motion features into the backbone network and concatenating them with the intermediate features extracted by the backbone network, including: Outputting the motion features extracted by the control network to the backbone network via the zero linear layer; The output of the zero linear layer is added to the output of the corresponding Transformer decoder layer in the backbone network; wherein the zero linear layer represents a linear layer whose parameters are initialized to zero.
3. The controllable music-driven three-dimensional dance movement generation model training method according to claim 1, characterized in that: The backbone network is a U-shaped structure, the number of Transformer decoder layers on the left and right arms of the U-shaped structure is the same, and the bottom layer has a Transformer decoder layer, and the method further includes: Input the music into the Transformer decoder layers on the left and right arms of the U-shaped structure, and input the noise motion data into the Transformer decoder layer of the left arm via the linear layer of the left arm; In each layer, the output of the left-arm Transformer decoder layer is concatenated with the output of the right-arm Transformer decoder layer of the next layer according to the feature dimension, and after the feature dimension size is adjusted by the linear layer of the current layer, it is input into the right-arm Transformer decoder layer of the current layer; finally, the right-arm linear layer outputs the predicted action data.
4. The controllable music-driven three-dimensional dance movement generation model training method according to claim 1, characterized in that: Constructing the loss of the predicted motion data and the three-dimensional dance motion data, comprising: Extracting predicted attribute information and real attribute information from the predicted motion data and the three-dimensional dance motion data respectively; the predicted attribute information includes predicted coordinates, predicted speed and predicted acceleration of human joints; the real attribute information includes real coordinates, real speed and real acceleration of human joints; Calculate the root mean square error of each attribute, and perform weighted summation of the root mean square error of each attribute to obtain the final loss function value.
5. The controllable music-driven three-dimensional dance movement generation model training method according to claim 1, characterized in that: Constructing the loss of the predicted motion data and the three-dimensional dance motion data generated by the control network, including: Extracting predicted attribute information and real attribute information from the predicted motion data generated by the control network and the three-dimensional dance motion data; the predicted attribute information includes predicted coordinates, predicted speed and predicted acceleration of human joints; the real attribute information includes real coordinates, real speed and real acceleration of human joints; Calculate the root mean square error of each attribute, and perform weighted summation of the root mean square error of each attribute to obtain the final loss function value.
6. A controllable music-driven three-dimensional dance movement generation method, characterized in that: The method comprises: Obtaining music for the three-dimensional dance movements to be generated and two-dimensional key point data for controlling the generated movements; The music and the two-dimensional key point data are input into a three-dimensional dance movement generation model trained by the controllable music-driven three-dimensional dance movement generation model training method according to any one of claims 1 to 5 to control the generation of three-dimensional dance movement data.
7. The controllable music-driven three-dimensional dance movement generation method according to claim 6, characterized in that: Inputting the music and the two-dimensional key point data into a three-dimensional dance movement generation model trained by the controllable music-driven three-dimensional dance movement generation model training method according to any one of claims 1 to 5, and controlling the generation of three-dimensional dance movement data, comprising: Inputting the music, the two-dimensional key point data and the random noise motion data generated based on Gaussian distribution into the three-dimensional dance motion generation model, and obtaining first predicted motion data after denoising the random noise motion data; adding noise to the first predicted motion data using a diffusion model to obtain first noise motion data; Inputting the music, the two-dimensional key point data and the first noise motion data into the three-dimensional dance motion generation model, and obtaining second predicted motion data after denoising the first noise motion data; Repeat the above steps for a preset number of times to obtain the final predicted action data.
8. The controllable music-driven three-dimensional dance movement generation method according to claim 6, characterized in that: The method further comprises: The three-dimensional dance movement data is converted into a three-dimensional modeling data format for rendering and display in three-dimensional modeling software.
9. A controllable music-driven three-dimensional dance movement generation device, comprising a processor, a memory, and a computer program / instruction stored in the memory, characterized in that: The processor is used to execute the computer program / instructions. When the computer program / instructions are executed, the device implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Music-driven dance generation method
CN110853670A
Systems, methods, kits, and apparatuses for using artificial intelligence for instructing smart machines in value chain networks
US20240144011A1
Cited By
Multi-modal driven human body action generation method based on large language model
CN120580356A