Method, device, equipment and storage medium for generating dance movements of virtual characters
By obtaining audio beat information and constructing a directed and authorized chart of dance movements, the problem of low efficiency in generating dance movements in traditional virtual characters is solved, and automated and intelligent dance movement generation is achieved.
Patent Information
- Application Number
- CN202210816238.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-07-12
AI Technical Summary
The traditional virtual character dance movement generation method requires manual participation, the process is complex and time-consuming, and the data needs to be re-recorded and converted each time a new dance is generated, which is inefficient.
By obtaining the beat information of the target audio, a directed and authorized image of the dance movement is constructed, and analyses and samples are swinged in the image to generate the dance movements of virtual characters.
Automatic and intelligent dance movement generation is realized, manual participation is reduced, and generation efficiency and intelligence are improved.
Smart Images

Figure CN115147519B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, computer equipment and storage medium for generating dance movements of a virtual character. Background Art
[0002] In the Metaverse and other fields, the character models of virtual characters have various actions, including continuous posture transitions when walking, turning movements, and waving movements. The character models should basically be able to perform the actions that people in the world have now, so that the virtual characters in the Metaverse can be more realistic and vivid, thus bringing a stronger sense of immersion. The above shows how important the various rich types of actions of virtual characters are in the Metaverse. Among them, the dance movements of virtual characters are also a very important application scenario in the application.
[0003] In the traditional scheme, the method of generating dance movements for virtual characters mainly involves first having a real person wear a motion capture device to perform a dance performance, using a professional motion capture device to record the dance movement data corresponding to the dance movements, and finally converting the recorded dance movement data into 3D movement data that can be used by the virtual character, and using the 3D movement data to drive the virtual character model to perform the corresponding dance movements.
[0004] However, the inventors have found that the above scheme requires manual participation, has a complicated process, and is time-consuming. In addition, each time a new dance needs to be generated, the dance movement data must be re-recorded and converted into 3D movement data that can be used by the virtual character. The level of intelligence is low. The above process results in a relatively low efficiency in constructing the dance movements of the virtual character model. Summary of the invention
[0005] The present application provides a method, device, computer equipment and storage medium for generating dance movements of a virtual character, so as to solve the technical problem of low efficiency in constructing dance movements of a virtual character model in traditional solutions.
[0006] A method for generating dance movements of a virtual character, comprising:
[0007] Acquire target audio, and extract beat information in the target audio, wherein the beat information includes the number of beat points and the beat landing time of all beat points;
[0008] Determine a walking starting point of a pre-constructed directed weighted graph of dance movements, wherein a node of the directed weighted graph of dance movements represents a quantitative feature of one of the dance movements, and dance movements at different nodes have different movement types;
[0009] Starting from the walking starting point, walking sampling is performed in the directed weighted graph of the dance action until the quantitative features of a plurality of target dance actions having the same number as the beat points are sampled, thereby obtaining the quantitative features of a plurality of sampled dance actions;
[0010] According to the beat landing time of each beat point and the sampling order of each sampled dance action in the multiple sampled dance actions, an action time is set for each sampled dance action, and the sampling order of each sampled dance action corresponds to the beat landing time of each beat point one by one;
[0011] The dance movements of the virtual character are generated according to the quantitative features and movement moments of all the sampled dance movements.
[0012] A device for generating dance movements of a virtual character, comprising:
[0013] An acquisition module, used to acquire target audio and extract beat information in the target audio, wherein the beat information includes the number of beat points and the beat landing time of all beat points;
[0014] A determination module is used to determine a walking starting point of a pre-constructed dance action directed weighted graph, wherein the nodes of the dance action directed weighted graph represent the quantitative characteristics of one target dance action, and the target dance actions at different nodes have different action types;
[0015] A walking module is used to start from the walking starting point and perform walking sampling in the directed weighted graph of the dance action until the quantitative features of a plurality of target dance actions having the same number as the beat points are sampled, thereby obtaining the quantitative features of a plurality of sampled dance actions;
[0016] A setting module, for setting a corresponding action moment for each sampled dance action according to the beat landing moment of each beat point and the sampling order of each dance action in the plurality of sampled dance actions, wherein the sampling order of the sampled dance actions corresponds to the beat landing moment of the beat point in a one-to-one manner;
[0017] The generation module is used to generate dance movements of the virtual character according to the quantitative features and movement moments of all the sampled dance movements.
[0018] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for generating dance movements of a virtual character are implemented.
[0019] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for generating dance movements of a virtual character are implemented.
[0020] In one of the schemes implemented by the above-mentioned method, device, computer equipment and storage medium for generating dance movements of virtual characters, the target audio is first obtained, and the beat information in the target audio is extracted, wherein the beat information includes the number of beat points contained in the target audio and the beat landing time of all beat points; the starting point of the walk of a pre-constructed directed weighted graph of dance movements is determined, the nodes of the directed weighted graph of dance movements are dance movements, and the movement types of dance movements at different nodes are different; starting from the starting point of the walk, walking sampling is performed in the directed weighted graph of dance movements until the quantitative features of multiple dance movements with the same number of beat points are sampled, and the quantitative features of multiple sampled dance movements are obtained; then, the dance movements of the virtual character are automatically generated according to the dance movements sampled from the directed weighted graph of dance movements. Compared with the traditional method of collecting and generating the dance movements of the virtual character through artificial wearable devices, the present application frees the cumbersome process of manual participation every time a new dance needs to be generated, and can automatically and intelligently generate dance movements according to the read target audio, thereby improving the efficiency of constructing the dance movements of the virtual character model, and the level of intelligence is also higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0022] Figure 1 It is a schematic diagram of an application environment of a method for generating dance movements of a virtual character in one embodiment of the present application;
[0023] Figure 2 It is a flow chart of a method for generating dance movements of a virtual character in one embodiment of the present application;
[0024] Figure 3 is a schematic diagram of a structure of a directed weighted graph of dance movements in one embodiment of the present application;
[0025] Figure 4 In one embodiment of the present application, a dance movement is constructed with a right to Figure 1 Flowchart diagram;
[0026] Figure 5 is a node diagram of key nodes of the human body used in an embodiment of the present application;
[0027] Figure 6 This is a schematic diagram of a process for training a target neural network model in one embodiment of the present application;
[0028] Figure 7 1. It is a schematic diagram of the network structure of the VQ-VAE network and the corresponding training process in one embodiment of the present application;
[0029] Figure 8 It is a structural schematic diagram of a device for generating dance movements of a virtual character in one embodiment of the present application;
[0030] Fig. 9 It is a structural diagram of a computer device in one embodiment of the present application. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0032] The method for generating dance movements of a virtual character provided in the embodiment of the present application can be applied in the following aspects: Figure 1 In an application environment, a terminal device or a server can obtain a target audio, and then extract the beat information in the target audio, wherein the beat information includes the number of beat points contained in the target audio and the beat landing time of all beat points, and according to the number of beat points, all required sampled dance movements are sampled from a pre-constructed directed weighted graph of dance movements, and finally the dance movements of the virtual character are generated according to the sampled dance movements. The terminal device may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server may be implemented with an independent server or a server cluster consisting of multiple servers.
[0033] It should be noted that the method for generating dance movements of virtual characters provided in the embodiments of the present application can be applied to the generation of dance movements of virtual characters in the Metaverse, including the generation of dance movements of virtual characters in virtual reality games, or the generation of dance movements of virtual characters in other non-game applications, and this application does not make any specific limitation.
[0034] In one embodiment, if Figure 2 As shown, a method for generating dance movements of a virtual character is provided, and the method is applied in Figure 1 The server in the example is used as an example to illustrate the process, including the following steps:
[0035] S10: Acquire target audio, and extract beat information in the target audio, wherein the beat information includes the number of beat points and the beat landing time of all beat points.
[0036] In the Metaverse and other fields, the character models of virtual characters have various actions, including continuous posture transitions when walking, turning movements, and waving movements. The character models should basically be able to perform the actions that people in the world have now, so that the virtual characters in the Metaverse can be more realistic and vivid, thus bringing a stronger sense of immersion. The above shows how important the various rich types of actions of virtual characters are in the Metaverse. Among them, the dance movements of virtual characters are also a very important application scenario in the application.
[0037] In an embodiment of the present application, a solution is provided for automatically generating a dance based on a directed weighted graph of dance movements. First, a target audio segment needs to be obtained. The target audio is the musical basis for subsequently generating dance movements corresponding to a virtual character. The target audio includes beat information. The beat information includes the number of beat points contained in the target audio and the beat landing times of all beat points. After the target audio is obtained, the beat information in the target audio is extracted. It can be understood that a music beat point is a unit for measuring audio rhythm. In the target audio, a series of beats with certain strengths and weaknesses are repeated at regular intervals. The beat landing time of the beat point records the time position of the beat point, that is, the duration of the beat point.
[0038] It should be noted that a beat tracking algorithm can be used to detect the music beat in the target audio, so as to obtain the beat landing time of all beat points of the target audio, and count the number of all beat points to obtain the number of beat points. The embodiment of the present application does not limit the beat tracking algorithm used. In some embodiments, the beat tracking algorithm can use a dynamic programming beat tracking algorithm or other beat tracking algorithms, which are not limited in the embodiment of the present application.
[0039] S20: Determine a walking starting point of a pre-constructed dance action directed weighted graph, wherein a node of the dance action directed weighted graph represents a quantitative feature of one target dance action, and target dance actions at different nodes have different action types.
[0040] In an embodiment of the present application, in addition to obtaining the target audio and extracting the beat information in the target audio, a directed weighted graph of dance movements is further obtained, and then a walking starting point of a pre-constructed directed weighted graph of dance movements is determined. The node of the directed weighted graph of dance movements represents one of the target dance movements, and the movement types of dance movements at different nodes are different.
[0041] It should be noted that a directed weighted graph is a storage structure, and the directed weighted graph of the present application is a directed weighted graph used to characterize the relationship between dance movements, which are collectively referred to as "directed weighted graphs of dance movements" in the embodiments of the present application. The directed weighted graph of dance movements includes multiple nodes, each node represents a different target dance movement, and each target dance movement corresponds to a quantitative feature, that is, different nodes correspond to different types of dance movements. The quantitative feature corresponding to the target dance movement characterizes the target dance movement, and the virtual character can be controlled to restore the corresponding target dance movement based on the quantitative feature. There is directionality between nodes, and the edge between nodes represents the transition from one target dance movement to another target dance movement. The value of the edge between nodes represents the probability of switching from one dance movement to another.
[0042] For example, Figure 3 As shown, Figure 3 is a schematic diagram of a structure of a directed weighted graph of dance movements in an embodiment of the present application. It should be noted that: Figure 3 This is just a few examples for the sake of illustration and is not intended to be limiting. Figure 3 In , five nodes are shown, including node 0, node 1, node 3, node 4, node 6 and node 9, wherein the specific numbers in the nodes represent node codes, each representing one of the target dance movements, and the edges between the nodes represent the switching probabilities between different target dance movements. For example, it can be seen that the switching probabilities from node 0 to nodes 4 and 6 are both 0.5, the switching probability from node 1 to nodes 0 and node 9 is 0.5, the switching probability from node 3 to node 0 is 1, the probability from node 9 to node 0 is also 1, and so on. It can be seen that the dance movement directed weighted graph is a pre-constructed directed weighted graph for characterizing the switching between different target dance movements.
[0043] In this step, the starting point of the walk of the pre-constructed dance movement directed weighted graph is determined, that is, a walk starting point is selected from all the nodes of the dance movement directed weighted graph as the basis for subsequent dance movement sampling.
[0044] It should also be noted that there is no time sequence restriction for step S10 and step S20.
[0045] S30: Starting from the walking starting point, walking sampling is performed in the directed weighted graph of the dance movements until quantitative features of a plurality of target dance movements having the same number as the beat points are sampled, thereby obtaining quantitative features of a plurality of sampled dance movements.
[0046] After determining the starting point of the walk in the directed weighted graph of the dance action, starting from the starting point of the walk, walk sampling is performed in the directed weighted graph of the dance action until the quantitative features of multiple target dance actions with the same number of beat points are sampled, and for ease of description and understanding, the target dance actions sampled from the directed weighted graph of the dance action are collectively referred to as sampled dance actions, and the corresponding quantitative features are referred to as quantitative features of the sampled dance actions. When walking sampling, the number of sampled dance actions sampled is the same as the number of beat points of the target audio. For example, if the number of beat points of the target audio is M, the quantitative features of M sampled dance actions are sampled from the directed weighted graph of the dance action.
[0047] As explained above, the dance action directed weighted graph is a directed weighted graph, that is, from a certain node, one can only move to another node according to the direction. For example, Figure 3 As shown, taking node 0 as an example, from node 0, one can only walk from node 0 to node 4 or node 6. Therefore, when sampling the directed weighted graph of dance movements, one can continuously sample according to the direction to obtain the required number of sampled dance movements.
[0048] S40: according to the beat landing time of each of the beat points and the sampling order of each of the multiple sampled dance actions, an action time is set for each of the sampled dance actions, and the sampling order of each of the sampled dance actions corresponds one-to-one to the beat landing time of each of the beat points.
[0049] S50: Generating dance movements of the virtual character according to the quantitative features and movement moments of all the sampled dance movements.
[0050] Starting from the walking starting point of the directed weighted graph of the dance action, walking sampling is performed in the directed weighted graph of the dance action until the quantitative features of multiple target dance actions with the same number of beat points are sampled, and after the quantitative features of multiple sampled dance actions are obtained, the corresponding action moment is set for each sampled dance action according to the beat landing moment of each beat point and the sampling order of each dance action in the multiple sampled dance actions. The sampling order of the sampled dance action corresponds to the beat landing moment of the beat point one by one. After setting the action moments corresponding to all the sampled dance actions, the dance actions of the virtual character can be generated according to the quantitative features and action moments of all the sampled dance actions, and even the subsequent dance actions of the virtual character can be controlled. It should be noted that since the action moments of all the sampled dance actions are known, it is known when to insert the required sampled dance actions and how many key dance actions to insert in the whole song. Through the quantitative features of the sampled dance actions, the virtual character can be controlled to show the corresponding dance actions, and finally the virtual character can be controlled to perform dancing behavior.
[0051] It can be understood that the number of sampled dance movements obtained from the directed weighted graph of dance movements corresponds to the number of beat points of the target audio, and the purpose is to set corresponding dance movements for each beat point of the target audio. Therefore, according to the beat landing moment of each beat point, the sampling order of each sampled dance movement is set to the corresponding action moment for each sampled dance movement. Exemplarily, assuming that the number of beat points is M, the number of sampled dance movements is also M, and the beat landing moments of the beat points are t0, t1, ..., tM-1 in sequence, and the sampling order of each dance movement in the multiple sampled dance movements is c0, c1, ..., cM-1, then the action moment of the sampled dance movement corresponding to c0 will be set to t0, and the action moment of the sampled dance movement corresponding to c1 will be set to t1, and so on, until all dance movements are set with corresponding action moments.
[0052] It can be seen from the above scheme that the embodiment of the present application provides a method for generating dance movements of a virtual character. Compared with the traditional method of collecting and generating dance movements of a virtual character through artificial wearable devices, it eliminates the cumbersome process of manual participation every time a new dance needs to be generated, and can automatically and intelligently generate dance movements according to the read target audio, thereby improving the efficiency of constructing the dance movements of the virtual character model and having a higher level of intelligence.
[0053] In the above-mentioned embodiment, it is mentioned that a directed weighted graph of dance movements will be pre-constructed. In one embodiment, a specific method for constructing a directed weighted graph of dance movements is provided, that is, Figure 4 As shown, before determining the node starting point of the pre-constructed dance action directed weighted graph, the method further includes the following steps S101-S105, wherein:
[0054] S101: Acquire a first dance video, where the first dance video includes a first dance picture and a first dance soundtrack.
[0055] In the process of constructing a directed weighted graph of dance movements, it is first necessary to obtain a first dance video, which is a 3D dance video. In a specific application scenario, the collected first dance video is a dance video shot from multiple angles. The first dance video contains dance pictures and music. In order to distinguish them from the dance videos of subsequent embodiments, the dance pictures and music contained in the first dance video are referred to as dance pictures and first music respectively. It is worth noting that multi-angle shooting is to ultimately generate the required 3D human body motion data.
[0056] S102: Analyze the first dance picture and the first dance soundtrack, extract dance movements at all beats of the first dance soundtrack, and obtain a target dance movement sequence.
[0057] like Figure 5 As shown, exemplarily, the dancer in the first dance video is divided into 25 key nodes, and the 25 key nodes of the dancer in the first dance video are analyzed to obtain action data of the 25 key nodes corresponding to each action. The above action data may specifically refer to the displacement vectors and rotation vectors corresponding to the 25 key nodes, that is, the displacement vector x and the rotation vector y of the key node, and the displacement vector x and the rotation vector y of the key node respectively represent the spatial position and the rotation angle of the key node. Exemplarily, the 25 key nodes include Node 0 (base of spine), Node 1 (mid-spine), Node 2 (neck), Node 3 (head), Node 4 (left shoulder), Node 5 (left elbow), Node 6 (left wrist), Node 7 (left hand), Node 8 (right shoulder), Node 9 (right elbow), Node 10 (right wrist), Node 11 (right hand), Node 12 (left hip), Node 13 (left knee), Node 14 (left ankle), Node 15 (left foot), Node 16 (right hip), Node 17 (right knee), Node 18 (right ankle), Node 19 (right foot), Node 20 (shoulder spine), Node 21 (left hand tip), Node 22 (left thumb), Node 23 (right hand tip) and Node 24 (right thumb).
[0058] It should be noted that in other applications, other key nodes or other numbers of key nodes may be used, and there is no specific limitation. For example, 18 key nodes may be used, etc. By selecting key nodes, data analysis work can be reduced while ensuring the generation of basic dance movements.
[0059] It should also be noted that, in some embodiments, an AI (artificial intelligence) image motion capture tool, or a developed AI motion capture model, can be used to analyze and process the first dance movement video to obtain the displacement vectors and rotation vectors of the key nodes of the dancer, without specific limitation.
[0060] In this step, the beat tracking algorithm can be used to detect the music beat in the target audio, so as to obtain the beat landing time of all the beat points of the first dance soundtrack, and then the corresponding image frame of the first dance picture is parsed according to the beat landing time of all the beat points of the first dance soundtrack, and the action data of the corresponding dancing figure in each image frame is obtained, so as to obtain the action data (i.e., position vector and rotation vector) corresponding to the dance action corresponding to all the beat points, and then obtain the target dance action sequence. In other words, the target dance action sequence contains action data corresponding to a series of dance actions.
[0061] It should be noted that the embodiments of the present application do not limit the beat tracking algorithm used. In some embodiments, the beat tracking algorithm may use a dynamic programming beat tracking algorithm or other beat tracking algorithms, which is not limited by the embodiments of the present application.
[0062] S103: Inputting the target dance action sequence into a pre-trained target neural network model, so that the target neural network model outputs the quantitative features and encoding of each target dance action in the dance action sequence.
[0063] After obtaining the target dance action sequence, the target dance action sequence can be used to input into a pre-trained target neural network model, so that the target neural network model can output the quantitative features and coding of each target dance action in the target dance action sequence. Among them, the target neural network model is based on a special neural network structure that is trained in advance using data, and the target neural network model is specifically configured to input a target dance action sequence, so that the target neural network model can input the quantitative features and coding of each dance action in the target dance action sequence. Through the target neural network model, on the one hand, the level of intelligence is improved, and the accuracy of the quantitative features of each target dance action output can also be guaranteed, providing the necessary basic conditions for the generation of subsequent dance actions.
[0064] S104: Counting the switching probabilities between different target dance movements in the target dance movement sequence.
[0065] It can be understood that the target dance action sequence generated above is obtained based on the data set. By analyzing the type and number of target dance action sequences, the switching probability of each different type of target dance action can be counted. The switching probability from a certain target dance action to another target dance action can be obtained by counting the number of times a certain target dance action is switched to another target dance action, and the proportion of all switching times. Accordingly, the switching probability between different target dance actions in the target dance action sequence can be obtained. For example, the target dance action sequence includes target dance action 1, target dance action 2, ..., target dance action S-1, and target dance action S. The switching probability of target dance action 1 and each dance action in target dance action 2-target dance action S is counted. Similarly, the switching probability of target dance action 2 and each target dance action in target dance action 1, target dance action 3-target dance action S is counted, and so on, until the switching probability between different target dance actions in the target dance action sequence is counted.
[0066] S105: constructing the dance movement directed weighted graph according to the switching probability between the different target dance movements and the encoding of each target dance movement, wherein the encoding of each node in the dance movement directed weighted graph corresponds to each target dance movement in the target dance movement sequence, and each target dance movement at each node in the dance movement directed weighted graph is represented by the quantitative feature corresponding to each target dance movement.
[0067] After counting the switching probabilities between different target dance movements in the target dance movement sequence, and after encoding each target dance movement, the dance movement directed weighted graph can be constructed. Specifically, the switching probabilities between target dance movements are used to determine the paths in the weighted directed graph, that is, the edges between nodes and the directions of the edges, and the final dance movement directed weighted graph can be constructed based on this, so that the encoding of each node in the dance movement directed weighted graph corresponds to each target dance movement in the target dance movement sequence, and each target dance movement of each node in the dance movement directed weighted graph is represented by the quantitative features corresponding to the each target dance movement. Exemplarily, the constructed dance movement directed weighted graph can be as follows Figure 3 As shown, no further detailed description is given here.
[0068] In this embodiment, a specific method for constructing a directed weighted graph of dance movements is proposed. The quantitative features and encoding of the dance movements are first obtained through deep learning, and the directed weighted graph of dance movements is confirmed based on the switching probability between different target dance movements. Since the directed weighted graph of dance movements is collected based on data of a real first dance video, the switching between real dance movements can be restored, thereby ensuring the continuity and naturalness of the dance movements of the subsequently generated virtual character.
[0069] It should be noted that the target neural network model is obtained by training with data in advance based on a special neural network structure, and in this application, the VQ-VAE network is used for training, and the trained VQ-VAE network model is used as the above target neural network model. Specifically, Figure 6 As shown, in one embodiment, the target neural network model is trained in the following manner:
[0070] S201: Acquire a second dance video, where the second dance video includes a second dance picture and a second dance music.
[0071] S202: Analyze the second dance soundtrack to extract the beat landing times of all beat points of the second dance soundtrack.
[0072] S203: extracting dance movements corresponding to all the beat points in the second dance soundtrack according to the beat landing moments of all the beat points in the second dance soundtrack, and obtaining a first training dance movement sequence.
[0073] The above S201-S203 process is similar to the process of analyzing the first dance video in the aforementioned embodiment, and will not be described in detail here. The difference is that the dance movement sequence obtained based on the first dance video in the aforementioned embodiment is to construct a directed weighted graph of dance movements, while the first training dance movement sequence extracted based on the second dance video in the embodiment of the present application is used to train the VQ-VAE network.
[0074] S204: Train the VQ-VAE network based on the first training dance action sequence until the loss value corresponding to the VQ-VAE network meets a preset condition, and use the VQ-VAE network whose loss value meets the preset condition as the target neural network model.
[0075] After obtaining the first training dance movement sequence, the VQ-VAE network can be trained based on the first training dance movement sequence until the loss value corresponding to the VQ-VAE network meets the preset conditions, and the VQ-VAE network whose loss value meets the preset conditions is used as the target neural network model.
[0076] It should be noted that, in some embodiments, the second dance video and the first dance video can be the same video, without specific limitation. Using the same data to train the target neural network and construct the directed weighted graph of dance movements can improve the overall processing efficiency and is more convenient.
[0077] Among them, Figure 7 As shown, Figure 7 The network structure diagram of the VQ-VAE network and the corresponding data processing process are shown in FIG. 1 , wherein the VQ-VAE network includes a dance action encoder E (Pose Encoder), a code book Z (Codebook Z) and a dance action decoder D (Pose Decoder D), wherein Pose represents a posture, that is, a dance action, and the VQ-VAE network is trained based on the first training dance action sequence until the loss value corresponding to the VQ-VAE network meets the preset conditions, and the detailed process of using the VQ-VAE network whose loss value meets the preset conditions as the target neural network model includes the following steps:
[0078] S2041: Input the first training dance action sequence into the encoder for encoding to obtain sequence encoding features.
[0079] S2042: Performing a quantization operation on the sequence encoding feature through the code book to obtain a sequence quantization feature.
[0080] S2043: Inputting the sequence quantization features into the decoder for decoding to obtain a second training dance action sequence.
[0081] For steps S2041-S2043, please continue to refer to Figure 7 As shown, Figure 7 P in the figure represents the first training dance action sequence. In the embodiment of the present application, the first training dance action sequence is mainly encoded and quantized to generate a code book Z of the dance action, wherein the code book Z is represented as N is the size of the codebook Z, and each Zi represents one of the dance moves. Specifically, the first training dance move sequence P is passed through the encoder E, and the encoder E encodes the dance move sequence P into a sequence encoding feature e, and then the sequence encoding feature e is quantized to obtain the sequence quantization feature e q Among them, the quantization rule is to select and e in the code book Z i The nearest feature, where e i Represents the encoding feature corresponding to one of the training dance movements i in the first training dance movement sequence.
[0082] It should be noted that, in some embodiments, the quantizing operation of the sequence coding features through the code book to obtain the sequence quantization features includes: selecting features that match the respective training dance movement coding features in the sequence coding features in the code book according to a preset calculation method to obtain the sequence quantization features;
[0083] The preset calculation method is: Among them, e i Represents the training dance action encoding features i, e q,i represents the training dance action quantization feature matched by the training dance action coding feature i, z represents the code sheet, z j Denotes the feature vector j in the codebook.
[0084] It can also be further seen from here that in the above formula, e q,i The q in represents the quantized operation, which indicates the quantized feature vector. The final sequence quantized feature e q Through the decoder D, it can be restored to the second training dance action sequence P * .
[0085] It can be understood that the encoder E can be implemented using a 1D temporal convolutional network, and the decoder D is configured accordingly, without specific limitation here.
[0086] S2044: Calculate the loss value using the following loss function.
[0087] S2045: The VQ-VAE network whose loss value after training meets the preset conditions is used as the VQ-VAE model.
[0088] When training the VQ-VAE network, the encoder D, decoder E, and codebook Z can be trained simultaneously. The loss function sampled during the training process is:
[0089] L VQ =L rec (P * ,P)+||sg[e]-e q ||+β||e-sg[e q ]||;
[0090] L VQ Represents the loss value, L rec (P * ,P) represents the reconstruction loss, ||sg[e]-e q || represents the code loss, β||e-sg[e q ]|| represents the commitment loss, the P represents the first training dance action sequence, the P * represents the second training dance action sequence, e represents the sequence encoding feature, sg[] indicates that the gradient is not updated during back propagation, and e q represents the quantitative characteristics of the sequence, and β represents a constant parameter.
[0091] It should be noted that β is a balance parameter and a constant. It can be set to 0.1 in the experiment. The specific training process is selected and is not limited here.
[0092] It can be understood that after the training is completed, each quantized feature of the code book Z in the VQ-VAE network represents one of the dance movements, and the quantized features and encoding of the corresponding dance movement can be output through the VQ-VAE network model later.
[0093] In combination with the above embodiment, in S103: that is, inputting the dance movement sequence into a pre-trained target neural network model, so that the target neural network model outputs the quantized features and encoding of each dance movement in the dance movement sequence, which becomes inputting the dance movement sequence into a pre-trained VQ-VAE network model, so that the VQ-VAE network outputs the quantized features and encoding of each dance movement in the dance movement sequence. The trained VQ-VAE network will use the code book Z to quantize the dance movement sequence. At this point, the key movements at each beat landing point in the first dance video are assigned a number M, M∈[0,N-1], and N is the size of the code book Z.
[0094] In this embodiment, the beat information of the second dance music in the second dance video is used to extract the first key dance action sequence and use it to train the VQ-VAE network. The training uses unlabeled data, which can greatly reduce the data annotation cost.
[0095] In one embodiment, in step S30, that is, starting from the walking starting point, walking sampling is performed in the directed weighted graph of the dance movement until the quantitative features of multiple target dance movements with the same number of beat points are sampled to obtain the quantitative features of multiple sampled dance movements, specifically including: starting from the walking starting point, random walking sampling is performed in the directed weighted graph of the dance movement, or walking sampling is performed according to a preset sampling rule until the quantitative features of multiple target dance movements with the same number of beat points are sampled to obtain the quantitative features of multiple sampled dance movements.
[0096] In this embodiment, the key dance movements are generated by random walking in a directed weighted graph of dance movements, which ensures the diversity of the generated dance movements, and because it is a directed weighted graph of dance movements constructed based on the present application, the switching of the dance movements finally generated will also be more natural. Of course, in addition to random walking, in some embodiments, walking sampling can also be performed according to preset sampling rules, which are not specifically limited. For example, the preset sampling rule can be performed in accordance with the switching probability. Specifically, if there are two directions, the one with a higher probability of walking first will be sampled along the corresponding edge, which is not specifically limited.
[0097] In one embodiment, step S20, i.e., determining the starting point of the walk of the pre-constructed dance action directed weighted graph, specifically includes the following steps:
[0098] S21: Determine the node starting point of the pre-constructed dance action directed weighted graph;
[0099] S22 uses the node starting point as the walking starting point of the directed weighted graph of the dance action.
[0100] In this embodiment, when generating the dance movements of the virtual character based on the target audio, the embodiment of the present application can first process the target audio to extract the beat information. A key dance movement should be generated at each beat point. Specifically, in the generated dance movement directed weighted graph, the movement starts from node 0 until enough dance movements are sampled. Since the dance movement directed weighted graph is constructed based on real dance movements, the movement starts from the node of the dance movement directed weighted graph, which can further ensure the naturalness of the generated dance movements.
[0101] It should be noted that in the embodiment of the present application, the process of the server generating the dance movements of the virtual character, the process of constructing the directed and weighted graph of the dance movements, and the process of training the target neural network model can be implemented by different device ends, for example, the same server, or different servers respectively, without specific limitation. When they are implemented by different servers respectively, the server for generating the dance movements of the virtual character only needs to construct the directed and weighted graph of the dance movements from other servers, and the server for constructing the directed and weighted graph of the dance movements only needs to train the target neural network model from other servers.
[0102] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0103] In one embodiment, a device for generating dance movements of a virtual character is provided, and the device for generating dance movements of a virtual character corresponds one-to-one to the method for generating dance movements of a virtual character in the above-mentioned embodiment. Figure 8 As shown, the dance movement generating device 10 of the virtual character includes an acquisition module 101, a determination module 102, a wandering module 103, a setting module 104 and a generating module 105. The detailed description of each functional module is as follows:
[0104] An acquisition module 101 is used to acquire a target audio and extract beat information in the target audio, wherein the beat information includes the number of beat points and the beat landing time of all beat points;
[0105] A determination module 102 is used to determine a starting point of a walk of a pre-constructed directed weighted graph of dance movements, wherein the nodes of the directed weighted graph of dance movements represent quantitative features of one target dance movement, and the target dance movements at different nodes have different movement types;
[0106] A walking module 103 is used to start from the walking starting point and perform walking sampling in the directed weighted graph of the dance movements until the quantitative features of a plurality of target dance movements having the same number as the beat points are sampled, thereby obtaining the quantitative features of a plurality of sampled dance movements;
[0107] A setting module 104 is used to set an action time for each sampled dance action according to the beat landing time of each beat point and the sampling order of each sampled dance action in the multiple sampled dance actions, and the sampling order of each sampled dance action corresponds to the beat landing time of each beat point one by one;
[0108] The generation module 105 is used to generate dance movements of the virtual character according to the quantitative features and movement moments of all the sampled dance movements.
[0109] In one embodiment, the dance movement generating device of the virtual character further includes a building module; wherein the building module is further used for:
[0110] Before determining the node starting point of the pre-constructed directed weighted graph of dance movements, obtaining a first dance video, wherein the first dance video includes a first dance picture and a first dance soundtrack;
[0111] Analyze the first dance picture and the first dance soundtrack, extract dance movements at all beats of the first dance soundtrack, and obtain a target dance movement sequence;
[0112] Inputting the target dance action sequence into a pre-trained target neural network model, so that the target neural network model outputs the quantitative features and encoding of each target dance action in the target dance action sequence;
[0113] Counting the switching probabilities between different target dance movements in the target dance movement sequence;
[0114] According to the switching probabilities between the different target dance movements and the encoding of each target dance movement, the dance movement directed weighted graph is constructed, the encoding of each node in the dance movement directed weighted graph corresponds to each target dance movement in the target dance movement sequence, and each target dance movement at each node in the dance movement directed weighted graph is represented by the quantitative features corresponding to each target dance movement.
[0115] In one embodiment, the dance movement generating device for a virtual character further includes a training module: the training module is used to:
[0116] Acquire a second dance video, wherein the second dance video includes a second dance picture and a second dance soundtrack;
[0117] Analyze the second dance soundtrack to extract the beat landing time of all beat points of the second dance soundtrack;
[0118] According to the beat landing time of all the beat points of the second dance soundtrack, dance movements corresponding to all the beat points are extracted from the second dance picture to obtain a first training dance movement sequence;
[0119] The VQ-VAE network is trained based on the first training dance action sequence until a loss value corresponding to the VQ-VAE network meets a preset condition, and the VQ-VAE network whose loss value meets the preset condition is used as the target neural network model.
[0120] In one embodiment, the VQ-VAE network includes an encoder, a codebook, and a decoder, and the training module is further used to:
[0121] Inputting the first training dance action sequence into the encoder for encoding to obtain sequence encoding features;
[0122] Performing a quantization operation on the sequence encoding feature through the code book to obtain a sequence quantization feature;
[0123] Inputting the sequence quantization features into the decoder for decoding to obtain a second training dance action sequence;
[0124] The loss value is calculated by the following loss function:
[0125] L VQ =L rec (P * ,P)+||sg[e]-e q ||+β||e-sg[e q ]||;
[0126] L VQ Represents the loss value, L rec (P * ,P) represents the reconstruction loss, ||sg[e]-e q || represents the code loss, β||e-sg[e q ]|| represents the commitment loss, the P represents the first training dance action sequence, the P * represents the second training dance action sequence, e represents the sequence encoding feature, sg[] indicates that the gradient is not updated during back propagation, and e q represents the quantitative characteristics of the sequence, and β represents a constant parameter;
[0127] The VQ-VAE network whose loss value after training meets the preset conditions is used as the VQ-VAE model.
[0128] In one embodiment, the training module is further used to:
[0129] Selecting features that match the respective training dance movement coding features in the sequence coding features in the code book according to a preset calculation method to obtain sequence quantization features;
[0130] The preset calculation method is: Among them, e i Represents the training dance action encoding features i, e q,i represents the training dance action quantization feature matched by the training dance action coding feature i, z represents the code sheet, z j Denotes the feature vector j in the codebook.
[0131] In one embodiment, the walking module 103 is specifically used to: start from the walking starting point, perform random walking sampling in the directed weighted graph of the dance movements, or perform walking sampling according to a preset sampling rule, until the quantitative features of multiple target dance movements with the same number of beat points are sampled, and the quantitative features of multiple sampled dance movements are obtained.
[0132] In one embodiment, the determination module 102 is specifically configured to determine a node starting point of a pre-constructed dance movement directed weighted graph; and use the node starting point as a walk starting point of the dance movement directed weighted graph.
[0133] For the specific definition of the virtual character dance movement generation device, please refer to the relevant definition of the virtual character dance movement generation method mentioned above, which will not be repeated here. Each module in the above-mentioned virtual character dance movement generation device can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0134] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig. 9 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The database of the computer device is used to store motion data or video data. The network interface of the computer device is used to communicate with an external capture device through a network connection. When the computer program is executed by the processor, a method for generating dance movements of a virtual character is implemented.
[0135] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:
[0136] Acquire target audio, and extract beat information in the target audio, wherein the beat information includes the number of beat points and the beat landing time of all beat points;
[0137] Determine a walking starting point of a pre-constructed dance action directed weighted graph, wherein a node of the dance action directed weighted graph represents a quantitative feature of one of the target dance actions, and the target dance actions at different nodes have different action types;
[0138] Starting from the walking starting point, walking sampling is performed in the directed weighted graph of the dance action until the quantitative features of a plurality of target dance actions having the same number as the beat points are sampled, thereby obtaining the quantitative features of a plurality of sampled dance actions;
[0139] According to the beat landing time of each beat point and the sampling order of each sampled dance action in the multiple sampled dance actions, an action time is set for each sampled dance action, and the sampling order of each sampled dance action corresponds to the beat landing time of each beat point one by one;
[0140] The dance movements of the virtual character are generated according to the quantitative features and movement moments of all the sampled dance movements.
[0141] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0142] Acquire target audio, and extract beat information in the target audio, wherein the beat information includes the number of beat points and the beat landing time of all beat points;
[0143] Determine a walking starting point of a pre-constructed dance action directed weighted graph, wherein a node of the dance action directed weighted graph represents a quantitative feature of one of the target dance actions, and the target dance actions at different nodes have different action types;
[0144] Starting from the walking starting point, walking sampling is performed in the directed weighted graph of the dance action until the quantitative features of a plurality of target dance actions having the same number as the beat points are sampled, thereby obtaining the quantitative features of a plurality of sampled dance actions;
[0145] According to the beat landing time of each beat point and the sampling order of each sampled dance action in the multiple sampled dance actions, an action time is set for each sampled dance action, and the sampling order of each sampled dance action corresponds to the beat landing time of each beat point one by one;
[0146] The dance movements of the virtual character are generated according to the quantitative features and movement moments of all the sampled dance movements.
[0147] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0148] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0149] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for generating dance movements of a virtual character, characterized in that: include: Acquire target audio, and extract beat information in the target audio, wherein the beat information includes the number of beat points and the beat landing time of all beat points; Determine a walking starting point of a pre-constructed dance action directed weighted graph, wherein a node of the dance action directed weighted graph represents a quantitative feature of one of the target dance actions, and the target dance actions at different nodes have different action types; Starting from the walking starting point, walking sampling is performed in the directed weighted graph of the dance action until the quantitative features of a plurality of target dance actions having the same number as the beat points are sampled, thereby obtaining the quantitative features of a plurality of sampled dance actions; According to the beat landing time of each beat point and the sampling order of each sampled dance action in the multiple sampled dance actions, an action time is set for each sampled dance action, and the sampling order of each sampled dance action corresponds to the beat landing time of each beat point one by one; Generating dance movements of a virtual character according to the quantitative features and movement moments of all the sampled dance movements; Before determining the node starting point of the pre-constructed dance action directed weighted graph, the method further includes: Acquire a first dance video, where the first dance video includes a first dance picture and a first dance soundtrack; Analyze the first dance picture and the first dance soundtrack, extract dance movements at all beats of the first dance soundtrack, and obtain a target dance movement sequence; Inputting the target dance action sequence into a pre-trained target neural network model, so that the target neural network model outputs the quantitative features and encoding of each target dance action in the target dance action sequence; Counting the switching probabilities between different target dance movements in the target dance movement sequence; According to the switching probabilities between the different target dance movements and the encoding of each target dance movement, the dance movement directed weighted graph is constructed, the encoding of each node in the dance movement directed weighted graph corresponds to each target dance movement in the target dance movement sequence, and each target dance movement at each node in the dance movement directed weighted graph is represented by the quantitative features corresponding to each target dance movement.
2. The method for generating dance movements of a virtual character according to claim 1, wherein: The target neural network model is trained in the following way: Acquire a second dance video, wherein the second dance video includes a second dance picture and a second dance soundtrack; Analyze the second dance soundtrack to extract the beat landing time of all beat points of the second dance soundtrack; According to the beat landing time of all the beat points of the second dance soundtrack, dance movements corresponding to all the beat points are extracted from the second dance picture to obtain a first training dance movement sequence; The VQ-VAE network is trained based on the first training dance action sequence until a loss value corresponding to the VQ-VAE network meets a preset condition, and the VQ-VAE network whose loss value meets the preset condition is used as the target neural network model.
3. The method for generating dance movements of a virtual character according to claim 2, wherein: The VQ-VAE network includes an encoder, a code book, and a decoder. The VQ-VAE network is trained based on the first training dance action sequence until a loss value corresponding to the VQ-VAE network meets a preset condition, and the VQ-VAE network whose loss value meets the preset condition is used as the target neural network model, including: Inputting the first training dance action sequence into the encoder for encoding to obtain sequence encoding features; Performing a quantization operation on the sequence encoding feature through the code book to obtain a sequence quantization feature; Inputting the sequence quantization features into the decoder for decoding to obtain a second training dance action sequence; The loss value is calculated by the following loss function: ; represents the loss value, represents the reconstruction loss, represents the code thinness loss, Indicates commitment loss, represents the first training dance action sequence, the represents the second training dance action sequence, represents the sequence encoding feature, Indicates that the gradient is not updated during back propagation. represents the quantitative characteristics of the sequence, Represents a constant parameter; The VQ-VAE network whose loss value meets the preset conditions after training is used as the VQ-VAE model.
4. The method for generating dance movements of a virtual character according to claim 3, wherein: The step of performing a quantization operation on the sequence encoding feature through the code book to obtain a sequence quantization feature includes: Selecting features that match the respective training dance movement coding features in the sequence coding features in the code book according to a preset calculation method to obtain sequence quantization features; The preset calculation method is: ,in, represents the training dance action encoding feature i, represents the training dance action quantization feature matched by the training dance action encoding feature i, Indicates that the code is thin, Denotes the feature vector j in the codebook.
5. The method for generating dance movements of a virtual character according to any one of claims 1 to 4, characterized in that: Starting from the walking starting point, walking sampling is performed in the directed weighted graph of the dance action until the quantitative features of multiple target dance actions with the same number of beat points are sampled, and the quantitative features of multiple sampled dance actions are obtained, including: Starting from the walking starting point, random walking sampling is performed in the directed weighted graph of the dance movements, or walking sampling is performed according to a preset sampling rule, until the quantitative features of multiple target dance movements with the same number of beat points are sampled, thereby obtaining the quantitative features of multiple sampled dance movements.
6. The method for generating dance movements of a virtual character according to any one of claims 1 to 4, characterized in that: The step of determining a starting point of walking the pre-constructed directed weighted graph of dance movements comprises: Determine the node starting point of the pre-constructed dance action directed weighted graph; The node starting point is used as the walking starting point of the directed weighted graph of the dance action.
7. A device for generating dance movements of a virtual character, characterized in that: include: An acquisition module, used to acquire target audio and extract beat information in the target audio, wherein the beat information includes the number of beat points and the beat landing time of all beat points; A determination module is used to determine a walking starting point of a pre-constructed dance action directed weighted graph, wherein the nodes of the dance action directed weighted graph represent the quantitative characteristics of one target dance action, and the target dance actions at different nodes have different action types; A walking module, used for starting from the walking starting point, walking sampling in the directed weighted graph of the dance action, until the quantitative features of a plurality of target dance actions having the same number as the beat points are sampled, thereby obtaining the quantitative features of a plurality of sampled dance actions; A setting module, for setting an action time for each sampled dance action according to the beat landing time of each beat point and the sampling order of each sampled dance action in the plurality of sampled dance actions, wherein the sampling order of each sampled dance action corresponds to the beat landing time of each beat point one by one; A generation module, used for generating dance movements of a virtual character according to the quantitative features and movement moments of all the sampled dance movements; The dance action generating device of the virtual character is also used for: obtaining a first dance video before determining the node starting point of the pre-constructed dance action directed weighted graph, wherein the first dance video includes a first dance picture and a first dance soundtrack; Analyze the first dance picture and the first dance soundtrack, extract dance movements at all beats of the first dance soundtrack, and obtain a target dance movement sequence; Inputting the target dance action sequence into a pre-trained target neural network model, so that the target neural network model outputs the quantitative features and encoding of each target dance action in the target dance action sequence; Counting the switching probabilities between different target dance movements in the target dance movement sequence; According to the switching probabilities between the different target dance movements and the encoding of each target dance movement, the dance movement directed weighted graph is constructed, the encoding of each node in the dance movement directed weighted graph corresponds to each target dance movement in the target dance movement sequence, and each target dance movement at each node in the dance movement directed weighted graph is represented by the quantitative features corresponding to each target dance movement.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for generating dance movements of a virtual character as claimed in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for generating dance movements of a virtual character as claimed in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Cboth generation method and device, equipment and storage medium
CN114386392A
Dance video generation method and device, and storage medium
CN114401439A