A method for generating dance poses based on a multi-feature fusion strategy

Through multi-feature fusion strategy and generative adversarial network, the problem of insufficient coordination of music and dance matching in the existing technology is solved, and a more natural and rich dance movement generation is achieved.

CN114998984BActive Publication Date: 2025-06-13SOUTHWEAT UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210458956.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-27
Publication Date
2025-06-13
Estimated Expiration
2042-04-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the matching coordination of music and dance and the authenticity of dance movements, especially in terms of movement transition and style coordination.

Method used

The music-generating dance posture method based on a multi-feature fusion strategy is adopted. Through feature extraction, feature fusion and posture generation steps, combining the Generative Adversarial Network (GAN) and the Autoencoder, the structure, beat and style characteristics are fused to generate natural and smooth dance postures.

Benefits of technology

It improves the coordination between movement and music in style and rhythm, enhances the richness and authenticity of dance movements, making the generated dance movements more smooth and natural, and consistent with the music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998984B_ABST
    Figure CN114998984B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating dance poses based on a multi-feature fusion strategy, including the steps of: feature extraction: preprocessing an audio file to obtain an audio sequence, and converting the audio sequence into feature data, which is composed of structural features, beat features, and style features; feature fusion: fusing the structural features, beat features, and style features to obtain a musical feature representation; pose generation: inputting the musical feature representation into a pose generator to obtain dance poses. The present invention can improve the coordination between movements and music in terms of style and rhythm, and enhance the richness of dance movements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer dance choreography, and particularly relates to a method for generating dance postures from music based on a multi-feature fusion strategy. Background Art

[0002] With the continuous development of music processing technology and motion capture technology, the technology of synthesizing dance movements based on music has gradually become a research hotspot in the fields of music understanding and dance synthesis. How to improve the matching between music and dance and the authenticity of synthesized dance is the key point of research.

[0003] Statistical models were the earliest work in this type of task. For example, actions were synthesized based on kernel-based probability distributions, but the disadvantage was the lack of action details. The action graph solved the problem of lacking action details in a non-parametric way. The action graph is a directed graph on an action dataset, where each node represents a posture and each edge represents the transition between two postures. Actions are generated by randomly walking on the graph. The disadvantage is the rationality of the generated transitions. Some methods solve this problem by parameterizing the transitions. The method based on kernel-based probability distributions lacks action details and appears as extremely rigid actions. The method of the action graph transforms the problem into finding the optimal path on the graph, solves the problem of lacking action details in a non-parametric way, and adding beat information to the action graph can synthesize rhythmic actions. However, it is difficult to have reasonable transitions between actions, and the actions are not coherent, appearing as a splicing of multiple segments of actions.

[0004] Nowadays, more often neural networks are used to generate 3D actions. An autoregressive model like RNN can theoretically generate infinite actions. Due to the problem of cumulative misalignment in itself, unnatural phenomena such as action stiffness and drift occur after several rounds of iteration. The staged neural network and its variants solve this problem by adjusting the network weights at each stage. Although the staged training of the network alleviates the stiffness problem by adjusting the network weights stage by stage, it cannot represent many types of actions, lacks the richness of represented actions, and the coordination between actions and music in terms of style and rhythm is not very ideal. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a method for generating dance postures from music based on a multi-feature fusion strategy, which improves the coordination between actions and music in terms of style and rhythm and enhances the richness of dance movements.

[0006] To achieve the above object, the technical solution adopted by the present invention is: A method for generating dance postures from music based on a multi-feature fusion strategy, comprising the steps:

[0007] Feature extraction: Preprocess the audio file to obtain an audio sequence, and convert the audio sequence into feature data, which consists of structural features, beat features, and style features;

[0008] Feature fusion: Fuse the structural features, beat features, and style features to obtain a musical feature representation;

[0009] Pose generation: Input the musical feature representation into a pose generator to obtain dance poses.

[0010] Furthermore, preprocessing the audio file to obtain an audio sequence includes the steps of:

[0011] Only sample the audio file to obtain a one-dimensional array representing the waveform;

[0012] Segment the one-dimensional array to obtain audio units corresponding to the duration of each frame;

[0013] Combine the obtained multiple audio units to obtain an audio sequence.

[0014] Furthermore, the feature extraction process consists of a structure extractor, a style extractor, and a beat extractor;

[0015] The structure extractor extracts a structural feature vector, the beat extractor extracts a beat feature vector, and the style extractor extracts a style feature vector. Finally, the three groups of obtained feature vectors are concatenated as the feature vector representing the music.

[0016] Furthermore, for the structure extractor, an audio encoder encodes the music file, and the vector obtained after encoding is processed using an attention mechanism after passing through an LSTM network to obtain a structural feature vector.

[0017] Furthermore, for the style extractor, an audio encoder encodes the music file, and the pre-trained music style extractor is directly used to extract the style features from the encoded vector to obtain a style feature vector.

[0018] Furthermore, the beat extractor uses an open-source tool to extract and obtain a beat feature vector.

[0019] Furthermore, in the feature fusion process, an autoencoder is used for fusion. The concatenated feature vector obtained in the previous stage is passed through an autoencoder-decoder, and the three features are fused to form a comprehensive feature representation as the feature representation of the music.

[0020] Furthermore, the pose generator uses a GAN-based image generation network, including a pose feature generator, a coherence discriminator, and a style discriminator;

[0021] Input the feature representation of music into the pose feature generator to output pose actions;

[0022] Input the pose actions into the coherence discriminator and the style discriminator respectively. The coherence discriminator and the style discriminator are used to constrain the generated dance poses, so that the generated actions are coherent and consistent with the music in style.

[0023] Furthermore, represent the pose actions with a sequence of skeletal key points, which are obtained by open-source tools to get the sequence shape, enabling better learning of the mapping relationship between data and actions and ignoring some irrelevant information such as the background and the human body.

[0024] Furthermore, the network generates data representing poses based on the GAN-based image generation network during the pose generation process. The loss in the pose feature generator consists of three parts, namely the adversarial loss, the joint-based reconstruction loss, and the feature matching loss.

[0025] Beneficial effects of adopting this technical solution:

[0026] The present invention proposes a method based on the generative adversarial network (GAN) and uses a multi-feature fusion strategy to improve the coordination between actions and music in style and rhythm during the process of generating dance from music, and enhance the richness of dance movements.

[0027] The present invention utilizes the multi-feature fusion strategy in combination with the generative adversarial network to make the generated dance movements more detailed and realistic. The task of generating dance from music is essentially a cross-domain generation task. Music and dance can be regarded as representations of feature data in different domains. Using the multi-feature fusion strategy and considering music style, beat, and structural information simultaneously makes the generated actions more smooth and natural and consistent with the music.

[0028] The present invention uses a neural network algorithm to implement the generation of human body pose data from audio data, completing the transformation from audition to vision.

[0029] The present invention represents human body postures with the skeletal diagrams of the human body and represents dance movements with a temporally continuous sequence of postures, improving the matching accuracy between dance movements and music.

[0030] The present invention preprocesses the music file representing sound using one-dimensional convolution to achieve feature extraction at the sound level. Based on the feature vector representing sound, feature extraction is performed in the music dimension. The model learns the mapping from audio data to image data based on the features representing music. In the music dimension, three important features in music theory, namely style features, beat features, and structural features, are considered simultaneously to guide the generation of dance poses from music. The multi-feature fusion strategy is used to extract music features from multiple dimensions of music, perform feature fusion on multiple features, and use the fused features to represent music features.

[0031] The present invention uses two discriminators, namely a continuity discriminator and a style discriminator, to constrain the generated pose sequences in the time dimension and the space dimension respectively, aiming to make the generated actions smooth and coordinated. Brief Description of the Drawings

[0032] Figure 1 It is a schematic diagram of the principle of a method for generating dance poses from music based on a multi-feature fusion strategy of the present invention;

[0033] Figure 2 It is a schematic diagram of the principle of preprocessing an audio file in an embodiment of the present invention. Detailed Embodiments

[0034] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings.

[0035] In this embodiment, the present invention proposes a method for generating dance poses from music based on a multi-feature fusion strategy, including the steps of:

[0036] Feature extraction: Preprocess the audio file to obtain an audio sequence, and convert the audio sequence into feature data, which consists of structural features, beat features and style features;

[0037] Feature fusion: Fuse the structural features, beat features and style features to obtain a music feature representation;

[0038] Pose generation: Input the music feature representation into a pose generator to obtain dance poses.

[0039] As an optimized solution of the above embodiment, preprocessing the audio file to obtain an audio sequence includes the steps of:

[0040] Only sample the audio file to obtain a one-dimensional array representing the waveform;

[0041] Segment the one-dimensional array to obtain audio units corresponding to the duration of each frame;

[0042] Combine the obtained multiple audio units to obtain an audio sequence.

[0043] As an optimized solution of the above embodiment, the feature extraction process consists of a structure extractor, a style extractor and a beat extractor;

[0044] The structure extractor extracts a structural feature vector, the beat extractor extracts a beat feature vector, the style extractor extracts a style feature vector, and finally the three groups of obtained feature vectors are concatenated as a feature vector representing music.

[0045] Among them, the structure extractor encodes a music file with an audio encoder, and the vector obtained after encoding is processed by an attention mechanism after passing through an LSTM network to obtain a structure feature vector.

[0046] Among them, the style extractor encodes a music file with an audio encoder, and directly uses a pre-trained music style extractor to extract style features from the encoded vector to obtain a style feature vector.

[0047] Among them, the beat extractor is extracted using an open-source tool to obtain a beat feature vector.

[0048] As an optimized solution of the above embodiment, in the feature fusion process, an autoencoder is used for fusion. The concatenated feature vectors obtained in the previous stage are passed through the self-encoder-decoder, and the three features are fused to form a comprehensive feature representation as the feature representation of the music.

[0049] As an optimized solution of the above embodiment, the pose generator adopts a GAN-based image generation network, including a pose feature generator, a coherence discriminator, and a style discriminator;

[0050] The feature representation of the music is input into the pose feature generator, and a pose action is output;

[0051] The pose action is respectively input into the coherence discriminator and the style discriminator, and the coherence discriminator and the style discriminator are used to constrain the generated dance poses. So that the generated actions are kept coherent and consistent with the music in style.

[0052] Preferably, the pose action is represented by a sequence of skeletal key points, which is obtained by an open-source tool to obtain a sequence shape; it can better learn the mapping relationship between data and actions and ignore some irrelevant information such as the background and the human body.

[0053] Preferably, the network generates data representing poses based on a GAN-based image generation network during the pose generation process, and the loss in the pose feature generator consists of three parts, namely the adversarial loss, the reconstruction loss based on joint points, and the feature matching loss.

[0054] 1. Adversarial loss L adv :

[0055] Ladv = E r [log D coh (S(r))] + E f,m [log[1 - D coh (S(f))] + E r,m [log D sty (r, m)] + E f [log[1 - D sty (f, m)];

[0056] Among them, S(·) represents sampling the sequence, r represents the real sequence, and f represents the generated sequence.

[0057] 2. Reconstruction loss Lrecon at the joint level:

[0058]

[0059] Among them, V represents the number of key points, and i represents the number of the joint coordinate

[0060] 3. Feature matching loss LF:

[0061]

[0062] Among them, Ф represents a pre-trained graph convolutional neural network, λ represents a hyperparameter, and j represents the number of layers of the pre-trained graph neural network.

[0063] Finally, the total loss function is as follows:

[0064] L = ω 1 L adv + ω 2 L recon + ω 3 L F ;

[0065] Among them, ω 1 , ω 2 , ω 3 represent the weights of each loss.

[0066] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for generating dance postures from music based on a multi-feature fusion strategy, characterized in that, it includes the steps: Feature extraction: Preprocess the audio file to obtain an audio sequence, and convert the audio sequence into feature data, which consists of structural features, beat features, and style features; The feature extraction process is composed of a structure extractor, a style extractor, and a beat extractor; the structure extractor extracts a structural feature vector, the beat extractor extracts a beat feature vector, and the style extractor extracts a style feature vector. Finally, the three groups of feature vectors obtained are concatenated as the feature vector representing the music; the style extractor encodes the music file with an audio encoder, and directly uses a pre-trained music style extractor to extract the style features from the encoded vector to obtain a style feature vector; the structure extractor encodes the music file with an audio encoder, and the vector obtained after encoding is processed by an attention mechanism after passing through an LSTM network to obtain a structural feature vector; the beat extractor is extracted using open-source tools to obtain a beat feature vector; Feature fusion: Fuse the structural features, beat features, and style features to obtain a music feature representation; Posture generation: Input the music feature representation into a posture generator to obtain dance postures.

2. A method for generating dance postures from music based on a multi-feature fusion strategy according to claim 1, characterized in that, Preprocessing the audio file to obtain an audio sequence includes the steps: Only sample the audio file to obtain a one-dimensional array representing the waveform; Segment the one-dimensional array to obtain audio units corresponding to the duration of each frame; Combine the obtained multiple audio units to obtain an audio sequence.

3. A method for generating dance postures from music based on a multi-feature fusion strategy according to claim 1, characterized in that, In the feature fusion process, an autoencoder is used for fusion. The concatenated feature vector obtained in the previous stage is passed through the autoencoder-decoder, and the three features are fused to form a comprehensive feature representation as the music feature representation.

4. A method for generating dance postures from music based on a multi-feature fusion strategy according to claim 1, characterized in that, The posture generator adopts a GAN-based image generation network, including a posture feature generator, a coherence discriminator, and a style discriminator; Input the music feature representation into the posture feature generator to output posture actions; Input the posture actions into the coherence discriminator and the style discriminator respectively, and the coherence discriminator and the style discriminator are used to constrain the generated dance postures.

5. A method for generating dance postures from music based on a multi-feature fusion strategy according to claim 4, characterized in that, The posture actions are represented by a sequence of skeletal key points, which are obtained by open-source tools to obtain the sequence shape.

6. A method for generating dance postures from music based on a multi-feature fusion strategy according to claim 4, characterized in that, The network generates data representing postures based on a GAN-based image generation network during the posture generation process. The loss in the posture feature generator consists of three parts, namely adversarial loss, reconstruction loss based on joint points, and feature matching loss.

Citation Information

Patent Citations

  • Dance animation processing method and device, electronic equipment and storage medium

    CN111179385A

  • Music-driven human skeleton dancing action generation system

    CN112700521A