Dance music generation method, device, terminal and readable storage medium

By performing feature extraction and convolutional neural network training on choreography music, dance music with emotion and expressions is generated, which solves the problem of poor dance effects in the existing technology without dance music, and realizes AI intelligent choreography, improving dance effects and saving costs.

CN114758636BActive Publication Date: 2025-08-22SHENZHEN TECHRISE ELECTRONICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210250022.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-14
Publication Date
2025-08-22
Estimated Expiration
2042-03-14

AI Technical Summary

Technical Problem

The existing dance music generation method is difficult to create dance movements that match the music that has not been danced, and the lack of dance expressions leads to stiff and mechanized dance effects and cannot meet the diverse music needs.

Method used

By extracting the choreography music, generating music feature vectors, and inputting a pre-trained dance music generation model, outputting target dance music containing dance movements and dance expressions, and training using a convolutional neural network model to generate dance movements with emotion.

Benefits of technology

It has achieved the creation of dance movements and expressions that match it for music that has not been danced, which has improved the dance effect and saved manpower and material resources. The generated dance music has soul and emotion, and is adapted to various music types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114758636B_ABST
    Figure CN114758636B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of computer technology and mainly provides a dance music generation method, device, terminal and readable storage medium. The present application extracts features from the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed, and then inputs the music feature vector into a pre-trained dance music generation model to obtain a target dance music including target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each of the target dance movements. Compared with merely matching stiff and mechanical dance movements with the music to be choreographed, the emotional dance music generated by the dance music generation method of the present application can effectively enhance the dance effect of the dance music and has higher use value. In addition, the present application can generate dance movements that match the music for unmatched music, thereby realizing AI intelligent choreography and saving manpower and material resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and in particular relates to a dance music generation method, device, terminal and readable storage medium. Background Art

[0002] At present, there are generally two ways to generate dance music. The first one comes from the choreographer's inspiration, and the choreographer choreographs according to the rhythm of the music; the second one is to use algorithms to perform three-dimensional modeling of existing dance movements, form a dance movement library, and output dance movements.

[0003] The first approach, however, is difficult to create a unified dance music library because each choreographer has a different understanding of music rhythm and dance moves, influenced by their own personal habits. This also requires significant manpower and resources. The second approach is limited to existing music and dance moves, and for music without matching dance moves to the music, there is a possibility of not being able to create dance moves that match the music. Furthermore, current dance music generation methods simply add dance moves to unmatched music without considering other factors that affect the dance effect, resulting in poor dance performance. Summary of the Invention

[0004] The present application provides a dance music generation method, device, terminal and readable storage medium, which can realize AI intelligent choreography and create dance movements that match the music for music that has not been matched with a dance, saving manpower and material resources. In addition, the generated dance movements can be matched with corresponding dance expressions to obtain dance music with emotion, thereby improving the dance effect of the dance music.

[0005] A first aspect of an embodiment of the present application provides a dance music generation method, comprising:

[0006] Get the music to be choreographed;

[0007] Performing feature extraction on the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed;

[0008] The music feature vector is input into a pre-trained dance music generation model to obtain a target dance music corresponding to the music to be choreographed output by the dance music generation model, wherein the target dance music includes target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each target dance movement.

[0009] A second aspect of the embodiments of the present application further provides a dance music generation device, comprising:

[0010] An acquisition unit, used to acquire the music to be choreographed;

[0011] A feature extraction unit, configured to extract features from the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed;

[0012] A dance music generation unit is used to input the music feature vector into a pre-trained dance music generation model to obtain a target dance music corresponding to the music to be choreographed output by the dance music generation model, wherein the target dance music includes target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each target dance movement.

[0013] A third aspect of an embodiment of the present application provides a terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the dance music generation method described in the first aspect.

[0014] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the dance music generation method described in the first aspect are implemented.

[0015] In the embodiment of the present application, feature extraction is performed on the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed, and then the music feature vector is input into a pre-trained dance music generation model to obtain a target dance music corresponding to the music to be choreographed output by the dance music generation model. Specifically, since the target dance music contains target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each of the target dance movements, that is, the dance movements generated by the present application have dance expressions that match them, the generated target dance music is an emotional and soulful dance music, which is different from simply matching the music to be choreographed with stiff and mechanical dance movements. Specifically, the dance music with emotion generated by the dance music generation method of the present application can effectively enhance the dance effect of the dance music and has a higher use value; in addition, since the music feature vectors corresponding to any type of music are input into the pre-trained dance music generation model, the present application can create dance music that matches it, without limiting the music feature vectors corresponding to the music input into the pre-trained dance music generation model to the music feature vectors corresponding to music that has been choreographed before. Therefore, there will be no situation where dance movements that match the music cannot be generated for music that has not been matched with a dance, thus realizing AI intelligent choreography and saving manpower and material resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A schematic diagram of the implementation flow of the dance music generation method provided in an embodiment of the present application.

[0017] Figure 2 A schematic diagram of the implementation flow of the training method for the dance music generation model provided in an embodiment of the present application.

[0018] Figure 3A schematic diagram of the implementation flow of the method for acquiring sample dance movements provided in an embodiment of the present application.

[0019] Figure 4 A schematic diagram of the implementation flow of the method for obtaining sample dance expressions provided in an embodiment of the present application.

[0020] Figure 5 A schematic structural diagram of a dance music generation device provided in an embodiment of the present application.

[0021] Figure 6 A schematic diagram of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0023] In practical applications, the dancer's emotions and expressions play a crucial role in the effectiveness of dance, not just the movements themselves. A set of movements paired with a rigid expression creates a distinctly different effect than a set of movements paired with a cheerful expression. However, most current choreography schemes simply provide rigid movements lacking expression, resulting in a dance that lacks soul, becoming overly rigid and mechanical.

[0024] The embodiments of the present application provide a dance music generation method, device, terminal and readable storage medium. By extracting features from the music to be choreographed, a music feature vector corresponding to the music to be choreographed is obtained. Then, the music feature vector is input into a pre-trained dance music generation model to obtain a target dance music corresponding to the music to be choreographed output by the dance music generation model. Specifically, since the target dance music contains target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each target dance movement, that is, the dance movements generated by the present application have dance expressions that match them, the generated target dance music is an emotional and soulful dance music, which is different from the target dance music that only matches the music to be choreographed with a dance expression. Compared with rigid and mechanical dance movements, the emotional dance music generated by the dance music generation method of the present application can effectively enhance the dance effect of the dance music and has higher use value. In addition, since the music feature vectors corresponding to any type of music are input into the pre-trained dance music generation model, the present application can create dance music that matches it, without limiting the music feature vectors corresponding to the music input into the pre-trained dance music generation model to the music feature vectors corresponding to the music that has been choreographed. Therefore, there will be no situation where dance movements that match the music cannot be generated for the music that has not been matched with the dance, thus realizing AI intelligent choreography and saving manpower and material resources.

[0025] In order to better illustrate the technical solution of the present application, examples are given below.

[0026] like Figure 1 The figure shows a flow chart of a dance music generation method according to an embodiment of the present application. The dance music generation method can be executed by a dance music generation device configured on a terminal. The terminal can be a smart terminal such as a mobile phone, tablet computer, or robot. The present application does not limit the terminal type.

[0027] Specifically, the dance music generation method provided in the embodiment of the present application can be implemented using the following steps 101 to 103:

[0028] Step 101: Obtain the music to be choreographed.

[0029] In the embodiments of the present application, the music to be choreographed refers to music for which dance movements need to be created. For example, the music to be choreographed can be music needed for dance teaching. The present application does not limit the source and type of the music to be choreographed.

[0030] Step 102: extract features from the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed.

[0031] Currently, music feature vectors corresponding to the music to be choreographed are generally obtained through manual labeling or by classifying music based on user comments and user listening history data.

[0032] Optionally, in an embodiment of the present application, in the above-mentioned process of extracting features of the music to be choreographed and obtaining the music feature vector corresponding to the music to be choreographed, one or more music features including Mel-scale Frequency Cepstral Coefficients (MFCC), chords, harmony, and rhythm corresponding to the music to be choreographed can be extracted, and a music feature vector corresponding to the music to be choreographed can be generated based on the music features.

[0033] For example, in the above process of extracting features from the music to be choreographed and obtaining the music feature vector corresponding to the music to be choreographed, the Mel-frequency cepstral coefficients corresponding to the music to be choreographed can be extracted to obtain the music feature vector corresponding to the music to be choreographed.

[0034] Specifically, in the process of extracting the Mel-frequency cepstral coefficients corresponding to the music to be choreographed and obtaining the music feature vector corresponding to the music to be choreographed, it can be achieved through the following steps: using an anti-aliasing filter with a bandwidth of 300-3400 Hz to pre-filter the music to be choreographed, then performing A / D conversion, pre-emphasis, framing, windowing, Fast Fourier Transformation (FFT), triangular window filtering, logarithm calculation, discrete cosine transform (DCT), and finally outputting the music feature vector corresponding to the music to be choreographed.

[0035] In an embodiment of the present application, since music features such as Mel-frequency cepstral coefficients, chords, harmony, and rhythm can all be used to characterize the music to be choreographed, after generating a music feature vector corresponding to the music to be choreographed based on the music features, the music feature vector can be input into a pre-trained dance music generation model for recognition, and the corresponding target dance music can be generated.

[0036] Optionally, in some embodiments of the present application, in the process of extracting features from the choreography music to obtain the music feature vector corresponding to the choreography music, the formula The discrete signals corresponding to the frequency components of the music to be choreographed are obtained by calculation; then, based on the discrete signals corresponding to the frequency components of the music to be choreographed, a music feature vector corresponding to the music to be choreographed is generated.

[0037] Among them, x(n) is the time domain signal corresponding to the music to be choreographed, W Nk [n] is the length N k The window function, S is the signal sampling frequency, δf is the frequency resolution, Q is a constant, and, f k represents the kth frequency component, and, f min is the lower limit of the frequency of the music to be choreographed, b is the number of frequency spectrum lines contained in one octave; k is an integer from 1 to M, and M is the number of filters used to filter the power spectrum of the music to be choreographed.

[0038] For example, in the above-mentioned process of generating the music feature vector corresponding to the music to be choreographed based on the discrete signals corresponding to the various frequency components of the music to be choreographed, when the discrete signals corresponding to the various frequency components of the music to be choreographed are x[1], x[2], x[3], ..., x[M], the music feature vector corresponding to the music to be choreographed can be expressed as F = (x[1], x[2], x[3], ..., x[M]).

[0039] It should be noted that the above is merely an example of how this application implements feature extraction of the music to be choreographed and obtains the music feature vector corresponding to the music to be choreographed, and is not intended to limit the scope of protection of this application. It is understood that in other embodiments of this application, other feature extraction methods may be used to obtain the music feature vector corresponding to the music to be choreographed. For example, the Librosa feature extraction tool may be used to implement feature extraction of the music to be choreographed and obtain the music feature vector corresponding to the music to be choreographed, but this application does not limit this.

[0040] Step 103: input the music feature vector into a pre-trained dance music generation model to obtain a target dance music corresponding to the music to be choreographed output by the dance music generation model.

[0041] The target dance music includes target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each target dance movement.

[0042] For example, the target dance music is a dance video, the background music of which is the music to be choreographed. The posture of the target object in each frame of the dance video and the character expression corresponding to each posture are the target dance movements corresponding to the music to be choreographed and the target dance expressions corresponding to each target dance movement.

[0043] Optionally, in an embodiment of the present application, the above-mentioned pre-trained dance music generation model can be a convolutional neural network model trained based on a large amount of sample dance music.

[0044] Optional, such as Figure 2 As shown, the dance music generation model can be generated by following steps 201 to 203.

[0045] Step 201: Obtain sample choreography music corresponding to the sample dance music, and obtain sample dance movements corresponding to the sample choreography music and sample dance expressions corresponding to each sample dance movement.

[0046] In an embodiment of the present application, the above-mentioned sample dance music can be a sample dance video, the background music of the sample dance video is the above-mentioned sample choreography music, and the posture of the sample object in each frame of the sample dance video and the character expression corresponding to each posture are the sample dance movements corresponding to the above-mentioned sample choreography music and the sample dance expressions corresponding to each sample dance movement.

[0047] In step 202, the music feature vector corresponding to the sample choreography music is input into the neural network model to be trained, and the neural network model outputs dance movements to be determined corresponding to the sample choreography music and dance expressions to be determined corresponding to each dance movement to be determined.

[0048] In the embodiment of the present application, the method for obtaining the music feature vector corresponding to the sample choreography music can refer to the description of the method for obtaining the music feature vector corresponding to the choreography music in the above step 102, and will not be repeated here.

[0049] Step 203, after adjusting the neural network model to be trained based on the differences between the to-be-determined dance movements corresponding to the sample choreography music and the sample dance movements corresponding to the sample choreography music, as well as the differences between the to-be-determined dance expressions corresponding to each to-be-determined dance movement and the sample dance expressions corresponding to the corresponding sample dance movement, returns to execute the steps of inputting the sample music feature vector corresponding to the sample choreography music into the neural network model to be trained, obtaining the to-be-determined dance movements corresponding to the sample choreography music output by the neural network model and the to-be-determined dance expressions corresponding to each to-be-determined dance movement, and subsequent steps, until the training of the neural network model to be trained is completed and the dance music generation model is obtained.

[0050] For example, when the difference between the to-be-determined dance movements corresponding to the sample choreography music and the sample dance movements corresponding to the sample choreography music, as well as the difference between the to-be-determined dance expressions corresponding to each to-be-determined dance movement and the sample dance expressions corresponding to the corresponding sample dance movement are all less than the preset difference value, the training of the neural network model to be trained is terminated to obtain a dance music generation model.

[0051] Among them, in some embodiments of the present application, such as Figure 3 As shown, in the above step 201, the process of obtaining sample dance movements corresponding to the sample choreography music can be implemented by adopting the following steps 301 to 303.

[0052] Step 301: extract the position coordinates of each joint feature point of the sample object in each frame of the sample image of the sample dance music.

[0053] In the embodiment of the present application, the above-mentioned position coordinates are coordinates under the same reference coordinate system, and can be two-dimensional coordinates or three-dimensional space coordinates, which is not limited in the present application.

[0054] Step 302 : Based on the position coordinates of each joint feature point of the sample object in each frame of the sample dance music, vectorize the corresponding joint feature points of the sample object to obtain the action vectors corresponding to each joint feature point of the sample object.

[0055] Since one frame of sample picture corresponds to a moment, after obtaining the position coordinates of each joint feature point of the sample object in each frame of sample picture of the sample dance music, it is equivalent to obtaining the position coordinates of each joint feature point of the sample object at different music playing moments. Therefore, based on the position coordinates of the corresponding joint feature points of the sample object in each frame of sample picture of the sample dance music, the corresponding joint feature points of the sample object can be vectorized and represented respectively to obtain the action vectors corresponding to each joint feature point of the sample object.

[0056] For example, after obtaining the position coordinates of the j-th joint feature point of the sample object in each frame of the sample image of the sample dance music, the obtained X position coordinates are vectorized to obtain the action vector A corresponding to the j-th joint feature point of the sample object. j .

[0057] Step 303: Combine the motion vectors corresponding to the joint feature points of the sample object to obtain a sample dance motion corresponding to the sample choreography music.

[0058] For example, the action vectors corresponding to the j joint feature points contained in the sample object are A1, A2, A3, ..., A j , then the sample dance movements corresponding to the sample choreography music can be expressed as P(A1, A2, A3, ..., A j ).

[0059] Optionally, in some embodiments of the present application, in the process of obtaining the sample dance movements corresponding to the sample choreography music, in addition to calculating based on the position coordinates of each joint feature point of the sample object in each frame sample image of the sample dance music in the same reference coordinate system, it can also be calculated based on the distance between the position of each joint feature point of the sample object in each frame sample image of the sample dance music and the reference position, for example, based on the Gaussian distance, Hamming distance, etc. between the position of each joint feature point of the sample object in each frame sample image of the sample dance music and the reference position, and the present application does not impose any restrictions on this.

[0060] Optionally, in some embodiments of the present application, such as Figure 4 As shown, in the process of obtaining the sample dance expression corresponding to the sample dance movement in the above step 201, the following steps 401 to 406 can be used to implement it.

[0061] Step 401: Obtain a face image of a sample subject in each frame of a sample dance music.

[0062] Step 402: Perform facial expression recognition on the above-mentioned facial image to obtain the facial expression of the sample object in each frame of the sample picture of the above-mentioned sample dance music.

[0063] Facial expressions are the result of one or more movements or states of facial muscles, and changes in facial movements reveal the dancer's emotional state. There are at least 21 facial expressions, including six basic ones: happiness, surprise, sadness, anger, disgust, and fear. There are also 15 compound expressions, such as surprise (happiness + surprise) and grief (sadness + anger), which are composed of these basic expressions.

[0064] As shown in Table 1 below, facial expressions are the result of facial movements or the state of facial muscles. Figure 3 The motion capture shown above cannot be obtained. The aforementioned motion capture mainly captures the changes in the key points of the human skeleton, rather than the changes in muscle movement. Therefore, the changes in muscle movement require separate expression recognition to obtain the facial expression of the sample subject in each sample image of the above dance music.

[0065] Table 1:

[0066]

[0067]

[0068] Specifically, in the process of performing expression recognition on the above-mentioned facial images to obtain the facial expressions of the sample objects in each frame of the sample dance music, the acquired facial images can be firstly aligned, and then the facial key points can be obtained to obtain the facial contour features, remove the information irrelevant to the face, and obtain the enhanced facial image through random enhancement processing; then, the enhanced facial image is subjected to feature extraction through a feature extraction algorithm to obtain the facial expression features, and finally the facial expression features are matched with the preset classification expressions to obtain the facial expressions of the sample objects in each frame of the sample dance music.

[0069] Step 403: performing frame processing on the sample choreography music corresponding to the sample dance music to obtain a music signal of each frame of the sample choreography music.

[0070] In the embodiment of the present application, the above-mentioned frame processing of the sample choreography music refers to segmenting the sample choreography music according to a preset time period or sampling number, that is, framing, to obtain a music signal of each frame of the sample choreography music.

[0071] Step 404 : obtaining the musical expression corresponding to each frame of the sample picture of the sample dance music based on the Mel spectrum of each frame of the music signal.

[0072] In an embodiment of the present application, before obtaining the musical expression corresponding to each frame of the sample picture of the sample dance music based on the Mel spectrum of each frame of the music signal, the neural network model can be trained in advance using the sample music with the musical expression labeled. After obtaining the network model for musical expression recognition, the Mel spectrum of each frame of the music signal is input into the network model respectively to obtain the musical expression corresponding to each frame of the sample picture of the sample dance music based on the Mel spectrum of each frame of the music signal.

[0073] Step 405 , compare the facial expression of the sample object in each frame of the sample dance music with the corresponding musical expression, and use the facial expression or musical expression corresponding to the target sample picture whose facial expression is consistent with the musical expression as the dance expression corresponding to the target sample picture.

[0074] In this application, when the facial expression of the sample object in the same frame sample picture is inconsistent with the corresponding musical expression, it means that there is a contradiction between the facial expression characteristics and the musical emotion characteristics, that is, the data is wrong. For example, the musical emotion characteristic is a happy expression, but the facial expression characteristic is a sad expression. The two are contradictory. Therefore, it is necessary to remove the data and only use the facial expression or musical expression corresponding to the target sample picture whose facial expression is consistent with the musical expression as the dance expression corresponding to the target sample picture to improve the accuracy of the dance expression.

[0075] Step 406 : combining the dance expressions corresponding to each target sample image frame to obtain a sample dance expression corresponding to the sample dance movement.

[0076] In the embodiment of the present application, by using the above Figure 3 and Figure 4 The sample dance movements corresponding to the sample choreography music and the sample dance expressions corresponding to each sample dance movement obtained in the manner shown are used to train the neural network model to be trained to obtain a dance music generation model, so that the dance music generation model can output target dance music including target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each target dance movement, thereby realizing AI intelligent choreography, being able to create dance movements that match the music for music that has not been matched with a dance, saving manpower and material resources, and also being able to match the generated dance movements with corresponding dance expressions to obtain dance music with emotion, thereby improving the dance effect of the dance music.

[0077] In the embodiments of the present application, the above-mentioned dance music generation method can be applied to terminals such as dancing robots, mobile phones, and computers, and can play an important role in application scenarios such as children learning to dance or adults learning to dance.

[0078] It should be noted that the above is merely an example of how to obtain the sample dance movements corresponding to the sample choreography music and the sample dance expressions corresponding to each sample dance movement. It can be understood that the sample dance movements corresponding to the above sample choreography music and the sample dance expressions corresponding to each sample dance movement can also be obtained through other methods, for example, through manual labeling, etc., and this application does not impose any restrictions on this.

[0079] It should also be noted that, for the sake of simplicity of description, the aforementioned method embodiments are all expressed as a series of action combinations. However, those skilled in the art should be aware that this application is not limited to the order of the actions described. In some embodiments of this application, certain steps may be performed in other orders.

[0080] For example, the above steps 401 and 402 may be performed after step 404, that is, steps 403 and 404 may be performed before step 401, and this application does not impose any limitation on this.

[0081] like Figure 5 , which is a schematic structural diagram of a dance music generating device 500 provided in an embodiment of the present application. The dance music generating device may include: an acquisition unit 501 , a feature extraction unit 502 and a dance music generating unit 503 .

[0082] An acquisition unit 501 is used to acquire music to be choreographed;

[0083] A feature extraction unit 502 is used to extract features from the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed;

[0084] The dance music generation unit 503 is used to input the music feature vector into a pre-trained dance music generation model to obtain a target dance music corresponding to the music to be choreographed output by the dance music generation model, wherein the target dance music includes target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each target dance movement.

[0085] Optionally, in some embodiments of the present application, the feature extraction unit 502 is used to extract one or more music features including Mel-cepstral coefficients, chords, harmony, and rhythm corresponding to the music to be choreographed, and generate a music feature vector corresponding to the music to be choreographed based on the music features.

[0086] Optionally, in some embodiments of the present application, the feature extraction unit 502 is further configured to:

[0087] Using the formula Calculate and obtain the discrete signals corresponding to the frequency components of the music to be choreographed; wherein x(n) is the time domain signal corresponding to the music to be choreographed, WNk [n] is the length N k The window function, S is the signal sampling frequency, δf is the frequency resolution, Q is a constant, and, f k represents the kth frequency component, and, f min is the lower frequency limit of the music to be choreographed, b is the number of frequency spectrum lines contained in one octave; k is an integer from 1 to M, and M is the number of filters used to filter the power spectrum of the music to be choreographed;

[0088] Based on the discrete signals corresponding to the frequency components of the music to be choreographed, a music feature vector corresponding to the music to be choreographed is generated.

[0089] Optionally, in some embodiments of the present application, the dance music generation device may further include a model training unit configured to:

[0090] Obtaining sample choreography music corresponding to the sample dance music, and obtaining sample dance movements corresponding to the sample choreography music and sample dance expressions corresponding to each of the sample dance movements;

[0091] Inputting the music feature vector corresponding to the sample choreography music into the neural network model to be trained, and obtaining the to-be-determined dance movements corresponding to the sample choreography music and the to-be-determined dance expressions corresponding to the to-be-determined dance movements output by the neural network model;

[0092] Based on the differences between the to-be-determined dance movements corresponding to the sample choreography music and the sample dance movements corresponding to the sample choreography music, as well as the differences between the to-be-determined dance expressions corresponding to each of the to-be-determined dance movements and the sample dance expressions corresponding to the corresponding sample dance movements, after adjusting the neural network model to be trained, the method returns to executing the steps of inputting the sample music feature vector corresponding to the sample choreography music into the neural network model to be trained, obtaining the to-be-determined dance movements corresponding to the sample choreography music and the to-be-determined dance expressions corresponding to each of the to-be-determined dance movements output by the neural network model, and subsequent steps, until the training of the neural network model to be trained is completed and the dance music generation model is obtained.

[0093] Optionally, in some embodiments of the present application, the above-mentioned model training unit is further used to:

[0094] Extracting the position coordinates of each joint feature point of the sample object in each frame of the sample picture of the sample dance music;

[0095] Based on the position coordinates of each joint feature point of the sample object in each frame of the sample picture of the sample dance music, the corresponding joint feature points of the sample object are respectively vectorized to obtain the action vectors corresponding to the joint feature points of the sample object;

[0096] The action vectors corresponding to the joint feature points of the sample object are combined to obtain the sample dance action corresponding to the sample choreography music.

[0097] Optionally, in some embodiments of the present application, the above-mentioned model training unit is further used to:

[0098] Obtaining a face image of a sample subject in each frame of a sample picture of the sample dance music;

[0099] Performing facial expression recognition on the facial image to obtain the facial expression of the sample subject in each frame of the sample dance music;

[0100] performing frame processing on the sample choreography music corresponding to the sample dance music to obtain a music signal of each frame of the sample choreography music;

[0101] Obtaining a musical expression corresponding to each frame of the sample picture of the sample dance music based on the Mel spectrum of each frame of the music signal;

[0102] Comparing the facial expression of the sample subject in each frame of the sample dance music with the corresponding musical expression, and taking the facial expression or musical expression corresponding to the target sample picture whose facial expression is consistent with the musical expression as the dance expression corresponding to the target sample picture;

[0103] The dance expressions corresponding to each frame of the target sample image are combined to obtain sample dance expressions corresponding to the sample dance movements.

[0104] It should be noted that, for ease and brevity of description, the specific operating process of the dance music generation device 500 described above can be referred to the description of the dance music generation method in the above embodiments, and will not be repeated here. Furthermore, it should be noted that the above embodiments can be combined with each other to obtain a variety of different embodiments, all of which fall within the scope of protection of this application.

[0105] like Figure 6 As shown, the embodiment of the present application also provides a terminal. The terminal can be configured with the dance music generation device shown in the above embodiments. Figure 6 As shown, the terminal 6 may include: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, the steps in the above-mentioned various dance music generation method embodiments are implemented, for example, Figure 1 Steps 101 to 103 are shown.

[0106] The processor 60 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0107] The memory 61 can be an internal storage unit of the terminal 6, such as a hard disk or memory. The memory 61 can also be an external storage device for the terminal 6, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the memory 61 can include both the internal storage unit of the terminal 6 and an external storage device. The memory 61 is used to store the aforementioned computer programs as well as other programs and data required by the terminal.

[0108] The computer program can be divided into one or more units, which are stored in the memory 61 and executed by the processor 60 to complete the present application. The one or more units can be a series of computer program instruction segments that can perform specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal that calls the camera. For example, the computer program can be divided into: an acquisition unit, a feature extraction unit, and a dance music generation unit, with specific functions as follows:

[0109] An acquisition unit, used to acquire the music to be choreographed;

[0110] A feature extraction unit, configured to extract features from the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed;

[0111] A dance music generation unit is used to input the music feature vector into a pre-trained dance music generation model to obtain a target dance music corresponding to the music to be choreographed output by the dance music generation model, wherein the target dance music includes target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each target dance movement.

[0112] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0113] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0114] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0115] In the embodiments provided in this application, it should be understood that the disclosed terminals and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical functional division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0116] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0117] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0118] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0119] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A dance music generation method, characterized in that: include: Get the music to be choreographed; Performing feature extraction on the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed; Inputting the music feature vector into a pre-trained dance music generation model to obtain a target dance music corresponding to the music to be choreographed, which is output by the dance music generation model, wherein the target dance music includes target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each target dance movement; the dance music generation model is trained using sample choreography music corresponding to the sample dance music, sample dance movements corresponding to the sample choreography music, and sample dance expressions corresponding to each sample dance movement; The step of obtaining the sample dance expression corresponding to the sample dance movement includes: Obtaining a face image of a sample subject in each frame of a sample picture of the sample dance music; Performing facial expression recognition on the facial image to obtain the facial expression of the sample subject in each frame of the sample dance music; performing frame processing on the sample choreography music corresponding to the sample dance music to obtain a music signal of each frame of the sample choreography music; Obtaining a musical expression corresponding to each frame of the sample picture of the sample dance music based on the Mel spectrum of each frame of the music signal; Comparing the facial expression of the sample subject in each frame of the sample dance music with the corresponding musical expression, and taking the facial expression or musical expression corresponding to the target sample picture whose facial expression is consistent with the musical expression as the dance expression corresponding to the target sample picture; The dance expressions corresponding to each frame of the target sample image are combined to obtain sample dance expressions corresponding to the sample dance movements.

2. The dance music generation method according to claim 1, wherein: The step of extracting features from the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed includes: One or more music features of the mel-cepstral coefficients, chords, harmony, and rhythm corresponding to the music to be choreographed are extracted, and a music feature vector corresponding to the music to be choreographed is generated based on the music features.

3. The dance music generation method according to claim 1, wherein: The step of extracting features from the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed includes: Using the formula Calculate and obtain the discrete signals corresponding to the frequency components of the music to be choreographed; wherein x(n) is the time domain signal corresponding to the music to be choreographed, The length is N k The window function, S is the signal sampling frequency, δf is the frequency resolution, Q is a constant, and, f k represents the kth frequency component, and, f min is the lower frequency limit of the music to be choreographed, b is the number of frequency spectrum lines contained in one octave; k is an integer from 1 to M, and M is the number of filters used to filter the power spectrum of the music to be choreographed; Based on the discrete signals corresponding to the frequency components of the music to be choreographed, a music feature vector corresponding to the music to be choreographed is generated.

4. The dance music generation method according to any one of claims 1 to 3, wherein: The steps of generating the dance music generation model include: Obtaining sample choreography music corresponding to the sample dance music, and obtaining sample dance movements corresponding to the sample choreography music and sample dance expressions corresponding to each of the sample dance movements; Inputting the music feature vector corresponding to the sample choreography music into the neural network model to be trained, and obtaining the to-be-determined dance movements corresponding to the sample choreography music and the to-be-determined dance expressions corresponding to the to-be-determined dance movements output by the neural network model; Based on the differences between the to-be-determined dance movements corresponding to the sample choreography music and the sample dance movements corresponding to the sample choreography music, as well as the differences between the to-be-determined dance expressions corresponding to each of the to-be-determined dance movements and the sample dance expressions corresponding to the corresponding sample dance movements, after adjusting the neural network model to be trained, the method returns to executing the steps of inputting the sample music feature vector corresponding to the sample choreography music into the neural network model to be trained, obtaining the to-be-determined dance movements corresponding to the sample choreography music and the to-be-determined dance expressions corresponding to each of the to-be-determined dance movements output by the neural network model, and subsequent steps, until the training of the neural network model to be trained is completed and the dance music generation model is obtained.

5. The dance music generation method according to claim 4, wherein: The obtaining of sample dance movements corresponding to the sample choreography music includes: Extracting the position coordinates of each joint feature point of the sample object in each frame of the sample picture of the sample dance music; Based on the position coordinates of each joint feature point of the sample object in each frame of the sample picture of the sample dance music, the corresponding joint feature points of the sample object are respectively vectorized to obtain the action vectors corresponding to the joint feature points of the sample object; The action vectors corresponding to the joint feature points of the sample object are combined to obtain the sample dance action corresponding to the sample choreography music.

6. A dance music generating device, characterized in that: include: An acquisition unit, used to acquire the music to be choreographed; A feature extraction unit, configured to extract features from the music to be choreographed to obtain a music feature vector corresponding to the music to be choreographed; a dance music generation unit, configured to input the music feature vector into a pre-trained dance music generation model to obtain a target dance music corresponding to the music to be choreographed, output by the dance music generation model, wherein the target dance music includes target dance movements corresponding to the music to be choreographed and target dance expressions corresponding to each target dance movement; the dance music generation model is trained using sample choreography music corresponding to the sample dance music, sample dance movements corresponding to the sample choreography music, and sample dance expressions corresponding to each sample dance movement; The dance music generation device also includes a model training unit, which is used to obtain the facial image of the sample object in each frame of the sample dance music when obtaining the sample dance expression corresponding to the sample dance movement; perform expression recognition on the facial image to obtain the facial expression of the sample object in each frame of the sample dance music; perform frame processing on the sample choreography music corresponding to the sample dance music to obtain the music signal of each frame of the sample choreography music; obtain the music expression corresponding to each frame of the sample picture of the sample dance music based on the Mel spectrum of each frame of the music signal; compare the facial expression of the sample object in each frame of the sample picture of the sample dance music with the corresponding music expression, and use the facial expression or music expression corresponding to the target sample picture with the consistent facial expression and music expression as the dance expression corresponding to the target sample picture; combine the dance expressions corresponding to each frame of the target sample picture to obtain the sample dance expression corresponding to the sample dance movement.

7. The dance music generating device according to claim 6, wherein: include: The feature extraction unit is used to extract one or more music features of Mel-cepstral coefficients, chords, harmony, and rhythm corresponding to the music to be choreographed, and generate a music feature vector corresponding to the music to be choreographed based on the music features.

8. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Music data processing method and device specific to humanoid robot

    CN106292423A

  • Music-driven dance generation method

    CN110853670A

  • Speech enhancement method based on constant constant frequency domain transformation

    CN111402909A