Animation generation method, apparatus, and device, and medium

By collecting user gestures and using animation generation models to generate matching music and action data, the problem of poor matching effects of dance movements and music in the prior art is solved, and the simplification of user interaction and the high-quality fit between actions and music is achieved.

WO2025124061A1PCT designated stage expired Publication Date: 2025-06-19MIGU CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/132201
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-11-15
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

When generating Dance Homo sapiens dance movements, the template matching method is limited to stored movements, resulting in poor matching effects between dance movements and music, and the inability to fully match the rhythm and emotions of each music. In real-time collection of user actions requires users to have a certain dance foundation.

Method used

By collecting user gesture actions in real time, extracting explicit feature vectors and encoding, generating implicit feature vectors, combining music feature extraction and action feature extraction, using pre-trained animation generation models to generate matching music data and action data, realizing the dual driving of music and action.

Benefits of technology

It simplifies the user interaction process, reduces the difficulty of user operation, and ensures the fit and matching between the actions and music, so that the generated actions not only conform to the user's gesture changes, but also match the performance of the music well, forming a smoother and more coordinated animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024132201_19062025_PF_FP_ABST
    Figure CN2024132201_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides an animation generation method, apparatus, and device, and a medium. The animation generation method comprises: collecting a target gesture action in real time; generating matching music data on the basis of a change condition of the target gesture action, and generating matching action data on the basis of the change condition of the target gesture action and the music data; and on the basis of the music data, generating a music to be played, and on the basis of the action data, controlling a target object to act.
Need to check novelty before this filing date? Find Prior Art

Description

Animation generation method, device, equipment and medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This disclosure claims priority to Chinese Patent Application No. 202311733510.5 filed in China on December 15, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of computer technology, and in particular to an animation generation method, apparatus, device, and medium. Background Art

[0004] With the development of metaverse technology, digital humans in the metaverse are increasingly being used in music videos, where digital humans dance to the music.

[0005] By capturing real-life movements or manually drawing key movement frames and saving them as a template library, the system automatically matches the optimal dance movement template to the input music rhythm and transmits it to the digital intelligent person for dancing. Alternatively, by collecting user movement data in real time and combining it with a given music rhythm for real-time processing and analysis, the system generates corresponding dance movement instructions based on the user's movement data and music rhythm, and transmits them to the digital intelligent person for dance performance.

[0006] However, the inventors discovered that existing technologies have at least the following problems: The template matching approach limits the digital AI's dance movement generation to the movements stored in the template, and the matching between dance movements and music is poor, failing to fully match the rhythm and emotion of each piece of music, resulting in a lack of coordination and fluidity between the dance movements and the music. Furthermore, the real-time user movement capture method requires the user to generate dance movements based on their actual movements captured by a camera, which requires a certain level of dance foundation. Summary of the Invention

[0007] The purpose of the embodiments of the present disclosure is to provide an animation generation method, device, equipment and medium, which realizes dual drive of music generation and target object action generation by obtaining user gesture actions, effectively simplifies the user interaction link, and makes the music and object actions more matched.

[0008] To achieve the above objectives, the present disclosure provides an animation generation method, comprising:

[0009] Collect target gestures in real time;

[0010] Generating matching music data according to the change of the target gesture action, and generating matching action data according to the change of the target gesture action and the music data;

[0011] The music to be played is generated according to the music data, and the action of the target object is controlled according to the action data.

[0012] As an improvement to the above solution, the generating of matching music data according to the change of the target gesture action, and the generating of matching action data according to the change of the target gesture action and the music data, include:

[0013] Extracting an explicit feature vector of the target gesture action;

[0014] Encoding the explicit feature vector to obtain an implicit feature vector of the target gesture action;

[0015] Performing music feature extraction on the implicit feature vector to obtain a music feature vector;

[0016] Predicting music data that matches the target gesture action based on the music feature vector;

[0017] Extracting motion features from the music feature vector to obtain a motion feature vector;

[0018] The action data matching the music data is predicted based on the action feature vector.

[0019] As an improvement to the above solution, the generating of matching music data according to the change of the target gesture action, and the generating of matching action data according to the change of the target gesture action and the music data, include:

[0020] Extracting an explicit feature vector of the target gesture action;

[0021] The explicit feature vector is input into a pre-built animation generation model for calculation to generate the music data and the action data; wherein the animation generation model is used to generate matching music data according to the target gesture action, and to generate matching action data according to the target gesture action and the music data.

[0022] As an improvement to the above solution, the animation generation model is trained through the following steps:

[0023] Acquire a training sample data set; wherein the training sample data set includes a plurality of sample element groups; wherein each of the sample element groups is composed of sample gestures, real music data, and real action data;

[0024] Constructing an initial machine learning model and initializing the parameters of the machine learning model;

[0025] Extracting an explicit feature vector of the sample gesture action, inputting the explicit feature vector into the machine learning model for calculation to generate corresponding predicted music data based on the sample gesture action, and generating corresponding predicted action data based on the sample gesture action and the predicted music data;

[0026] generating a target loss function based on the real music data, the predicted music data, the real action data, and the predicted action data;

[0027] According to the target loss function, the parameters of the machine learning model are updated to obtain a trained animation generation model.

[0028] As an improvement to the above solution, the step of obtaining a training sample data set includes:

[0029] Acquire a first video containing background music information and character action information, and divide the first video into a plurality of first video frames;

[0030] Selecting each of the first video frames in sequence, extracting the character's skeletal key points from the character motion information of the currently selected first video frame to generate the real motion data; and extracting an audio spectrum from the background music information of the currently selected first video frame to generate the real music data;

[0031] Acquire a second video containing continuous hand gestures, and divide the second video into a plurality of second video frames;

[0032] Decomposing and analyzing the gestures in each second video frame from multiple feature dimensions, and matching corresponding sample gestures to the real music data and the real motion data generated for the same first video frame based on the key music information and key motion information represented by each feature dimension of all the gestures, to generate the sample element group;

[0033] The training sample data set is constructed based on the generated several sample element groups.

[0034] As an improvement to the above solution, the characteristic dimensions include amplitude strength, speed frequency, movement direction and hand shape;

[0035] The key music information represented by the amplitude and strength of the gesture action includes the volume and strength of the music data, and the key action information represented includes the action range and strength of the action data;

[0036] The key music data represented by the speed frequency of the gesture action includes the rhythm and speed of the music data, and the key action information represented includes the rhythm of the action data;

[0037] The key music information represented by the movement direction of the gesture action includes the pitch of the musical notes of the music data, and the key action information represented includes the spatial position and movement direction of the action data;

[0038] The key music data represented by the hand shape of the gesture action includes the note effects of the music data, and the key action information represented includes the action content of the action data.

[0039] As an improvement to the above solution, generating a target loss function based on the real music data, the predicted music data, the real action data, and the predicted action data includes:

[0040] generating a first constrained loss function according to the relationship between the real music data and the predicted music data;

[0041] generating a second constraint loss function according to a relationship between the real motion data and the predicted motion data;

[0042] According to the first constraint loss function, the second constraint loss function and the corresponding preset weight coefficients, a weighted summation algorithm is adopted to generate a target loss function.

[0043] The present disclosure also provides an animation generation device, including:

[0044] Data acquisition module, used to collect target gestures in real time;

[0045] a data calculation module, configured to generate matching music data according to the change of the target gesture action, and generate matching action data according to the change of the target gesture action and the music data;

[0046] The data control module is used to generate music to be played according to the music data and control the target object action according to the action data.

[0047] An embodiment of the present disclosure also provides an animation generation device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the animation generation method as described in any one of the above items is implemented.

[0048] An embodiment of the present disclosure further provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute any one of the above-described animation generation methods.

[0049] Compared with the prior art, the animation generation method, device, equipment and medium disclosed in this disclosure have the following beneficial effects:

[0050] First, gestures that are easier to operate, easier to capture, and interactive are used as input features of the model to achieve dual drive of music generation and action generation. Compared with the method of collecting user body movements to generate the actions of the target object, it can simplify the user interaction link and reduce the difficulty of user operation.

[0051] Second, the disclosed embodiment analyzes gesture action features to obtain music data, and further obtains action data based on the supervision of both music data features and gesture action features. Compared to directly generating action data using gesture data alone, the disclosed embodiment uses the characteristic distribution of music combined with gesture features to derive action features after the gesture data characterizes the music features, which can better ensure that the action data and music data have high-quality fit and matching when used in combination output. The effective fusion of the three features of gesture, music, and action strengthens the relationship between them, ensuring a natural match between action and music beats. The generated action not only conforms to the user's gesture change trend and meets the user's personalized design needs, but also matches the performance of the music very well, making the produced action and music more smooth and coordinated.

[0052] Third, the disclosed embodiments enable users to drive the generation of background music and the movement of target objects in real time through gesture interaction in the real world, enabling users to better integrate into animation scenes, enhance the expressiveness of virtual space, improve the performance effect of animation, enable users to better immerse themselves in the interactive experience, and meet the needs of different users. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] FIG1 is a flow chart of a method for generating an animation provided by an embodiment of the present disclosure;

[0054] FIG2 is a flow chart of another preferred method for generating an animation in an embodiment of the present disclosure;

[0055] FIG3 is a flow chart of another preferred method for generating an animation in an embodiment of the present disclosure;

[0056] FIG4 is a schematic diagram of the principle of the animation generation model in an embodiment of the present disclosure;

[0057] FIG5 is a schematic diagram of a process for training an animation generation model in an embodiment of the present disclosure;

[0058] FIG6 is a schematic structural diagram of an animation generating device provided by an embodiment of the present disclosure;

[0059] FIG7 is a schematic structural diagram of an animation generation device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0061] In the description of this application, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application.

[0062] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include one or more of such features. Throughout this application, unless otherwise specified, "plurality" means two or more.

[0063] In the description of this application, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.

[0064] Referring to FIG1 , which is a flow chart of a method for generating an animation provided by an embodiment of the present disclosure, the present disclosure provides a method for generating an animation, which is specifically performed through steps S11 to S13:

[0065] S11, real-time acquisition of target gestures;

[0066] S12, generating matching music data according to the change of the target gesture action, and generating matching action data according to the change of the target gesture action and the music data;

[0067] S13. Generate music to be played according to the music data, and control the action of the target object according to the action data.

[0068] In an embodiment of the present disclosure, in response to a preset animation generation control instruction, target gesture movements are collected in real time to facilitate subsequent processing and analysis, generate music data and movement data, and generate animation based on the music data and movement data; wherein, the animation includes music and the movement of the target object.

[0069] It should be noted that the animation generated by the embodiments of the present disclosure can be a two-dimensional video or a three-dimensional virtual scene, such as a metaverse scene. When the animation is in the form of a video, the target object is the person in the video; when the animation is in the form of a virtual scene, the target object is the digital human in the virtual scene. The form of the movement can be dance, gymnastics, etc., without specific limitation here.

[0070] In an optional implementation manner, the real-time acquisition of target gestures includes:

[0071] The user's gesture is captured by a preset camera device, and gesture data within a number of unit time intervals is collected to generate the target gesture.

[0072] Specifically, the target gesture action of the embodiment of the present disclosure can be obtained by collecting user actions in real time. The user sets up a corresponding camera device, such as a camera or mobile phone, and performs continuous gestures in front of the camera device in real time. The camera device captures the gestures in real time to generate the target gesture action.

[0073] In another optional implementation manner, the real-time acquisition of target gestures includes:

[0074] The target gesture action is generated by obtaining a pre-collected continuous gesture video, dividing the continuous gesture video into unit frames, and obtaining gesture data within a plurality of unit time intervals.

[0075] Specifically, the target gesture in the disclosed embodiments can be extracted from a pre-recorded video. A user uses a common video capture device, such as a camera or mobile phone, to capture a continuous gesture video. The continuous gesture is then divided into frames, breaking it down into gesture data within a number of discrete time intervals.

[0076] It should be noted that the unit time interval is consistent with the frame time interval of the gesture action during model training. For example, the value is 30ms-100ms (that is, 10 to 30 frames per second).

[0077] Preferably, step S12, i.e., generating matching music data according to the change of the target gesture action, and generating matching action data according to the change of the target gesture action and the music data, includes:

[0078] Extracting features of the target gesture action to obtain an explicit feature vector;

[0079] The explicit feature vector is input into a pre-built animation generation model for calculation to generate the music data and the action data; wherein the animation generation model is used to generate matching music data according to the target gesture action, and to generate matching action data according to the target gesture action and the music data.

[0080] Specifically, after acquiring the target gesture action, the display feature vector of the target gesture action is extracted, including the steady-state position of the gesture, finger strength, finger movement trend and finger change frequency, etc. These explicit feature vectors are input into the preset animation generation model for calculation and analysis. In the embodiment of the present disclosure, an animation generation model is pre-trained, the input value of the model is the gesture action, and the output value is music data and action data. The animation generation model can generate matching music data according to the input gesture action, and after generating the music data, it generates matching action data according to the input gesture action and the generated music data.

[0081] The trained model is used to extract and analyze features of multiple continuous target gesture actions to generate corresponding discrete music data and action data. The discrete music data and action data are merged separately to obtain corresponding continuous music data and continuous action data.

[0082] Finally, the music data and the action data are integrated, the corresponding music is played in the animation scene, and the target object in the animation scene is driven to perform personalized actions with the generated music to form an action performance.

[0083] The present invention provides a method for generating an animation, which has the following beneficial effects:

[0084] First, gestures that are easier to operate, easier to capture, and interactive are used as input features of the model to achieve dual drive of music generation and action generation. Compared with the method of collecting user body movements to generate the actions of the target object, it can simplify the user interaction link and reduce the difficulty of user operation.

[0085] Second, the disclosed embodiment analyzes gesture action features to obtain music data, and further obtains action data based on the supervision of both music data features and gesture action features. Compared to directly generating action data using gesture data alone, the disclosed embodiment uses the characteristic distribution of music combined with gesture features to derive action features after the gesture data characterizes the music features, which can better ensure that the action data and music data have high-quality fit and matching when used in combination output. The effective fusion of the three features of gesture, music, and action strengthens the relationship between them, ensuring a natural match between action and music beats. The generated action not only conforms to the user's gesture change trend and meets the user's personalized design needs, but also matches the performance of the music very well, making the produced action and music more smooth and coordinated.

[0086] Third, the disclosed embodiments enable users to drive the generation of background music and the movement of target objects in real time through gesture interaction in the real world, enabling users to better integrate into animation scenes, enhance the expressiveness of virtual space, improve the performance effect of animation, enable users to better immerse themselves in the interactive experience, and meet the needs of different users.

[0087] As a preferred embodiment, the present disclosure further implements the above embodiment. Referring to FIG2 , which is a flow chart of another preferred animation generation method in the present disclosure, the present disclosure further optimizes the analysis and processing process of the target gesture action. Step S12, i.e., generating matching music data based on the change of the target gesture action, and generating matching action data based on the change of the target gesture action and the music data, includes:

[0088] S121, extracting an explicit feature vector of the target gesture action;

[0089] S122, encoding the explicit feature vector to obtain an implicit feature vector of the target gesture action;

[0090] S123, performing music feature extraction on the implicit feature vector to obtain a music feature vector;

[0091] S124, predicting music data that matches the target gesture action based on the music feature vector;

[0092] S125, extracting motion features from the music feature vector to obtain a motion feature vector;

[0093] S126. Predicting action data that matches the music data based on the action feature vector.

[0094] In the disclosed embodiment, display feature extraction is performed on the target gesture action collected in real time to obtain a display feature vector. Feature encoding is performed on the explicit feature vector to obtain the preceding implicit feature vector h_head of the user gesture. The implicit feature vector h_head of the target gesture action is then processed, and the corresponding predicted value mel_m' of the music spectrum data is output. The implicit feature vector h_head of the target gesture action and the music spectrum data mel_m' are then processed, and the corresponding predicted value y_d' of the action data is output.

[0095] Specifically, the embodiment of the present disclosure uses the animation generation model to perform the above steps S122 to S126. See Figures 3 and 4. Figure 3 is a flow chart of another preferred animation generation method in the embodiment of the present disclosure, and Figure 4 is a schematic diagram of the principle of the animation generation model in the embodiment of the present disclosure. The animation generation model includes a pre-encoder, a music distribution mapper Dis_m, a music decoder D_m, an action distribution mapper Dis_d, and an action decoder D_d. Then step 12 is specifically as follows:

[0096] S121, extracting an explicit feature vector of the target gesture action;

[0097] S122′, inputting the explicit feature vector of the target gesture action into the pre-encoder for encoding to obtain the implicit feature vector of the target gesture action;

[0098] S123′, inputting the implicit feature vector into the music distribution mapper to extract music features to obtain a music feature vector;

[0099] S124′, inputting the music feature vector into the music decoder for decoding, and predicting music data matching the target gesture action;

[0100] S125′, inputting the music feature vector into the action distribution mapper to extract action features to obtain an action feature vector;

[0101] S126': input the action feature vector into the action decoder for decoding, and predict action data that matches the music data.

[0102] In the disclosed embodiment, the explicit feature vector of the target gesture is input into the pre-encoder Eh for feature encoding, obtaining the pre-implicit feature vector h_head of the user gesture. Next, the implicit feature vector h_head is fed into the music distribution mapper Dis_m, which resamples the implicit feature vector, maps it to the latent space distribution generated by the music, and outputs the corresponding music feature vector h_m. The music feature vector h_m is then fed into the music decoder D_m for decoding, obtaining the predicted value mel_m' of the music spectrum data, i.e., the music data. In addition, the music feature vector h_m is sent to the action distribution mapper Dis_d, and the action distribution mapper Dis_d resamples the music feature vector h_m, maps it to the latent space distribution generated by the action, and outputs the corresponding action feature vector h_d; the action feature vector h_d is sent to the action decoder D_d for decoding to obtain the corresponding action data prediction value y_d', that is, the action data, so that according to the music data and the action data, the target object in the animation scene is driven to perform personalized action performances with the generated music.

[0103] By adopting the technical means of the embodiments of the present disclosure, by constructing a model including a pre-encoder, a music distribution mapper, a music decoder, an action distribution mapper and an action decoder, an end-to-end generation algorithm for implicit feature extraction of gesture actions, music latent space modeling, and action latent space modeling is proposed. While decoupling music and action, the correlation between the two is strengthened, making the music and action more matched.

[0104] As a preferred implementation, the embodiment of the present disclosure is further implemented on the basis of any of the above embodiments, and the embodiment of the present disclosure further optimizes the construction and training process of the animation generation model.

[0105] Specifically, referring to FIG5 , which is a flow chart of training an animation generation model in an embodiment of the present disclosure, the animation generation model is trained through the following steps S21 to S25:

[0106] S21, obtaining a training sample data set; wherein the training sample data set includes a plurality of sample element groups; wherein each of the sample element groups is composed of sample gestures, real music data, and real action data;

[0107] S22. Construct an initial machine learning model and initialize the parameters of the machine learning model;

[0108] S23. Extracting an explicit feature vector of the sample gesture action, inputting the explicit feature vector into the machine learning model for calculation to generate corresponding predicted music data based on the sample gesture action, and generating corresponding predicted action data based on the sample gesture action and the predicted music data;

[0109] S24, generating a target loss function based on the real music data, the predicted music data, the real action data, and the predicted action data;

[0110] S25. Update the parameters of the machine learning model according to the target loss function to obtain a trained animation generation model.

[0111] In the embodiment of the present disclosure, a training sample data set is first acquired, and the training sample data set includes several sample element groups (gi, y_mi, y_di). The elements in each of the sample element groups include a sample gesture action gi, real music data y_mi and real action data y_di, indicating that a sample gesture action gi corresponds to the action data y_di and music data y_mi at the i-th moment.

[0112] Optionally, in order to better allow audio data to participate in training, the music data y_m is converted into the frequency domain space based on Fourier transform to obtain the corresponding mel spectrum data mel_m. Therefore, the sample element group is expressed as (gi, mel_mi, y_di), which represents a sample gesture action gi corresponding to the action data y_di at the i-th moment and the mel spectrum graph data mel_mi of the music data.

[0113] Furthermore, an initial machine learning model is constructed and its parameters are initialized. Model training is performed based on each sample element group in the training sample dataset. For each sample element group i, features are extracted from the sample gesture action gi to obtain an explicit feature vector, which is input into the machine learning model to generate corresponding predicted music data mel_m'i and predicted action data y_d'i.

[0114] Preferably, the machine learning model includes a pre-encoder Eh, a music distribution mapper Dis_m, a music decoder D_m, an action distribution mapper Dis_d, and an action decoder D_d. Then step S23 is specifically as follows:

[0115] Extracting an explicit feature vector of the sample gesture action, inputting the explicit feature vector of the sample gesture action into the pre-encoder for encoding, and obtaining an implicit feature vector of the sample gesture action;

[0116] Inputting the implicit feature vector of the sample gesture action into the music distribution mapper for calculation to obtain the music feature vector of the sample gesture action;

[0117] Inputting the music feature vector of the sample gesture action into the music decoder for decoding to obtain the predicted music data;

[0118] Inputting the music feature vector of the sample gesture action into the action distribution mapper for calculation to obtain the action feature vector of the sample gesture action;

[0119] The action feature vector of the sample gesture action is input into the action decoder for decoding to obtain the predicted action data.

[0120] Furthermore, according to the real music data mel_mi, the predicted music data mel_m'i, the real action data y_di and the predicted action data y_d'i, a target loss function L is generated. final . By judging whether the target loss function reaches the preset convergence condition; when the target loss function does not reach the preset convergence condition, the parameters in the machine learning model, including the parameters of the pre-encoder Eh, the music distribution mapper Dis_m, the music decoder D_m, the action distribution mapper Dis_d and the action decoder D_d, are updated, and the updated machine learning model is used to re-analyze and calculate another sample element group i+1 in the training sample data set, and the target loss function is calculated again. In this way, the parameter settings of the machine learning model are continuously adjusted to continuously reduce the target loss function until the value of the target loss function tends to be minimized, that is, when the preset convergence condition is reached, the trained animation generation model is obtained.

[0121] Optionally, the parameter updating method of the machine learning model is: calculating the gradient of the loss function and using the gradient descent method to update the parameters in the machine learning model.

[0122] As a preferred embodiment, step S24, i.e., generating a target loss function based on the real music data, the predicted music data, the real action data, and the predicted action data, includes steps S241 to S243:

[0123] S241, generating a first constrained loss function according to the relationship between the real music data and the predicted music data;

[0124] S242: generating a second constraint loss function according to the relationship between the real motion data and the predicted motion data;

[0125] S243. Generate a target loss function using a weighted summation algorithm based on the first constraint loss function, the second constraint loss function, and corresponding preset weight coefficients.

[0126] Specifically, the first constraint loss function is:

[0127] The second constraint loss function is:

[0128] The objective loss function is:

[0129] L final =αL m +(1-α)L d ;

[0130] Among them, α is the weight coefficient corresponding to the first constraint loss function, and 1-α is the weight coefficient corresponding to the second constraint loss function.

[0131] By adopting the technical means of the embodiments of the present disclosure, the animation generation model is constructed and trained in a supervised manner, thereby improving the matching between gesture actions and the generated music data and action data.

[0132] As a preferred embodiment, the disclosed embodiment further optimizes the means for obtaining the training sample data set. Step S21, i.e., obtaining the training sample data set, includes steps S211 to S215:

[0133] S211, obtaining a first video including background music information and character action information, and dividing the first video into a plurality of first video frames;

[0134] S212, sequentially selecting each of the first video frames, extracting the character's skeletal key points from the character action information of the currently selected first video frame to generate the real action data; and extracting an audio spectrum from the background music information of the currently selected first video frame to generate the real music data;

[0135] S213: Acquire a second video containing continuous hand gestures, and divide the second video into a plurality of second video frames;

[0136] S214: Decompose and analyze the gestures in each second video frame from multiple feature dimensions, and match corresponding sample gestures to the real music data and the real motion data generated for the same first video frame based on the key music information and key motion information represented by each feature dimension of all the gestures, to generate the sample element group;

[0137] S215: Construct the training sample data set according to the generated sample element groups.

[0138] In an embodiment of the present disclosure, at least one first video containing background music information and character movement information is obtained, and at least one second video containing continuous hand gestures is obtained. The first video and the second video are respectively frame-processed to obtain a plurality of first video frames and second video frames.

[0139] Obtain real action data: extract n skeletal key points in each time series from the first video, and define and describe these skeletal key points as the real labels y_d of the action data in model training.

[0140] Obtain real music data: Extract the audio spectrum from the background music of the first video as the real label y_m of the music data in model training. At the same time, in order to better allow the audio data to participate in training, convert it into the frequency domain space based on Fourier transform to obtain the corresponding Mel spectrum data mel_m.

[0141] Obtaining sample gestures: The gestures extracted from the second video are disassembled and labeled. The disassembled feature dimensions include: the amplitude, strength, speed, frequency, movement direction, and / or hand shape of the gesture at each time point. Different feature dimensions are used to represent different key music information of the music data and different key action information of the action data. According to the corresponding representation relationship, the acquired real music data and real action data are labeled to obtain a sample element group (gi, mel_mi, y_di), which represents the action data y_di and the mel spectrum graph data mel_mi of the music data corresponding to a sample gesture gi at the i-th moment.

[0142] Preferably, the characteristic dimensions include amplitude strength, speed frequency, movement direction and hand shape.

[0143] The key music information represented by the amplitude and strength of the gesture action includes the volume and strength of the music data, and the key action information represented includes the action range and strength of the action data; the key music data represented by the speed and frequency of the gesture action includes the rhythm and speed of the music data, and the key action information represented includes the rhythm of the action data; the key music information represented by the moving direction of the gesture action includes the pitch of the notes of the music data, and the key action information represented includes the spatial position and moving direction of the action data; the key music data represented by the hand shape of the gesture action includes the note effect of the music data, and the key action information represented includes the action content of the action data.

[0144] As an example, the correspondence between gesture actions and key music information of music data is as follows:

[0145] Regarding amplitude and force: The amplitude and force of gestures can be mapped to the volume and strength of the music. Larger amplitude and force gestures can correspond to higher volume and stronger musical performances, while smaller amplitude and force gestures can correspond to lower volume and softer musical performances.

[0146] Regarding speed and frequency: The speed and frequency of gestures can be mapped to the rhythm and tempo of the music. Fast and continuous gestures can correspond to fast-paced and active music, while slow and intermittent gestures can correspond to slow-paced and calm music.

[0147] Regarding movement direction: The vertical position and direction of a gesture can be mapped to the pitch of musical notes. A higher position of the gesture can correspond to a higher note, while a lower position of the gesture can correspond to a lower note. The up and down movement of the gesture can correspond to the rising and falling of the note.

[0148] Regarding hand morphology: The shape of the hand gestures, the posture and movement of the fingers can be mapped to musical note effects. For example, stroking the fingers can produce continuous note effects, while bending and twisting the hand gestures can produce note variations and glissando effects.

[0149] The correspondence between gesture actions and key action information of action data is as follows:

[0150] Regarding amplitude and strength: The amplitude of a gesture can be mapped to the range and strength of the motion data. Larger gestures can correspond to a greater range and strength within the action, while smaller gestures correspond to light or subtle movements within the action.

[0151] Regarding speed and frequency: the speed of gestures can be mapped to the rhythm of the motion data. Fast gestures can be mapped to fast-paced movements in the motion, while slow gestures can be mapped to slow-paced movements in the motion.

[0152] Regarding movement direction: The position and direction of a gesture can be mapped to the position and movement of the action in space. The up and down position of the gesture can be mapped to the height of the body part in the action, while the left and right position can be mapped to the lateral movement in the action.

[0153] Regarding hand shape: The shape of the gesture and the posture of the fingers can be mapped to different actions. For example, the opening and closing of the gesture can be mapped to the expansion and contraction of the action, and the bending and straightening of the fingers can be mapped to the gesture expression of the action.

[0154] Therefore, based on the above correspondence, the acquired real music data and real action data are analyzed and labeled, so as to match the corresponding sample gestures and construct the training sample data set.

[0155] By adopting the technical means of the embodiments of the present disclosure, during the model training process, the key music features and key action features represented by the gesture movements are labeled and matched with the action data and music data extracted from the video to obtain a training sample data set, which effectively improves the accuracy of the construction and training of the animation generation model.

[0156] 6 is a schematic diagram of the structure of an animation generation device provided by an embodiment of the present disclosure. The present disclosure provides an animation generation device 30, including:

[0157] Data acquisition module 31, used for real-time acquisition of target gestures;

[0158] a data calculation module 32 for generating matching music data according to the change of the target gesture action, and generating matching action data according to the change of the target gesture action and the music data;

[0159] The data control module 33 is used to generate music to be played according to the music data and control the action of the target object according to the action data.

[0160] The present invention provides an animation generation device with the following advantages:

[0161] First, gestures that are easier to operate, easier to capture, and interactive are used as input features of the model to achieve dual drive of music generation and action generation. Compared with the method of collecting user body movements to generate the actions of the target object, it can simplify the user interaction link and reduce the difficulty of user operation.

[0162] Second, the disclosed embodiment analyzes gesture action features to obtain music data, and further obtains action data based on the supervision of both music data features and gesture action features. Compared to directly generating action data using gesture data alone, the disclosed embodiment uses the characteristic distribution of music combined with gesture features to derive action features after the gesture data characterizes the music features, which can better ensure that the action data and music data have high-quality fit and matching when used in combination output. The effective fusion of the three features of gesture, music, and action strengthens the relationship between them, ensuring a natural match between action and music beats. The generated action not only conforms to the user's gesture change trend and meets the user's personalized design needs, but also matches the performance of the music very well, making the produced action and music more smooth and coordinated.

[0163] Third, the disclosed embodiments enable users to drive the generation of background music and the movement of target objects in real time through gesture interaction in the real world, enabling users to better integrate into animation scenes, enhance the expressiveness of virtual space, improve the performance effect of animation, enable users to better immerse themselves in the interactive experience, and meet the needs of different users.

[0164] As a preferred embodiment, the data calculation module 32 is specifically configured to:

[0165] Extracting an explicit feature vector of the target gesture action;

[0166] Encoding the explicit feature vector to obtain an implicit feature vector of the target gesture action;

[0167] Performing music feature extraction on the implicit feature vector to obtain a music feature vector;

[0168] Predicting music data that matches the target gesture action based on the music feature vector;

[0169] Extracting motion features from the music feature vector to obtain a motion feature vector;

[0170] The action data matching the music data is predicted based on the action feature vector.

[0171] As a preferred embodiment, the data calculation module 32 is specifically configured to:

[0172] Extracting an explicit feature vector of the target gesture action;

[0173] The explicit feature vector is input into a pre-built animation generation model for calculation to generate the music data and the action data; wherein the animation generation model is used to generate matching music data according to the target gesture action, and to generate matching action data according to the target gesture action and the music data.

[0174] As a preferred embodiment, the animation generation model is trained by the following steps:

[0175] Acquire a training sample data set; wherein the training sample data set includes a plurality of sample element groups; wherein each of the sample element groups is composed of sample gestures, real music data, and real action data;

[0176] Constructing an initial machine learning model and initializing the parameters of the machine learning model;

[0177] Extracting an explicit feature vector of the sample gesture action, inputting the explicit feature vector into the machine learning model for calculation to generate corresponding predicted music data based on the sample gesture action, and generating corresponding predicted action data based on the sample gesture action and the predicted music data;

[0178] generating a target loss function based on the real music data, the predicted music data, the real action data, and the predicted action data;

[0179] According to the target loss function, the parameters of the machine learning model are updated to obtain a trained animation generation model.

[0180] Preferably, the acquiring of a training sample data set includes:

[0181] Acquire a first video containing background music information and character action information, and divide the first video into a plurality of first video frames;

[0182] Selecting each of the first video frames in sequence, extracting the character's skeletal key points from the character motion information of the currently selected first video frame to generate the real motion data; and extracting an audio spectrum from the background music information of the currently selected first video frame to generate the real music data;

[0183] Acquire a second video containing continuous hand gestures, and divide the second video into a plurality of second video frames;

[0184] Decomposing and analyzing the gestures in each second video frame from multiple feature dimensions, and matching corresponding sample gestures to the real music data and the real motion data generated for the same first video frame based on the key music information and key motion information represented by each feature dimension of all the gestures, to generate the sample element group;

[0185] The training sample data set is constructed based on the generated several sample element groups.

[0186] Preferably, the characteristic dimensions include amplitude, speed, frequency, movement direction and / or hand shape;

[0187] The key music information represented by the amplitude and strength of the gesture action includes the volume and strength of the music data, and the key action information represented includes the action range and strength of the action data;

[0188] The key music data represented by the speed frequency of the gesture action includes the rhythm and speed of the music data, and the key action information represented includes the rhythm of the action data;

[0189] The key music information represented by the movement direction of the gesture action includes the pitch of the musical notes of the music data, and the key action information represented includes the spatial position and movement direction of the action data;

[0190] The key music data represented by the hand shape of the gesture action includes the note effects of the music data, and the key action information represented includes the action content of the action data.

[0191] Preferably, generating a target loss function based on the real music data, the predicted music data, the real action data and the predicted action data includes:

[0192] generating a first constrained loss function according to the relationship between the real music data and the predicted music data;

[0193] generating a second constraint loss function according to a relationship between the real motion data and the predicted motion data;

[0194] According to the first constraint loss function, the second constraint loss function and the corresponding preset weight coefficients, a weighted summation algorithm is adopted to generate a target loss function.

[0195] It should be noted that the animation generation device provided in the embodiment of the present disclosure is used to execute all the process steps of the animation generation method in the above embodiment. The working principles and beneficial effects of the two correspond one to one, and thus will not be described in detail.

[0196] Refer to Figure 7, which is a structural diagram of an animation generation device provided by an embodiment of the present disclosure. The embodiment of the present disclosure provides an animation generation device 40, including a processor 41, a memory 42, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the animation generation method described in any one of the above embodiments.

[0197] It should be noted that the animation generation device provided in the embodiment of the present disclosure is used to execute all the process steps of the animation generation method in the above embodiment. The working principles and beneficial effects of the two correspond one to one, and thus will not be described in detail.

[0198] An embodiment of the present disclosure further provides a computer-readable storage medium, which includes a stored computer program. When the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the animation generation method as described in any one of the above embodiments.

[0199] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0200] The above is a preferred embodiment of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present disclosure. These improvements and modifications are also considered to be within the scope of protection of the present disclosure.

Claims

1. A method for generating an animation, comprising: Collect target gestures in real time; Generate matching music data according to the change of the target gesture action, and generate matching action data according to the change of the target gesture action and the music data; The music to be played is generated according to the music data, and the action of the target object is controlled according to the action data.

2. The method for generating an animation according to claim 1, wherein: The generating of matching music data according to the change of the target gesture action, and generating matching action data according to the change of the target gesture action and the music data, includes: Extracting an explicit feature vector of the target gesture action; Encoding the explicit feature vector to obtain an implicit feature vector of the target gesture action; Performing music feature extraction on the implicit feature vector to obtain a music feature vector; Predicting music data matching the target gesture action according to the music feature vector; Extracting action features from the music feature vector to obtain an action feature vector; The action data matching the music data is predicted based on the action feature vector.

3. The method for generating an animation according to claim 1, wherein: The generating of matching music data according to the change of the target gesture action, and generating matching action data according to the change of the target gesture action and the music data, includes: Extracting an explicit feature vector of the target gesture action; The explicit feature vector is input into a pre-built animation generation model for calculation to generate the music data and the action data; wherein the animation generation model is used to generate matching music data according to the target gesture action, and to generate matching action data according to the target gesture action and the music data.

4. The method for generating an animation according to claim 3, wherein: The animation generation model is trained by the following steps: Acquire a training sample data set; wherein the training sample data set includes a plurality of sample element groups; wherein each of the sample element groups is composed of sample gesture actions, real music data and real action data; Constructing an initial machine learning model and initializing parameters of the machine learning model; Extracting an explicit feature vector of the sample gesture action, inputting the explicit feature vector into the machine learning model for calculation, so as to generate corresponding predicted music data according to the sample gesture action, and generating corresponding predicted action data according to the sample gesture action and the predicted music data; Generate a target loss function according to the real music data, the predicted music data, the real action data and the predicted action data; According to the target loss function, the parameters of the machine learning model are updated to obtain a trained animation generation model.

5. The method for generating an animation according to claim 4, wherein: The step of obtaining a training sample data set includes: Acquire a first video including background music information and character action information, and divide the first video into a plurality of first video frames; Selecting each of the first video frames in turn, extracting the skeleton key points of the character from the character action information of the currently selected first video frame to generate the real action data; and extracting the audio spectrum from the background music information of the currently selected first video frame to generate the real music data; Acquire a second video containing continuous gesture actions, and divide the second video into a plurality of second video frames; Decomposing and analyzing the gesture actions in each of the second video frames from multiple feature dimensions, and matching the corresponding sample gesture actions to the real music data and the real action data generated for the same first video frame according to the key music information and key action information represented by each feature dimension of all the gesture actions, so as to generate the sample element group; The training sample data set is constructed based on the generated plurality of sample element groups.

6. The method for generating an animation according to claim 5, wherein: The characteristic dimensions include amplitude, speed, frequency, movement direction and / or hand shape; The key music information represented by the amplitude and strength of the gesture action includes the volume and strength of the music data, and the key action information represented includes the action range and strength of the action data; The key music data represented by the speed frequency of the gesture action includes the rhythm and speed of the music data, and the key action information represented includes the rhythm of the action data; The key music information represented by the moving direction of the gesture action includes the pitch of the musical notes of the music data, and the key action information represented includes the spatial position and moving direction of the action data; The key music data represented by the hand shape of the gesture action includes the note effects of the music data, and the key action information represented includes the action content of the action data.

7. The method for generating an animation according to claim 5, wherein: The generating a target loss function according to the real music data, the predicted music data, the real action data and the predicted action data comprises: Generate a first constraint loss function according to the relationship between the real music data and the predicted music data; generating a second constraint loss function according to the relationship between the real action data and the predicted action data; According to the first constraint loss function, the second constraint loss function and the corresponding preset weight coefficient, a weighted summation algorithm is adopted to generate a target loss function.

8. An animation generating device, comprising: Data acquisition module, used to collect target gestures in real time; A data calculation module, used to generate matching music data according to the change of the target gesture action, and to generate matching action data according to the change of the target gesture action and the music data; The data control module is used to generate music to be played according to the music data and control the action of the target object according to the action data.

9. An animation generation device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the animation generation method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, the computer-readable storage medium comprising a stored computer program, wherein: When the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the animation generation method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for generating music

    CN109413351A

  • Digital human animation evaluation optimization method and device, equipment, medium and product

    CN115620096A

  • Digital human gesture generation method and device, equipment and storage medium

    CN116524074A

  • Animation control information construction method and device, equipment, medium and product

    CN116543077A

  • Animation generation method and device, equipment and medium

    CN117710536A