Digital human animation evaluation optimization method and device, equipment, medium and product

By generating and pushing multiple digital human dance animations in the live broadcast room, and using user evaluation data to select the best motion generation model and samples, the problem of data scarcity is solved, and the quality of digital human animations and user experience are improved.

CN115620096BActive Publication Date: 2026-04-28GUANGZHOU HUANJU SHIDAI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU HUANJU SHIDAI INFORMATION TECH CO LTD
Filing Date
2022-10-24
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, due to the scarcity of data samples required for training digital humans and the high production costs, motion generation models often fail to achieve good results, frequently exhibiting motion anomalies such as clipping, static images, and unsmooth movements.

Method used

By acquiring the audio data feature sequence of the live broadcast room, a dance animation is generated using multiple converged motion generation models and simultaneously pushed to users in the live broadcast room. User evaluation data is received, and the best dance animation and motion generation model are selected. The audio feature sequence and motion sequence information of the best dance animation are expanded into samples in the training dataset, and the motion generation model is iteratively trained.

Benefits of technology

The training data quality of the motion generation model was improved, the generated dance animations were more in line with user perception, the user experience was enhanced, and high-quality data samples were dynamically generated through user evaluations, which enriched the training dataset and improved the model's generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620096B_ABST
    Figure CN115620096B_ABST
Patent Text Reader

Abstract

The application relates to a digital human animation evaluation optimization method and device, equipment, medium and product, the method comprising: acquiring audio feature sequences of audio data of on-air music in a live broadcast room, using a plurality of converged action generation models to respectively generate action sequence information corresponding to the audio feature sequences, and obtaining action sequence information corresponding to each of a plurality of digital humans; applying each action sequence information to the corresponding digital human, synchronously pushing corresponding dance animation to the live broadcast room for playing, and receiving user evaluation data of the dance animation; determining the best dance animation and the best action generation model according to the user evaluation data; and expanding the audio feature sequences and the action sequence information corresponding to the best dance animation into data samples in a training data set of the best action generation model. The application can determine high-quality dance animation, thereby improving the inference ability of the action generation model, saving the continuous evolution cost of the model, and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to digital human virtual technology, and more particularly to a digital human animation evaluation and optimization method, apparatus, equipment, medium, and product thereof. Background Technology

[0002] With the rise of the "metaverse," the introduction of "digital humans" into live streaming scenarios has become a standard feature of live streaming applications, with major manufacturers and live streaming platforms successively launching virtual "digital human" functions. Live streaming using "digital humans" has become a popular choice for many streamers. In live streaming scenarios, "digital humans" offer many new possibilities, with using external materials such as sound, motion, and music to drive their movement becoming a novel and popular feature.

[0003] Currently, solutions for driving "digital humans" using external materials typically employ machine learning to model a motion generation model. This model is then trained to convergence using data samples, making it suitable for generating corresponding motion sequence information for the target digital human. This motion sequence information is then applied to the digital human and rendered to generate the corresponding animation. The motion sequence information includes motion control data corresponding to each image frame in the "digital human's" motion images, forming information frames corresponding to each image frame. The motion control data in each information frame corresponds to the control data provided by each skeletal node of the digital human. By using the control data of each skeletal node to instruct the drawing of the "digital human's" posture, and by continuously drawing multiple image frames, the animation of the digital human can be obtained.

[0004] In practice, the scarcity of data samples required for training digital humans and the high cost of sample production often lead to motion generation models failing to meet expectations, resulting in anomalies such as clipping, static images, and unsmooth movements. How to comprehensively improve the information quality at each stage of the motion generation model application process to enhance digital human animation warrants in-depth research. Summary of the Invention

[0005] The purpose of this application is to solve the above-mentioned problems by providing a digital human animation evaluation and optimization method, as well as corresponding devices, equipment, non-volatile readable storage media, and computer program products.

[0006] According to one aspect of this application, a method for evaluating and optimizing digital human animation is provided, comprising the following steps:

[0007] The audio feature sequence of the music playing in the live room is obtained, and multiple converged action generation models are used to generate action sequence information corresponding to the audio feature sequence, thereby obtaining the corresponding action sequence information of multiple digital humans.

[0008] Each of the aforementioned action sequence information is applied to its corresponding digital human to generate a corresponding dance animation for each of the multiple digital humans, which is then synchronously pushed to the live streaming room for playback. User evaluation data applied to the dance animation in the live streaming room is received.

[0009] Based on the user evaluation data, determine the best dance animation among the dance animations and the best motion generation model among the converged motion generation models;

[0010] The audio feature sequence and motion sequence information corresponding to the best dance animation are amplified into data samples in the training dataset of the best motion generation model.

[0011] Optionally, generating corresponding dance animations for each of the multiple digital humans and simultaneously pushing them to the live stream for playback includes:

[0012] Initialize the poses and movements of multiple digital humans;

[0013] Establish a unified coordinate system among the multiple digital humans;

[0014] Based on the unified coordinate system, the sequence information of each action is applied to its corresponding digital human, so that the dance animation generated by each digital human keeps the actions synchronized.

[0015] Optionally, after generating the corresponding dance animations for each of the multiple digital humans and synchronously pushing them to the live stream for playback, the process includes:

[0016] The graphical user interface of the live broadcast room displays dance animations corresponding to the multiple digital humans, and the dance animations are synchronized with the music being played.

[0017] The graphical user interface displays multiple scoring controls corresponding to the multiple dance animations, which are used to obtain the corresponding subjective scores of users for each dance animation.

[0018] The interaction information generated by the users in the live broadcast room interacting in the graphical user interface is obtained, and the corresponding user evaluation data of each digital human is determined according to the interaction information of each digital human. The interaction information includes any one or more of the following: subjective score, bullet screen text, and chat area text.

[0019] Optionally, determining the best dance animation in the dance animation and the best motion generation model in the converged motion generation models based on the user evaluation data includes:

[0020] The corresponding rating score for each dance animation is determined based on the user review data for each dance animation.

[0021] The standard deviation was determined based on the corresponding evaluation scores of all dance animations;

[0022] Dance animations with evaluation scores higher than the standard deviation were selected as the best dance animations.

[0023] The converged motion generation model that produces the most and best dance animations is selected as the best motion generation model.

[0024] Optionally, the rating score for each dance animation can be determined based on user review data, including:

[0025] Input the text type data from the user review data corresponding to each dance animation into a preset rating prediction model to predict the first rating. The text type data includes any one or more of the following: bullet screen text and chat area text.

[0026] The subjective score from the user review data corresponding to each dance animation is used as the second rating.

[0027] The first and second scores of each dance animation are weighted and combined to obtain the evaluation score of the dance animation.

[0028] Optionally, after amplifying the audio feature sequence and motion sequence information corresponding to the best dance animation into data samples in the training dataset of the best motion generation model, the process includes:

[0029] The training of the action generation model is restarted using data samples from the training dataset, and the action generation model is trained to a convergent state, which serves as an upgraded action generation model.

[0030] The upgraded motion generation model is configured as a callable service of the live streaming room, used to generate corresponding motion sequence information for the audio feature sequence of the audio data of the music playing in the live streaming room, and to generate the dance animation of the digital human in the live streaming room based on the motion sequence information.

[0031] Optionally, restarting the training of the action generation model using data samples from the training dataset includes:

[0032] A single data sample is retrieved from the training dataset, the audio feature sequence in the data sample is determined as a training sample, and the action sequence information corresponding to the audio data is determined as a supervision label;

[0033] The training samples are input into the motion generation model to predict motion sequence information used to control the digital human to generate motion images;

[0034] Based on the supervised labels, the loss value of the action sequence information obtained by the action generation model is calculated. Based on the loss value, a decision is made on whether to continue iterative training of the action generation model until it reaches a convergence state and constitutes an upgraded action generation model.

[0035] According to another aspect of this application, a digital human animation evaluation and optimization apparatus is provided, comprising:

[0036] The action generation module is configured to acquire the audio feature sequence of the audio data of the music playing in the live room, and use multiple converged action generation models to generate action sequence information corresponding to the audio feature sequence, thereby obtaining the corresponding action sequence information of multiple digital humans.

[0037] The animation push module is configured to apply each of the action sequence information to its corresponding digital human, generate the corresponding dance animations of the multiple digital humans, and push them to the live broadcast room for playback, and receive user evaluation data in the live broadcast room that affects the dance animations.

[0038] The evaluation processing module is configured to determine the best dance animation in the dance animation and the best motion generation model in the converged motion generation model based on the user evaluation data.

[0039] The sample augmentation module is configured to augment the audio feature sequence and motion sequence information corresponding to the best dance animation into data samples in the training dataset of the best motion generation model.

[0040] According to another aspect of this application, a digital human animation evaluation and optimization device is provided, including a central processing unit and a memory, wherein the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the digital human animation evaluation and optimization method described in this application.

[0041] According to another aspect of this application, a non-volatile readable storage medium is provided, which stores a computer program implemented according to the digital human animation evaluation and optimization method in the form of computer-readable instructions, wherein the computer program, when invoked by a computer, executes the steps included in the method.

[0042] According to another aspect of this application, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in any embodiment of this application.

[0043] Compared with existing technologies, this application has several technological advantages, including but not limited to:

[0044] First, this application prepares multiple converged motion generation models. These models are used to generate motion sequence information for their corresponding digital humans, and dance animations are generated for each digital human based on the motion sequence information. These dance animations are then simultaneously displayed in a live broadcast room. User evaluation data for each dance animation is obtained through the live broadcast room. The best dance animation and the best motion generation model are selected based on the user evaluation data. The audio feature sequence of the audio data corresponding to the best dance animation and the motion sequence information generated therefrom are constructed as data samples and amplified into the training dataset used by the best motion generation model. Since the training dataset can train the best motion generation model, it indicates that it has excellent sample quality. On this basis, the data samples generated by the user evaluation selection are further amplified, which further improves the overall sample quality of the training dataset. Thus, the amplified training dataset can be used to retrain the motion generation model, iteratively obtaining an upgraded version of the motion generation model. This process is repeated to continuously improve the quality of the generated digital human dance animations.

[0045] Secondly, given the scarcity of digital human-related data samples, this application first uses an initially trained, converged motion generation model to generate dance animations for digital humans. Then, based on user evaluation data of these dance animations, it selects the best dance animation from each one and creates data samples accordingly. Thus, at a lower cost, it dynamically generates a massive amount of high-quality data samples using user evaluation data. These data samples are diverse due to differences in audio data, motion generation models, and user evaluation data, exhibiting excellent feature generalization effects. This enriches the training dataset required for training the motion generation model and makes it more suitable for retraining the motion generation model.

[0046] Furthermore, this application obtains dance animations of multiple digital humans on the backend, while displaying the corresponding dance animations of multiple digital humans on the graphical user interface of the live broadcast room on the frontend. This allows users to compare multiple dance animations and provide valuable user evaluation data after comparison. This evaluation data is essentially subjective evaluation data, unlike simple data cleaning through algorithms. With the help of subjective evaluation data, the selected dance animations are of higher quality, and therefore the audio feature sequences and motion sequence information of their corresponding audio data are of higher quality. The motion generation model trained in this way, when used to produce motion information sequences and generate dance animations, will inevitably produce dance animations that are more in line with user perception and improve user experience. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a schematic diagram of the network architecture of an exemplary application environment for this application;

[0049] Figure 2 This is a schematic diagram illustrating an exemplary digital human and the distribution of its skeletal key points according to this application;

[0050] Figure 3 This is a flowchart illustrating one embodiment of the digital human animation evaluation and optimization method of this application.

[0051] Figure 4 This is a schematic diagram illustrating the process of controlling the synchronous generation of dance animations of multiple digital humans in an embodiment of this application;

[0052] Figure 5 This is a schematic diagram illustrating the initial pose of an exemplary digital human according to this application;

[0053] Figure 6 This is a schematic diagram illustrating the process of a terminal device displaying dance animation and obtaining user evaluation data in an embodiment of this application;

[0054] Figure 7 This is a schematic diagram illustrating the process of determining the best dance animation and the best motion generation model in the embodiments of this application;

[0055] Figure 8 This is a schematic diagram illustrating the process of determining the evaluation score of a dance animation in an embodiment of this application;

[0056] Figure 9 This is a flowchart illustrating the optimization of the action generation model in an embodiment of this application;

[0057] Figure 10 This is a flowchart illustrating the upgrade training process of an exemplary action generation model in an embodiment of this application.

[0058] Figure 11 This is a schematic diagram of the digital human animation evaluation and optimization device of this application;

[0059] Figure 12 This is a schematic diagram of the structure of a digital human animation evaluation and optimization device used in this application. Detailed Implementation

[0060] Please see Figure 1This application describes an exemplary application scenario using a network architecture including a terminal device 80, a media server 81, and an application server 82. The application server 82 can be used to deploy a live streaming service, and the media server 81 can be used to generate a video stream corresponding to a digital human based on action sequence information. The action sequence information can be generated by calling a trained and converged action generation model, inputting an audio feature sequence of audio data, and then extracting semantic features from the audio feature sequence and inferring its generation. When a user accesses the live streaming service provided by the application server 82 from their terminal device 80, enters the corresponding live stream room, and begins to control the output video stream corresponding to the digital human, they can trigger the generation of action sequence information for controlling the digital human's movement. For example, by playing music, providing a currently playing song to determine its audio data, obtaining the audio feature sequence corresponding to the audio data, and then calling the action generation model to generate the action sequence information based on the audio feature sequence. Then, the media server 81 generates a corresponding dance animation based on the action sequence information, and, under the coordination of the application server 82, pushes the dance animation to the live stream room as a video stream, thereby providing a live streaming service for the digital human.

[0061] The action sequence information can be generated by a pre-trained action generation model that has reached a convergent state, based on audio feature sequences. The action generation model can be a machine learning model or a deep learning model, which is pre-trained using data samples from a corresponding training dataset. These data samples include training samples and corresponding supervision labels. The training samples can be audio feature sequences, and the supervision labels can be action sequence information. Each action generation model can specifically generate its corresponding action sequence information for a given digital human.

[0062] The motion sequence information includes information frames corresponding to each image frame in the digital human dance animation to be generated. Each information frame stores control data describing the digital human in the corresponding image frame within the dance animation. This control data is described by motion vectors or static position data of the digital human's corresponding preset skeletal key points. The specific data format can be standardized to adapt to the output format of the motion generation model or arbitrarily preset. For example, in one embodiment, when the motion sequence information represents each information frame, it can provide the offset corresponding to the rotation values ​​of each preset skeletal key point on the three axes of the image coordinate system. In another embodiment, when the motion sequence information represents each information frame, it can provide the position coordinates of each preset skeletal key point in the image coordinate system. Regardless of the form used, the control data of all skeletal key points can be represented as vectors in one information frame. Thus, the entire motion sequence information is described as a matrix with time sequence as the row index. It can be seen that as long as the internal data structure of the motion sequence information is defined in advance, the motion sequence information can be constructed using any parsable data organization method.

[0063] The key skeletal features of the digital human, such as Figure 2 As shown in the example, these key points are mainly distributed in the various skeletal joints of the digital human's body and limbs. The displacement of these key skeletal points in the image coordinate system can usually drive the corresponding body parts of the digital human to produce corresponding motion effects. By progressively adjusting the positional information of the same set of key skeletal points in multiple information frames, the corresponding image frames that produce gradual changes in motion of the digital human can be controlled accordingly. By playing the video stream composed of these image frames in an orderly manner, the motion effects of the digital human body parts corresponding to the corresponding key skeletal points can be visually presented.

[0064] The digital human can be pre-modeled so that, according to the position information of each skeletal key point in the information frame, the image frame position and viewpoint of the body parts corresponding to each skeletal key point can be adjusted. This allows for the subsequent generation of corresponding digital human image frames through rendering, and the dance animation can be constructed from these sequentially generated image frames. Furthermore, depending on actual needs, corresponding background images or other foreground images can be added when generating the digital human image frames. Those skilled in the art can implement this flexibly.

[0065] Please see Figure 3 According to a digital human animation evaluation and optimization method provided in this application, one embodiment includes the following steps:

[0066] Step S1100: Obtain the audio feature sequence of the audio data of the music playing in the live room, and use multiple converged action generation models to generate action sequence information corresponding to the audio feature sequence, thereby obtaining the corresponding action sequence information of multiple digital humans.

[0067] In an exemplary application scenario, to control the generation of dance animations from multiple digital humans, a motion generation model is trained to convergence for each digital human. This model generates motion sequence information for the corresponding digital human. Using control data from each information frame in the motion sequence information, the digital human is controlled to set the position and viewpoint of its skeletal key points in three-dimensional space frame by frame and render the corresponding image frames. Organizing these image frames sequentially yields the corresponding dance animation. All converged motion generation models can be configured as callable services in a live streaming service, allowing the media server to invoke one or more converged motion generation models based on the settings and triggers of each live streaming room. These models are then used to generate motion sequence information corresponding to multiple digital humans based on the music playing in the live streaming room.

[0068] In one embodiment, when the live stream is open, the broadcaster can specify a piece of music for playback, which will become the playing music when it is played. Whether the user-specified playing music is played from a file on the broadcaster's local machine or from cloud storage on the server, the audio data of the playing music will be pushed to the live stream via the media server to reach all viewers in the live stream.

[0069] When the media server obtains the audio data of the playing music, it can perform audio preprocessing algorithms to extract the audio feature sequences. For example, the audio data can be windowed and framed to obtain multiple data frames. A Fast Fourier Transform (FFT) is then performed on each data frame to transform it from the time domain to the frequency domain. A Mel-scale filter is then applied to each data frame, followed by a logarithmic transform to obtain the Mel spectrum. Audio features are extracted from the Mel spectrum to obtain an audio feature sequence. This audio feature sequence can be provided to the action generation model for inference to obtain the corresponding digital human's action sequence information.

[0070] A piece of music is usually several minutes long. To address this, the audio data of the music being played can be acquired according to a certain preset duration, and the audio feature sequences of each segment of the audio data can be extracted and input into the corresponding action generation model in order to generate the action sequence information of the music being played segment by segment.

[0071] In other embodiments, the features of the Mel spectrum can be replaced by CQT features, Chroma features, temporal features, etc., of the audio data of the playing audio to obtain the audio feature sequence of the audio data. In short, as long as the corresponding audio features can effectively describe the style-invariant features in the audio data of the playing music, especially the rhythm information therein, they can be used as the audio features required by this application to construct the corresponding audio feature sequence.

[0072] The audio feature sequence extracted from the audio data of the playing music is further synchronously input into multiple motion generation models pre-trained to convergence in this application. Each motion generation model, based on its pre-trained reasoning ability, infers and generates motion sequence information for controlling the digital human to generate motion images based on the audio feature sequence.

[0073] The action sequence information stores multiple information frames corresponding to a video generated by the digital human. Each information frame corresponds to the position and viewpoint control data of key skeletal points in the three-dimensional model of the digital human. For example... Figure 2 As an example, a 3D model of a digital human has 24 skeletal key points. Therefore, the action sequence information stores the control data corresponding to these 24 skeletal key points in different image frames of the video stream. The image information set consisting of the control data of all 24 skeletal key points corresponding to each image frame is regarded as an information frame. Usually, there is a one-to-one correspondence between information frames and image frames according to the time sequence.

[0074] Since the digital human is a 3D model obtained through 3D modeling, the control data can generally be represented in a 3D coordinate system. This allows the control data representation to be quickly mapped to the 3D model of the digital human, facilitating the actual image transformation and adjustment of the corresponding skeletal key points in the 3D model of the digital human.

[0075] In one embodiment, the control data can be described by the relative position information between skeletal keypoints. For example, taking a certain skeletal keypoint on the body of the digital human's 3D model as a reference point, the position information of the reference point is represented as [0,0,0]. The position information of other skeletal keypoints can be represented as [Δx_(i,j),Δy_(i,j),Δz_(i,j)], where Δ represents the coaxial relative offset, i represents the sequence number of the information frame, and j represents the sequence number of the skeletal keypoint. In another embodiment, the position information of each skeletal keypoint can also be described by the absolute position information of a 3D coordinate system. For example, the position information of a skeletal keypoint can be represented as [x_(i,j),y_(i,j),z_(i,j)].

[0076] In practice, any method can be used to represent the internal data structure of the action sequence information. By using a pre-defined data structure, the action generation model can represent the control data of all skeletal keypoints in all information frames of the action sequence information, thereby representing the corresponding position and viewpoint information, which is equivalent to determining the corresponding posture information. Regardless of the method used to represent the control data of the skeletal keypoints, as long as it can be parsed and utilized by this application, it is acceptable.

[0077] Step S1200: Apply each of the action sequence information to its corresponding digital human to generate the corresponding dance animations of the multiple digital humans and push them to the live broadcast room for playback, and receive user evaluation data in the live broadcast room that affects the dance animations;

[0078] After the media server obtains the corresponding motion sequence information for each motion generation model, it can apply one of the motion sequence information to each digital human. This allows each digital human to change the position and perspective of its skeletal key points according to the control data of each information frame in its corresponding motion sequence information, thereby generating multiple image frames to present different movement postures. These image frames are then synthesized sequentially to form a complete dance animation. In one embodiment, when generating the image frames corresponding to the various information frames, background images and / or decorative images can be added to the digital human to make the scene more realistic.

[0079] In one embodiment, when generating the dance animation, a media server can synthesize the audio data of the playing music into it, and ensure that the audio and dance animation are synchronized in rhythm according to the temporal correspondence between the action sequence information and the audio data, thereby obtaining a corresponding live stream. The live stream is then pushed to the live room of the broadcaster. After the viewers in the live room obtain the live stream, the player in the live room parses and plays it, thereby obtaining an audio and video playback effect where the music and the digital human's movements are synchronized, thus realizing a virtual live streaming service.

[0080] In another embodiment, the rhythm synchronization information between the dance animation and the audio data of the playing music, the audio data, and each dance animation can be synchronously pushed to the terminal device of the audience user in the live broadcast room, and the terminal device can synchronously play the audio data and each dance animation according to the rhythm synchronization information.

[0081] In another embodiment, whether on a media server or a terminal device, when processing the display relationship of multiple digital human dance animations generated based on the same audio data, each dance animation can be set up and displayed side by side in different windows of the same layout, so that viewers can intuitively compare the imaging quality between different dance animations.

[0082] It's easy to understand that when the dance animation is pushed to the live stream, the visual effect of multiple digital humans simultaneously performing the dance animation will appear on the terminal device. Based on this, user evaluation data from each viewer in the live stream can be monitored for their interaction with any one of the dance animations. This user evaluation data can be extracted from user chat text generated in the live stream's chat flow or obtained using a dedicated rating control. In short, on the server side, whether it's the media server or the application server, the evaluation score for each dance animation can be quantified based on the user evaluation data, allowing for the selection of the best animation based on its score.

[0083] Step S1300: Determine the best dance animation in the dance animation and the best motion generation model in the converged motion generation model based on the user evaluation data;

[0084] For dance animations of multiple digital humans generated from a single audio data, user evaluation data generated for each dance animation in the live broadcast room is first obtained. Then, based on statistical methods, the evaluation score of the dance animation is determined according to the multiple user evaluation data of the same dance animation. This realizes the quantification of the user's subjective feeling for the corresponding dance animation by the evaluation score, that is, the quality level of the corresponding dance animation is measured by the evaluation score. The higher the evaluation score, the better the quality of the dance animation; otherwise, the lower the evaluation score, the worse the quality of the corresponding dance animation.

[0085] In some embodiments, the dance animation with the highest evaluation score among multiple digital human dance animations generated from a single audio data set can theoretically be determined as the best dance animation. In other embodiments, a threshold can be preset corresponding to the evaluation score. For multiple digital human dance animations generated from a single audio data set, if any one of the evaluation scores is higher than the threshold, the corresponding dance animation is determined as the best dance animation. This approach is applicable when more than three dance animations are generated from a single audio data set, allowing multiple best dance animations to be selected.

[0086] In some embodiments, during the complete playback of the entire piece of music, steps S1100 and S1200 can be executed cyclically for different segments of audio data to generate all the dance animations of all digital humans and determine their corresponding user evaluation data. Then, in conjunction with this step, the evaluation scores of the corresponding dance animations are filtered using the threshold to centrally determine the best dance animations that are higher than the threshold from all the dance animations.

[0087] In other embodiments, when music is played in a large number of live streaming rooms on the live streaming platform, steps S1100 and S1200 are executed cyclically to generate a large number of dance animations, and user evaluation data corresponding to each dance animation is obtained, and its corresponding evaluation score is further determined. Then, all the dance animations generated corresponding to multiple live streaming rooms and multiple playing music can be periodically evaluated, and their corresponding evaluation scores can be determined based on the user evaluation data of each dance animation. Then, the best dance animations that are higher than the threshold can be selected using the evaluation scores.

[0088] It's easy to understand that the converged motion generation model that produces the best dance animation can be identified as the optimal motion generation model. The reason for identifying the optimal motion generation model is that the training dataset on which the optimal motion generation model relies usually has a better composition of data samples. Subsequently, new data samples can be added based on this training dataset to iteratively upgrade the motion generation model.

[0089] In one embodiment, within the same dance animation set, the proportion of dance animations generated by each motion generation model that are determined to be the best dance animation can be statistically analyzed. When a certain motion generation model obtains the highest proportion of best dance animations, it can be determined as the best motion generation model. The scope of the dance animation set can typically include one dance animation for each of multiple digital humans generated from a single audio data set; or it can include multiple dance animations for multiple digital humans generated from multiple audio data sets of a playing song. Of course, the scope of the dance animation set can also be expanded to a massive number of dance animations generated from multiple playing songs, and the specific implementation can be flexible.

[0090] As can be seen, based on user review data of dance animations, the best dance animation can be selected from those generated by the converged motion generation models. Furthermore, based on the proportion of the best dance animations produced by each motion generation model, the best motion generation model can be determined.

[0091] Step S1400: Expand the audio feature sequence and motion sequence information corresponding to the best dance animation into data samples in the training dataset of the best motion generation model.

[0092] The best dance animation obtained in this application is a high-quality dance animation determined based on user evaluation data corresponding to the subjective evaluation of the audience. The high-quality dance animation can guide the training of the motion generation model. Therefore, the audio feature sequences corresponding to all the best dance animations determined in this application and the motion sequence information generated by the motion generation model based on the audio feature sequences can be constructed into new data samples. These data samples can be added and stored in the training dataset on which the best motion generation model determined in this application depends for training, thereby expanding the training dataset. It's easy to understand that since the best motion generation model is the model that produces the highest proportion of the best dance animations, the data samples in its training dataset are relatively high-quality. Furthermore, data samples corresponding to the best dance animations produced by various motion generation models are added to expand the dataset. These best dance animations may originate from different motion generation models and have different characteristics. Therefore, the data samples in the expanded training dataset can generalize the diverse features between audio feature sequences and their corresponding motion sequence information. Moreover, it performs better in reflecting the mapping relationship between audio feature sequences and motion sequence information. Consequently, the expanded training dataset makes it easier to retrain the motion generation model to convergence, and the converged motion generation model gains stronger reasoning ability, enabling it to generate higher-quality motion sequence information based on audio feature sequences, thereby producing higher-quality dance animations.

[0093] As can be seen from the above embodiments, this application has multiple technical advantages, including but not limited to:

[0094] First, this application prepares multiple converged motion generation models. These models are used to generate motion sequence information for their corresponding digital humans, and dance animations are generated for each digital human based on the motion sequence information. These dance animations are then simultaneously displayed in a live broadcast room. User evaluation data for each dance animation is obtained through the live broadcast room. The best dance animation and the best motion generation model are selected based on the user evaluation data. The audio feature sequence of the audio data corresponding to the best dance animation and the motion sequence information generated therefrom are constructed as data samples and amplified into the training dataset used by the best motion generation model. Since the training dataset can train the best motion generation model, it indicates that it has excellent sample quality. On this basis, the data samples generated by the user evaluation selection are further amplified, which further improves the overall sample quality of the training dataset. Thus, the amplified training dataset can be used to retrain the motion generation model, iteratively obtaining an upgraded version of the motion generation model. This process is repeated to continuously improve the quality of the generated digital human dance animations.

[0095] Secondly, given the scarcity of digital human-related data samples, this application first uses an initially trained, converged motion generation model to generate dance animations for digital humans. Then, based on user evaluation data of these dance animations, it selects the best dance animation from each one and creates data samples accordingly. Thus, at a lower cost, it dynamically generates a massive amount of high-quality data samples using user evaluation data. These data samples are diverse due to differences in audio data, motion generation models, and user evaluation data, exhibiting excellent feature generalization effects. This enriches the training dataset required for training the motion generation model and makes it more suitable for retraining the motion generation model.

[0096] Furthermore, this application obtains dance animations of multiple digital humans on the backend, while displaying the corresponding dance animations of multiple digital humans on the graphical user interface of the live broadcast room on the frontend. This allows users to compare multiple dance animations and provide valuable user evaluation data after comparison. This evaluation data is essentially subjective evaluation data, unlike simple data cleaning through algorithms. With the help of subjective evaluation data, the selected dance animations are of higher quality, and therefore the audio feature sequences and motion sequence information of their corresponding audio data are of higher quality. The motion generation model trained in this way, when used to produce motion information sequences and generate dance animations, will inevitably produce dance animations that are more in line with user perception and improve user experience.

[0097] Based on any embodiment of this application, please refer to Figure 4 The process includes generating dance animations for each of the multiple digital humans and simultaneously pushing them to the live stream for playback, including:

[0098] Step S1210: Initialize the poses of multiple digital humans;

[0099] To facilitate comparison of dance animations by multiple digital humans generated from the same audio data, the postures of each digital human can be coordinated and controlled before generating the corresponding dance animations, ensuring that the dance animations of each digital human have a posture synchronization effect.

[0100] Specifically, the 3D models of each digital human can be uniformly set to the same standardized pose, for example... Figure 5 The T-pose shown can be any other pose; you can set it flexibly for actual use.

[0101] Step S1220: Establish a unified coordinate system among the multiple digital humans;

[0102] Each digital character in this application has the same skeletal keypoint layout, such as Figure 2 As shown. Based on this, it can be done according to... Figure 2 The distribution relationship shown in the figure establishes a unified coordinate system so that all 3D models of digital humans can be manipulated in a standardized manner according to the unified coordinate system to achieve posture adjustment.

[0103] For example, such as Figure 2 As shown, we can first take a plane with the dantian bone point 0, the left leg bone point 1, and the right leg bone point 2 of the digital human, and find the normal vector of this plane as the x-axis. Then, take a vector with the dantian bone point 0 and the pharyngeal bone point 12 as the y-axis. Then, take a plane with the x and y axes, and find the normal vector of this plane as the z-axis. Thus, we obtain a unified coordinate system with the origin corresponding to the dantian bone point.

[0104] Step S1230: Based on the unified coordinate system, apply the action sequence information of each action to its corresponding digital human so that the dance animation generated by each digital human keeps the action synchronized.

[0105] For the multiple digital humans described in this application, the aforementioned unified coordinate system can be used as the standard coordinate system for posture control. During the application of the corresponding action sequence information for each digital human, based on the control data of each information frame in the action sequence information of each digital human, and using the unified coordinate system as a reference, the posture of the digital human can be adjusted to match the control data of the information frame. This allows for rendering of the digital human and generation of corresponding image frames. Each image frame in the dance animation of multiple digital humans is prepared according to this process, ultimately achieving synchronization of the movements in the dance animation of all the digital humans.

[0106] For example, when the control data of the information frame in the action sequence information is represented as a rotation amount, the corresponding digital human pose can be generated by applying the following formula using the relevant values ​​in the control data of each information frame: Digital human 3D model skeleton rotation vector = initial rotation vector * rotation vector around x-axis * rotation vector around y-axis * rotation vector around z-axis.

[0107] As can be seen from the above embodiments, when applying the motion sequence information generated by each motion generation model to each digital human, a unified coordinate system is used to implement synchronous control of the postures of multiple digital humans, thereby generating corresponding dance animations for each digital human. This ensures that each digital human maintains motion synchronization with each other in its dance animation, thus ensuring that users can obtain a consistent visual experience and that the phenomenon of motion asynchrony in each dance animation will not occur due to hardware factors such as network conditions and device memory.

[0108] Based on any embodiment of this application, please refer to Figure 6After generating the corresponding dance animations for each of the multiple digital humans and simultaneously pushing them to the live stream for playback, the process includes:

[0109] Step S2100: Display the dance animations corresponding to the multiple digital humans in the graphical user interface of the live broadcast room, and keep the dance animations in rhythm with the music being played.

[0110] After the dance animations of each digital human are pushed to the live broadcast room, the media server actually pushes the live stream containing the dance animations to the terminal devices of each viewer user in the live broadcast room.

[0111] After receiving the live stream, the terminal device obtains the dance animation from it. In one embodiment, if the dance animation is a composite of the dance animations of multiple digital humans, and may also include music requested by the broadcaster, the terminal device only needs to play the live stream directly. When the terminal device plays the live stream, the dance animation, which includes the dance animations of multiple digital humans, is displayed in the corresponding graphical user interface of the live stream. Furthermore, since the media server has pre-processed the state synchronization relationship between the various digital humans, the movements of the dance animations of the various digital humans obtained in the graphical user interface are necessarily synchronized when playing the live stream, and the dance movements of the digital humans in each dance animation are also synchronized in rhythm with the music being played.

[0112] In another embodiment, if the dance animations of the various digital humans are independent of the audio data of the playing music, the media server will also synchronously push the rhythm information corresponding to the audio data. Therefore, based on the rhythm information, the terminal device can control the synchronized playback of the dance animations of the various digital humans and the audio data, thereby ensuring that the movement states of the dance animations of the various digital humans are rhythmically synchronized with the playing music. The rhythm information can represent the timestamps corresponding to each drumbeat in the audio data of the playing music, so that the terminal device can maintain the synchronization relationship through the correspondence between the playback time of the playing music and the timestamps, and the time correspondence between the timestamps and the image frames of the dance animations of the various digital humans.

[0113] Step S2200: Display multiple scoring controls on the graphical user interface corresponding to the multiple dance animations, for obtaining the corresponding user subjective scores for each dance animation;

[0114] To facilitate the acquisition of at least some user evaluation data, when the dance animation is displayed on the graphical user interface, a corresponding rating control can be displayed for each dance animation. The rating control can be used to input a specific rating value, or to input a rating level that corresponds to a specific rating value, or it can simply be used to allow the user to select one of the multiple dance animations as the best one, as long as the backend can correspond the selected and unselected to the corresponding rating values, such as 1 and 0. It can be seen that there are many feasible ways to implement the rating control, and its specific presentation can be flexibly designed.

[0115] After the terminal device displays the multiple rating controls in the graphical user interface, the user can give a rating value corresponding to one or more dance animations according to their own wishes through the corresponding rating controls. The rating value is actually the user's subjective score for each dance animation after watching and comparing all the dance animations displayed on the same screen. As the name suggests, the user's subjective score is the user's subjective rating, which represents the user's subjective viewing experience. Compared with objective evaluation indicators, it can better reflect the viewing experience of the dance animation for the audience.

[0116] Step S2300: Obtain the interaction information generated by the users in the live broadcast room interacting in the graphical user interface, and determine the corresponding user evaluation data of each digital human based on the interaction information of each digital human. The interaction information includes any one or more of the following: subjective score, bullet screen text, and chat area text.

[0117] The user review data for each dance animation is not limited to subjective scores obtained through the corresponding rating controls. Users can also provide other types of user review data in various other forms within the graphical user interface of the live stream.

[0118] In one embodiment, users may submit their evaluation information for one or more dance animations via bullet comments or chat. Therefore, the terminal device can process the bullet comment text and / or chat text, extract the content containing the evaluation information for a certain dance animation, and identify it as other user evaluation data for the corresponding dance animation.

[0119] In the terminal device, the information generated by the user's interaction with the graphical user interface of the live broadcast room is the interactive information. This includes subjective scores generated by the user through the evaluation controls, as well as bullet screen text, chat area text, etc., all of which belong to the interactive information. It also includes other types of interactive information, such as information generated by gift-giving events. It is easy to understand that, given the numerous types of information, the terminal device primarily applies the user evaluation data requirements of the various embodiments of this application, extracting the corresponding interactive information from the interactive information to determine the corresponding user evaluation data. This user evaluation data is submitted by the terminal device to the server for further processing in order to achieve the purpose of selecting the best dance animation in this application.

[0120] As can be seen from the above embodiments, this application can obtain user evaluation data of each dance animation from the terminal device in a variety of ways, including quantitative subjective scores or non-quantitative text information. In this way, it can achieve deep integration with the online live streaming business, so that viewers can watch multiple digital human synchronized dance animations in the live streaming room, thus innovating the business model; and obtain the corresponding user evaluation data of each dance animation, so as to achieve the purpose of selecting the best dance animation, thus achieving two goals at once.

[0121] Based on any embodiment of this application, please refer to Figure 7 Determining the best dance animation and the best motion generation model among the converged motion generation models based on the user evaluation data includes:

[0122] Step S1310: Determine the corresponding evaluation score for each dance animation based on the user evaluation data corresponding to each dance animation.

[0123] In one embodiment, quantifiable subjective scores from user review data corresponding to each dance animation can be used to determine the uniqueness score of each dance animation among multiple dance animations pushed in the same group.

[0124] Taking the subjective scores determined by users in the user review data as an example, for a dance animation, there may be 30 subjective scores generated by 30 users. In this case, we only need to simply average these 30 subjective scores to obtain the average value, which can be used as the final evaluation score of the dance animation.

[0125] In another embodiment, binary subjective scores from user review data corresponding to each dance animation can be used. For example, when a user approves and selects a dance animation, the subjective score of that dance animation is assigned a value of 1; otherwise, it is 0. In this case, the number of users who approve the same dance animation can be used to determine the evaluation score of that dance animation. Specifically, the number of approvals obtained by each dance animation in this case can be counted, and this number of approvals can be used as the unique evaluation score of that dance animation.

[0126] In other embodiments, other artificial intelligence methods can be further used to determine other forms of quantitative scores based on the bullet screen text, chat area text, etc. in the user evaluation data, so as to jointly determine the evaluation score corresponding to each dance animation.

[0127] Step S1320: Determine the standard deviation based on the corresponding evaluation scores of all dance animations;

[0128] Following step S1310, for a large number of dance animations generated successively, each dance animation can obtain its corresponding evaluation score. In order to facilitate the generation of the best dance animation based on statistics, for each dance animation, its standard deviation is first determined, that is, the average value of the evaluation scores of all dance animations generated by all motion generation models is first calculated, and then the variance of each dance animation is calculated with the average value. Based on the obtained variance, the standard deviation is calculated with the evaluation scores of each dance animation and the variance, which is used to measure the dispersion of the evaluation scores of each dance animation.

[0129] Step S1330: Select the dance animation with an evaluation score higher than the standard deviation as the best dance animation;

[0130] Based on the established standard deviation, it can be used as a threshold to select the best from the entire pool of dance animations. Specifically, the evaluation score of each dance animation can be compared with the standard deviation. When the evaluation score is higher than the standard deviation, the corresponding dance animation can be determined as the best; otherwise, it is not the best. Therefore, from the massive amount of dance animations with user evaluation data, multiple best dance animations can be identified, and these best dance animations are determined through user evaluation.

[0131] Step S1340: Select the converged motion generation model that produces the most and best dance animations as the best motion generation model.

[0132] The several optimal dance animations selected from a vast pool of dance animations may have been generated from motion sequence information produced by different converged motion generation models. Therefore, it is necessary to determine the source models of the motion sequence information corresponding to these optimal dance animations, and then determine the total number of optimal dance animations generated by each motion generation model. The motion generation model that produces the optimal dance animation with the largest output can be identified as the optimal motion generation model, indicating that this optimal model can obtain higher-quality motion sequence information and generate higher-quality dance animations. Furthermore, the superior performance of the optimal motion generation model is due to the higher-quality data samples in the training dataset used for its training. Therefore, the audio data and motion sequence information corresponding to the optimal dance animations can be further used to expand the data samples in the training dataset to train even better motion generation models.

[0133] As can be seen from the above embodiments, this application can quickly determine the best dance animation based on the user evaluation data obtained from each dance animation and by applying statistical principles. Furthermore, based on the statistical number of each motion generation model corresponding to the best dance animation, the best motion generation model is selected from all converged motion generation models. The data samples corresponding to the best dance animation are used to expand the training dataset on which the best motion generation model depends, thus efficiently and quickly preparing for subsequent upgrades and iterations of the motion generation model.

[0134] Based on any embodiment of this application, please refer to Figure 8 Based on user review data for each dance animation, a corresponding rating score was determined for each dance animation, including:

[0135] Step S1311: Input the text type data in the user evaluation data corresponding to each dance animation into the preset rating prediction model to predict the first rating. The text type data includes any one or more of the following: bullet screen text and chat area text.

[0136] A rating prediction model can be pre-trained to predict the confidence level of the classification corresponding to the positive sample based on the text type data, and this confidence level is used as the rating corresponding to the text type data. The text type data can be bullet screen text, chat text, or a combination of both.

[0137] In an exemplary model architecture, the rating prediction model includes a text feature extraction model and a classifier. The text feature extraction model is implemented using encoders such as RNN, LSTM, and BERT to extract deep semantic information from the encoded information of text-type data. The classifier is responsible for mapping the deep semantic information to a preset classification space to obtain the confidence score corresponding to each classification in the classification space. The classification space can be a multi-class space or a binary classification space, in which there exists a classification supervised by positive samples. The confidence score obtained by this classification can be used to represent the rating that the corresponding text-type data can obtain.

[0138] During the training phase of the rating prediction model, manually prepared training samples can be used for training. These training samples consist of text from users evaluating the quality of a dance animation. By manually interpreting these training samples, it is determined whether the user's text is a positive or negative review. Text corresponding to positive reviews is used as positive samples, and text corresponding to negative reviews is used as negative samples. When the training samples are input into the rating prediction model for training, corresponding supervision is implemented for both positive and negative samples. Through iterative training, the rating prediction model is trained to a convergent state. Thus, the rating prediction model can express its rating for the corresponding text type data through the confidence score of its classification corresponding to the positive sample.

[0139] For each dance animation, text data can be extracted from the user review data corresponding to each user, and input into the rating prediction model to predict its corresponding rating. Then, the average of the ratings corresponding to all users of the dance animation is calculated as the first rating. Thus, each dance animation can obtain a corresponding first rating.

[0140] Step S1312: Obtain the subjective score from the user review data corresponding to each dance animation as the second score;

[0141] As mentioned earlier, user review data can also include subjective scores corresponding to user reviews. Similarly, for the subjective scores of all users corresponding to the same dance animation, the average of these scores can be used to obtain the second rating for each dance animation.

[0142] Step S1313: Weighted fusion of the first and second scores of each dance animation to obtain the evaluation score of the dance animation.

[0143] At this point, each dance animation has received its first and second scores. By weighting and combining these two scores, the corresponding evaluation score for each dance animation can be obtained. The weights required for weighting the first and second scores can be normalized weights, which can be flexibly set by those skilled in the art.

[0144] As can be seen from the above embodiments, when determining the evaluation score of each dance animation, text-based data is used to determine the first score, and subjective scores are used to determine the second score. Since text-based data has the characteristic of being relatively comprehensive in meaning, and subjective scores have the characteristic of being quantifiable, weighting and integrating the scores obtained from these two parts to determine the evaluation score of each dance animation is more scientific, comprehensive and accurate.

[0145] Based on any embodiment of this application, please refer to Figure 9 After amplifying the audio feature sequence and motion sequence information corresponding to the best dance animation into data samples in the training dataset of the best motion generation model, the process includes:

[0146] Step S3100: Restart the training of the action generation model using data samples from the training dataset, and train the action generation model to a convergent state, which serves as an upgraded action generation model;

[0147] The various converged action generation models in this application typically use different training datasets during their respective training phases to differentiate them. In this application, the training dataset used by the best action generation model is considered the optimal training dataset. Adding data samples obtained by expanding the dataset according to this application to the training dataset used by the optimal action generation model is for the purpose of using this training dataset to iteratively upgrade the action generation model.

[0148] Therefore, after the training dataset is sufficiently expanded with newly added data samples, the training task of the motion generation model described in this application can be restarted. In the training task, each data sample in the training dataset is sequentially called to iteratively train the motion generation model until it converges, becoming an upgraded motion generation model. Because the data samples in the training dataset used by the upgraded motion generation model are constructed through user evaluation and selection, it is expected to improve the inference ability of the motion generation model, and the generated motion sequence information is expected to further present high-quality dance animation.

[0149] Step S3200: Configure the upgraded motion generation model as a callable service of the live broadcast room, which is used to generate corresponding motion sequence information for the audio feature sequence of the audio data of the music playing in the live broadcast room, and generate the dance animation of the digital human in the live broadcast room according to the motion sequence information.

[0150] After obtaining the upgraded motion generation model, the old motion generation model, originally configured for use in the live streaming room to generate motion sequence information based on the audio feature sequence of the playing music, no longer needs to be used. Therefore, the upgraded motion generation model can be configured as a callable service for the live streaming room. When the live streaming room requests music, its audio feature sequence can be extracted from the audio data of the playing music. Then, the upgraded motion generation model is called to generate corresponding motion sequence information based on the audio feature sequence. Based on the motion sequence information, the corresponding digital human is controlled to change posture to generate a corresponding dance animation. Finally, the dance animation is pushed to the live streaming room to achieve a digital human performance effect.

[0151] As can be seen from the above embodiments, this application can realize the iterative upgrading of the motion generation model. The training dataset can be iteratively upgraded using the converged motion generation model, and the motion generation model can be iteratively upgraded again using the training data with expanded data samples. This iterative upgrade continuously improves the quality of the digital human dance animation in the live broadcast room, thereby continuously improving the user experience and service experience.

[0152] Based on any embodiment of this application, please refer to Figure 10 Restarting the training of the action generation model using data samples from the training dataset includes:

[0153] Step S3210: Call a single data sample from the training dataset, determine the audio feature sequence in the data sample as a training sample, and determine the action sequence information corresponding to the audio data as a supervision label;

[0154] When iterative training of the action generation model of this application is required, a single data sample is retrieved from the training dataset for each iteration. As mentioned earlier, the data sample contains audio feature sequences and their corresponding action sequence information. The audio feature sequences are obtained by extracting audio features from audio data, while the action sequence information corresponds to the audio feature sequences. In the training dataset, the action sequence information in some of the expanded data samples is actually generated by inferring the audio feature sequences using an older version of the action generation model. Since the task of the action generation model is to generate corresponding action sequence information based on the audio feature sequences, the audio feature sequences in the data samples can be used as training samples as input, while the action sequence information in the data samples can be used as supervision labels to supervise the model output.

[0155] Step S3220: Input the training samples into the motion generation model to predict the motion sequence information used to control the digital human to generate motion images;

[0156] During training, the training samples used in the current iteration are input into the action generation model. The action generation model then infers based on its modeling reasoning function and predicts the corresponding action sequence information. This action sequence information is predicted by the model in real time and can theoretically be used to control the digital human to generate motion images.

[0157] Step S3230: Calculate the loss value of the action sequence information obtained by the action generation model based on the supervision label, and decide whether to continue iterative training of the action generation model based on the loss value until it reaches a convergence state and constitutes an upgraded action generation model.

[0158] Based on the supervision label, and according to the loss function used by the action generation model, the loss value of the action sequence information predicted by the model relative to the supervision label can be calculated. The goal of model training is to minimize the loss value, for example, to 0. Therefore, a threshold is usually set to compare whether the loss value reaches the threshold. When the threshold is reached, it indicates that the action generation model has reached convergence, and training can be terminated to obtain an upgraded action generation model. When the threshold is not reached, it indicates that the model has not reached convergence, and gradient updates need to be performed on the action generation model according to the loss value. The weight parameters of each link are corrected through backpropagation to make it closer to convergence. Then, the next iteration continues from step S3210, and so on, until the action generation model is retrained to convergence and an upgraded action generation model is obtained.

[0159] As can be seen from the above embodiments, the motion generation model of this application is upgraded by retraining with a training dataset that has been expanded with data samples. Since the data samples in the training dataset have been expanded with high-quality data samples, the sample features have been further generalized. Therefore, it is easier to train the motion generation model to convergence quickly. Moreover, with the help of data samples with generalized features, the upgraded motion generation model is expected to have stronger reasoning ability. Based on the motion sequence information it predicts, the generated digital human dance animation will be of higher quality.

[0160] Please see Figure 11According to one aspect of this application, a digital human animation evaluation and optimization device includes a motion generation module 1100, an animation push module 1200, an evaluation processing module 1300, and a sample augmentation module 1400. The motion generation module 1100 is configured to acquire audio feature sequences of audio data from music playing in a live stream, and generate motion sequence information corresponding to the audio feature sequences using multiple converged motion generation models, thereby obtaining motion sequence information for each of the multiple digital humans. The animation push module 1200 is configured to apply each motion sequence information to its corresponding digital human, generating dance animations for each of the multiple digital humans, which are then synchronously pushed to the live stream for playback, and to receive user evaluation data from the live stream that affects the dance animations. The evaluation processing module 1300 is configured to determine the best dance animation and the best motion generation model among the converged motion generation models based on the user evaluation data. The sample augmentation module 1400 is configured to augment the audio feature sequences and motion sequence information corresponding to the best dance animation into data samples in the training dataset of the best motion generation model.

[0161] Based on any embodiment of this application, the animation push module 1200 includes: a posture initialization unit, configured to initialize the action postures of multiple digital humans; a coordinate system unit, configured to establish a unified coordinate system among the multiple digital humans; and a synchronization control unit, configured to apply the action sequence information of each action to its corresponding digital human based on the unified coordinate system, so that the dance animations generated by each digital human keep the actions synchronized.

[0162] Based on any embodiment of this application, the animation push module 1200 further includes: an animation display module, configured to display dance animations corresponding to the plurality of digital humans in the graphical user interface of the live broadcast room, and to keep the dance animations rhythmically synchronized with the music being played; a control display module, configured to display a plurality of rating controls in the graphical user interface corresponding to the plurality of dance animations, for obtaining the corresponding user subjective scores of each dance animation; and a data processing module, configured to obtain the interaction information generated by the interaction of users in the graphical user interface of the live broadcast room, and to determine the corresponding user evaluation data of each digital human based on the interaction information of each digital human, wherein the interaction information includes any one or more of the following: the subjective scores, bullet screen text, and chat area text.

[0163] Based on any embodiment of this application, the evaluation processing module 1300 includes: a scoring determination unit, configured to determine the evaluation score corresponding to each dance animation based on the user evaluation data corresponding to each dance animation; a benchmark calibration unit, configured to determine the standard deviation based on the evaluation scores of all dance animations; an animation optimization unit, configured to select the dance animation with an evaluation score higher than the standard deviation as the best dance animation; and a model optimization unit, configured to select the converged motion generation model that produces the most best dance animations as the best motion generation model.

[0164] Based on any embodiment of this application, the scoring determination unit includes: a first scoring submodule, configured to input text type data from user evaluation data corresponding to each dance animation into a preset scoring prediction model to predict a first score, wherein the text type data includes any one or more of bullet screen text and chat area text; a second scoring submodule, configured to obtain the subjective score from user evaluation data corresponding to each dance animation as a second score; and a scoring fusion submodule, configured to weightedly fuse the first score and the second score of each dance animation to obtain the evaluation score of the dance animation.

[0165] Based on any embodiment of this application, the sample augmentation module 1400 further includes: a training restart module, configured to restart the training of the action generation model using data samples from the training dataset, training the action generation model to a convergent state, and serving as an upgraded action generation model; and a service restart module, configured to configure the upgraded action generation model as a callable service of the live streaming room, used to generate corresponding action sequence information for the audio feature sequence of the audio data of the music playing in the live streaming room, and to generate a dance animation of the digital human in the live streaming room based on the action sequence information.

[0166] Based on any embodiment of this application, the restart training module includes: a sample retrieval unit, configured to retrieve a single data sample from the training dataset, determine the audio feature sequence in the data sample as a training sample, and determine the action sequence information corresponding to the audio data as a supervision label; a model prediction unit, configured to input the training sample into the action generation model, and predict the action sequence information used to control the digital human to generate motion images; and an iterative decision unit, configured to calculate the loss value of the action sequence information obtained by the action generation model based on the supervision label, and decide whether to continue iteratively training the action generation model according to the loss value until it reaches a convergence state and constitutes an upgraded action generation model.

[0167] Another embodiment of this application provides a digital human animation evaluation and optimization device. For example... Figure 12The diagram shows the internal structure of a digital human animation evaluation and optimization device. This device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable, non-volatile storage medium stores an operating system, a database, and computer-readable instructions. The database stores information sequences, and when executed by the processor, these computer-readable instructions enable the processor to implement a digital human animation evaluation and optimization method.

[0168] The processor of the digital human animation evaluation and optimization device provides computing and control capabilities, supporting the operation of the entire device. The device's memory can store computer-readable instructions, which, when executed by the processor, cause the processor to perform the digital human animation evaluation and optimization method of this application. The device's network interface is used for communication with a terminal.

[0169] Those skilled in the art will understand that Figure 12 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the digital human animation evaluation and optimization device to which the present application is applied. The specific digital human animation evaluation and optimization device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0170] In this embodiment, the processor is used to execute... Figure 11 The specific functions of each module are described, and the memory stores the program code and various data required to execute the above modules or sub-modules. The network interface is used to realize data transmission between user terminals or servers. The non-volatile readable storage medium in this embodiment stores the program code and data required to execute all modules in the digital human animation evaluation and optimization device of this application, and the server can call the server's program code and data to execute the functions of all modules.

[0171] This application also provides a non-volatile readable storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the digital human animation evaluation and optimization method of any embodiment of this application.

[0172] This application also provides a computer program product, including a computer program / instructions that, when executed by one or more processors, implement the steps of the method described in any embodiment of this application.

[0173] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM).

[0174] In summary, this application can optimize the dance animation obtained by the motion generation model by collecting user evaluation data from the live broadcast room. The motion sequence information of the identified high-quality dance animation and the audio feature sequence that generated the motion sequence information are constructed as data samples for iterative training of the motion generation model, thereby improving the inference ability of the motion generation model, saving the continuous evolution cost of the motion generation model, and improving the user experience of the digital human's dance animation.

Claims

1. A method for evaluating and optimizing digital human animation, characterized in that, include: The audio feature sequence of the music playing in the live room is obtained, and multiple converged action generation models are used to generate action sequence information corresponding to the audio feature sequence, thereby obtaining the corresponding action sequence information of multiple digital humans. Each of the aforementioned action sequence information is applied to its corresponding digital human to generate dance animations for each of the multiple digital humans, which are then synchronously pushed to the live stream for playback. User evaluation data applied to the dance animations in the live stream is received. This process includes: initializing the action postures of the multiple digital humans; establishing a unified coordinate system among the multiple digital humans; and, based on the unified coordinate system, applying each action sequence information to its corresponding digital human to keep the dance animations generated by each digital human synchronized. Determining the best dance animation and the best motion generation model among the converged motion generation models based on the user evaluation data includes: determining the evaluation score corresponding to each dance animation based on the user evaluation data corresponding to each dance animation; determining the standard deviation based on the evaluation scores of all dance animations; selecting the dance animation with an evaluation score higher than the standard deviation as the best dance animation; and selecting the converged motion generation model that produces the most best dance animations as the best motion generation model. The audio feature sequence and motion sequence information corresponding to the best dance animation are amplified into data samples in the training dataset of the best motion generation model.

2. The digital human animation evaluation and optimization method according to claim 1, characterized in that, After generating the corresponding dance animations for each of the multiple digital humans and synchronously pushing them to the live stream for playback, the process includes: The graphical user interface of the live broadcast room displays dance animations corresponding to the multiple digital humans, and the dance animations are synchronized with the music being played. The graphical user interface displays multiple scoring controls corresponding to the multiple dance animations, which are used to obtain the corresponding subjective scores of users for each dance animation. The interaction information generated by the users in the live broadcast room interacting in the graphical user interface is obtained, and the corresponding user evaluation data of each digital human is determined according to the interaction information of each digital human. The interaction information includes any one or more of the following: subjective score, bullet screen text, and chat area text.

3. The digital human animation evaluation and optimization method according to claim 1, characterized in that, The rating scores for each dance animation were determined based on user review data, including: Input the text type data from the user review data corresponding to each dance animation into a preset rating prediction model to predict the first rating. The text type data includes any one or more of the following: bullet screen text and chat area text. The subjective score from the user review data corresponding to each dance animation is used as the second rating. The first and second scores of each dance animation are weighted and combined to obtain the evaluation score of the dance animation.

4. The digital human animation evaluation and optimization method according to any one of claims 1 to 3, characterized in that, After amplifying the audio feature sequence and motion sequence information corresponding to the optimal dance animation into data samples in the training dataset of the optimal motion generation model, the process includes: The training of the action generation model is restarted using data samples from the training dataset, and the action generation model is trained to a convergent state, which serves as an upgraded action generation model. The upgraded motion generation model is configured as a callable service of the live streaming room, used to generate corresponding motion sequence information for the audio feature sequence of the audio data of the music playing in the live streaming room, and to generate the dance animation of the digital human in the live streaming room based on the motion sequence information.

5. The digital human animation evaluation and optimization method according to claim 4, characterized in that, Restarting the training of the action generation model using data samples from the training dataset includes: A single data sample is retrieved from the training dataset, the audio feature sequence in the data sample is determined as a training sample, and the action sequence information corresponding to the audio data is determined as a supervision label; The training samples are input into the motion generation model to predict motion sequence information used to control the digital human to generate motion images; Based on the supervised labels, the loss value of the action sequence information obtained by the action generation model is calculated. Based on the loss value, a decision is made on whether to continue iteratively training the action generation model until it reaches a convergence state and constitutes an upgraded action generation model.

6. A digital human animation evaluation and optimization device, characterized in that, include: The action generation module is configured to acquire the audio feature sequence of the audio data of the music playing in the live room, and use multiple converged action generation models to generate action sequence information corresponding to the audio feature sequence, thereby obtaining the corresponding action sequence information of multiple digital humans. The animation push module is configured to apply each of the aforementioned action sequence information to its corresponding digital human, generating dance animations for each of the multiple digital humans and synchronously pushing them to the live stream for playback. It also receives user evaluation data from the live stream that affects the dance animations. The module includes: a posture initialization unit, configured to initialize the action postures of the multiple digital humans; a coordinate system unit, configured to establish a unified coordinate system among the multiple digital humans; and a synchronization control unit, configured to apply each action sequence information to its corresponding digital human based on the unified coordinate system, ensuring that the dance animations generated by each digital human remain synchronized. The evaluation processing module is configured to determine the best dance animation and the best motion generation model among the converged motion generation models based on the user evaluation data. This includes: a scoring determination unit, configured to determine the corresponding evaluation score for each dance animation based on the user evaluation data; a benchmark calibration unit, configured to determine the standard deviation based on the corresponding evaluation scores of all dance animations; an animation optimization unit, configured to select the dance animation with an evaluation score higher than the standard deviation as the best dance animation; and a model optimization unit, configured to select the converged motion generation model that produces the most best dance animations as the best motion generation model. The sample augmentation module is configured to augment the audio feature sequence and motion sequence information corresponding to the best dance animation into data samples in the training dataset of the best motion generation model.

7. The digital human animation evaluation and optimization device according to claim 6, characterized in that, This device also includes: An animation display module is configured to display dance animations corresponding to the multiple digital humans in the graphical user interface of the live broadcast room, and to keep the dance animations in rhythm synchronization with the music being played. The control display module is configured to display multiple scoring controls on the graphical user interface corresponding to the multiple dance animations, for obtaining the corresponding subjective scores of each dance animation. The data processing module is configured to acquire the interaction information generated by the users in the live broadcast room interacting in the graphical user interface, and determine the corresponding user evaluation data of each digital human based on the interaction information of each digital human. The interaction information includes any one or more of the following: subjective score, bullet screen text, and chat area text.

8. The digital human animation evaluation and optimization device according to claim 6 or 7, characterized in that, The scoring determination unit includes: The first rating submodule is set to input the text type data from the user evaluation data corresponding to each dance animation into a preset rating prediction model to predict the first rating. The text type data includes any one or more of the following: bullet screen text and chat area text. The second scoring submodule is set to obtain the subjective score from the user review data corresponding to each dance animation as the second score; The scoring fusion submodule is configured to weightedly fuse the first and second scores of each dance animation to obtain the evaluation score for that dance animation.

9. A digital human animation evaluation and optimization device, comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 5.

10. A non-volatile readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 5, which, when invoked by a computer, executes the steps included in the corresponding method.

Citation Information

Patent Citations

  • Method and system for generating character model dance animation

    CN112330779A

  • KR20200014510A