Action Data Generation Method, Apparatus, Device, Storage Medium, and Program Product
By using the diffusion model to generate matching skeleton action sequences, the problem of low matching between digital human body movements and speech output text is solved, and the user's immersion experience is improved.
Patent Information
- Application Number
- CN202510260633.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-06
AI Technical Summary
In the prior art, the matching degree of digital human body movements and text output through voice is poor, which affects the user's immersion experience.
By obtaining the audio characteristics and corresponding text of the target speech, the text segments that need to be performed synchronously based on semantic understanding are determined, and a matching skeleton action sequence is generated using the diffusion model to improve the matching degree of limb movements and speech content.
It improves the matching degree between digital human body movements and voice content, and enhances the immersive experience of users.
Smart Images

Figure CN119741405B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, storage medium, and program product for generating motion data. Background Art
[0002] Currently, in some human-computer interaction scenarios, a digital human interacts with a user through voice. To enhance the user's immersive experience, the digital human not only needs smooth and natural lip synchronization but also needs to be paired with context-appropriate body movements.
[0003] Currently, usually, the text that the digital human needs to output through voice is obtained, and based on this text, a skeleton motion sequence (representing the positions of the skeleton points of the digital human at different times) is searched in the motion database. When the digital human outputs the text through voice, the digital human is driven to perform body movements based on the searched skeleton motion sequence. However, this body movement control method has the problem that the matching degree between the body movements of the digital human and the text output through voice is relatively poor. Summary of the Invention
[0004] In view of the above problems, this application provides a method, apparatus, device, storage medium, and program product for generating motion data to improve the matching degree between the body movements of the digital human and the text output through voice. The specific solutions are as follows:
[0005] In a first aspect of this application, a method for generating motion data is provided, including:
[0006] Obtain the audio features of the target voice and the text corresponding to the target voice; the audio features include at least one of the following: the rhythm feature of the target voice, and the hidden layer features obtained by performing hidden layer feature extraction on the target voice;
[0007] Based on the semantic understanding of the text, determine the target text segments in the text that require the digital human to synchronously perform body movements, the categories of body movements corresponding to each target text segment, and the position encodings of each action frame in the to-be-generated skeleton motion sequence corresponding to each target text segment; the skeleton motion sequence corresponding to each target text segment represents the positions of each skeleton point at different times during the process of the digital human outputting the target text segment through voice.
[0008] For each target text segment, at least use the audio features, the category of the body movement corresponding to this target text segment, and the position encodings of each action frame corresponding to this target text segment as the control conditions of the diffusion model, and generate the skeleton motion sequence corresponding to this target text segment through the diffusion model.
[0009] In a possible implementation, based on the semantic understanding of the text, determine the target text segments in the text that require the digital human to synchronously execute limb actions. The categories of limb actions corresponding to each target text segment include:
[0010] Process the text through a large model to insert marker tags into the text; each group of marker tags represents the start and end positions of the target text segments in the text that require the digital human to synchronously execute limb actions, and the category of limb actions corresponding to each target text segment; the categories of limb actions include: intentions and the corresponding limb actions.
[0011] In a possible implementation, based on the semantic understanding of the text, determine the target text segments in the text that require the digital human to synchronously execute limb actions. The categories of limb actions corresponding to each target text segment include:
[0012] Process the text through a large model to insert marker tags into the text; each group of marker tags represents the start and end positions of the target text segments in the text that require the digital human to synchronously execute limb actions, and the category of limb actions corresponding to each target text segment; the categories of limb actions include: intentions and the corresponding limb actions;
[0013] Perform label correction on the text after inserting marker tags to delete abnormal marker tags or correct relevant information for abnormal marker tags.
[0014] In a possible implementation, the performing label correction on the text after inserting marker tags includes: performing label correction on the text after inserting marker tags by using at least one of the following methods:
[0015] Compare the text after removing the marker tags with the text corresponding to the target speech. If the comparison result indicates inconsistency, discard the text after inserting the marker tags and use the text corresponding to the target speech;
[0016] Perform a normative judgment on each group of marker tags and delete the groups of marker tags that do not meet the preset specifications;
[0017] Based on the start and end positions of each character in each target text segment in the target speech, determine the duration of the skeleton action sequence corresponding to each target text segment; if the duration of the skeleton action sequence corresponding to any target text segment does not meet the target duration range of the limb action represented by the marker tag marking the any target text segment, update the duration of the skeleton action sequence corresponding to the any target text segment to a duration within the target duration range.
[0018] In a possible implementation, determining the position encodings of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment includes:
[0019] Determining the number of action frames in the to-be-generated skeleton action sequence corresponding to each target text segment according to the start and end positions of each character in the target text segment in the target speech and a preset frame rate;
[0020] Obtaining the position encodings of each action frame in the skeleton action sequence according to the positions of each action frame in the skeleton action sequence.
[0021] In a possible implementation, obtaining the audio features of the target speech includes:
[0022] Extracting the hidden layer features of the target speech by a general speech pre-training model;
[0023] And / or, extracting the audio start point and beat information of the audio signal from the target speech through a speech signal processing library as the rhythm features of the target speech.
[0024] In a possible implementation, generating at least two initial skeleton action sequences corresponding to the target text segment through the diffusion model; among adjacent two initial skeleton action sequences, the last m action frames in the previous initial skeleton action sequence and the first m action frames in the subsequent initial skeleton action sequence correspond to the same sub-skeleton action sequence in the skeleton action sequence corresponding to the target text segment, and the sub-skeleton action sequence includes m action frames;
[0025] Performing weighted summation on the m action frames corresponding to the same sub-skeleton action sequence in adjacent two initial skeleton action sequences to obtain m target action frames corresponding to the same sub-skeleton action sequence; wherein, the weights of the last m action frames in the previous initial skeleton action sequence gradually decrease, and the weights of the first m action frames in the subsequent initial skeleton action sequence gradually increase.
[0026] In a possible implementation, it further includes at least one of the following:
[0027] Obtaining the input reference person image, and extracting the skeleton features from the reference person image;
[0028] Obtaining the person style information;
[0029] Correspondingly, the control conditions of the diffusion model further include:
[0030] At least one of the skeleton features and the person style information.
[0031] In a possible implementation, feature extraction is performed on the reference person image to obtain skeleton features, including:
[0032] Skeleton point extraction is performed on the reference person image to obtain an initial skeleton point set; the initial skeleton point set is translated and / or scaled to obtain a target skeleton point set; feature extraction is performed on the target skeleton point set to obtain the skeleton features;
[0033] Alternatively, skeleton point extraction is performed on the reference person image through a pose estimation model to obtain an initial skeleton point set; the initial skeleton point set is translated and / or scaled to obtain a target skeleton point set; feature extraction is performed on the target skeleton point set to obtain a first feature; the second feature output by a preset intermediate layer of the pose estimation model and the first feature constitute the skeleton features.
[0034] In a possible implementation, the diffusion model is trained with a number of skeleton action sequences as the training sample set and the target feature corresponding to each skeleton action sequence as the control condition of the diffusion model;
[0035] The different skeleton action sequences in the training sample set are obtained by performing skeleton point extraction on the video frames in different video segments; a limb action is synchronously displayed by the person in each video segment when outputting speech;
[0036] The target feature corresponding to each skeleton action sequence in the training sample set at least includes: the audio feature of the speech synchronized with the video segment corresponding to the skeleton action sequence, the category of the limb action in the video segment corresponding to the skeleton action sequence, and the position encoding of each video frame in the video segment corresponding to the skeleton action sequence.
[0037] In a possible implementation, during the training process of the diffusion model, some training samples are randomly selected, and the category and / or position encoding of the limb action in the target feature corresponding to the randomly selected training samples are set to zero.
[0038] In a possible implementation, in the skeleton action sequence output by the diffusion model, each action frame includes a human skeleton point set and a hand key point set;
[0039] During the training process of the diffusion model, the parameters of the diffusion model are updated based on the difference between the noise point set sequence predicted by the diffusion model and the noise point set sequence input to the diffusion model;
[0040] Wherein, when calculating the difference between the noise point set sequences predicted by the diffusion model and the noise point sets at the same positions of the noise point set sequence input to the diffusion model, calculate the overlapping degree of the two hands in the video frame corresponding to the noise point set at the same position; if the overlapping degree is greater than the threshold, when calculating the difference between the noise point sets at the same position, set the difference of the hand region to zero.
[0041] The second aspect of the present application provides an action data generation device, including:
[0042] An acquisition module, configured to acquire the audio feature of the target speech and the text corresponding to the target speech; the audio feature includes at least one of the following: the rhythm feature of the target speech, and the hidden layer feature obtained by performing hidden layer feature extraction on the target speech;
[0043] A determination module, configured to determine, based on the semantic understanding of the text, the target text segments in the text that require the digital human to synchronously execute limb actions, the category of the limb actions corresponding to each target text segment, and the position encoding of each action frame in the skeleton action sequence to be generated corresponding to each target text segment; the skeleton action sequence corresponding to each target text segment represents the positions of each skeleton point at different moments when the digital human outputs the target text segment in speech.
[0044] A generation module, configured to, for each target text segment, use at least the audio feature, the category of the limb actions corresponding to the target text segment, and the position encoding of each action frame corresponding to the target text segment as the control conditions of the diffusion model, and generate a skeleton action sequence corresponding to the target text segment through the diffusion model.
[0045] The third aspect of the present application provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement the action data generation method in the first aspect or any implementation manner of the first aspect.
[0046] The fourth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:
[0047] The memory is used to store a computer program;
[0048] The processor is configured to execute the computer program so that the electronic device can implement the action data generation method in the first aspect or any implementation manner of the first aspect.
[0049] A fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the action data generation method according to the first aspect or any implementation manner of the first aspect.
[0050] With the above technical solutions, the action data generation method, device, equipment, storage medium and program product provided by the present application obtain the audio features of the target voice and the text corresponding to the target voice; the audio features include at least one of the following: the rhythm feature of the target voice, and the hidden layer features obtained by performing hidden layer feature extraction on the target voice; based on the semantic understanding of the text, determine the target text segments in the text that require the digital human to synchronously execute limb actions, the categories of limb actions corresponding to each target text segment, and the position encodings of each action frame in the skeleton action sequence to be generated corresponding to each target text segment; the skeleton action sequence corresponding to each target text segment represents the positions of each skeleton point at different moments when the digital human outputs the target text segment by voice; for each target text segment, at least use the audio features, the category of the limb action corresponding to the target text segment, and the position encodings of each action frame corresponding to the target text segment as the control conditions of the diffusion model, and generate the skeleton action sequence corresponding to the target text segment through the diffusion model. The present application proposes an action data generation scheme based on joint perception of voice and semantics. By at least using the audio features of the target voice, the categories of limb actions corresponding to the target text segments determined based on the semantics of the text corresponding to the target voice, and the position encodings of each action frame in the skeleton action sequence to be generated as the control conditions of the diffusion model, the diffusion model is guided to generate action data that matches the control conditions, that is, the skeleton action sequence, improving the matching degree between the skeleton action sequence and the voice content, and thus improving the matching degree between the limb actions of the digital human driven by the skeleton action sequence and the voice content. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.
[0052] Figure 1 It is an exemplary diagram of the key points of the human skeleton provided by the present application;
[0053] Figure 2 It is another exemplary diagram of the key points of the human skeleton provided by the present application;
[0054] Figure 3 It is a flowchart of an implementation of the action data generation method provided by the present application;
[0055] Figure 4 A flowchart of an implementation for providing the position encoding of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment in the present application;
[0056] Figure 5 A schematic structural diagram of an action data generation device provided in the present application;
[0057] Figure 6 A schematic structural diagram of an electronic device provided in the present application. Detailed implementation manners
[0058] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments part of the present application are only used to explain the specific embodiments of the present application, rather than intended to limit the present application.
[0059] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.
[0060] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0061] To better understand the present application, some concepts will be described first.
[0062] Action data, that is, the skeleton action sequence, can reflect the movement postures and action changes of the active object (such as a person or an animal). Taking the active object as a human being as an example, the skeleton action sequence refers to a sequence composed of a series of action frames arranged in chronological order, and each action frame is composed of the coordinates (two-dimensional coordinates or three-dimensional coordinates) of at least one skeleton key point (also called a skeleton point or a joint point) of the human body. Each human body has multiple skeleton key points, and the skeleton key points of each human body in each action frame can reflect the instantaneous movement posture (or called instantaneous action) of a human body. If an action frame includes the skeleton key points of N (N is an integer greater than 0) human bodies, then this action frame represents the instantaneous movement postures of N human bodies.
[0063] Such asFigure 1 As shown in the figure, it is an example diagram of the key points of the human skeleton provided by the embodiment of the present application. In this example, a human body is represented by 18 key points of the skeleton (numbered 0-17 in sequence). Based on this, if a motion frame includes the key points of N human skeletons, then there are N×18 key points of the skeleton in this motion frame, or there are N sets of key points of the skeleton in this motion frame, and each set of key points of the skeleton is composed of 18 key points of one human body.
[0064] As Figure 2 shown in the figure, it is another example diagram of the key points of the human skeleton provided by the embodiment of the present application. In this example, a human body is represented by 17 key points of the skeleton (numbered 0-16 in sequence). Based on this, if a motion frame includes the key points of N human skeletons, then there are N×17 key points of the skeleton in this motion frame, or there are N sets of key points of the skeleton in this motion frame, and each set of key points of the skeleton is composed of 17 key points of one human body.
[0065] That is to say, if a skeleton motion sequence includes T motion frames, each motion frame includes the key points of N human skeletons, and each human body corresponds to M key points of the skeleton, then the t (t = 1, 2, 3,..., T) -th motion frame in this skeleton motion sequence is composed of the coordinates of N×M key points of the skeleton, representing the instantaneous motion of N human bodies at the t -th moment. The M key points of the skeleton corresponding to the n (n = 1, 2, 3,..., N) -th human body in the t -th motion frame represent the instantaneous motion of the n -th human body at the t -th moment; the coordinates of the m (m = 1, 2, 3,..., M) -th key point of the n -th human body in different motion frames may be the same or different. Based on this, each motion frame can be regarded as a skeleton diagram, and the pixel values at the positions of the N×M key points of the skeleton in this skeleton diagram are different from the pixel values at other positions.
[0066] A Digital Human is a virtual character that simulates human characteristics such as appearance, behavior, and emotion through computer technology. Digital Humans usually have rich information processing capabilities, simulation capabilities, and learning capabilities, and can provide intelligent customized services according to people's needs. Digital Humans can also be called virtual humans, digital virtual humans, virtual digital humans, etc.
[0067] The following explains the solution of the present application.
[0068] As Figure 3 shown in the figure, it is a flowchart of an implementation of the method for generating motion data provided by the embodiment of the present application, which may include:
[0069] Step S301: Obtain the audio features of the target voice and the text corresponding to the target voice.
[0070] The audio features include at least one of the following: the rhythm feature of the target speech, and the hidden layer feature obtained by performing hidden layer feature extraction on the target speech.
[0071] Optionally, the rhythm feature of the target speech may include at least one of: audio onsets and beats information. The rhythm feature can be used as a shallow feature of the target speech.
[0072] The hidden layer feature obtained by performing hidden layer feature extraction on the target speech can be used as a high-level feature of the target speech.
[0073] The text corresponding to the target speech (for the convenience of description and distinction, denoted as text) can be the text obtained by performing speech recognition on the target speech;
[0074] Alternatively, the target speech is a speech signal obtained by performing text-to-speech (TTS) on the text text.
[0075] Alternatively, the target speech is the speech signal of the reader collected when the reader reads the text text.
[0076] Step S302: Based on the semantic understanding of the text, determine the target text segments in the text that require the digital human to synchronously execute limb movements, the categories of limb movements corresponding to each target text segment, and the position encodings of each action frame in the skeleton action sequence to be generated corresponding to each target text segment.
[0077] The skeleton action sequence to be generated corresponding to each target text segment represents the positions of each skeleton point at different moments when the digital human outputs the target text segment by voice. This skeleton action sequence reflects the limb movements of the digital human when outputting the target text segment by voice.
[0078] In this application, according to the semantics of the text corresponding to the target speech, the text segments (denoted as target text segments) in the text that require the digital human to synchronously execute limb movements, the categories of limb movements to be synchronously executed, and the position encodings of each action frame in the skeleton action sequence corresponding to the target text segments are determined.
[0079] The position encoding of each action frame in the skeleton action sequence represents the position of this action frame in the skeleton action sequence.
[0080] Step S303: For each target text segment, at least use the above audio features, the category of the limb actions corresponding to the target text segment, and the position encodings of the respective action frames corresponding to the target text segment as the control conditions for the diffusion model, and generate a skeleton action sequence corresponding to the target text segment through the diffusion model. The generated skeleton action sequence is used to drive the digital human to perform the limb actions corresponding to the target text segment when the digital human outputs the target text segment in speech. If the limb actions corresponding to the target text segment are presented through video segments, each action frame is used to generate a video frame in the video segment corresponding to the target text segment.
[0081] A noise sequence of a preset length sampled randomly and the control conditions can be input into the diffusion model to obtain a skeleton action sequence of the preset length that the diffusion model denoises the noise sequence based on the control conditions. Among them, the control conditions at least include the above audio features, the category of the limb actions corresponding to the target text segment, and the position encodings of the respective action frames corresponding to the target text segment.
[0082] Among them, the noise sequence is composed of multiple noise frames, and the data volume of each noise frame in the noise sequence is the same as that of the action frame. For example, if there are two-dimensional coordinates of N×M skeleton key points in an action frame, then there are N×M×2 data in the action frame, and the data volume of each noise frame is also N×M×2.
[0083] In the case where the length of the skeleton action sequence to be generated corresponding to the target text segment is greater than the preset length, the skeleton action sequence corresponding to the target text segment can be generated in segments, that is, a skeleton action sequence of the preset length is generated each time until all the generated skeleton action sequences of the preset length can be combined to obtain the skeleton action sequence to be generated corresponding to the target text segment.
[0084] As an example, assume that the length of the noise sequence is 10. If the length of the skeleton action sequence to be generated corresponding to the target text segment is 15 (that is, the skeleton action sequence includes 15 action frames), then the skeleton action sequence needs to be generated in segments.
[0085] For example, first generate the first 10 action frames (i.e., the 1st - 10th action frames) in the skeleton action sequence for the first time, and generate the last 5 action frames (i.e., the 11th - 15th action frames) for the second time. In this way, the action frames generated twice are directly spliced to obtain the skeleton action sequence corresponding to the target text segment.
[0086] For another example, first generate the first 10 action frames in the skeleton action sequence (i.e., the 1st - 10th action frames), and then generate the subsequent 8 action frames the second time (i.e., the 8th - 15th action frames). That is, there is an overlapping window of 3 frames in the skeleton action sequences generated twice (i.e., the 8th - 10th action frames). In this way, when obtaining the skeleton action sequence corresponding to the target text segment based on the skeleton action sequences generated twice, the 1st - 7th action frames of the skeleton action sequence corresponding to the target text segment are the first 7 action frames among the 10 action frames generated the first time, the 11th - 15th action frames are the last 5 action frames among the 8 action frames generated the second time, and the 8th - 10th action frames are obtained by fusing the overlapping action frames in the overlapping window. In this way, the limb actions represented by the skeleton action sequence corresponding to the target text segment are more coherent.
[0087] As an example, assume that the length of the noise sequence is J, that is, the noise sequence includes J noise frames. If any continuous J action frames in the skeleton action sequence corresponding to the target text segment are to be generated, then for the jth (j = 1, 2, 3, ……, J) noise frame in the noise sequence of length J, the diffusion model can denoise the jth noise frame based on the control condition corresponding to the jth noise frame to obtain the action frame corresponding to the jth noise frame. Among them, the control condition corresponding to the jth noise frame at least includes: the above - mentioned audio features, the category of the limb actions corresponding to the target text segment, and the position encoding of the jth action frame in the above - mentioned continuous J action frames. Based on this, the action frame corresponding to the jth noise frame is the jth action frame in the above - mentioned continuous J action frames.
[0088] For example, assume that the length of the noise sequence is 10 and the length of the skeleton action sequence corresponding to the target text segment is 15. Then a noise sequence of length 10 can be randomly sampled first (for the convenience of description, denoted as the first noise sequence). For the jth (j = 1, 2, 3, ……, 10) noise frame in the first noise sequence, the diffusion model can denoise the jth noise frame based on the control condition corresponding to the jth noise frame to obtain the jth action frame in the skeleton action sequence corresponding to the target text segment. Among them, the control condition corresponding to the jth noise frame at least includes: the above - mentioned audio features, the category of the limb actions corresponding to the target text segment, and the position encoding of the jth action frame in the skeleton action sequence.
[0089] After generating the 1st to 10th action frames, it is necessary to generate the 8th to 15th action frames. At this time, a noise sequence of length 10 can be randomly sampled again (for the convenience of description, denoted as the second noise sequence). For the jth (j = 1, 2, 3, ……, 10) noise frame in the second noise sequence, the diffusion model can denoise the jth noise frame based on the control condition corresponding to the jth noise frame to obtain the (j + 7)th action frame in the skeleton action sequence corresponding to the target text segment. Among them, the control condition corresponding to the jth noise frame at least includes: the above audio features, the category of the limb actions corresponding to the target text segment, and the position encoding of the (j + 7)th action frame in the skeleton action sequence. The control conditions corresponding to the 8th to 10th noise frames in the second noise sequence are the same.
[0090] The category of limb actions corresponding to each target text segment can be represented by one-hot encoding. For example, assuming there are C categories of limb actions, the category of limb actions corresponding to each target text segment can be represented by a vector of length C. The elements at different positions in the vector correspond to different categories of limb actions. Among them, the value at the position corresponding to the category of limb actions corresponding to the target text segment is 1, and the values at other positions are 0.
[0091] The action data generation method provided in the embodiments of the present application proposes an action data generation scheme based on joint perception of speech and semantics. By using at least the audio features of the target speech, the category of limb actions corresponding to the target text segment determined based on the semantics of the text corresponding to the target speech, and the position encoding of each action frame in the skeleton action sequence to be generated as the control conditions of the diffusion model, the diffusion model is guided to generate action data that matches the control conditions, that is, the skeleton action sequence, improving the matching degree between the skeleton action sequence and the speech content, and thus improving the matching degree between the limb actions of the digital human driven by the skeleton action sequence and the speech content.
[0092] In an optional embodiment, based on the semantic understanding of the text, to determine the target text segment in the text that requires the digital human to synchronously execute limb actions, one implementation manner of the category of limb actions corresponding to each target text segment can be:
[0093] Process the text through a large model to insert marker tags in the text; each group of marker tags represents the start and end positions of the target text segment in the text that requires the digital human to synchronously execute limb actions, and the category of limb actions corresponding to each target text segment; the category of limb actions includes: intention and the corresponding limb actions.
[0094] Optionally, the large model can be a Large Language Model (LLM) or a Multimodal Large Language Model (MLLM).
[0095] A large language model refers to a model that can process large-scale natural language data. It is usually based on deep learning techniques, especially the Transformer architecture, and has powerful text understanding and generation capabilities. Of course, the network structure of the large language model is not limited to the Transformer architecture, and other network structures can also be used. This application does not make specific limitations in this regard.
[0096] A multimodal large model is a deep learning model that combines a large language model and a large vision model (or other types of models, such as audio models), and can process and understand various types of data, such as text, images, audio, and video, etc.
[0097] That is to say, through the semantic understanding of the text by the large model in this application, the start and end positions of the target text segment that requires the digital human to synchronously execute limb movements in the text, the intention of the target text segment, and the corresponding limb movements of the target text segment are determined.
[0098] Optionally, "intention-action" labels (for the convenience of narration and distinction, denoted as action labels) that conform to the application scenario can be designed in advance. The action labels represent the categories of limb movements. For example, greeting-wave, emphasizing-pointing at number one, emphasizing-raising hand, praising-giving a thumbs up, etc. Taking the action label "greeting-wave" as an example, "greeting" is the intention expressed by the text, and "wave" is the limb movement that should be executed for this intention. To facilitate the marking of the text, different placeholders can be used to represent different action labels, that is, different placeholders represent different categories of limb movements. For example, use <unused1>Indicates "Greeting - Waving", using <unused2>Indicates "farewell - waving", using <unused3>Denotes "Emphasis - More than number one", using <unused4>Indicates "emphasis - raising hand", using <unused5>It means "praise - giving a thumbs up", etc.
[0099] When inserting marker tags in the text, in order to represent the start and end positions of the target text segment, the start tag can be inserted after the placeholder. <start>And termination tag <end>, used to clarify the start and end times of limb movements. That is to say, the marker tags inserted in the text appear in pairs.
[0100] For example, assume the text corresponding to the target voice is "Friends, I believe you will surely give a thumbs up to this city like me and inadvertently find that touch. Goodbye, looking forward to meeting you in Anhui", and an example of inserting marker tags in this text can be:
[0101] Friends, I believe you will surely give a thumbs up to this city like me <unused5> <start>The city gives a thumbs up, <unused5> <end>And inadvertently find that touch. <unused2> <start>Goodbye, <unused2> <end>Looking forward to meeting you in Anhui.
[0102] In the above example, <unused5> <start>And <unused5> <end>This set of markup tags characterizes that the intention of the target text segment "The city gives a thumbs up" is to praise, and the corresponding body movement is to give a thumbs up; the starting position of the target text segment is the character "city", and the ending position of the target text segment is the character "finger"; similarly, <unused2> <start>and <unused2> <end>This set of markup tags characterizes that the intention of the target text segment "Goodbye" is to bid farewell, and the corresponding physical action is waving; the starting position of the target text segment is the character "zai", and the ending position of the target text segment is the character "le".
[0103] Optionally, the text corresponding to the target speech can be added to a preset prompt template to obtain a prompt, and the prompt is input into the large model to obtain the text after the insertion of the markup tags output by the large model based on the prompt. Among them, the prompt instructs the large model to predict the category of the physical action corresponding to the person when speaking, which is achieved by inserting action tags that conform to its semantics into the given text (i.e., the text added to the prompt template). The prompt also provides a list of action tags and instructs the large model to use the specified starting tag when inserting the action tags (for example, <start>) indicates the start of an action, using the specified termination tag (e.g., <end>Indicates the end of an action, etc.
[0104] As an example, the prompt template can be:
[0105] "... Predict the corresponding body movements of a person while speaking, and achieve this by inserting action tags that conform to the semantics in the given text. The action tags include, for greeting: <unused1>, indicating farewell <unused2>……Requirements: <start>Indicates the start of an action, <end>Indicates the end of the action...
[0106] Given text: ".
[0107] The above prompt template is only for illustrative purposes, to illustrate the structure and some content of the prompt template. "..." represents some omitted content.
[0108] When adding the text corresponding to the target speech to the prompt template, the text corresponding to the target speech is added after the "Given text".
[0109] For example, adding the text "Friends, I believe you will surely give a thumbs up to this city just like me and inadvertently find that touch. Goodbye and look forward to meeting you in Anhui" to the prompt template results in the following prompt:
[0110] "... Predict the corresponding body movements of the person when speaking, by inserting action tags that conform to the semantics in the given text. Action tags include, for greeting: <unused1>, indicating farewell <unused2>……Requirements: <start>Indicates the start of an action, <end>Indicates the end of the action...
[0111] Given text: "Friends, I believe you will surely give a thumbs up to this city like me and inadvertently find that touch. Goodbye, looking forward to meeting you in Anhui."
[0112] Optionally, the large model is trained by using text samples as training samples and the text after inserting marker tags corresponding to the text samples as sample labels in a supervised fine-tuning (SFT) manner for the general large model. The text samples are text samples of the application scenario. By performing supervised fine-tuning on the general large model, the accuracy and task adaptability of the large model in predicting intent-actions that conform to the context semantics from the text can be improved.
[0113] When training the general large model, add the text sample to the prompt template to obtain the prompt corresponding to the text sample, input the prompt corresponding to the text sample into the general large model to obtain the text after inserting marker tags output by the general large model, and update the parameters of the general large model with the goal that the text after inserting marker tags output by the large model approaches the sample label corresponding to the text sample.
[0114] In an optional embodiment, when determining the position encoding of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment, the position encoding of each action frame in the to-be-generated skeleton action sequence corresponding to the target text segment can be determined based on the alignment relationship between the determined target text segment and the target speech.
[0115] Optionally, a flowchart of an implementation for determining the position encoding of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment is as Figure 4 shown and may include:
[0116] Step S401: Determine the number of action frames in the to-be-generated skeleton action sequence corresponding to each target text segment according to the start and end positions of each character in each target text segment in the target speech and the preset frame rate.
[0117] The forced alignment (FA) technique can be used to align the target speech with the corresponding text word by word, that is, generate the start timestamp and end timestamp of each character in the text in the target speech.
[0118] The preset frame rate refers to the frame rate of the to-be-generated skeleton action sequence, and its essence is the frame rate of the video matched with the target speech. The video matched with the target speech is the video when the digital human outputs the text corresponding to the target speech in voice.
[0119] For each target text segment, the duration of the speech segment corresponding to the target text segment can be determined according to the start timestamp of the first character in the target text segment in the target speech and the end timestamp of the last character in the target text segment in the target speech (i.e., the duration from the start timestamp of the first character in the target text segment in the target speech to the end timestamp of the last character in the target text segment in the target speech). Multiply this duration by the preset frame rate to obtain the number of action frames in the to-be-generated skeleton action sequence corresponding to the target text segment.
[0120] Optionally, the duration of the speech segment corresponding to the target text segment can also be determined in the following way: According to the start timestamp and end timestamp of each character in the target text segment in the target speech, determine the duration of the speech segment corresponding to each character in the target text segment, and sum up the durations of the speech segments corresponding to the respective characters in the target text segment to obtain the duration of the speech segment corresponding to the target text segment.
[0121] Step S402: Obtain the position encoding of each action frame in the skeleton action sequence according to the position of each action frame in the skeleton action sequence.
[0122] Optionally, assume that the position encoding is a d-dimensional vector, and the number of action frames in the to-be-generated skeleton action sequence corresponding to a certain target text segment is P. For each action frame, according to the position of the action frame in the skeleton action sequence (denoted as pos, pos = 0, 1, 2, 3, ……, P - 1), and the position of each element in the position encoding corresponding to the action frame (a d-dimensional vector) in the position encoding, determine the value of each element in the position encoding corresponding to the action frame.
[0123] As an example, for the k-th (k = 0, 1, 2, 3, ……, d) element in the position encoding of each action frame, the value of the k-th element can be determined in the following way.
[0124] (1)
[0125] Where is the value of the k-th element in the position encoding corresponding to the pos-th action frame.
[0126] In an optional embodiment, since the marked tags are inserted into the text by the large model and errors may occur, to further improve the accuracy of the action data, after inserting the marked tags into the text, before determining the position encodings of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment, the text after inserting the marked tags can be corrected for tags first, so as to delete abnormal marked tags or correct relevant information for abnormal marked tags, and then based on the corrected text after inserting the marked tags, determine the target text segments and the position encodings of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment.
[0127] Optionally, an implementation manner of correcting the text after inserting the marked tags provided by this application can be:
[0128] Compare the text after removing the marked tags with the text corresponding to the target speech. If the comparison result indicates that the two are inconsistent, it means that when the large model inserts tags into the text, it modifies the text corresponding to the target speech. Such a modification may change the semantics of the text corresponding to the target speech. Therefore, discard the text after inserting the marked tags, that is, consider that there is no text segment in the text corresponding to the target speech that requires synchronized limb actions, and directly use the text corresponding to the target speech for subsequent processing. Since there is no target text segment, the above process of determining the position encodings of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment is not executed.
[0129] If the comparison result indicates that the two are consistent, it means that when the large model inserts tags into the text, it does not modify the text corresponding to the target speech, and then execute the above process of determining the position encodings of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment.
[0130] Optionally, another implementation manner of correcting the text after inserting the marked tags provided by this application can be:
[0131] Judge the normativity of each group of marked tags, delete each group of marked tags that do not meet the preset norms, and only retain each group of marked tags that meet the preset norms.
[0132] As an example, the preset norms may include but are not limited to the following three items:
[0133] Belong to a preset set of action categories. Based on this, for any set of markup tags, if the action category in this set of markup tags does not belong to the preset set of action categories, it is determined that this set of markup tags does not conform to the preset specification. Specifically, for any set of markup tags, if the placeholder in this set of markup tags belongs to the preset set of placeholders, it is determined that the limb action category represented by the placeholder in this set of markup tags belongs to the preset set of limb action categories and conforms to the pre-examination specification; if the placeholder in this set of markup tags does not belong to the preset set of placeholders, it is determined that the limb action category represented by the placeholder in this set of markup tags does not belong to the preset set of limb action categories and does not conform to the preset specification.
[0134] The placeholders are paired. Based on this, for any set of markup tags, if the placeholders in this set of markup tags are not paired, it is determined that this set of markup tags does not conform to the preset specification.
[0135] The start tag and the end tag are paired. Based on this, for any set of markup tags, if the start tag and the end tag in this set of markup tags are not paired (that is, there is only a start tag, or, there is only an end tag, or, the start tag is incorrect, or, the end tag is incorrect, etc.), it is determined that this set of markup tags does not conform to the preset specification.
[0136] Optionally, another implementation manner provided by this application for correcting the tags of the text after inserting the markup tags may be:
[0137] Based on the start and end positions of each character in each target text segment in the target speech, determine the duration of the skeleton action sequence corresponding to each target text segment.
[0138] If the duration of the skeleton action sequence corresponding to any target text segment does not meet the target duration range of the limb action represented by the markup tag for marking this any target text segment, update the duration of the skeleton action sequence corresponding to this any target text segment to a duration within the above target duration range.
[0139] Among them, the target duration range of any limb action can be obtained by statistically analyzing the duration of this any limb action in a number of real videos.
[0140] Optionally, when updating the duration of the skeleton action sequence corresponding to this any target text segment to a duration within the above duration range, a duration can be randomly sampled within the above target duration range as the duration of the skeleton action sequence corresponding to this any target text segment.
[0141] Optionally, another implementation manner provided by this application for correcting the tags of the text after inserting the markup tags may be:
[0142] Compare the text after removing the markup tags with the text corresponding to the target speech. If the comparison result indicates that the two are inconsistent, discard the text after inserting the markup tags, that is, use the text corresponding to the target speech for subsequent processing, and do not perform the process of determining the position encoding of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment. For the specific implementation method, refer to the foregoing embodiments and will not be elaborated here.
[0143] If the comparison result indicates that the two are consistent, then perform a normative judgment on each group of markup tags, delete each group of markup tags that do not meet the preset specifications, and only retain each group of markup tags that meet the preset specifications. For the specific implementation method, refer to the foregoing embodiments and will not be elaborated here.
[0144] Optionally, another implementation manner for correcting the tags of the text after inserting the markup tags provided by this application may be:
[0145] Compare the text after removing the markup tags with the text corresponding to the target speech. If the comparison result indicates that the two are inconsistent, discard the text after inserting the markup tags, that is, use the text corresponding to the target speech for subsequent processing, and do not perform the process of determining the position encoding of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment. For the specific implementation method, refer to the foregoing embodiments and will not be elaborated here.
[0146] If the comparison result indicates that the two are consistent, then determine the duration of the skeleton action sequence corresponding to each target text segment based on the start and end positions of each character in each target text segment in the target speech; if the duration of the skeleton action sequence corresponding to any target text segment does not meet the target duration range of the limb action represented by the markup tag for marking the any target text segment, update the duration of the skeleton action sequence corresponding to the any target text segment to a duration within the above target duration range. For the specific implementation method, refer to the foregoing embodiments and will not be elaborated here.
[0147] Optionally, another implementation manner for correcting the tags of the text after inserting the markup tags provided by this application may be:
[0148] Perform a normative judgment on each group of markup tags, delete each group of markup tags that do not meet the preset specifications, and only retain each group of markup tags that meet the preset specifications. For the specific implementation method, refer to the foregoing embodiments and will not be elaborated here.
[0149] Based on the start and end positions of each character in each remaining target text segment in the target speech, determine the duration of the skeleton action sequence corresponding to each remaining target text segment; if, among the remaining target text segments, the duration of the skeleton action sequence corresponding to any target text segment does not meet the target duration range of the limb action represented by the marking label for marking that any target text segment, update the duration of the skeleton action sequence corresponding to that any target text segment to a duration within the above-mentioned target duration range. For the specific implementation method, refer to the foregoing embodiments and will not be elaborated here.
[0150] Optionally, another implementation method for correcting the labels of the text after inserting the marking labels provided in this application may be:
[0151] Compare the text after removing the marking labels of the text after inserting the marking labels with the text corresponding to the target speech. If the comparison result indicates that the two are inconsistent, discard the text after inserting the marking labels, that is, use the text corresponding to the target speech for subsequent post-processing, and do not perform the process of determining the position encoding of each action frame in the skeleton action sequence to be generated corresponding to each target text segment. For the specific implementation method, refer to the foregoing embodiments and will not be elaborated here.
[0152] If the comparison result indicates that the two are consistent, then perform a normative judgment on each group of marking labels, delete each group of marking labels that do not meet the preset specifications, and only retain each group of marking labels that meet the preset specifications. For the specific implementation method, refer to the foregoing embodiments and will not be elaborated here.
[0153] Based on the start and end positions of each character in each remaining target text segment in the target speech, determine the duration of the skeleton action sequence corresponding to each remaining target text segment; if, among the remaining target text segments, the duration of the skeleton action sequence corresponding to any target text segment does not meet the target duration range of the limb action represented by the marking label for marking that any target text segment, update the duration of the skeleton action sequence corresponding to that any target text segment to a duration within the above-mentioned target duration range. For the specific implementation method, refer to the foregoing embodiments and will not be elaborated here.
[0154] In an optional embodiment, one implementation method for obtaining the audio features of the target speech may be:
[0155] Extract the hidden layer features of the target speech through a general speech pre-training model. That is to say, this application can use the hidden layer features of the target speech as the audio features of the target speech.
[0156] Optionally, the general speech pre-training model may include but is not limited to any one of the following: WavLM model, Hubert model, wav2vec model, etc.
[0157] As an example, if the speech pre-training model adopts an encoder-decoder architecture, such as the WavLM model, the audio features of the target speech can be the hidden layer features output by the decoding module (decoder) of the speech pre-training model.
[0158] In an optional embodiment, one implementation manner of obtaining the audio features of the target speech can be:
[0159] Extract the onset points (onsets) and beat information (beats) of the audio signal from the target speech through a speech signal processing library as the rhythm features of the target speech. That is to say, the present application can use the rhythm features of the target speech as the audio features of the target speech.
[0160] Optionally, the speech signal processing library can include but is not limited to at least one of the following: librosa, Audacity, etc.
[0161] In an optional embodiment, one implementation manner of obtaining the audio features of the target speech can be:
[0162] Extract the hidden layer features of the target speech through a general speech pre-training model.
[0163] Extract the onset points (onsets) and beat information (beats) of the audio signal from the target speech through a speech signal processing library as the rhythm features of the target speech.
[0164] That is to say, the present application can use the hidden layer features and rhythm features of the target speech as the audio features of the target speech.
[0165] In an optional embodiment, the diffusion model generates a skeleton action sequence including J action frames each time. For any target text segment, if the number of action frames included in the to-be-generated skeleton action sequence corresponding to the any target text segment is less than or equal to J, when generating the skeleton action sequence corresponding to the any target text segment through the diffusion model, only generate one skeleton action sequence including J action frames corresponding to the target text segment through the diffusion model. In this case, if the to-be-generated skeleton action sequence corresponding to the any target text segment includes F action frames, where F < J, the control conditions and noise frames are the same when the diffusion model generates the last J - F + 1 action frames.
[0166] If the number of action frames included in the to-be-generated skeleton action sequence corresponding to any one of the target text segments is greater than J, when generating the skeleton action sequence corresponding to any one of the target text segments through the diffusion model, it is necessary to generate at least two skeleton action sequences including J action frames corresponding to the target text segment through the diffusion model. For the convenience of narration and distinction, the at least two skeleton action sequences including J action frames are denoted as at least two initial skeleton action sequences. The at least two initial skeleton action sequences satisfy the following conditions:
[0167] Among two adjacent initial skeleton action sequences, the last m action frames in the previous initial skeleton action sequence correspond to the same sub-skeleton action sequence in the skeleton action sequence corresponding to the target text segment as the first m action frames in the subsequent initial skeleton action sequence, and the sub-skeleton action sequence includes m action frames.
[0168] As an example, assume that the diffusion model generates a skeleton action sequence including 10 action frames each time (that is, the diffusion model generates a skeleton action sequence with a length of 10 each time). If the length of the to-be-generated skeleton action sequence corresponding to the target text segment is 15 (that is, the skeleton action sequence includes 15 action frames), then it is necessary to generate the skeleton action sequence in segments.
[0169] For example, first generate the first 10 action frames (that is, the 1st - 10th action frames) in the to-be-generated skeleton action sequence corresponding to the target text segment for the first time, and generate the last 8 action frames (that is, the 8th - 15th action frames) for the second time. That is, there is an overlapping window of 3 (that is, m = 3) frames in the skeleton action sequences generated twice (that is, the 8th - 10th action frames in the skeleton action sequence corresponding to the target text segment). In this way, when obtaining the skeleton action sequence corresponding to the target text segment based on the skeleton action sequences generated twice, the 1st - 7th action frames in the skeleton action sequence corresponding to the target text segment are the first 7 action frames among the 10 action frames generated for the first time, the 11th - 15th action frames are the last 5 action frames among the 8 action frames generated for the second time, and the 8th - 10th action frames are obtained by fusing the overlapping action frames in the overlapping window. In this way, the limb actions represented by the skeleton action sequence corresponding to the target text segment are more coherent and smooth.
[0170] Perform weighted summation on the m action frames corresponding to the same sub-skeleton action sequence in two adjacent initial skeleton action sequences to obtain m target action frames corresponding to the same sub-skeleton action sequence; among them, the weights of the last m skeleton actions in the previous initial skeleton action sequence gradually decrease, and the weights of the first m skeleton actions in the subsequent initial skeleton action sequence gradually increase.
[0171] In the above example, perform weighted summation on the 8th action frame in the previous initial skeleton action sequence and the 1st action frame in the subsequent initial skeleton action sequence to obtain the 8th action frame in the skeleton action sequence corresponding to the target text segment.
[0172] Weight-sum the 9th action frame in the previous initial skeleton action sequence and the 2nd action frame in the subsequent initial skeleton action sequence to obtain the 9th action frame in the skeleton action sequence corresponding to the target text segment.
[0173] Weight-sum the 10th action frame in the previous initial skeleton action sequence and the 3rd action frame in the subsequent initial skeleton action sequence to obtain the 10th action frame in the skeleton action sequence corresponding to the target text segment.
[0174] Among them, in the previous initial skeleton action sequence, the weight of the 8th action frame is greater than the weight of the 9th action frame, and the weight of the 9th action frame is greater than the weight of the 10th action frame; in the subsequent initial skeleton action sequence, the weight of the 1st action frame is less than the weight of the 2nd action frame, and the weight of the 2nd action frame is less than the weight of the 3rd action frame.
[0175] In an optional embodiment, the action data generation method provided by the embodiments of the present application may further include:
[0176] Obtain the input reference person image, perform feature extraction on the reference person image to obtain skeleton features.
[0177] The reference person image may be a single half-body photo or a full-body front photo of a person input by the user.
[0178] Optionally, skeleton point extraction may be performed on the reference person image to obtain an initial skeleton point set; translation and / or scaling may be performed on the initial skeleton point set to obtain a target skeleton point set; feature extraction may be performed on the target skeleton point set to obtain skeleton features.
[0179] As an example, a pre-trained pose estimation model may be used to process the reference person image to extract the coordinates of multiple skeleton key points . As an example, assuming the skeleton structure shown in Figure 1 , the coordinates of 18 skeleton key points numbered from 0 to 17 can be extracted. Assuming the skeleton structure shown in Figure 2 , the coordinates of 17 skeleton key points numbered from 0 to 16 can be extracted.
[0180] The pose estimation model may include, but is not limited to, any one of the following models: DWPose, RTMPose, DeepPose, Stacked Hourglass, CPN, HRNet, etc.
[0181] As an example, the main skeleton point set in the initial skeleton point set may be used (for example, Figure 1 In the shown skeletal structure, the skeletal key points numbered 0, 1, 2, 5, 8, 11, 14, 15, 16, and 17 are the main skeletal key points. The positional relationship between these main skeletal key points and the main skeletal point set in the reference skeletal diagram, as well as the difference between the relative positional relationship between the main skeletal key point set in the initial skeletal point set and the relative positional relationship between the main skeletal point set in the reference skeletal diagram, are used to translate and / or scale the initial skeletal point set so that the position and size of the skeletal structure formed by the initial skeletal point set in the skeletal diagram are the same as or similar to the position and size of the skeletal structure in the reference skeletal diagram in the reference skeletal diagram.
[0182] The coordinates of each skeletal key point in the target skeletal point set can be concatenated in order to form a vector (the dimension of this vector is 2M×1), which is used as the skeletal feature.
[0183] Alternatively, the coordinates of each skeletal key point in the target skeletal point set can be concatenated in order to form a vector, and this vector is linearly transformed to obtain the skeletal feature.
[0184] Optionally, a pre-trained pose estimation model can be used to extract skeletal key points from the reference person image to obtain an initial skeletal point set; the initial skeletal point set is translated and / or scaled to obtain a target skeletal point set; feature extraction is performed on the target skeletal point set to obtain a first feature; a second feature output by a preset intermediate layer of the pose estimation model is obtained; the first feature and the second feature constitute the skeletal feature.
[0185] As an example, the coordinates of each skeletal key point in the target skeletal point set can be concatenated in order to form a vector (the dimension of this vector is 2M×1), which is used as the first feature.
[0186] Alternatively, the coordinates of each skeletal key point in the target skeletal point set can be concatenated in order to form a vector, and this vector is linearly transformed to obtain the first feature.
[0187] As an example, if the pose estimation model adopts an encoder-decoder architecture, such as the DWPose model, the second feature can be the hidden layer feature output by a certain layer (which can be the last layer of the decoder or other layers of the decoder) of the decoding module (decoder) of the pose estimation model.
[0188] In the case of obtaining the skeletal feature, the control condition of the diffusion model can also include: the skeletal feature. When the skeletal feature is included in the control condition, the generated skeletal action sequence matches the skeletal structure of the person in the reference person image.
[0189] In an optional embodiment, the action data generation method provided by the embodiments of the present application may further include:
[0190] Obtain the character style information. The character style information can be input by the user or be the default character style information of the system.
[0191] Correspondingly, the control conditions of the diffusion model can also include: the character style information. When the character style information is included in the control conditions, the style of the limb movements represented by the generated skeleton action sequence matches the obtained character style information.
[0192] In an alternative embodiment, the method for generating action data provided by the embodiments of the present application may further include:
[0193] Obtain the input reference character image, extract features from the reference character image to obtain skeleton features.
[0194] Obtain the character style information.
[0195] Correspondingly, the control conditions of the diffusion model can also include: the skeleton features and the character style information.
[0196] In an alternative embodiment, the diffusion model is trained with a number of skeleton action sequences as the training sample set, and the target features corresponding to each skeleton action sequence are used as the control conditions of the diffusion model;
[0197] Among them, different skeleton action sequences in the training sample set are obtained by extracting skeleton points from video frames in different video segments; a limb movement is synchronously shown by the character in each video segment when outputting speech.
[0198] The target features corresponding to each skeleton action sequence in the training sample set at least include: the audio features of the speech synchronized with the video segment corresponding to the skeleton action sequence, the category of the limb movements in the video segment corresponding to the skeleton action sequence, and the position encodings of each video frame in the video segment corresponding to the skeleton action sequence.
[0199] Optionally, the target features corresponding to each skeleton action sequence in the training sample set may further include at least one of the following: the skeleton features of the reference character image, the target character style information.
[0200] The process of training the diffusion model based on the control conditions can refer to existing solutions and will not be elaborated here.
[0201] In an alternative embodiment, during the training process of the diffusion model, some training samples are randomly selected, and the category and / or position encoding of the limb movements in the target features corresponding to the randomly selected training samples are set to zero.
[0202] By performing a drop operation on some control conditions (that is, setting the category and / or position encoding of the limb movements in the target features corresponding to the randomly selected training samples to zero), the generalization ability of the diffusion model can be enhanced, overfitting can be prevented, the model robustness can be improved, and feature diversity can be promoted.
[0203] In an optional embodiment, each skeleton action sequence in the training sample set includes a set of human skeleton points (for example, Figure 1 or Figure 2 the set of skeleton points shown) and a set of hand key points. Correspondingly, in the skeleton action sequence output by the diffusion model, each action frame includes a set of human skeleton points and a set of hand key points.
[0204] As an example, there are 21 key points in the set of hand key points of the left hand, and there are also 21 key points in the set of hand key points of the right hand.
[0205] During the training process of the diffusion model, the parameters of the diffusion model are updated based on the difference between the noise point set sequence predicted by the diffusion model and the noise point set sequence input to the diffusion model.
[0206] Among them, when calculating the difference between the noise point sets at the same position in the noise point set sequence predicted by the diffusion model and the noise point set sequence input to the diffusion model, calculate the overlapping degree of the two hands in the video frame corresponding to the noise point sets at the same position; if the overlapping degree is greater than the threshold, it is considered that the overlapping degree of the two hands is high and the posture is complex, then when calculating the difference between the noise point sets at the same position, set the difference in the hand area to zero.
[0207] For example, when calculating the difference between the noise point sets at the j-th position in the noise point set sequence predicted by the diffusion model and the noise point set sequence input to the diffusion model, calculate the overlapping degree of the two hands in the video frame corresponding to the noise point sets at the j-th position; if the overlapping degree is greater than the threshold, it is considered that the overlapping degree of the two hands is high and the posture is complex, then when calculating the difference between the noise point sets at the j-th position, set the difference in the hand area to zero.
[0208] Among them, in any video frame, the overlapping degree of the two hands can be determined by the following method:
[0209] Determine the minimum bounding rectangle of the left hand (for the convenience of description and distinction, denoted as the first minimum bounding rectangle) and the minimum bounding rectangle of the right hand (for the convenience of description and distinction, denoted as the second minimum bounding rectangle) in the any video frame; calculate the intersection over union (IoU) of the first minimum bounding rectangle and the second minimum bounding rectangle as the overlapping degree of the two hands.
[0210] During the training of the diffusion model, by setting the differences in the hand regions when the hands are complexly overlapped to zero, the interference of the hands can be avoided and the robustness of the model can be improved.
[0211] Corresponding to the method embodiment, an embodiment of the present application provides an action data generation device. A schematic structural diagram of the action data generation device provided by the embodiment of the present application is as Figure 5 shown, and may include:
[0212] An acquisition module 501, a determination module 502 and a generation module 503;
[0213] Among them, the acquisition module 501 is used to acquire the audio features of the target speech and the text corresponding to the target speech; the audio features include at least one of the following: the rhythm features of the target speech, and the hidden layer features obtained by performing hidden layer feature extraction on the target speech;
[0214] The determination module 502 is used to determine, based on the semantic understanding of the text, the target text segments in the text that require the digital human to synchronously execute limb actions, the categories of limb actions corresponding to each target text segment, and the position encodings of each action frame in the skeleton action sequence to be generated corresponding to each target text segment; the skeleton action sequence corresponding to each target text segment represents the positions of each skeleton point at different moments when the digital human outputs the target text segment in speech.
[0215] The generation module 503 is used to, for each target text segment, use at least the audio features, the category of the limb action corresponding to the target text segment, and the position encodings of each action frame corresponding to the target text segment as control conditions of the diffusion model, and generate a skeleton action sequence corresponding to the target text segment through the diffusion model.
[0216] The action data generation device provided by the embodiment of the present application proposes an action data generation scheme based on joint perception of speech and semantics. By using at least the audio features of the target speech, the categories of limb actions corresponding to the target text segments determined based on the semantics of the text corresponding to the target speech, and the position encodings of each action frame in the skeleton action sequence to be generated as control conditions of the diffusion model, the diffusion model is guided to generate action data that matches the control conditions, that is, the skeleton action sequence, improving the matching degree between the skeleton action sequence and the speech content, and thus improving the matching degree between the limb actions of the digital human driven by the skeleton action sequence and the speech content.
[0217] In an optional embodiment, when the determination module 502 determines the target text segments in the text that require the digital human to synchronously execute limb actions and the categories of limb actions corresponding to each target text segment based on the semantic understanding of the text, it is used for:
[0218] Process the text through a large model to insert marker tags into the text; each set of marker tags represents the start and end positions of a target text segment in the text that requires the digital human to synchronously perform limb actions, and the limb action category corresponding to each target text segment; the limb action category includes: an intention and the corresponding limb action.
[0219] In an alternative embodiment, the determination module 502 determines, based on the semantic understanding of the text, the target text segments in the text that require the digital human to synchronously perform limb actions, and the categories of the limb actions corresponding to each target text segment, including:
[0220] Process the text through a large model to insert marker tags into the text; each set of marker tags represents the start and end positions of a target text segment in the text that requires the digital human to synchronously perform limb actions, and the limb action category corresponding to each target text segment; the limb action category includes: an intention and the corresponding limb action;
[0221] Perform label correction on the text after inserting the marker tags to delete abnormal marker tags or correct relevant information for abnormal marker tags.
[0222] In an alternative embodiment, when the determination module 502 performs label correction on the text after inserting the marker tags, it is used to: perform label correction on the text after inserting the marker tags by using at least one of the following methods:
[0223] Compare the text after removing the marker tags with the text corresponding to the target speech. If the comparison result indicates inconsistency, discard the text after inserting the marker tags and use the text corresponding to the target speech;
[0224] Perform a normative judgment on each set of marker tags and delete each set of marker tags that does not conform to the preset specification;
[0225] Based on the start and end positions of each character in each target text segment in the target speech, determine the duration of the skeleton action sequence corresponding to each target text segment; if the duration of the skeleton action sequence corresponding to any target text segment does not meet the target duration range of the limb action represented by the marker tags marking the any target text segment, update the duration of the skeleton action sequence corresponding to the any target text segment to a duration within the target duration range.
[0226] In an alternative embodiment, when the determination module 502 determines the position encoding of each action frame in the to-be-generated skeleton action sequence corresponding to each target text segment, it is used to:
[0227] Determine the number of action frames in the to-be-generated skeleton action sequence corresponding to each target text segment according to the start and end positions of each character in the target text segment in the target speech and a preset frame rate;
[0228] Obtain the position encoding of each action frame in the skeleton action sequence according to the position of each action frame in the skeleton action sequence.
[0229] In an optional embodiment, when the obtaining module 501 obtains the audio features of the target speech, it is used for:
[0230] Extract features from the target speech through a general speech pre-training model to obtain the hidden layer features of the target speech;
[0231] And / or, extract the audio start point and beat information of the audio signal from the target speech through a speech signal processing library as the rhythm features of the target speech.
[0232] In an optional embodiment, the generating module 503 generates at least two initial skeleton action sequences corresponding to the target text segment through the diffusion model; among adjacent two initial skeleton action sequences, the last m action frames in the previous initial skeleton action sequence and the first m action frames in the subsequent initial skeleton action sequence correspond to the same sub-skeleton action sequence in the skeleton action sequence corresponding to the target text segment, and the sub-skeleton action sequence includes m action frames;
[0233] The generating module 503 is further configured to perform weighted summation on the m action frames corresponding to the same sub-skeleton action sequence in adjacent two initial skeleton action sequences to obtain m target action frames corresponding to the same sub-skeleton action sequence; among them, the weights of the last m action frames in the previous initial skeleton action sequence gradually decrease, and the weights of the first m action frames in the subsequent initial skeleton action sequence gradually increase.
[0234] In an optional embodiment, the obtaining module 501 is further configured to perform at least one of the following:
[0235] Obtain the input reference person image, extract features from the reference person image to obtain skeleton features;
[0236] Obtain the person style information;
[0237] Correspondingly, the control conditions of the diffusion model further include:
[0238] At least one of the skeleton features and the person style information.
[0239] In an optional embodiment, when the obtaining module 501 extracts features from the reference person image to obtain skeleton features, it is used for:
[0240] Extract the skeleton points from the reference person image to obtain an initial skeleton point set; perform translation and / or scaling on the initial skeleton point set to obtain a target skeleton point set; perform feature extraction on the target skeleton point set to obtain the skeleton feature;
[0241] Alternatively, extract the skeleton points from the reference person image through a pose estimation model to obtain an initial skeleton point set; perform translation and / or scaling on the initial skeleton point set to obtain a target skeleton point set; perform feature extraction on the target skeleton point set to obtain a first feature; the second feature output by a preset intermediate layer of the pose estimation model and the first feature constitute the skeleton feature.
[0242] In an optional embodiment, the diffusion model is trained with a number of skeleton action sequences as the training sample set and the target feature corresponding to each skeleton action sequence as the control condition of the diffusion model;
[0243] The different skeleton action sequences in the training sample set are obtained by extracting skeleton points from the video frames in different video segments; a limb action is synchronously displayed by the person in each video segment when outputting speech;
[0244] The target feature corresponding to each skeleton action sequence in the training sample set at least includes: the audio feature of the speech synchronized with the video segment corresponding to the skeleton action sequence, the category of the limb action in the video segment corresponding to the skeleton action sequence, and the position encoding of each video frame in the video segment corresponding to the skeleton action sequence.
[0245] In an optional embodiment, the action data generation device further includes a training module for training the diffusion model;
[0246] During the training process of the diffusion model by the training module, some training samples are randomly selected, and the category and / or position encoding of the limb action in the target feature corresponding to the randomly selected training samples are set to zero.
[0247] In an optional embodiment, the action data generation device further includes a training module for training the diffusion model; in the skeleton action sequence output by the diffusion model, each action frame includes a human skeleton point set and a hand key point set;
[0248] During the training process of the diffusion model by the training module, the parameters of the diffusion model are updated based on the difference between the noise point set sequence predicted by the diffusion model and the noise point set sequence input to the diffusion model;
[0249] Among them, when calculating the difference between the noise point set sequences predicted by the diffusion model and the noise point sets at the same positions of the noise point set sequence input into the diffusion model, calculate the overlapping degree of the two hands in the video frames corresponding to the noise point sets at the same positions; if the overlapping degree is greater than the threshold, when calculating the difference between the noise point sets at the same positions, set the difference in the hand region to zero.
[0250] An electronic device is also provided in an embodiment of the present application. Refer to Figure 6 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the embodiment of the present application. The electronic device in the embodiment of the present application can be a terminal device (such as, a vehicle-mounted computer, a large-screen device, a smart home, a mobile phone, a tablet computer, a laptop computer, a desktop computer, etc.), or a server (it can be a single server, or a server cluster, or it can be a cloud server, etc.). Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present application.
[0251] As Figure 6 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0252] Generally, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.
[0253] An embodiment of the present application also provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any of the action data generation methods provided in the embodiments of the present application.
[0254] In an embodiment of the present application, a computer-readable storage medium is further provided. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the action data generation methods provided in the embodiments of the present application.
[0255] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in the present application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0256] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for the present application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0257] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. Professional technicians can use different methods to implement the described functions for each specific solution, but such implementation should not be considered to exceed the scope of the present application.
[0258] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions may be transmitted from a website, a computer, a training device, or a data center to another website, a computer, a training device, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0259] The embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. For the same or similar parts among the embodiments, reference may be made to each other.
[0260] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.< / end> < / start> < / end> < / start> < / end> < / start> < / end> < / unused2> < / start> < / unused2> < / end> < / unused5> < / start> < / unused5> < / end> < / unused2> < / start> < / unused2> < / end> < / unused5> < / start> < / unused5> < / end> < / start>
Claims
1. A method for generating motion data, characterized in that: include: Obtaining audio features of a target speech and text corresponding to the target speech; the audio features include at least one of the following: rhythm features of the target speech, and hidden features obtained by performing hidden feature extraction on the target speech; Based on the semantic understanding of the text, determine the target text segment in the text that requires the digital human to synchronously perform body movements, the category of body movements corresponding to each target text segment, and the position codes of each action frame in the skeleton action sequence to be generated corresponding to each target text segment; the skeleton action sequence corresponding to each target text segment represents the positions of each skeleton point at different times during the process of the digital human outputting the target text segment in voice; For each target text segment, at least the audio feature, the category of the body movement corresponding to the target text segment, and the position encoding of each action frame corresponding to the target text segment are used as control conditions of the diffusion model, and a skeleton action sequence corresponding to the target text segment is generated through the diffusion model; The diffusion model is obtained by training with a number of skeleton action sequences as training sample sets and with the target features corresponding to each skeleton action sequence as the control conditions of the diffusion model; The different skeleton action sequences in the training sample set are obtained by extracting skeleton points from video frames in different video clips; the characters in each video clip synchronously display a body movement when outputting speech; The target features corresponding to each skeleton action sequence in the training sample set include at least: audio features of speech synchronized with the video clip corresponding to the skeleton action sequence, categories of body movements in the video clip corresponding to the skeleton action sequence, and position codes of each video frame in the video clip corresponding to the skeleton action sequence.
2. The method according to claim 1, characterized in that Based on the semantic understanding of the text, a target text segment in the text that requires the digital human to synchronously perform body movements is determined. The category of body movements corresponding to each target text segment includes: The text is processed by a large model to insert mark tags into the text; each group of mark tags represents the start and end positions of the target text segments in the text that require the digital human to synchronously perform physical movements, as well as the physical movement category corresponding to each target text segment; the physical movement category includes: intention and the physical movement corresponding to the intention.
3. The method according to claim 1, characterized in that Based on the semantic understanding of the text, a target text segment in the text that requires the digital human to synchronously perform body movements is determined. The category of body movements corresponding to each target text segment includes: The text is processed by a large model to insert marking tags into the text; each group of marking tags represents the start and end positions of the target text segments in the text that require the digital human to synchronously perform body movements, and the body movement category corresponding to each target text segment; the body movement category includes: intention and the body movement corresponding to the intention; The text after the mark label is inserted is corrected to delete the abnormal mark label or correct the relevant information for the abnormal mark label.
4. The method according to claim 3, characterized in that The tag correction of the text after the inserted mark tag includes: using at least one of the following methods to correct the text after the inserted mark tag: The text after the inserted mark tag is removed and compared with the text corresponding to the target speech. If the comparison result is inconsistent, the text after the inserted mark tag is discarded and the text corresponding to the target speech is used; Conduct normative judgment on each group of mark labels, and delete each group of mark labels that do not meet the preset norms; Based on the start and end positions of each character in each target text segment in the target speech, the duration of the skeleton action sequence corresponding to each target text segment is determined; if the duration of the skeleton action sequence corresponding to any target text segment does not meet the target duration range of the limb action represented by the marking tag used to mark the any target text segment, the duration of the skeleton action sequence corresponding to the any target text segment is updated to a duration within the target duration range.
5. The method according to any one of claims 2 to 4, characterized in that: Determine the position encoding of each action frame in the skeleton action sequence to be generated corresponding to each target text segment, including: Determine the number of action frames in the skeleton action sequence to be generated corresponding to each target text segment according to the start and end positions of each character in each target text segment in the target speech and a preset frame rate; According to the position of each action frame in the skeleton action sequence, the position code of each action frame in the skeleton action sequence is obtained.
6. The method according to claim 1, characterized in that The step of obtaining the audio features of the target speech comprises: Extracting features of the target speech using a general speech pre-training model to obtain hidden features of the target speech; And / or, extracting the audio starting point and beat information of the audio signal from the target speech through a speech signal processing library as the rhythm feature of the target speech.
7. The method according to claim 1, characterized in that At least two initial skeleton action sequences corresponding to the target text segment are generated by the diffusion model; wherein, in two adjacent initial skeleton action sequences, the last m action frames in the previous initial skeleton action sequence and the first m action frames in the next initial skeleton action sequence correspond to the same sub-skeleton action sequence in the skeleton action sequence corresponding to the target text segment, and the sub-skeleton action sequence includes m action frames; The m action frames corresponding to the same sub-skeleton action sequence in two adjacent initial skeleton action sequences are weighted and summed to obtain m target action frames corresponding to the same sub-skeleton action sequence; among them, the weights of the last m action frames in the previous initial skeleton action sequence gradually decrease, and the weights of the first m action frames in the next initial skeleton action sequence gradually increase.
8. The method according to claim 1, characterized in that Also includes at least one of the following: Obtaining an input reference person image, performing feature extraction on the reference person image, and obtaining skeleton features; Get character style information; Accordingly, the control conditions of the diffusion model also include: At least one of the skeleton feature and the character style information.
9. The method according to claim 8, characterized in that Extracting features from the reference person image to obtain skeleton features includes: Extracting skeleton points from the reference character image to obtain an initial skeleton point set; translating and / or scaling the initial skeleton point set to obtain a target skeleton point set; extracting features from the target skeleton point set to obtain the skeleton features; Alternatively, skeleton points are extracted from the reference character image through a posture estimation model to obtain an initial skeleton point set; the initial skeleton point set is translated and / or scaled to obtain a target skeleton point set; feature extraction is performed on the target skeleton point set to obtain a first feature; a second feature output by a preset intermediate layer of the posture estimation model and the first feature constitute the skeleton feature.
10. The method according to claim 1, characterized in that During the training of the diffusion model, some training samples are randomly selected, and the category and / or position codes of the limb movements in the target features corresponding to the randomly selected training samples are set to zero.
11. The method according to claim 1, characterized in that: In the skeleton action sequence output by the diffusion model, each action frame includes a human skeleton point set and a hand key point set; In the process of training the diffusion model, updating the parameters of the diffusion model based on the difference between the noise point set sequence predicted by the diffusion model and the noise point set sequence input to the diffusion model; Wherein, when calculating the difference between the noise point set sequence predicted by the diffusion model and the noise point set at the same position of the noise point set sequence input into the diffusion model, the degree of overlap of the two hands in the video frame corresponding to the noise point set at the same position is calculated; if the degree of overlap is greater than a threshold, when calculating the difference of the noise point set at the same position, the difference in the hand area is set to zero.
12. A motion data generating device, characterized in that: include: An acquisition module is used to obtain audio features of a target speech and text corresponding to the target speech; the audio features include at least one of the following: rhythm features of the target speech, and hidden features obtained by performing hidden feature extraction on the target speech; A determination module is used to determine, based on the semantic understanding of the text, a target text segment in the text that requires the digital human to synchronously perform a physical action, a category of the physical action corresponding to each target text segment, and a position code of each action frame in a skeleton action sequence to be generated corresponding to each target text segment; the skeleton action sequence corresponding to each target text segment represents the position of each skeleton point at different times during the process of the digital human outputting the target text segment in voice; A generation module, for corresponding to each target text segment, using at least the audio feature, the category of the body movement corresponding to the target text segment, and the position encoding of each action frame corresponding to the target text segment as control conditions of a diffusion model, and generating a skeleton action sequence corresponding to the target text segment through the diffusion model; The diffusion model is obtained by training with a number of skeleton action sequences as training sample sets and with the target features corresponding to each skeleton action sequence as the control conditions of the diffusion model; The different skeleton action sequences in the training sample set are obtained by extracting skeleton points from video frames in different video clips; the characters in each video clip synchronously display a body movement when outputting speech; The target features corresponding to each skeleton action sequence in the training sample set include at least: audio features of speech synchronized with the video clip corresponding to the skeleton action sequence, categories of body movements in the video clip corresponding to the skeleton action sequence, and position codes of each video frame in the video clip corresponding to the skeleton action sequence.
13. A computer program product, characterized in that It comprises computer-readable instructions, and when the computer-readable instructions are executed on an electronic device, the electronic device implements the motion data generating method as claimed in any one of claims 1 to 11.
14. An electronic device, characterized in that: The electronic device comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the motion data generating method as described in any one of claims 1 to 11.
15. A computer storage medium, characterized in that The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the action data generation method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Virtual object control method and device, computer equipment and storage medium
CN118466887A
Motion video generation method, related device and medium
CN119229218A