Personalized whole body action generation method based on voice input

By constructing methods that separate quantize potential spaces and integrate audio features, the problem of unnatural movements in the prior art is solved, and more accurate and natural whole-body movement generation is achieved.

CN120339475APending Publication Date: 2025-07-18TIANXIANG RUIYI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510406940.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When generating human gestures in the whole body, it is difficult for the prior art to accurately capture semantic information and rhythm information of speech content, resulting in unnatural movements and under-exploration and utilization of potential features of gestures and audio data, affecting the fidelity and accuracy of movements.

Method used

The separation quantization latent space construction method based on a variational autoencoder is adopted to extract the characteristic representation of the head and body respectively, and the audio features are fused through the time cross attention and content self-attention mechanism to generate natural and smooth whole-body movements.

Benefits of technology

Improve the accuracy and naturalness of action generation, ensure the coordination and consistency of head and body movements, and the generated actions are more in line with the audio content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339475A_ABST
    Figure CN120339475A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized whole body action generation method based on voice input, and the method comprises the steps: obtaining body parameters suitable for a human body model, and extracting the audio features of a user; based on the body parameters, aiming at the head model and the body model, separately constructing separation quantitative potential spaces from the two variational auto-encoders; rhythm and text content are extracted from audio features of the user, and appropriate feature representations fused with audio content and rhythm are generated for the head model and the body model respectively; the mask posture is processed, effective body prompt information is coded, and the audio features and the body prompt information are selectively fused through time cross attention, so that reconstruction of the mask posture is realized; and respectively decoding the motion information of the head and the body, estimating global translation, and generating final whole body motion. According to the method, data features are fully mined and utilized to improve the accuracy and naturalness of action generation, so that the generated action is more consistent with audio content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital humans, and particularly relates to a personalized full-body motion generation method based on voice input. Background Art

[0002] With the development of virtual reality (VR), augmented reality (AR) and human-computer interaction technologies, generating human body motions that match the audio content has become a popular research direction. Traditional methods often focus on simple lip synchronization or rule-based body motion generation. For example, early technologies generated simple body motions through predefined audio-pose mapping rules. This method appears too rigid when faced with complex speech content and diverse human motion requirements, and cannot generate natural and fluent full-body motions that conform to semantics. The existing technologies mainly have the following two problems.

[0003] One is the lack of sufficient motion realism. When existing technologies generate full-body human gestures, it is often difficult to accurately capture the semantic information and rhythm information contained in the speech content, resulting in a mismatch between the generated motions and the audio content, and the motions looking unnatural. For example, in a speech segment expressing strong emotions, the generated body posture may not accurately convey this emotional intensity. Many methods have defects in handling the coordinated motions of the hands, head and the whole body, and situations such as the disconnection between hand and body motions and unreasonable head postures will occur, which seriously affects the realism of the generated motions.

[0004] The other is the lack of effective utilization of data. Some existing algorithms do not fully explore and utilize the potential features in gesture data and audio data. For example, the spatial and temporal information in gesture data is not effectively extracted and fused, resulting in the loss of important details during the motion generation process and affecting the accuracy of the motions. At the same time, for audio data, only some basic features such as frequency and amplitude are often simply extracted, without deeply exploring the semantic and rhythm patterns therein, and these rich audio information cannot be converted into meaningful body motions. Summary of the Invention

[0005] The present invention aims to fully explore and utilize data features to improve the accuracy and naturalness of motion generation.

[0006] According to the above objective, in the first aspect of the present invention, a personalized full-body motion generation method based on voice input is provided, and the method includes:

[0007] Obtain body parameters applicable to the human body model and extract the audio features of the user;

[0008] Based on the body parameters, for the head model and the body model, respectively construct separate quantization latent spaces from two variational autoencoders;

[0009] Extract the rhythm and text content from the user's described audio features, and generate appropriate feature representations that fuse the audio content and rhythm for the head model and the body model respectively;

[0010] Process the masked pose, encode the effective body cue information, and selectively fuse the audio features and the body cue information through cross - attention over time to achieve the reconstruction of the masked pose;

[0011] Decode the action information of the head and the body respectively, and estimate the global translation to generate the final full - body action.

[0012] Furthermore, obtaining the body parameters applicable to the human model includes:

[0013] Given the marker positions and marker - position offsets, calculate the potential markers by combining the differentiable surface point - mapping function and the vertex normal function, and determine the initial parameters of the body by minimizing a series of loss terms, including: shape parameters, pose parameters, and transformation parameters;

[0014] According to the human physiological characteristics, use the Kolmogorov - Smirnov test to analyze the data distribution and adjust the initial parameters;

[0015] Convert the Blendshape weights of the face into parameters available for the FLAME model.

[0016] Furthermore, obtaining the optimized VQ - VAEs model includes:

[0017] Encode the processed data;

[0018] Find the closest quantization vectors;

[0019] Calculate the VQ - VAE loss;

[0020] Optimize the VQ - VAEs to obtain the latent space.

[0021] Furthermore, the generation of appropriate feature representations that fuse the audio content and rhythm includes:

[0022] Extract the rhythm feature R from the speech audio 1:T , and use the pre - trained embedding feature from the transcript as the content feature C 1:T ;

[0023] Use a temporal convolutional network and a linear mapping to align the feature encodings of the rhythm and the content;

[0024] Through the formula F 1:T =α×R 1:T +(1 - α)×C1:T Fuse the rhythm features and content features to generate the comprehensive audio feature F 1:T , where α is a dynamic weight that is adaptively adjusted according to the importance of the rhythm features and content features.

[0025] Furthermore, the reconstruction of the masked pose includes:[[]]

[0026] Process the masked tokens and replace them with the learned masked embeddings;

[0027] Use a spatial convolutional encoder to summarize the spatial information and compress the spatial features;

[0028] Encode the body cue information through a temporal cross-attention mechanism, and the formula is where H ∈ R T×512 , representing the encoded body cue information, and p t is the sum of the learned speaker embedding and the position pose-related information PPE;

[0030] Adopt a Transformer decoder based on temporal cross-attention to fuse the body cue information and audio features, and the formula is And minimize the Manhattan distance in the latent space to achieve the reconstruction of the masked pose.

[0031] Furthermore, decode the action information of the head and body respectively, and estimate the global translation, including:[[]]

[0032] Concatenate the body cue information H with the audio features to decode the head action information;

[0033] Obtain through a vector quantization decoder VQ-Decoder and then use a pre-trained global motion predictor to estimate the global translation, where is the body motion information, is the estimated global translation,

[0034] Furthermore, the conversion of the Blendshape weights of the face into parameters available for the FLAME model includes:[[]]

[0035] Given the face Blendshape weight information B obtained from the original data ARKit ∈ R T×51 , through a handcrafted Blendshape template V template ∈ R 52 , using the formula Drive the FLAME topology vertices, and then optimize the transformation matrix W ∈ R by minimizing 51×103 so as to map B ARKit to B Flame ∈ R T×(100+3) , that is, the parameter space of the FLAME model, and obtain the parameters related to facial features.

[0036] In the second aspect of the present invention, there is provided a computer-readable storage medium on which a computer program is stored, wherein when the computer program is executed by a processor, the personalized full-body motion generation method based on voice input described in the first aspect above is implemented.

[0037] In the third aspect of the present invention, there is further provided an electronic device for personalized full-body motion generation based on voice input, which includes: a processor and a memory connected to the processor; the memory stores instructions executable by the processor, and when the instructions are executed by the processor, the full-body motion generation method described in the first aspect above is implemented.

[0038] Compared with the prior art, the personalized full-body motion generation method based on voice input disclosed in the embodiments of the present invention aims to fully exploit and utilize data features to improve the accuracy and naturalness of motion generation. By using technologies such as spatial convolution and temporal self-attention, spatial and temporal features in the motion pose data of training data are comprehensively extracted, providing richer detailed information for motion generation and avoiding the loss of key information. The present invention not only focuses on the basic features of audio, but also deeply explores information such as semantics, rhythm, and prosody therein, and incorporates this information into the motion generation process through a reasonable fusion strategy, making the generated motion more conform to the audio content. By designing effective audio feature extraction and fusion mechanisms, such as content and rhythm self-attention mechanisms, semantic and rhythm information in speech is accurately captured and transformed into body postures highly matching the semantics. By separately modeling the head and the body and considering their mutual relationship during the motion generation process, the limb coordination problem is solved, ensuring that the generated full-body motion is natural and smooth. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic flowchart of the personalized full-body motion generation method based on voice input in the embodiments of the present invention.

[0040] Figure 2 It is a schematic flowchart of body parameter processing in the embodiments of the present invention.

[0041] Figure 3 It is a schematic flowchart of processing the VQ-VAEs model in the embodiments of the present invention.

[0042] Figure 4Schematic diagram of the process of fusing audio features in the embodiments of the present invention.

[0043] Figure 5 Schematic diagram of the process of processing mask postures in the embodiments of the present invention.

[0044] Figure 6 Schematic diagram of the structure of an electronic device for generating personalized full-body movements based on voice input in the embodiments of the present invention. Detailed implementation manners

[0045] The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and cannot be used to limit the protection scope of the present invention. For example, certain terms are used in the specification and claims to refer to specific components. Those skilled in the art should understand that hardware or software manufacturers may use different terms to refer to the same component. The specification and claims do not use the difference in names as a way to distinguish components, but use the difference in functions of components as the criterion for distinction. The subsequent description in the specification is a preferred implementation manner for implementing the present invention, but the description is for the purpose of illustrating the general principles of the present invention and is not used to limit the scope of the present invention. The protection scope of the present invention shall be subject to what is defined by the appended claims.

[0046] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0047] Embodiment 1

[0048] Please refer to Figures 1 to 5 As shown, Embodiment 1 of the present invention provides a method for generating personalized full-body movements based on voice input. The method includes:

[0049] Step S1: Obtain body parameters applicable to the human model and extract the audio features of the user;

[0050] Step S2: Based on the body parameters, for the head model and the body model, respectively construct separate quantization latent spaces from two variational autoencoders, and an optimized VQ-VAEs model;

[0051] Step S3: Extract the rhythm and text content from the extracted audio features of the user, and generate feature representations that are suitable and fused with audio content and rhythm for the head model and the body model respectively;

[0052] Step S4: Process the mask posture, encode effective body prompt information, and selectively fuse the audio features and the body prompt information through temporal cross-attention to achieve the reconstruction of the mask posture;

[0053] Step S5: Decode the motion information of the head and the body respectively, and estimate the global translation to generate the final full-body movement.

[0054] Next, each step of the present invention will be described in detail.

[0055] In step S1, accurate and applicable body parameters related to body shape, posture, transformation parameters, and facial features are obtained from the action and posture dataset for training. At the same time, the voice audio data for training is prepared to lay a solid foundation for subsequent model training, and finally, the processed posture data and voice audio data are output.

[0056] Specifically, obtaining the body parameters applicable to the human body model includes the following process:

[0057] Initialization of body parameters: Given the marker positions and marker position offsets, potential markers are calculated by combining the differentiable surface vertex mapping function and the vertex normal function, and the initial parameters of the body, including shape parameters, posture parameters, and transformation parameters, are determined by minimizing a series of loss terms. Specifically, with the captured marker positions X ∈ R T×K×3 , where T represents the time dimension to reflect the changes of actions in the time series, K is the number of markers to identify different position points, and this data completely describes the position information in the three-dimensional space. At the same time, the predefined marker position offset Y ∈ R K×3 , which is mainly used to adjust the marker positions to enable the model to better adapt to the actual situation, and combined with the differentiable surface vertex mapping function and the vertex normal function to preliminarily determine the body shape posture ω ∈ R T ×55×3 and transformation parameter τ ∈ R T×3 , where represents the parameter set of the body shape, ω represents the parameters of the posture in the time and space dimensions, and τ is the parameter related to transformation. Finally, the potential marker for each frame is calculated through the formula to provide a data basis for subsequent optimization, where is the user-defined vertex-to-marker function. In the entire action generation process, subsequent optimization of body parameters requires knowing the difference between the currently estimated body part positions (i.e., potential markers) and the actual observed positions in order to adjust through the loss function to make the parameters more reasonable. The body parameters are preliminarily determined by minimizing a series of loss terms, and these loss terms include the data term surface distance posture and shape priors and and marker initialization regularization ε I the velocity constant term The overall objective function is where The potential markers for each calculated frame is a set of parameters for body shape, Ω represents the set of pose parameters ω, Υ is the set of corresponding transformation parameters, and λ ω is the weight of each loss term, used to balance the influence of different loss terms on the overall objective. For example, the data term Here is a function related to the simulated markers. This data term is used to measure the distance between the simulated markers and the observed markers. When refining the body parameters, the initialized parameters need to be adjusted, and the calculated can reflect the fitting degree of the model to the data under the current parameters, providing a reference for parameter refinement. In addition, when training VQ-VAEs, the optimized body parameters (affected by ) are used as input features and mapped to the latent space through the encoder network, making the representation in the latent space more consistent with the actual body poses.

[0058] Refinement of body parameters: According to human physiological characteristics, the Kolmogorov-Smirnov test is used to analyze the data distribution, and the initialized parameters (body shape parameters pose parameters ω and transformation parameters τ) are adjusted. Specifically, according to human physiological characteristics, such as the rule that the length of the neck and head is about of the total body length, and the fingers except the thumb should not bend backward under normal circumstances, etc., and using the Kolmogorov-Smirnov test to analyze the data distribution, the initialized parameters are adjusted to make the parameters more reasonable.

[0059] Conversion of Blendshape weights to FLAME parameters: It is independent of the processing of the aforementioned body parameters. It mainly converts the Blendshape weights of the face into parameters available for the FLAME model. First, given the 51-dimensional Blendshape weight information B ARKit ∈R T×51 obtained from the original motion capture camera stored action data, and then, through the handcrafted Blendshape template V template ∈R 52 , using the formula to drive the FLAME topology vertices, and then minimizing to optimize the transformation matrix W ∈ R 51×103 , so as to map B ARKit to B Flame ∈R T×(100+3), namely the parameter space of the FLAME model, to obtain parameters related to facial features. After converting Blendshape weights into FLAME parameters, the rich parameter space of the FLAME model can be utilized to achieve more refined and accurate control of facial expressions, enabling the generation of more natural and realistic facial animations.

[0060] The processed body pose data is used to construct the latent space of the head and body during the stage of constructing the combined discrete head and body prior, while the speech audio data is used to extract audio features during the rhythm self-attention stage.

[0061] In step S2, a combined discrete head and body prior will be constructed based on Vector Quantized Variational Autoencoders (VQ-VAEs). This stage is dedicated to constructing separate quantized latent spaces for the head and body. By optimizing the VQ-VAEs (Variational Autoencoders), the latent representations of the head and body poses are learned, and finally, the quantized latent spaces for the head and body are obtained respectively. and the optimized VQ-VAEs model, where, q h represents the quantized latent vector of the head, which is the discrete representation obtained in the head quantized latent space by encoding the head-related pose data G h ∈R T×106 through the encoding process of VQ-VAEs, containing the key features of head movement and posture. q b represents the quantized latent vector of the body, corresponding to the pose data G of the body (as a whole) b ∈R T×(55×6+4+3) The discrete representation in the body quantized latent space after being encoded by VQ-VAEs encompasses the important features of the movement and posture of each part of the body. Separating the quantized spaces of the head and body to form two independent latent spaces allows for independent modeling, analysis, and operation of the head and body. For example, when generating images or 3D models, the states of the head and body can be adjusted independently without interference from each other.

[0062] First, for the head G h ∈R T×106 and the body (as a whole) G b ∈R T×(55×6+4+3) of the pose data, separate quantized latent spaces from two VQ-VAEs are constructed respectively. The latent spaces of the head and body provide the basic feature representations for the subsequent rhythm self-attention stage, masked audio pose stage, and head and body action decoding stage. In these stages, the model will be based on these latent representations, that is, the vector representations (q h and q b) to perform more complex calculations and processing by combining information such as audio features. The spatially encoded body cue information H∈R is obtained by refining the spatially encoded features through the temporal cross-attention mechanism (without forward feedback) TCA. T×512 The generation of the body cue information H here is closely related to the latent spatial representation of the body, because the initial input data spatial masked pose G is associated with the construction of the body latent spatial representation and combined with audio features in the subsequent decoding stage.

[0063] Next, by jointly optimizing a series of loss terms, the VQ-VAEs can accurately represent the features of the head and body in the latent space. The process of obtaining the latent space includes the following steps: for the processed pose data obtained in the dataset construction stage, specifically divided into head pose data G h ∈R T×106 and body (overall) pose data G b ∈R T×(55×6+4+3) perform encoding; find the closest quantization vector; calculate the VQ-VAE loss; optimize the VQ-VAEs to obtain the latent space.

[0064] Schematically, in the above optimization measures, specifically, through the formula The quantization vector q h closest to the encoded input data Z (Z is the corresponding G encoded with a time window size w = 1, for the head it is G b and for the body it is G i . The overall loss function is where is the reconstruction loss, used to measure the difference between the generated gesture and the original gesture G; and are the velocity and acceleration losses, which help the model learn natural motion patterns; is the commitment loss, used to constrain the relationship between the encoded vector Z and the quantization vector q. By minimizing VQ–VAEs

[0066] can learn accurate and reasonable latent representations of the head and body postures, providing more reliable basic features for the subsequent rhythm self-attention stage and masked audio pose stage.

[0067] In step S3, enter the rhythm self-attention stage to generate a feature representation that fuses the content and rhythm of the audio and is appropriate, including the following process:

[0068] S31. Extract the rhythm feature R 1:T from the speech audio, and use the pre-trained embedding feature from the transcript as the content feature C1:T ;

[0069] S32. Use a temporal convolutional network (TCN) and linear mapping to encode and align the rhythm features and content features so that they can be better processed by the model;

[0070] S33. Through the formula F 1:T = α × R 1:T + (1 - α) × C 1:T Fuse the rhythm features and content features to generate comprehensive facial features F that integrate speech rhythm and content features 1:T , which are used as the input for the final decoder and to generate facial expressions related to the speech content. α is a dynamic weight that is adaptively adjusted according to the importance of the rhythm features and content features.

[0071] In this stage, the rhythm (onset and amplitude) is mainly extracted from the speech audio A as the explicit audio rhythm R 1:T , while the pre-trained embedding E from the transcript is used as the content C 1:T , and the two are fused to generate feature representations for the head and body that match the audio content and rhythm respectively, and finally the fused facial features F 1:T are obtained, which encode the prosodic and semantic information of the speech.

[0072] The fused facial features F 1:T are used to fuse with the body cue information in the masked audio pose stage to guide the reconstruction of the latent space and the generation of poses. In the head and body motion decoding stage, the audio features are combined with the body-related information for the final motion information decoding.

[0073] In step S4, the reconstruction of the masked pose will be performed. This stage focuses on processing the masked pose, encoding effective body cue information, and selectively fusing the audio and body cues through temporal cross-attention to achieve the reconstruction of the masked pose, and finally obtaining the reconstructed masked pose information and the encoded body cue information. For a given temporal and spatial masked pose Since the 0 value may still represent a specific motion pose, the learned masked embedding e mask ∈

[0074] R 256 is used to replace the masked token to prepare for subsequent processing. The reconstructed masked pose information and the encoded body cue information are used for the decoding of the head motion information in the head and body motion decoding stage. At the same time, the body cue information also provides an important basis for estimating the global translation.

[0075] Among them, the reconstruction of the masked pose specifically includes the following process:

[0076] The mask mark is processed and replaced with the learned mask embedding to prepare for subsequent processing; the mask mark after replacement contains e mask The data will be input into the spatial convolution encoder.

[0077] The spatial convolutional encoder summarizes the spatial information and compresses the spatial features into R T×256 ; Use a spatial convolutional encoder to summarize spatial information and mask pose data from space The representative information that can reflect the spatial characteristics of body posture is extracted, and the spatial characteristics are compressed to R T×256 , extracting spatial information and compressing dimensions.

[0078] The TCA encodes body cue information through the temporal cross attention mechanism (without forward feedback), and the formula is Among them, H∈R T×512 , represents the encoded body cue information, p t is the sum of the learned speaker embedding and position and posture related information PPE, Represents the spatial masking posture The result after spatial convolutional coding. First, the mask mark is processed and replaced with the learned e mask ∈R 256 Mask embedding, followed by a spatial convolutional encoder on the processed spatial mask pose data containing mask embedding The spatial convolution encoder slides the convolution kernel on the data, performs weighted summation and other operations on the local area, summarizes the spatial information, and compresses the spatial features from the original 337 dimensions to 256 dimensions to obtain the compressed spatial feature representation. This result is

[0079] The Transformer decoder based on temporal cross attention is used to fuse body prompt information and audio features. The decoding formula of the body is: The decoding formula for the header is And minimize the L1 distance in the latent space To achieve the reconstruction of the mask pose, the L1 distance refers to the Manhattan distance, which refers to the sum of the absolute differences between two points in all dimensions.

[0080] In step S5, the head and body movements and global translation decoding are implemented. Based on the information obtained in the previous step (F 1:T ,H,p t) Decode the action information of the head and body respectively, estimate the global translation, and finally obtain the head facial expression, body action information and global translation information. These information together constitute a complete description of the full-body human gesture. Obtain the body motion information from the decoder VQ-Decoder After that, use the pre-trained global motion predictor Estimate the global translation, is the body motion information, is the estimated global translation, is the pre-trained global motion predictor.

[0081] Decoding the action information of the head and body and estimating the global translation include the following processes:

[0082] Since the head and body movements are relatively independent, first connect the body cue information H with the audio feature to decode the head action information;

[0083] Obtain through the vector quantization decoder VQ-Decoder After that, use the pre-trained global motion predictor Estimate the global translation, where, is the body motion information, is the estimated global translation, is the pre-trained global motion predictor.

[0084] The facial expression generation is mainly generated in the decoding stage of the model. Considering the weak relationship between the face and body movements, in the face decoding step, directly connect the masked body cue with the audio feature, and like the body decoding process, perform the final decoding of the facial latent feature to generate the facial expression.

[0085] The user uploads an audio file. The audio is noise-free and requires pure human voice. The audio generates full-body actions including facial expressions and body movements through the method of the present invention.

[0086] Embodiment 2

[0087] Please refer to Figure 6, Embodiment 2 of the present invention provides an electronic device for processing data information. The electronic device includes at least one processor 201 and at least one memory 202; wherein, the processor 201 and the memory 202 are directly connected to each other, or communicate with each other through a communication interface 203, or are electrically connected through one or more communication buses or signal lines to achieve data transmission or interaction; the memory 202 stores program instructions executable by the processor 201, and the processor 201 calls the program instructions to execute a personalized full-body motion generation method based on voice input disclosed in Embodiment 1. For example, it realizes: obtaining body parameters applicable to the human model and extracting the audio features of the user; based on the body parameters, respectively constructing separate quantization latent spaces from two variational autoencoders for the head model and the body model; extracting rhythm and text content from the user's audio features, and generating appropriate feature representations fused with audio and rhythm for the head model and the body model respectively; processing the masked pose, encoding effective body prompt information, and selectively fusing the audio features and the body prompt information through temporal cross-attention to realize the reconstruction of the masked pose; respectively decoding the motion information of the head and the body, and estimating the global translation to generate the final full-body motion.

[0088] Among them, the memory 202 can be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable read-only memory (EPROM), electrically erasable read-only memory (EEPROM), etc.

[0089] The processor 201 can be an integrated circuit chip with signal processing capabilities. The processor 201 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0090] It can be understood that Figure 6 The structure shown is only schematic, and the electronic device may also include more or fewer components than those shown Figure 6 in it, or have a different configuration from that shown Figure 6 in it. Figure 6 Each component shown in it can be implemented by hardware, software or a combination thereof.

[0091] Embodiment 3

[0092] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor 201, it implements the personalized full-body motion generation method based on voice input in the first embodiment. For example, it realizes: obtaining body parameters applicable to a human body model and extracting audio features of a user; respectively constructing separate quantization latent spaces from two variational autoencoders for a head model and a body model based on the body parameters; extracting rhythm and text content from the audio features of the user, and generating appropriate feature representations fused with audio and rhythm for the head model and the body model respectively; processing the masked pose, encoding effective body prompt information, and selectively fusing audio features and body prompt information through temporal cross-attention to realize the reconstruction of the masked pose; respectively decoding the motion information of the head and the body, and estimating the global translation to generate the final full-body motion.

[0093] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0094] It should be noted that for those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

[0095] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0096] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A personalized full-body motion generation method based on voice input, characterized in that The method includes: Obtaining body parameters applicable to a human body model and extracting audio features of a user; Based on the body parameters, separately constructing a separate quantization latent space from two variational autoencoders for a head model and a body model; Extracting rhythm and text content from the audio features of the user, and generating appropriate feature representations for the head model and the body model respectively, which are fused with audio content and rhythm; Processing the masked pose, encoding effective body prompt information, and selectively fusing audio features and body prompt information through temporal cross-attention to achieve the reconstruction of the masked pose; Separately decoding the action information of the head and the body, and estimating the global translation to generate the final full-body action.

2. The action generation method according to claim 1, wherein Obtaining body parameters applicable to a human body model includes: Given marker positions and marker position offsets, calculating potential markers by combining a differentiable surface point mapping function and a vertex normal function, and determining the initial parameters of the body by minimizing a series of loss terms, including: shape parameters, pose parameters, and transformation parameters; According to human physiological characteristics, using the Kolmogorov-Smirnov test to analyze the data distribution and adjusting the initial parameters; Converting the Blendshape weights of the face into parameters available for the FLAME model.

3. The action generation method according to claim 2, characterized in that, Obtaining the latent space includes: Encoding the processed body pose data; Finding the quantization vector closest thereto; Calculating the VQ-VAE loss; Optimizing the VQ-VAEs to obtain the latent space.

4. The action generation method according to claim 3, wherein The generating appropriate feature representations fused with audio content and rhythm includes: Extract the rhythm feature R from the speech audio 1:T and use the pre-trained embedding feature from the transcript as the content feature C 1:T ; Using a temporal convolutional network and a linear mapping to align the feature encodings of rhythm and content; Through formula F 1:T = α × R 1:T + (1 - α) × C 1:T Fuse the rhythm feature and the content feature to generate the comprehensive audio feature F 1:T , where α is the dynamic weight and is adaptively adjusted according to the importance of the rhythm feature and the content feature.

5. The action generation method according to claim 4, wherein The reconstruction of the masked pose includes: Processing the masked markers and replacing them with learned masked embeddings; Using a spatial convolutional encoder to summarize spatial information and compressing the spatial features; Encoding body cue information through a temporal cross-attention mechanism, the formula is where \(H\in\mathbb{R}\) T×512 , representing the encoded body cue information, \(p\) t is the sum of the learned speaker embedding and the position-pose related information PPE, represents the result after spatially convolving and encoding the spatial masked pose ; The Transformer decoder based on temporal cross-attention is used to fuse the body cue information and audio features, and the formula is and minimize the Manhattan distance in the latent space to achieve the reconstruction of the masked pose.

6. The action generation method according to claim 5, wherein Separately decoding the action information of the head and the body and estimating the global translation includes: Connecting the body prompt information H with the audio features and decoding the head action information; Obtained by the vector quantization decoder VQ - Decoder After that, use the pre - trained global motion predictor To estimate the global translation, where Is the body motion information, Is the estimated global translation, Is the pre - trained global motion predictor.

7. The action generation method according to claim 2, wherein The converting the Blendshape weights of the face into parameters available for the FLAME model includes: Given the facial Blendshape weight information B obtained from the original data ARKit ∈R T×51 , through the handcrafted Blendshape template V template ∈R 52 , using the formula to drive the FLAME topology vertices, and then by minimizing to optimize the transformation matrix W ∈ R 51×103 , thereby mapping B ARKit to B Flame ∈R T×(100+3) , that is, the parameter space of the FLAME model, to obtain the parameters related to facial features.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the personalized full-body action generation method according to any one of claims 1 to 7 above.

9. An electronic device for personalized full-body motion generation based on voice input, characterized in that, Including: A processor and a memory connected to the processor; The memory stores instructions executable by the processor, and when the instructions are executed by the processor, they implement the personalized full-body action generation method according to any one of claims 1 to 7 above.