Visual and motion joint representation method and system for human action generation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-08-11
AI Technical Summary
[0007]鉴于上述,为解决现有纯音频驱动方法中存在的动作生成过度平滑、缺乏三维空间几何先验以及跨模态语义对齐能力不足的问题
[0023]与现有技术相比,本发明具有的有益效果至少包括:
Smart Images

Figure CN121505186B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and computer graphics technology, specifically relating to a method and system for visual and motion joint representation of human motion generation. Background Technology
[0002] With the rapid development of industries such as virtual reality, digital entertainment, and virtual anchors, high-fidelity and highly natural 3D digital human driving technology has become a focus of industry attention. In human-computer interaction scenarios, how to automatically generate body movements that are semantically consistent with the speech content, rhythmically coordinated, and expressive, based solely on the input audio signal, is one of the core technical challenges in improving the intelligence and immersive experience of digital humans.
[0003] Currently, the mainstream technologies in the field of audio-driven human motion generation focus on data-driven deep learning methods. These methods train neural networks to directly map one-dimensional acoustic features (such as Mel spectrum and prosodic features) to three-dimensional parametric human body models (such as SMPL and SMPL-X) coefficients or skeletal keypoint sequences. While these methods address the issue of motion continuity to some extent, the mapping from audio to motion is inherently a complex one-to-many nondeterministic problem. Existing models often tend to output the average of multiple reasonable motion paths, resulting in generated motion sequences that are too small in amplitude, overly smooth, lack spatial variation and detail, and have insufficient expressiveness.
[0004] Furthermore, a fundamental flaw of pure audio-driven technology is that audio signals inherently lack three-dimensional spatial geometric information. When a model relies solely on acoustic features, it must infer complex three-dimensional human postures, spatial displacements, and limb structural relationships from these one-dimensional temporal signals. This makes it difficult for the model to generate actions with precise entity representation (such as specific gestures expressing size or direction) or that conform to complex environmental contexts, severely limiting the spatial accuracy and naturalness of the actions.
[0005] In recent years, multimodal visual large models (such as Wan-S2V) have made breakthrough progress in the field of video generation, demonstrating a powerful ability to generate high-quality, semantically consistent video clips based on audio. These pre-trained large models actually contain rich visual prior knowledge of human pose structure, scene understanding, and spatiotemporal evolution. However, these models typically output two-dimensional pixel video streams directly and cannot directly extract and convert them into accurate and actuated 3D human motion coefficients.
[0006] Therefore, designing an effective mechanism to inject these powerful visual prior knowledge as guiding signals into the audio-action generation pipeline, and solving the problems of lack of spatial sense and weak action semantics in pure audio-driven models, is a technical challenge that urgently needs to be overcome in this field. Summary of the Invention
[0007] In view of the above, to address the problems of overly smooth motion generation, lack of three-dimensional spatial geometric priors, and insufficient cross-modal semantic alignment capabilities in existing pure audio-driven methods, the purpose of this invention is to provide a visual and action joint representation method and apparatus for human motion generation. By introducing high-level visual features extracted from a pre-trained multimodal large model as guiding signals and constructing a joint representation space for visual and action alignment, it is able to generate high-fidelity human motion coefficient sequences with accurate spatial structure, rich semantics, and high consistency with audio rhythm.
[0008] To achieve the above-mentioned objectives, an embodiment provides a visual and motion joint representation method for human motion generation, comprising the following steps: Visual feature vectors containing human pose structure and spatial geometry information are extracted from target videos using a pre-trained multimodal large model. Using visual feature vectors as guiding information, the visual and action joint representation module aligns visual features and action features in a unified latent space based on the guiding information to obtain joint representations, and generates action feature sequences based on the joint representations. An action decoder is used to decode the action feature sequence to generate a sequence of human action coefficients that are consistent with the semantics and rhythm of the target audio content.
[0009] Preferably, the pre-trained multimodal large model adopts a pre-trained Wan-S2V model, wherein the Wan-S2V model includes a variational autoencoder component and a DiT module; During feature extraction, the audio feature sequence corresponding to the target video is injected into the DiT module as a condition, and the latent space features of the output layer of the DiT module are extracted as visual feature vectors during the forward inference process.
[0010] Preferably, the vision and motion joint representation module includes a vision and motion joint encoder and a motion feature decoder. The vision feature vector is input as guiding information to the vision and motion joint encoder to achieve alignment of vision features and motion features in a unified latent space, generating a vision and motion joint representation containing human posture structure and spatial geometric information. This joint representation generates a motion feature sequence containing rich motion priors through the motion feature decoder.
[0011] Preferably, the visual and motion joint representation module is trained before application. The training architecture includes a video encoder, a motion encoder, a visual and motion joint encoder, a visual feature decoder, and a motion feature decoder. Paired monocular video sequences and human motion coefficient sequences are processed by the video encoder and motion encoder to extract visual embeddings and motion embeddings, respectively. The joint encoder maps the visual embeddings and motion embeddings to a unified latent space through a cross-modal attention mechanism to obtain a joint representation. The joint representation Z is decoded by the visual feature decoder and motion feature decoder to generate visual features and motion features, respectively.
[0012] Preferably, based on the aforementioned architecture, a generative alignment strategy based on a diffusion model is adopted, utilizing action diffusion loss. With visual diffusion loss Optimize the vision and motion co-encoder; Action diffusion loss The ability of the joint characterization Z to generate action modes is used to constrain the ability of the characterization Z to generate action modes, and it is expressed as follows:
[0013] in Indicates the action characteristics in the spreading step The noise representation below, To add Gaussian noise, For prediction noise in the action feature decoder, Expressing expectations, Represents the square of the L2 norm; Visual diffusion loss The ability of the joint representation to generate visual modalities is used to constrain the ability of the joint representation to do so, and it is expressed as follows:
[0014] in Indicates latent visual features in the diffusion step The noise representation below, The number of potential visual features in a single frame. i Indicates latent visual feature index, To jointly characterize the corresponding local latent representation in Z, Indicating targeting Added Gaussian noise, This represents the prediction noise in the visual feature decoder.
[0015] Preferably, the paired monocular video sequences and human motion coefficient sequences used during the training of the visual and motion joint representation module are constructed in the following manner: To address the issue of missing motion coefficients in monocular videos, the Aios backbone network is used to predict the SMPL-X 3D human body parameters in video frames. For the problem of motion coefficients existing but lacking corresponding monocular videos, a parametric rendering method based on the SMPL-X model is adopted. Through differentiable rendering technology based on the graphics pipeline, motion coefficients are mapped to corresponding monocular human videos, thus constructing data pairs that are consistent with vision and motion.
[0016] Preferably, the action decoder employs a UNet module with a stream matching training strategy and an ordinary differential equation numerical integrator, which can convert the input action feature sequence into a human action coefficient sequence.
[0017] Preferably, the action decoder also needs to be trained before being used. During training, an action encoder and action decoder are introduced to form a training framework. The training adopts two stages. The first stage trains the action encoder and action decoder with an RVQVAE structure for each body part. To achieve multi-level residual quantization, the model employs... Cascaded quantizers, training loss It consists of two parts: reconstruction loss and commitment loss, specifically expressed as follows:
[0018] in, Represents the actual sequence of actions. This represents the sequence of actions for reconstruction, the first item. The first term represents the absolute difference, used to constrain the accuracy of the reconstruction; the second term represents the commitment loss, used to constrain the stability of the quantization process. This represents the weighting parameter corresponding to the committed loss. Indicates the relationship with RVQ The residual levels are accumulated. Indicates the first The output vector of the layer quantizer Indicates the first The nearest neighbor discrete vector retrieved from the layer codebook. Indicates the first Weighting coefficients for layer residual quantization This represents a continuous sequence of representations output by the motion encoder. This represents a discrete codebook representation sequence obtained through vector quantization. This indicates that gradient propagation is blocked during backpropagation. Indicates the mean squared error; The second stage freezes the action encoder and replaces the action decoder with the action decoder from the UNet module for training, with the training loss... Represented as:
[0019] in, This represents the initial noisy action sequence. Represents the target action sequence. Represents a time variable. Indicates time Between and Interpolation states between them This represents the velocity field predicted by the UNet module. For UNet parameters, Indicates in Calculate the expectation on the joint distribution.
[0020] To achieve the above-mentioned objectives, the embodiments also provide a visual and motion joint representation system for human motion generation, comprising: Feature extraction module: used to extract visual feature vectors containing human pose structure and spatial geometric information from target videos using a pre-trained multimodal large model; Joint Representation Mapping Module: Used to align visual features and action features in a unified latent space using visual feature vectors as guiding information, and obtain joint representations based on the guiding information from the visual and action joint representation module, and generate action feature sequences based on the joint representations; Action generation module: Used to decode action feature sequences using an action decoder to generate a sequence of human action coefficients that are consistent with the semantics and rhythm of the target audio content.
[0021] To achieve the above-mentioned objectives, the embodiments also provide a computing device, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned visual and motion joint representation method for human motion generation.
[0022] To achieve the above-mentioned objectives, the embodiments also provide a computer-readable storage medium storing a program that, when executed by a processor, implements the above-mentioned visual and motion joint representation method for human motion generation.
[0023] Compared with the prior art, the beneficial effects of the present invention include at least the following: This invention utilizes visual feature vectors extracted from a pre-trained multimodal large model as guidance, providing crucial spatial geometric constraints for action generation. This effectively overcomes the limitations of purely audio-driven methods, such as lack of spatial awareness and unreasonable limb structure. Simultaneously, by constructing a joint representation of vision and action in a unified latent space, the model can accurately capture the deep correlation between audio rhythm and limb dynamics, significantly improving the over-smoothing phenomenon of traditional methods and enhancing semantic alignment accuracy. Furthermore, the bidirectional data augmentation process based on parameter estimation and differentiable rendering designed in this invention constructs high-quality aligned data at low cost, effectively solving the problem of scarce cross-modal training data. Combined with a stream matching strategy and a part-specific RVQVAE encoding architecture, this further ensures that the generated human action sequences possess high diversity while achieving structural rationality and natural fluency. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of the overall visual and motion joint representation method for human motion generation provided in the embodiment; Figure 2 This is a schematic diagram of the overall network architecture for the joint visual and motion representation of human motion generation provided in the embodiment; Figure 3 This is a schematic diagram of the training architecture of the vision and action joint representation module provided in the embodiment; Figure 4 This is a schematic diagram of the action codec structure based on RVQ-VAE and stream matching strategy provided in the embodiment; Figure 5 This is a structural block diagram of the visual and motion joint representation system for human motion generation provided in the embodiment. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of this invention.
[0027] The inventive concept of this invention is as follows: Addressing the problems of excessively smooth motion, lack of three-dimensional spatial geometric priors, and inaccurate cross-modal semantic alignment in existing pure audio-driven human motion generation technologies, this invention proposes a modeling scheme enhanced with visual features. This scheme breaks away from the traditional paradigm of relying solely on one-dimensional audio signals for motion inference, innovatively utilizing a pre-trained multimodal large model to extract visual feature vectors containing rich implicit posture structure and spatial geometric information as guidance. By constructing a joint representation space of vision and motion, the visual modality of monocular video is deeply aligned with the parameterized human motion modality, thereby ensuring that the final generated motion coefficient sequence is not only rhythmically synchronized with the audio but also possesses the characteristics of realistic human movement in terms of spatial structure and semantic expressiveness.
[0028] like Figure 1 and Figure 2 As shown in the embodiment, a visual and motion joint representation method for human motion generation includes the following steps: S1 uses a pre-trained multimodal large model to extract visual feature vectors containing human posture structure and spatial geometric information from the target video.
[0029] In the example, the input duration is first... The target audio signal (such as a waveform file) Preprocessing is performed. Specifically, an audio feature extractor (such as the pre-trained Wav2Vec2.0) is used to convert the original waveform into an audio feature sequence. ,in For sequence length, For audio feature dimensions.
[0030] Specifically, the pre-trained multimodal large model uses a pre-trained Wan-S2V model as the backbone network for visual feature extraction. The Wan-S2V model internally includes a variational autoencoder component and a DiT (Diffusion Transformer) module. The DiT module is used to perform temporal modeling and cross-modal interaction modeling on the audio feature vectors and the visual feature vectors to extract visual feature vectors that jointly contain temporal information under multimodal conditions. In implementation, the aforementioned audio feature sequence... As a conditional input, the latent space features of the DiT module in the Wan-S2V model are used as initial noise latent variables. Unlike conventional video generation tasks that require complete denoising sampling to generate pixel-level videos, this step aims to obtain high-level visual semantic representations. Therefore, during the model's forward inference process, the latent space features of the DiT module's output layer are extracted as visual feature vectors. .in, Indicates the temporal length of the corresponding video frame (aligned with the audio duration). Indicates the number of feature channels. This represents the potential spatial resolution after downsampling.
[0031] This visual feature vector It not only encodes temporal information synchronized with the audio rhythm, but also implicitly includes the human body's posture structure, limb geometric relationships, and movement trends in three-dimensional space through priors learned from large-scale video data by the Wan-S2V model, thus providing a strong guiding signal with spatial geometric information for subsequent joint representation mapping.
[0032] S2 uses visual feature vectors as guiding information and utilizes the visual and action joint representation module to align visual features and action features in a unified latent space based on the guiding information to obtain joint representations. Based on the joint representations, action feature sequences are generated.
[0033] In this embodiment, the vision and action joint representation module includes a vision and action joint encoder and an action feature decoder, which converts the visual feature vectors... As guiding information input to the vision-motion joint encoder, visual and motion features are aligned in a unified latent space to generate a joint visual and motion representation that includes human posture structure and spatial geometry information. This joint characterization An action feature sequence containing rich motion priors is generated using an action feature decoder. This is used for generating subsequent actions.
[0034] The visual and action joint representation module is trained before being applied, and the specific training architecture is as follows: Figure 3 As shown, the training architecture includes a video encoder, an action encoder, a vision and action joint encoder, a visual feature decoder, and an action feature decoder. The training architecture aims to establish a shared unified latent space in the vision and action joint encoder, such that visual features from monocular videos and action features from human body parameters are highly consistent in semantics and geometry within this unified latent space.
[0035] Specifically, during the training phase, pairs of monocular video sequences are input. and human motion coefficient sequence A video encoder is used to extract data from a monocular video sequence. Extracting visual embeddings ; A motion encoder is used to obtain human motion coefficient sequences Extracting Action Embedding The joint encoder utilizes a cross-modal attention mechanism. and Mapped to dimension The unified latent space is jointly characterized. The joint representation Z is decoded by a visual feature decoder and an action feature decoder to generate visual and action features, respectively. The visual and action feature decoders use the same diffusion model structure, but the output visual and action features differ in dimensionality.
[0036] To ensure this joint characterization It possesses the ability to reconstruct both visual appearance and action structure. The implementation example employs a generative alignment strategy based on a diffusion model, utilizing action diffusion loss. With visual diffusion loss Optimize the vision and motion co-encoder.
[0037] Among them, action diffusion loss The ability of the joint characterization Z to generate action modes is used to constrain the ability of the characterization Z to generate action modes, and it is expressed as follows:
[0038] in Indicates the action characteristics in the spreading step The noise representation below, To add Gaussian noise, For prediction noise in the action feature decoder, Expressing expectations, Represents the square of the L2 norm; Visual diffusion loss The ability of the joint representation to generate visual modalities is used to constrain the ability of the joint representation to do so, and it is expressed as follows:
[0039] in Indicates latent visual features in the diffusion step The noise representation below, The number of potential visual features in a single frame. i Indicates latent visual feature index, To jointly characterize the corresponding local latent representation in Z, Indicating targeting Added Gaussian noise, This represents the prediction noise in the visual feature decoder.
[0040] By minimizing the joint loss A well-trained vision-action joint encoder can process only the input visual feature vector. The accurate mapping is a joint representation Z containing rich motion priors.
[0041] To support efficient training of the joint visual and action representation module, and addressing the pain point that monocular videos and action coefficients often appear in pairs in existing data, this invention also designs a bidirectional data augmentation process based on the Aios backbone network and parametric rendering, specifically including the following two aspects: On the one hand, for data scenarios where monocular videos exist but corresponding motion coefficients are missing: the embodiment uses the Aios backbone network as the core pose estimator to predict the SMPL-X 3D human body parameters in video frames. The SMPL-X parameters specifically include shape parameters, body pose parameters, hand pose parameters, chin pose parameters, facial expression parameters, and camera parameters. To address the jitter and inconsistency issues that may arise from frame-by-frame prediction, the embodiment employs the following optimization strategies: First, the shape parameters and camera intrinsic parameters predicted in the first video frame are globally fixed to ensure consistency between the human body shape and imaging geometry throughout the entire video sequence; second, temporal smoothing (Gaussian smoothing) is applied to the body and local pose parameters to mitigate noise interference from single-frame prediction; finally, all parameters are integrated to form a complete, continuous sequence of human motion coefficients that is visually highly aligned with the original video.
[0042] On the other hand, for data scenarios where motion coefficients exist but corresponding monocular videos are lacking: This embodiment utilizes the generative characteristics of the SMPL-X model and employs a differentiable rendering technique based on the graphics pipeline to map existing human motion coefficients to corresponding monocular human videos. Specifically, the motion coefficients are input into the SMPL-X model to generate a 3D mesh, and combined with randomly sampled background images and texture lighting conditions, realistic 2D human motion videos are rendered. In this way, a synthetic data pair with strict consistency between vision and motion is constructed, greatly enriching the diversity of training samples for the joint representation module.
[0043] S3 uses an action decoder to decode the action feature sequence and generate a sequence of human action coefficients that are consistent with the semantics and rhythm of the target audio content.
[0044] like Figure 4 As shown, the action decoder uses a UNet module with a stream matching training strategy and an ordinary differential equation (ODE) numerical integrator, which can convert the input action feature sequence into a human action coefficient sequence.
[0045] In the reasoning phase of action generation, the first step is to use the standard Gaussian distribution. Mid-sample initial noise Using the ODE numerical integrator, the action feature sequence output in step S2 is... and time variables As a conditional input, in time During the process of moving from 0 to 1, the UNet module adjusts according to the current noise level. and time variables and in action feature sequence Mapping features after the time-series mapping layer As a guide, predict the velocity field of the current state. The ODE numerical integrator is based on the velocity field. Update the noise state, thus gradually reducing the initial noise. Transform into target action The final output sequence of human motion coefficients includes body posture parameters, hand posture parameters, and facial parameters. These parameters are sequential in time, conform to human anatomical structure in space, and are highly aligned with the semantic rhythm of the input audio, thus completing a high-quality modeling from audio to human motion coefficients.
[0046] The action decoder also needs to be trained before it can be used. During training, the action encoder and action decoder are introduced to form a training framework, such as... Figure 4 As shown, the motion encoder adopts an RVQVAE architecture for body part segmentation, specifically dividing the body parts into face, upper body, lower body, and hands. The training process consists of two phases. The first phase trains the motion encoder and motion decoder based on the RVQVAE structure for each body part. To achieve multi-level residual quantization, the model employs... Cascaded quantizers, training loss It consists of two parts: reconstruction loss and commitment loss, specifically expressed as follows:
[0047] in, Represents the actual sequence of actions. This represents the sequence of actions for reconstruction, the first item. The first term represents the absolute difference, used to constrain the accuracy of the reconstruction; the second term represents the commitment loss, used to constrain the stability of the quantization process. This represents the weighting parameter corresponding to the committed loss. Indicates the relationship with RVQ The residual levels are accumulated. Indicates the first The output vector of the layer quantizer Indicates the first The nearest neighbor discrete vector retrieved from the layer codebook. Indicates the first Weighting coefficients for layer residual quantization This represents a continuous sequence of representations output by the motion encoder. This represents a discrete codebook representation sequence obtained through vector quantization. This indicates that gradient propagation is blocked during backpropagation. Indicates the mean squared error; The second stage freezes the action encoder and replaces the action decoder with the action decoder from the UNet module for training, with the training loss... Represented as:
[0048] in, This represents the initial noisy action sequence. Represents the target action sequence. Represents a time variable. Indicates time Between and Interpolation states between them This represents the velocity field predicted by the UNet module. For UNet parameters, Indicates in Calculate the expectation on the joint distribution.
[0049] After building and training the model through the above steps, it can pass Figure 1 The flowchart shown takes the target audio as input and outputs human motion coefficients that match the rhythm and cadence of the audio.
[0050] Based on the above inventive concept, such as Figure 5 As shown, the embodiment also provides a visual and motion joint representation system 70 for human motion generation. At the hardware level, the system includes a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they implement the aforementioned modeling method. At the logical function level, the system 70 specifically includes: a feature extraction module 71, used to receive target audio and extract visual feature vectors containing human posture structure and spatial geometric information from the target video using a pre-trained multimodal large model; a joint representation mapping module 72, used to use the visual feature vectors as guiding information, and utilize the visual and motion joint representation module to align visual features and motion features in a unified latent space based on the guiding information to obtain a joint representation, and generate a motion feature sequence based on the joint representation; and a motion generation module 73, used to decode the motion feature sequence using a motion decoder to generate a human motion coefficient sequence consistent with the semantics and rhythm of the target audio content.
[0051] It should be noted that the visual and motion joint representation system for human motion generation provided in the above embodiments should be illustrated using the above-described functional module division when modeling from audio to human motion coefficients. The functions can be assigned to different functional modules as needed, i.e., the internal structure of the terminal or server can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the visual and motion joint representation system for human motion generation provided in the above embodiments and the visual and motion joint representation method embodiments for human motion generation belong to the same concept. Their specific implementation process is detailed in the visual and motion joint representation method embodiments for human motion generation, and will not be repeated here.
[0052] Based on the same inventive concept, the embodiments also provide a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements the above-described visual and motion joint representation method for human motion generation, specifically including the following steps: S1, using a pre-trained multimodal large model to extract visual feature vectors containing human posture structure and spatial geometric information from the target video; S2 uses visual feature vectors as guiding information and utilizes the visual and action joint representation module to align visual features and action features in a unified latent space based on the guiding information to obtain joint representations. Based on the joint representations, action feature sequences are generated. S3 uses an action decoder to decode the action feature sequence and generate a sequence of human action coefficients that are consistent with the semantics and rhythm of the target audio content.
[0053] In this embodiment, the computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data.
[0054] The specific embodiments described above illustrate the technical solution and beneficial effects of the present invention in detail. It should be understood that the above description is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for joint visual and motion representation of human motion generation, characterized in that, Includes the following steps: Visual feature vectors containing human pose structure and spatial geometric information are extracted from target videos using a pre-trained multimodal large model. The pre-trained multimodal large model adopts a pre-trained Wan-S2V model, which includes a variational autoencoder component and a DiT module. During feature extraction, the audio feature sequence corresponding to the target video is injected into the DiT module as a condition, and the latent space features of the output layer of the DiT module are extracted as visual feature vectors during the forward inference process. The training process includes a visual and action joint representation module comprising a visual and action joint encoder and an action feature decoder. The training architecture includes a video encoder, an action encoder, a visual and action joint encoder, a visual feature decoder, and an action feature decoder. Paired monocular video sequences and human motion coefficient sequences are processed by the video encoder and action encoder to extract visual embeddings and action embeddings, respectively. The joint encoder maps the visual embeddings and action embeddings to a unified latent space through a cross-modal attention mechanism to obtain a joint representation. The joint representation is then decoded by the visual feature decoder and action feature decoder to generate visual features and action features, respectively. Based on this architecture, a generative algorithm based on a diffusion model is employed. The alignment strategy optimizes the visual and action joint encoder through action diffusion loss and visual diffusion loss. Using visual feature vectors as guiding information, the visual and action joint representation module aligns visual features and action features in a unified latent space based on the guiding information to obtain a joint representation. Based on the joint representation, an action feature sequence is generated. The process includes: inputting visual feature vectors as guiding information into the visual and action joint encoder, aligning visual features and action features in a unified latent space, generating a joint representation of visual and action features containing human posture structure and spatial geometric information, and generating an action feature sequence containing rich motion priors through the action feature decoder. The action decoder, which includes a UNet module employing a stream matching training strategy and an ordinary differential equation numerical integrator, decodes the action feature sequence to generate a human action coefficient sequence consistent with the semantics and rhythm of the target audio content. Specifically, this includes: using an ODE numerical integrator, the action feature sequence output in step S2 is decoded to generate a sequence of human action coefficients consistent with the semantics and rhythm of the target audio content. and time variables As a conditional input, in time During the process of moving from 0 to 1, the UNet module adjusts according to the current noise level. and time variables and in the action feature sequence Mapping features after the time-series mapping layer As a guide, predict the velocity field of the current state. The ODE numerical integrator is based on the velocity field. Update the noise state, thus gradually reducing the initial noise. Transform into target action This is the final output sequence of human motion coefficients; The motion decoder also needs to be trained before it can be used. During training, a training framework is formed by the motion encoder and motion decoder. The training consists of two phases. The first phase trains the motion encoder and motion decoder using an RVQVAE structure for each body part. The cascaded quantizer has a training loss consisting of reconstruction loss and commitment loss. In the second stage, the action encoder is frozen, and the action decoder is replaced with the action decoder of the UNet module for training.
2. The visual and motion joint representation method for human motion generation according to claim 1, characterized in that, Action diffusion loss The ability of the joint characterization Z to generate action modes is used to constrain the ability of the characterization Z to generate action modes, and it is expressed as follows: in Indicates the action characteristics in the spreading step The noise representation below, To add Gaussian noise, For prediction noise in the action feature decoder, Expressing expectations, Represents the square of the L2 norm; Visual diffusion loss The ability of the joint representation to generate visual modalities is used to constrain the ability of the joint representation to do so, and it is expressed as follows: in Indicates latent visual features in the diffusion step The noise representation below, The number of potential visual features in a single frame. i Represents the latent visual feature index. To jointly characterize the corresponding local latent representation in Z, Indicating targeting Added Gaussian noise, This represents the prediction noise in the visual feature decoder.
3. The visual and motion joint representation method for human motion generation according to claim 1, characterized in that, The paired monocular video sequences and human motion coefficient sequences used during the training of the vision and motion joint representation module are constructed in the following manner: To address the issue of missing motion coefficients in monocular videos, the Aios backbone network is used to predict the SMPL-X 3D human body parameters in video frames. For the problem of motion coefficients existing but lacking corresponding monocular videos, a parametric rendering method based on the SMPL-X model is adopted. Through differentiable rendering technology based on the graphics pipeline, motion coefficients are mapped to corresponding monocular human videos, thus constructing data pairs that are consistent with vision and motion.
4. The visual and motion joint representation method for human motion generation according to claim 1, characterized in that, Training loss Specifically, it is expressed as follows: in, Represents the actual sequence of actions. This represents the sequence of actions for reconstruction, the first item. The first term represents the absolute difference, used to constrain the accuracy of the reconstruction; the second term represents the commitment loss, used to constrain the stability of the quantization process, where... This represents the weighting parameter corresponding to the committed loss. Indicates the relationship with RVQ The residual levels are accumulated. Indicates the first The output vector of the layer quantizer Indicates the first The nearest neighbor discrete vector retrieved from the layer codebook. Indicates the first Weighting coefficients for layer residual quantization This represents a continuous sequence of representations output by the motion encoder. This represents a discrete codebook representation sequence obtained through vector quantization. This indicates that gradient propagation is blocked during backpropagation. Indicates the mean squared error; Training loss in the second phase Represented as: in, This represents the initial noisy action sequence. Represents the target action sequence. Represents a time variable. Indicates time Between and Interpolation states between them This represents the velocity field predicted by the UNet module. For UNet parameters, Indicates in Calculate the expectation on the joint distribution.
5. A computer-readable storage medium, characterized in that, It stores a program that, when executed by a processor, implements the visual and motion joint representation method for human motion generation as described in any one of claims 1-4.
Citation Information
Patent Citations
Action generation method and device based on multi-modal pre-training, robot and medium
CN120663303A
Children story video generation method and system based on AI
CN120812370A
Three-dimensional digital human generation method and system capable of voice interaction
CN120931773A
Video generation method and device, intelligent agent, electronic equipment and storage medium
CN121000951A
Robot action generation method and device
CN121061872A