A dance action generation method based on a five-dimensional quadratic kernel model

By using a five-dimensional quadratic kernel model and diffusion iterative transformation technology, combined with inertial motion capture and Jukebox encoder, the problems of rhythm and emotional expression in dance movement generation are solved, achieving high-quality dance movement generation suitable for digital dance art creation.

CN122492980APending Publication Date: 2026-07-31TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2026-04-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively capture the subtle rhythms, fluid spatiotemporal dynamics, and expressive body language of dance movements. They also have limited cross-modal alignment capabilities between music and dance, resulting in generated dance movements that are inconsistent with the emotional expression of the music, lack artistic resonance, and fail to reach the artistic level of professional dance performances.

Method used

A five-dimensional quadratic kernel modeling method is adopted, combined with diffusion iterative transformation (DIT) technology and quadratic kernel parameter representation. Dance data is collected through an inertial motion capture system to construct a three-dimensional human geometric representation. Modeling is performed based on a spatiotemporally continuous symbolic distance field. Audio features are extracted by combining Jukebox encoder to construct a diffusion Transformer generation architecture. Multi-objective collaborative optimization is carried out to generate high-quality dance movement sequences.

Benefits of technology

It achieves the preservation of dance movement details and the enhancement of fluidity, the precise mapping of musical emotional characteristics with stylized dance movements, and the coordinated and consistent generation of dance movements with musical emotional expression, reaching the artistic level of professional dance performance, and is suitable for the intelligent creation of digital dance art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492980A_ABST
    Figure CN122492980A_ABST
Patent Text Reader

Abstract

This invention proposes a method for generating dance movements using a five-dimensional quadratic kernel model. The method involves collecting full-body motion data from dancers to construct a three-dimensional geometric representation of the human body; modeling a five-dimensional quadratic kernel hybrid model based on a spatiotemporally continuous symbolic distance field, and extracting the five-dimensional quadratic kernel parameters as motion representations; using an encoder to extract rhythm, melody, and emotional features from the audio; injecting audio features through adaptive layer normalization to learn the spatiotemporal distribution mapping of the five-dimensional quadratic kernel parameters; reconstructing the dynamic symbolic distance field based on the generated parameters to calculate the training loss; achieving multi-objective collaborative optimization; inputting music in segments, training the network to output five-dimensional quadratic kernel parameters and reconstructing the symbolic distance field; and obtaining the dance movement sequence through isosurface extraction and post-processing. This invention combines the five-dimensional quadratic kernel model with DiT technology, designing a multi-stage progressive generation process and a cross-modal alignment algorithm, which can be applied to fields such as digital human performance and intelligent dance teaching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of 3D vision, computer graphics, and 3D model reconstruction, and in particular to a method for generating dance movements using a five-dimensional quadratic kernel model. Background Technology

[0002] Currently, AI-based dance motion generation technology still faces numerous technical challenges in practical applications, especially when dealing with highly artistic and expressive dance movements, where existing solutions exhibit significant limitations. Specifically, the following three key technical bottlenecks exist in this field:

[0003] Firstly, in dance movement modeling, traditional parametric methods (such as those based on parametric human models like SMPL) primarily focus on the geometric representation of human posture, making it difficult to effectively capture the subtle rhythms, fluid spatiotemporal dynamics, and expressive body language unique to dance. These methods typically simplify human movement into discrete joint trajectories, ignoring the continuity, rhythm, and artistry inherent in dance movements, resulting in generated movements lacking the rhythmic beauty and dynamic expressiveness characteristic of dance.

[0004] Secondly, in terms of music-dance cross-modal alignment, existing models have limited ability to understand music signals, making it difficult to deeply explore the multi-layered emotional information contained in the audio. Although some methods attempt to establish a connection between music and dance through simple rhythmic features or spectral analysis, these methods cannot accurately capture the subtle emotional changes, melodic fluctuations, and dynamic layers in the music, resulting in a significant inconsistency between the generated dance movements and the emotional expression of the music, making it difficult to achieve true artistic resonance.

[0005] Finally, regarding the quality of generation and artistic expression, due to the combined effect of the aforementioned technological limitations, current AI-generated dance movements often exhibit mechanical, repetitive, and emotionally shallow characteristics. The generated results fall short of the artistic standards required for professional dance performances and fail to meet the high standards of naturalness, expressiveness, and artistic appeal demanded in practical applications, severely hindering the in-depth application of AI technology in dance creation, teaching, and performance. Summary of the Invention

[0006] The present invention aims to at least partially solve one of the technical problems in the related art.

[0007] Therefore, the first objective of this invention is to propose a dance motion generation method based on a five-dimensional quadratic kernel model, which innovatively combines diffusion iterative transformation (DIT) technology with quadratic kernel parameter representation: DIT technology effectively solves the problem of detail preservation in long sequence motions through its progressive generation characteristics, while quadratic kernel parameter modeling enables a multi-scale accurate description of the dynamic geometric features of the human body.

[0008] Another objective of this invention is to propose a dance motion generation system based on a five-dimensional quadratic kernel model.

[0009] To achieve the above objectives, a first aspect of the present invention proposes a method for generating dance movements using a five-dimensional quadratic kernel model, comprising: S1, based on the full-body motion data of professional dancers dancing to music collected by the inertial motion capture system, outputs a three-dimensional human geometric representation of each frame of motion constructed based on the depth symbolic distance field algorithm. S2, input the three-dimensional geometric representation, and perform time-segmented modeling based on the spatiotemporally continuous symbolic distance field, outputting the five-dimensional quadratic kernel hybrid model parameters as a motion representation; S3, for the audio signal of the music played during the dancer's dance, output an audio sequence containing rhythm, melody and emotional information based on the Jukebox encoder; S4. Based on the audio features of the audio sequence and the parameters of the five-dimensional quadratic kernel hybrid model represented by the movement sequence of the dancer dancing to the music, a generative architecture based on diffusion Transformer is constructed. The audio semantic features are injected into the denoising network through adaptive layer normalization, and the edge distribution and conditional probability mapping of the five-dimensional quadratic kernel parameters in the spatiotemporal field are learned. S5. Reconstruct the continuous dynamic symbolic distance field based on the generated five-dimensional quadratic kernel parameters, and calculate the loss function value during the training process based on the reconstruction result. S6. During the diffusion training process, input the loss function value, introduce inverse kinematics loss, rhythm alignment loss and style consistency loss to perform multi-objective collaborative optimization, and output the trained conditional generation network. S7, the music signal to be generated for the dance is segmented, and audio features are extracted using the method described in step S3; the corresponding audio features are input into the trained conditional generation network obtained in step S6 for inference to obtain a five-dimensional quadratic kernel parameter sequence; based on the five-dimensional quadratic kernel parameter sequence, a continuous dynamic symbolic distance field is reconstructed, and the final dance movement sequence is output through isosurface extraction and post-processing.

[0010] In one embodiment of the present invention, the human body SMPL parameterized template captured in a single frame is characterized as a symbolic distance field SDF.

[0011] In one embodiment of the present invention, the music and the dance sequence in S1 are in one-to-one correspondence.

[0012] In one embodiment of the present invention, actions in consecutive frames are timed. The Segmentation; for modeling each action sequence segment, step S2 specifically includes the following steps: Step S21, five-dimensional quadratic kernel modeling, transforms the symbol distance field data of multiple frames of this time series into... N Modeling is done using a 5x5 matrix, meaning the number of rows is equal to the 3D spatial resolution. n and time frame count T product N=n × T The data matrix has 5 columns; the data is randomly divided into... K Given a set of groups, calculate the mean vector for each group. Step S22, for N Data and K Calculate the mean square error for each mean vector, assign the data point to the group with the smallest mean square error, and iteratively calculate the mean; repeat this step 14 times to obtain the final grouping. K The dataset yields the mean vector, covariance matrix, and weights. ; Step S23: Reconstruct the formula using a five-dimensional quadratic kernel gate function. Reconstructing symbolic distance field data ,in Represents a quadratic kernel gate function, This represents the conditional mean of the second-order kernel.

[0013] In one embodiment of the present invention, the motion generation is driven by the music accompanying the dancer, and step S3 specifically includes the following steps: Step S31: Divide the input audio into fixed duration segments. The A segment of seconds, for each audio segment A i The Jukebox model with frozen weights was used as the audio encoder to extract latent features Z. i =Jukebox(A i )∈R Ta×Dz ,in T a The number of time steps for an audio segment. D z Jukebox feature dimensions; Step S32: Adapt high-dimensional features to the conditional space of the diffusion model through a learnable projection layer. Use a classifier-independent guidance strategy and set empty features with a 10% probability during training so that the model can learn both conditional and unconditional generation.

[0014] In one embodiment of the present invention, step S4 specifically includes the following steps: Step S41: Synchronously segment the dance movement data into segments of Ta seconds, with each segment represented by a five-dimensional quadratic kernel parameter Θ. i , i This is the segment number; for segments shorter than Ta seconds, the last frame action is used to fill the gap. Step S42: Construct a denoising network based on the DiT block of the diffusion transformer. f θ It utilizes its built-in adaptive layer normalization unit to inject audio features and time step information into the latent space, thereby achieving global correlation modeling and denoising prediction of human geometric sequences. Step S43, using audio-action clip pairs ( C i ,Θ i Using the training samples, optimize and simplify the diffusion loss. L simple =E [|| - f θ (Θ t , t , C i )|| 2 ].

[0015] In one embodiment of the present invention, step S5 specifically includes the following steps: Step S51: Construct a five-dimensional quadratic kernel gate function The formula for obtaining it is as follows:

[0016] in It is the weight of the j-th array; Based on location x , y , z and time variables t The secondary nuclear edge distribution, in which The expression for the five-dimensional quadratic kernel is: f ( x , y , z , t , w ),set up ,So f ( x , y , z , t , w ) expressed as :

[0017] Where is the mean vector of the j-th group of data; Let be the covariance matrix of the j-th data set; where, It is a variable in k sets of data x , y , z , t , w The random variables of the j-th data group, the mean, covariance matrix, and weights of the j-th data group have all been obtained in S2 and are known quantities. Then:

[0018] Step S52: Construct the conditional mean function, and obtain the formula as follows:

[0019] in, Let x, y, z, t, w be random variables of the j-th data set. The mean vector of the j-th data set obtained from S3 is: The mean vector of the position variables is The covariance matrix of the j-th data group obtained from S3 is: , It is a submatrix of the covariance matrix, where the first 4×4 dimension matrix is ​​the covariance matrix of the position variables: ; Step S53: Construct the reconstructed symbol distance field, and obtain the formula as follows: , in To reconstruct the symbolic distance field.

[0020] In one embodiment of the present invention, step S6 specifically includes the following steps: Step S61: Based on the human template symbol distance field reconstructed in step S5, extract the joint positions. The joint positions of the parameterized human model SMPL that originally captured the actions in the sequence. J ori ( i Calculate the mean square error loss ; Step S62: Extract the center of gravity velocity curve from the reconstructed action symbol distance field sequence. v ( t ), and detect its local extreme points as the rhythm of the movement. t p Music beat sequences extracted from audio features tb Perform dynamic time warping (DTW) alignment ; Step S63: Use a pre-trained dance movement style classifier C style Style scoring is performed on the reconstructed action sequences, and negative log-likelihood loss is calculated to enhance style preservation. ; Step S64: Combine the above loss with the diffusion-based loss to form the total loss function: L total = L simple + l 1 L 1+ l 2 L 2+ l 3 L 3, in l 1, l 2, l 3 represents the loss coefficient.

[0021] In one embodiment of the present invention, step S7, which obtains the final dance movement sequence through isosurface extraction and post-processing, specifically includes the following steps: Step S71: For each frame of the reconstructed symbolic distance field, for each edge, if the symbols of the two connected vertices are different, interpolate on this edge, calculate the position where the interpolation value is 0, and place the vertex. Step S72: For each frame of the reconstructed symbolic distance field, construct triangular patches between vertices.

[0022] The method of this invention not only breaks through the limitations of traditional methods in terms of motion fluency and style fidelity, but also achieves for the first time a precise mapping between musical emotional characteristics and stylized dance movements through an innovative cross-modal attention mechanism, providing a new technical paradigm for the intelligent creation of digital dance art.

[0023] This invention also proposes a dance movement generation system based on a five-dimensional quadratic kernel model, comprising: The motion data acquisition and geometric representation construction module is used to represent the three-dimensional geometry of each frame of motion by collecting SMPL data of professional dancers dancing to music based on the inertial motion capture system and using the depth symbolic distance field algorithm. The five-dimensional quadratic kernel parameter extraction module is used to input the three-dimensional geometry, perform time-segmented modeling based on the spatiotemporally continuous symbolic distance field, and output the five-dimensional quadratic kernel hybrid model parameters as a motion representation. The audio feature extraction module is used to output an audio sequence containing rhythm, melody and emotional information based on the Jukebox encoder for the audio signal of the music played during the dancer's dance. The Conditional Generative Network Training Module is used to construct a generative architecture based on the diffusion Transformer, which is based on the audio features of the audio sequence and the parameters of the five-dimensional quadratic kernel hybrid model represented by the dance sequence of the dancer dancing to the music. The audio semantic features are injected into the denoising network through adaptive layer normalization, and the edge distribution and conditional probability mapping of the five-dimensional quadratic kernel parameters in the spatiotemporal field are learned. The differentiable reconstruction and loss calculation module is used to reconstruct the continuous dynamic symbolic distance field based on the generated five-dimensional quadratic kernel parameters, and to calculate the loss function value during the training process based on the reconstruction result. The multi-objective collaborative optimization module is used to input the loss function value during the diffusion training process, introduce inverse kinematics loss, rhythm alignment loss and style consistency loss to perform multi-objective collaborative optimization, and output the trained conditional generation network. The dance movement sequence generation module is used to input new music signals of the dance to be generated into the trained conditional generation network in segments for inference and generation, and output a five-dimensional quadratic kernel parameter sequence. After segmented reconstruction of the continuous dynamic symbolic distance field, the final dance movement sequence is output through isosurface extraction and post-processing.

[0024] In summary, the beneficial effects of the present invention are as follows: 1. This invention uses five-dimensional quadratic kernel correlation statistics to perform dynamic joint modeling of human body templates, effectively utilizing the theoretical characteristics of the five-dimensional quadratic kernel model to establish a correlation between multi-dimensional human dynamic data in time and space.

[0025] 2. This invention uses the mean vector, covariance matrix and weights of a five-dimensional quadratic kernel model as the basis for expression, which can effectively learn the dynamic local features of the three-dimensional human body template action sequence, and make the understanding of human body action have spatiotemporal inductive generalization.

[0026] 3. This invention uses dynamic multi-plane decomposition for parameter optimization, which can solve the difficulty of iterative calculation of high-dimensional data and optimize data decomposition.

[0027] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0028] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1This is a flowchart of a dance motion generation method based on a five-dimensional quadratic kernel model according to an embodiment of the present invention; Figure 2 This is a structural diagram of a dance motion generation system based on a five-dimensional quadratic kernel model according to an embodiment of the present invention. Detailed Implementation

[0029] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0031] The following describes, with reference to the accompanying drawings, a method and system for generating dance movements using a five-dimensional quadratic kernel model, according to an embodiment of the present invention.

[0032] Figure 1 This is a flowchart illustrating a method for generating dance movements using a five-dimensional quadratic kernel model, according to an embodiment of the present invention. Figure 1 As shown, a method for generating dance movements using a five-dimensional quadratic kernel model includes the following steps: S1, based on the full-body motion data of professional dancers dancing to music collected by the inertial motion capture system, outputs a three-dimensional human geometric representation of each frame of motion constructed based on the depth symbolic distance field algorithm. It is understood that this invention converts SMPL data into a symbolic distance field using the following method: By establishing a mapping from spatial sampling points to the surface of the SMPL parameterized model through Mesh to Symbolic Distance Field (Mesh2SDF), the shortest directional distance from points to the human body mesh is calculated, thereby transforming the discrete skeleton-driven human body model into a continuous differentiable symbolic distance field (SDF) expression.

[0033] Specifically, this method first generates a regular or adaptive set of spatial sampling points in three-dimensional space. Then, it uses the SMPL (Skinned Multi-Person Linear Model) parameterized human body model to generate a corresponding triangular mesh representation. This mesh is jointly controlled by the shape parameter β and the pose parameter θ, which can accurately describe the human body geometry under different body shapes and poses. By constructing an efficient spatial acceleration data structure (such as a BVH tree or KD-Tree), it provides a fast nearest-point search capability for subsequent distance calculations, laying the foundation for realizing the conversion from discrete meshes to continuous fields.

[0034] In the specific implementation process, for each spatial sampling point, the system quickly locates the nearest triangular facet of the SMPL human body mesh through the spatial acceleration structure and accurately calculates the shortest Euclidean distance from the sampling point to the triangular facet. Simultaneously, the system uses ray casting combined with mesh topology information to determine the interior-exterior relationship of the sampling point relative to the human body mesh, thereby determining the sign attribute of the distance (negative for interior, positive for exterior). After organizing the distance values ​​and sign information of all sampling points into a three-dimensional scalar field, the discrete sampling results are further transformed into a continuous, smooth signed distance field function through continuous processing methods such as trilinear interpolation or radial basis functions. This function supports accurate querying and gradient calculation at any spatial location.

[0035] Using the above method, the originally discrete skeleton-driven human mesh model was successfully transformed into a continuous, differentiable symbolic distance field representation. This SDF representation exhibits good geometric continuity and mathematical differentiability, enabling end-to-end gradient propagation through automatic differentiation frameworks (such as PyTorch or TensorFlow), and supporting direct optimization of the human shape parameter β and pose parameter θ. This continuous field representation demonstrates significant advantages in applications such as 3D human reconstruction, physical simulation, collision detection, and inverse dynamics, providing a powerful mathematical tool and representational foundation for deep learning-based human modeling and animation generation.

[0036] S2, input the three-dimensional geometric representation, and perform time-segmented modeling based on the spatiotemporally continuous symbolic distance field, outputting the five-dimensional quadratic kernel hybrid model parameters as a motion representation; It is understood that the preferred embodiment of this invention is a five-dimensional quadratic kernel hybrid model: the five-dimensional model can jointly model the time dimension, that is, a set of spatial three-dimensional positions and time dimensions have corresponding signed distance field values. Specifically: It enables spatiotemporal joint modeling, completely encoding the dynamic human movement process into a unified mathematical representation. Specifically, this five-dimensional model consists of three-dimensional spatial coordinates (x, y, z), a time dimension (t), and the corresponding symbolic distance field value w (SDF value), forming a five-dimensional mapping relationship of (x, y, z, t, w). This design allows the system to simultaneously capture the geometric shape of the human body in space and its dynamic changes in the time dimension within a single continuous field representation, avoiding the problem of separating spatial and temporal information in traditional methods.

[0037] By explicitly incorporating the temporal dimension into the modeling process of the symbolic distance field, the five-dimensional model can naturally express the continuity and temporal consistency of human motion. In practical applications, given any spatial location and time point, the system can directly query the SDF value corresponding to that spatiotemporal point, thereby achieving a precise geometric description of the human body at any time and any location. This spatiotemporal joint representation is particularly suitable for tasks such as dynamic human reconstruction, motion prediction, and temporal action generation, effectively capturing the geometric continuity and physical plausibility of the motion process.

[0038] Furthermore, the differentiability of the five-dimensional model allows the time dimension to participate in the end-to-end gradient optimization process. Through an automatic differentiation framework, the system can not only optimize geometric accuracy in the spatial dimension but also learn and optimize motion patterns in the temporal dimension, achieving deep modeling of human motion dynamics. This unified spatiotemporal representation provides a powerful mathematical tool for deep learning-based dynamic human body analysis, significantly improving the model's performance in complex temporal tasks.

[0039] S3, for the audio signal of the music played during the dancer's dance, output an audio sequence containing rhythm, melody and emotional information based on the Jukebox encoder; Understandably, Jukebox is a primitive audio generation model based on a multi-level vector quantization autoencoder (VQ-VAE) and an autoregressive Transformer, capable of learning and extracting latent musical representations containing melody, rhythm, and timbre features.

[0040] In its implementation, the Jukebox model employs a hybrid architecture combining a multi-level Vector Quantized Variational Autoencoder (VQ-VAE) with an autoregressive Transformer. The multi-level VQ-VAE progressively compresses the original audio waveform into discrete latent representations at different time scales through multiple layers of encoders. Each layer corresponds to a different granularity of musical structure: the bottom-level encoder captures high-frequency details such as timbre and transient features; the middle-level encoder extracts rhythmic patterns and beat information; and the top-level encoder learns macroscopic musical semantics such as melody contours and harmonic structures. Through vector quantization, continuous latent vectors are mapped onto a predefined discrete codebook, forming a compact and information-rich sequence of discrete representations.

[0041] In the audio feature extraction stage, the raw audio signals recorded during the dancers' performance are input into the Jukebox encoder. After multi-level feature extraction and quantization processing, the output is a discrete latent sequence containing multi-layered musical information. This sequence not only contains precise rhythmic temporal information (such as beat position, rhythm density, and tempo changes), but also encodes rich melodic features (such as pitch contours, interval relationships, and tonal structure) and deep emotional semantic information (such as musical emotional intensity, emotional polarity, and dynamic changes). These multi-dimensional audio representations can be precisely time-aligned with the dance movement sequence, providing a crucial feature foundation for subsequent tasks such as dance-music synchronization analysis, motion generation, and style transfer.

[0042] S4. Based on the audio features of the audio sequence and the parameters of the five-dimensional quadratic kernel hybrid model represented by the movement sequence of the dancer dancing to the music, a generative architecture based on diffusion Transformer is constructed. The audio semantic features are injected into the denoising network through adaptive layer normalization, and the edge distribution and conditional probability mapping of the five-dimensional quadratic kernel parameters in the spatiotemporal field are learned. S5. Reconstruct the continuous dynamic symbolic distance field based on the generated five-dimensional quadratic kernel parameters, and calculate the loss function value during the training process based on the reconstruction result. S6. During the diffusion training process, input the loss function value, introduce inverse kinematics loss, rhythm alignment loss and style consistency loss to perform multi-objective collaborative optimization, and output the trained conditional generation network. S7, the music signal to be generated for the dance is segmented, and audio features are extracted using the method described in step S3; the corresponding audio features are input into the trained conditional generation network obtained in step S6 for inference to obtain a five-dimensional quadratic kernel parameter sequence; based on the five-dimensional quadratic kernel parameter sequence, a continuous dynamic symbolic distance field is reconstructed, and the final dance movement sequence is output through isosurface extraction and post-processing.

[0043] Specifically, the dance motion generation method based on the five-dimensional quadratic kernel model of the present invention is implemented through the following steps: Full-body motion data of dancers are collected using an inertial motion capture system, and a three-dimensional human geometric representation of each frame of motion is constructed using a depth symbolic distance field algorithm; a five-dimensional quadratic kernel hybrid model is modeled in time segments based on a spatiotemporally continuous symbolic distance field, and the five-dimensional quadratic kernel parameters controlling the dynamic symbolic distance field are extracted as motion representations; this representation can model spatial and temporal information together with human motion, capturing motion details of different parts of the human body, as well as dynamic changes at different time scales; audio features containing rhythm, melody, and emotional information are extracted based on a Jukebox encoder; and a diffusion-based Transformer (DiT) model is constructed. The generative architecture injects audio semantic features into the denoising network through adaptive layer normalization (adaLN-Zero) to learn the edge distribution and conditional probability mapping of five-dimensional quadratic kernel parameters in the spatiotemporal field. Based on the generated five-dimensional quadratic kernel parameters, a continuous dynamic symbolic distance field is reconstructed to generate the loss function during training. In the diffusion training, inverse kinematics loss, rhythm alignment loss, and style consistency loss are introduced to achieve multi-objective collaborative optimization through differentiable reconstruction. The music is input in segments, and through the trained network, the five-dimensional quadratic kernel parameters are output and the continuous dynamic symbolic distance field is reconstructed segment by segment. The final dance movement sequence is obtained through isosurface extraction and post-processing.

[0044] Furthermore, the human SMPL parameterized template captured in a single frame is characterized as a symbolic distance field (SDF).

[0045] Furthermore, the music and the dancers' movement sequences in S1 are in one-to-one correspondence.

[0046] Furthermore, assign time to continuous frame actions. The Segmentation. For modeling each action sequence segment, step S2 specifically includes the following steps: Step S21, five-dimensional quadratic kernel modeling, transforms the symbol distance field data of multiple frames of this time series into... N Modeling is done using a 5x5 matrix, meaning the number of rows is equal to the 3D spatial resolution. n and time frame count T product N=n × T The data matrix has 5 columns. The data is randomly divided into... K Given a set of groups, calculate the mean vector for each group. Step S22, for N Data and K Calculate the mean squared error for each mean vector, then group the data point to the group with the smallest mean squared error, and iteratively calculate the mean. Repeat this step 14 times to obtain the final grouping. KThe dataset yields the mean vector, covariance matrix, and weights. ; Step S23: Reconstruct the formula using a five-dimensional quadratic kernel gate function. Reconstructing symbolic distance field data ,in Represents a quadratic kernel gate function, This represents the conditional mean of the second-order kernel.

[0047] Furthermore, the movement generation is driven by the music accompanying the dancer, and step S3 specifically includes the following steps: Step S31: Divide the input audio into fixed duration segments. The A segment of seconds, for each audio segment A i The Jukebox model with frozen weights was used as the audio encoder to extract latent features Z. i =Jukebox(A i )∈R Ta×Dz ,in T a The number of time steps for an audio segment. D z Jukebox feature dimensions; Step S32: Adapt high-dimensional features to the conditional space of the diffusion model through a learnable projection layer. Use a classifier-independent guidance strategy and set empty features with a 10% probability during training so that the model can learn both conditional and unconditional generation.

[0048] Furthermore, step S4 specifically includes the following steps: Step S41: Synchronously segment the dance movement data into segments of Ta seconds, with each segment represented by a five-dimensional quadratic kernel parameter Θ. i , i This is the segment number. For segments shorter than Ta seconds, the last frame's action is used to fill the gap. Step S42: Construct a denoising network based on the DiT block of the diffusion transformer. f θ It utilizes its built-in adaptive layer normalization unit to inject audio features and time step information into the latent space, thereby achieving global correlation modeling and denoising prediction of human geometric sequences. Step S43, using audio-action clip pairs ( C i ,Θ i Using the training samples, optimize and simplify the diffusion loss. L simple =E [|| - f θ (Θ t ,t , C i )|| 2 ].

[0049] Furthermore, step S5 specifically includes the following steps: Step S51: Construct a five-dimensional quadratic kernel gate function The formula for obtaining it is as follows: , in It is the weight of the j-th array; Based on location x , y , z and time variables t The secondary nuclear edge distribution, in which The expression for the five-dimensional quadratic kernel is: f ( x , y , z , t , w ),set up ,So f ( x , y , z , t , w ) expressed as :

[0050] Where is the mean vector of the j-th group of data; Let be the covariance matrix of the j-th data set. It is a variable in k sets of data x , y , z , t , w The random variables of the j-th data group, the mean, covariance matrix, and weights of the j-th data group have all been obtained in S2 and are known quantities. Then:

[0051] Step S52: Construct the conditional mean function, and obtain the formula as follows:

[0052] in, Let x, y, z, t, w be random variables of the j-th data set. The mean vector of the j-th data set obtained from S3 is: The mean vector of the position variables is The covariance matrix of the j-th data group obtained from S3 is: , It is a submatrix of the covariance matrix, where the first 4×4 dimension matrix is ​​the covariance matrix of the position variables: .

[0053] Step S53: Construct the reconstructed symbol distance field, and obtain the formula as follows: , in To reconstruct the symbolic distance field.

[0054] Furthermore, step S6 specifically includes the following steps: Step S61: Based on the human template symbol distance field reconstructed in step S5, extract the joint positions. The joint positions of the parameterized human model SMPL that originally captured the actions in the sequence. J ori ( i Calculate the mean square error loss ; Step S62: Extract the center of gravity velocity curve from the reconstructed action symbol distance field sequence. v ( t ), and detect its local extreme points as the rhythm of the movement. t p Music beat sequences extracted from audio features t b Perform dynamic time warping (DTW) alignment ; Step S63: Use a pre-trained dance movement style classifier C style Style scoring is performed on the reconstructed action sequences, and negative log-likelihood loss is calculated to enhance style preservation. ; Step S64: Combine the above loss with the diffusion-based loss to form the total loss function: L total = L simple + l 1 L 1+ l 2 L 2+ l 3 L 3 in l 1, l 2, l 3 represents the loss coefficient.

[0055] Furthermore, step S7 obtains the final dance movement sequence through isosurface extraction and post-processing, specifically including the following steps: Step S71: For each frame of the reconstructed symbolic distance field, for each edge, if the symbols of the two connected vertices are different, interpolate on this edge, calculate the position where the interpolation value is 0, and place the vertex. Step S72: For each frame of the reconstructed symbolic distance field, construct triangular patches between vertices.

[0056] This invention presents a dance motion generation method based on a five-dimensional quadratic kernel model. It acquires full-body motion data of dancers using an inertial motion capture system and constructs a three-dimensional geometric representation of the human body for each frame of motion using a depth symbolic distance field algorithm. A five-dimensional quadratic kernel hybrid model is created by segmenting the spatiotemporally continuous symbolic distance field into time periods, and the five-dimensional quadratic kernel parameters controlling the dynamic symbolic distance field are extracted as motion representations. This representation can model spatial and temporal information together with human motion, capturing motion details of different parts of the body and dynamic changes at different time scales. An audio sequence containing rhythm, melody, and emotional information is extracted based on a Jukebox encoder. A diffusion-based Transformer (DiT) model is then constructed. The generative architecture injects audio semantic features into the denoising network through adaptive layer normalization (adaLN-Zero), learning the edge distribution and conditional probability mapping of five-dimensional quadratic kernel parameters in the spatiotemporal field. Based on the generated five-dimensional quadratic kernel parameters, a continuous dynamic symbolic distance field is reconstructed to generate the loss function during training. Inverse kinematics loss, rhythm alignment loss, and style consistency loss are introduced in diffusion training, achieving multi-objective collaborative optimization through differentiable reconstruction. Segmented input music is processed by the trained network, outputting five-dimensional quadratic kernel parameters and reconstructing the continuous dynamic symbolic distance field segment by segment. The final dance sequence is obtained through isosurface extraction and post-processing. A dual optimization framework combining DIT technology and quadratic kernel parameters is proposed, and a multi-stage progressive generation process is designed to ensure motion quality. A cross-modal alignment algorithm specifically for dance is developed, applicable to fields such as digital human performance and intelligent dance teaching.

[0057] To implement the method of the above embodiments, such as Figure 2 As shown, this invention proposes a dance motion generation system 10 based on a five-dimensional quadratic kernel model, comprising: The motion data acquisition and geometric representation construction module 100 is used to represent the three-dimensional geometry of each frame of motion by acquiring the full-body motion data of professional dancers dancing to music based on the inertial motion capture system and using the depth symbolic distance field algorithm. The five-dimensional quadratic kernel parameter extraction module 200 is used to input the three-dimensional geometry, perform time-segmented modeling based on the spatiotemporally continuous symbolic distance field, and output the five-dimensional quadratic kernel hybrid model parameters as a motion representation. The audio feature extraction module 300 is used to output an audio sequence containing rhythm, melody and emotional information based on the Jukebox encoder for the audio signal of the music played during the dancer's dance. The Conditional Generative Network Training Module 400 is used to construct a generative architecture based on diffusion Transformer based on the audio features of the audio sequence and the parameters of the five-dimensional quadratic kernel hybrid model represented by the dance sequence of the dancer dancing to music. The audio semantic features are injected into the denoising network through adaptive layer normalization, and the edge distribution and conditional probability mapping of the five-dimensional quadratic kernel parameters in the spatiotemporal field are learned. Differentiable reconstruction and loss calculation module 500 is used to reconstruct a continuous dynamic symbolic distance field based on the generated five-dimensional quadratic kernel parameters, and to calculate the loss function value during the training process based on the reconstruction result; The multi-objective collaborative optimization module 600 is used to input the loss function value during the diffusion training process, introduce inverse kinematics loss, rhythm alignment loss and style consistency loss to perform multi-objective collaborative optimization, and output the trained conditional generation network. The dance movement sequence generation module 700 is used to input new music signals of the dance to be generated into the trained conditional generation network for inference generation in segments, and output a five-dimensional quadratic kernel parameter sequence. After segmented reconstruction of the continuous dynamic symbol distance field, the final dance movement sequence is output through isosurface extraction and post-processing.

[0058] The system of this invention proposes a dual optimization framework combining DIT technology and secondary kernel parameters, designs a multi-stage progressive generation process to ensure motion quality, develops a cross-modal alignment algorithm specifically for dance, and can be applied to fields such as digital human performance and intelligent dance teaching.

[0059] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0060] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A method for generating dance motion by a five-dimensional quadratic kernel model, characterized in that, Includes the following steps: S1, based on the full-body motion data of professional dancers dancing to music collected by the inertial motion capture system, outputs a three-dimensional human geometric representation of each frame of motion constructed based on the depth symbolic distance field algorithm. S2, input the three-dimensional geometric representation, and perform time-segmented modeling based on the spatiotemporally continuous symbolic distance field, outputting the five-dimensional quadratic kernel hybrid model parameters as a motion representation; S3, for the audio signal of the music played during the dancer's dance, output an audio sequence containing rhythm, melody and emotional information based on the Jukebox encoder; S4. Based on the audio features of the audio sequence and the parameters of the five-dimensional quadratic kernel hybrid model represented by the movement sequence of the dancer dancing to the music, a generative architecture based on diffusion Transformer is constructed. The audio semantic features are injected into the denoising network through adaptive layer normalization, and the edge distribution and conditional probability mapping of the five-dimensional quadratic kernel parameters in the spatiotemporal field are learned. S5. Reconstruct the continuous dynamic symbolic distance field based on the generated five-dimensional quadratic kernel parameters, and calculate the loss function value during the training process based on the reconstruction result. S6. During the diffusion training process, input the loss function value, introduce inverse kinematics loss, rhythm alignment loss and style consistency loss to perform multi-objective collaborative optimization, and output the trained conditional generation network. S7, the music signal to be generated for the dance is segmented, and audio features are extracted using the method described in step S3; The corresponding audio features are input into the trained conditional generation network obtained in step S6 for inference to obtain a five-dimensional quadratic kernel parameter sequence; based on the five-dimensional quadratic kernel parameter sequence, a continuous dynamic symbolic distance field is reconstructed, and the final dance movement sequence is output through isosurface extraction and post-processing. 2.The dance move generation method of the five-dimensional quadratic kernel modeling according to claim 1, wherein, The human body SMPL parameterized template captured in a single frame is characterized as a symbolic distance field (SDF). 3.The dance move generation method of the five-dimensional quadratic kernel modeling according to claim 1, wherein, The music and the dancers' movement sequences in S1 are in one-to-one correspondence.

4. The dance motion generation method based on the five-dimensional quadratic kernel model as described in claim 1, characterized in that, Time stamping consecutive frame actions Ta Segmenting; for each segment action sequence modeling, the step S2 specifically comprises the following steps: Step S21, five-dimensional quadratic kernel modeling, transforms the symbol distance field data of multiple frames of this time series into... N Modeling is done using a 5x5 matrix, meaning the number of rows is equal to the 3D spatial resolution. n and time frame count T product N=n × T The data matrix has 5 columns; the data is randomly divided into... K Given a set of groups, calculate the mean vector for each group. Step S22, for N Data and K Calculate the mean square error for each mean vector, assign the data point to the group with the smallest mean square error, and iteratively calculate the mean; repeat this step 14 times to obtain the final grouping. K The dataset yields the mean vector, covariance matrix, and weights. ; Step S23: Reconstruct the formula using a five-dimensional quadratic kernel gate function. Reconstructing symbolic distance field data ,in Represents a quadratic kernel gate function, This represents the conditional mean of the second-order kernel.

5. The dance motion generation method based on the five-dimensional quadratic kernel model as described in claim 1, characterized in that, The movement is generated based on the music accompanying the dancer, and step S3 specifically includes the following steps: Step S31: Divide the input audio into fixed duration segments. Ta A segment of seconds, for each audio segment A i The Jukebox model with frozen weights was used as the audio encoder to extract latent features Z. i =Jukebox(A i )∈R Ta×Dz ,in T a The number of time steps for an audio segment. D z Jukebox feature dimensions; Step S32: Adapt high-dimensional features to the conditional space of the diffusion model through a learnable projection layer. Use a classifier-independent guidance strategy and set empty features with a 10% probability during training so that the model can learn both conditional and unconditional generation.

6. The dance motion generation method based on the five-dimensional quadratic kernel model as described in claim 1, characterized in that, Step S4 specifically includes the following steps: Step S41: Synchronously segment the dance movement data into segments of Ta seconds, with each segment represented by a five-dimensional quadratic kernel parameter Θ. i , i This is the segment number; for segments shorter than Ta seconds, the last frame action is used to fill the gap. Step S42: Construct a denoising network based on the DiT block of the diffusion transformer. f θ It utilizes its built-in adaptive layer normalization unit to inject audio features and time step information into the latent space, thereby achieving global correlation modeling and denoising prediction of human geometric sequences. Step S43, using audio-action clip pairs ( C i ,Θ i Using the training samples, optimize and simplify the diffusion loss. L simple =E [|| - f θ (Θ t , t , C i )|| 2 ].

7. The action generation method as described in claim 1, characterized in that, Step S5 specifically includes the following steps: Step S51: Construct a five-dimensional quadratic kernel gate function The formula for obtaining it is as follows: in It is the weight of the j-th array; Based on location x , y , z and time variables t The secondary nuclear edge distribution, in which The expression for the five-dimensional quadratic kernel is: f ( x , y , z , t , w ),set up ,So f ( x , y , z , t , w ) expressed as : Where is the mean vector of the j-th group of data; Let be the covariance matrix of the j-th data set; where, It is a variable in k sets of data x , y , z , t , w The random variables of the j-th data group, the mean, covariance matrix, and weights of the j-th data group have all been obtained in S2 and are known quantities. Then: Step S52: Construct the conditional mean function, and obtain the formula as follows: in, Let x, y, z, t, w be random variables of the j-th data set. The mean vector of the j-th data set obtained from S3 is: The mean vector of the position variables is The covariance matrix of the j-th data group obtained from S3 is: , It is a submatrix of the covariance matrix, where the first 4×4 dimension matrix is ​​the covariance matrix of the position variables: ; Step S53: Construct the reconstructed symbol distance field, and obtain the formula as follows: , in To reconstruct the symbolic distance field.

8. The dance motion generation method based on the five-dimensional quadratic kernel model as described in claim 1, characterized in that, Step S6 specifically includes the following steps: Step S61: Based on the human template symbol distance field reconstructed in step S5, extract the joint positions. The joint positions of the parameterized human model SMPL that originally captured the actions in the sequence. J ori ( i Calculate the mean square error loss ; Step S62: Extract the center of gravity velocity curve from the reconstructed action symbol distance field sequence. v ( t ), and detect its local extreme points as the rhythm of the movement. t p Music beat sequences extracted from audio features t b Perform dynamic time warping (DTW) alignment ; Step S63: Use a pre-trained dance movement style classifier C style Style scoring is performed on the reconstructed action sequences, and negative log-likelihood loss is calculated to enhance style preservation. ; Step S64: Combine the above loss with the diffusion-based loss to form the total loss function: L total = L simple + λ 1 L 1+ λ 2 L 2+ λ 3 L 3, in λ 1, λ 2, λ 3 represents the loss coefficient.

9. The dance motion generation method based on the five-dimensional quadratic kernel model as described in claim 1, characterized in that, Step S7 obtains the final dance movement sequence through isosurface extraction and post-processing, specifically including the following steps: Step S71: For each frame of the reconstructed symbolic distance field, for each edge, if the symbols of the two connected vertices are different, interpolate on this edge, calculate the position where the interpolation value is 0, and place the vertex. Step S72: For each frame of the reconstructed symbolic distance field, construct triangular patches between vertices.

10. A dance motion generation system based on a five-dimensional quadratic kernel model, characterized in that, include: The motion data acquisition and geometric representation construction module is used to represent the three-dimensional geometry of each frame of motion by collecting SMPL data of professional dancers dancing to music based on the inertial motion capture system and using the depth symbolic distance field algorithm. The five-dimensional quadratic kernel parameter extraction module is used to input the three-dimensional geometry, perform time-segmented modeling based on the spatiotemporally continuous symbolic distance field, and output the five-dimensional quadratic kernel hybrid model parameters as a motion representation. The audio feature extraction module is used to output an audio sequence containing rhythm, melody and emotional information based on the Jukebox encoder for the audio signal of the music played during the dancer's dance. The Conditional Generative Network Training Module is used to construct a generative architecture based on the diffusion Transformer, which is based on the audio features of the audio sequence and the parameters of the five-dimensional quadratic kernel hybrid model represented by the dance sequence of the dancer dancing to the music. The audio semantic features are injected into the denoising network through adaptive layer normalization, and the edge distribution and conditional probability mapping of the five-dimensional quadratic kernel parameters in the spatiotemporal field are learned. The differentiable reconstruction and loss calculation module is used to reconstruct the continuous dynamic symbolic distance field based on the generated five-dimensional quadratic kernel parameters, and to calculate the loss function value during the training process based on the reconstruction result. The multi-objective collaborative optimization module is used to input the loss function value during the diffusion training process, introduce inverse kinematics loss, rhythm alignment loss and style consistency loss to perform multi-objective collaborative optimization, and output the trained conditional generation network. The dance movement sequence generation module is used to input new music signals of the dance to be generated into the trained conditional generation network in segments for inference and generation, and output a five-dimensional quadratic kernel parameter sequence. After segmented reconstruction of the continuous dynamic symbolic distance field, the final dance movement sequence is output through isosurface extraction and post-processing.