Real-time dance movement generation system based on AI music rhythm
Through dynamic optimal transmission and deep learning technology, the synchronization relationship between dance movements and music rhythm is optimized, and the problem of insufficient synchronization between dance movements and music rhythm in the existing technology is solved, and high-precision alignment and generation of movement diversity is achieved.
Patent Information
- Application Number
- CN202510260788.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, dance movements and music rhythms are not synchronized enough in complex dynamic scenarios, and the diversity and fluency of movement generation are poor, making it difficult to meet users' needs for personalized and diversified dances.
Using a technical solution combining dynamic optimal transmission and deep learning, the synchronization relationship between dance movements and music rhythm is optimized through the music rhythm extraction module, dance movement generation module, dynamic synchronization module and action adjustment module, and the synchronization relationship between dance movements and music is achieved with high-precision alignment of movements and music.
It improves the synchronization accuracy and fluency of dance movements and music rhythms, enhances the diversity and expressiveness of the movements, and solves the problem of action and music mismatch in complex rhythms and dynamic changing scenarios.
Smart Images

Figure CN120199210A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence and multimedia interaction, and specifically to a real-time dance movement generation system based on AI music rhythm. Background Art
[0002] With the rapid development of artificial intelligence technology, the integration of music and dance is gradually becoming a research hotspot in the fields of virtual entertainment, stage performance, and education. By using artificial intelligence algorithms to generate dance movements that match the music rhythm, not only can the complexity of traditional manual choreography be reduced, but also technical support can be provided for virtual idols, intelligent performance systems, and immersive entertainment experiences. Currently, AI-based music and dance generation systems are gradually realizing the transformation from simple action imitation to highly personalized generation.
[0003] In the prior art, fixed action templates or rule-based methods are usually adopted to generate dance movements that match the music rhythm. Some methods attempt to synchronize the music rhythm and dance movements through time-sliding windows or dynamic time warping (DTW). However, these methods often rely on preset action templates and lack the ability to adapt to the diversity and complexity of music. At the same time, the flexibility and expressiveness of action generation are limited, making it difficult to meet the needs of users for personalized and diverse dances. In addition, in the process of synchronizing action generation with the music rhythm, the prior art usually cannot well solve complex rhythm and dynamic change scenarios, resulting in a mismatch problem between the generated dance movements and the music rhythm.
[0004] Aiming at the problem of insufficient synchronization between dance movements and music rhythm in the prior art, the present invention fundamentally optimizes the accuracy and fluency of action generation by introducing a technical solution combining dynamic optimal transport and deep learning, enabling the dance movements to naturally and real-time match complex music rhythms, and providing an efficient and flexible solution for various application scenarios. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a real-time dance movement generation system based on AI music rhythm, which solves the problems of insufficient synchronization between dance movements and music rhythm in complex dynamic scenarios, and poor diversity and fluency of action generation in the prior art.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A real-time dance movement generation system based on AI music rhythm, comprising:
[0007] A music rhythm extraction module, configured to extract the beat time and intensity features of the input music and generate a music rhythm distribution;
[0008] The dance movement generation module generates an initial dance movement distribution according to the extracted music rhythm distribution, and is used to generate the initial dance movement distribution based on a pre-trained deep learning model;
[0009] The dynamic synchronization module is used to calculate the mapping relationship between the music rhythm distribution and the dance movement distribution through the dynamic optimal transport method;
[0010] The movement adjustment module is used to adjust the time series of the initial dance movements according to the mapping relationship and perform smoothing processing;
[0011] The movement display module is used to map the adjusted dance movement sequence to an avatar or a 3D skeleton model and play it in real-time synchronization with the music rhythm.
[0012] Preferably, the music rhythm distribution extracted by the music rhythm extraction module includes multiple beat time points and corresponding beat intensities, and each beat intensity is obtained from the music signal through short-time Fourier transform and a dynamic beat tracker.
[0013] Preferably, the dance movement generation module uses a pre-trained deep learning model to generate an initial dance movement distribution. The dance movement distribution includes multiple movement occurrence time points and corresponding movement intensities, and the generated movement time points are initially aligned with the rhythm characteristics of the input music.
[0014] Preferably, the dynamic synchronization module calculates the mapping relationship between the music rhythm distribution and the dance movement distribution through the dynamic optimal transport method. The mapping relationship is realized by minimizing the transport cost function, and the transport cost function represents the squared error of the offset between the music rhythm time points and the dance movement time points.
[0015] Preferably, the transport density constructed by the dynamic synchronization module satisfies the following constraints:
[0016] The sum of the edges of the transport density is respectively consistent with the sum of the edges of the music rhythm distribution and the dance movement distribution;
[0017] The transport density is always non-negative.
[0018] Preferably, the movement adjustment module adjusts the dance movement time series through the transport density, and the adjusted time points satisfy the physical consistency of movement generation through smoothing processing. The smoothing processing is calculated according to the time inertia factor of the movement.
[0019] Preferably, the dynamic synchronization module iteratively solves the optimal transport density through the Sinkhorn-Knopp algorithm. The algorithm takes the initial distribution parameters as input and finally obtains the mapping probability from the music rhythm time points to the dance movement time points through multiple rounds of normalization calculations.
[0020] Preferably, the action adjustment module superimposes the feature intensity of the music rhythm time point on the corresponding dance action time point through a mapping relationship.
[0021] Preferably, the system uses a sliding window mechanism to segment the input music. Each segmented music corresponds to independently calculated music rhythm distribution and dance action distribution, and the transmission density is quickly updated through incremental optimization between sliding windows.
[0022] Preferably, the action display module realizes the output of dance action animation synchronized with the music rhythm by mapping the adjusted dance action time series to the joint trajectory of the virtual image or the 3D bone model.
[0023] The present invention provides a real-time dance action generation system based on AI music rhythm.
[0024] It has the following beneficial effects:
[0025] 1. The present invention adopts a technical solution that combines dynamic optimal transport and variational optimization. By optimizing the synchronization relationship between the music rhythm distribution and the dance action distribution, it achieves the technical effect of highly accurate alignment of actions and music rhythm. Compared with the technical solutions using simple time sliding windows or dynamic time warping methods in the prior art, it solves the problem of action and music mismatch in complex music rhythm change scenarios, making the dance actions more natural and fluent.
[0026] 2. The present invention realizes the preliminary generation of the dance action distribution by introducing a dance action generation model based on multi-modal features and using deep learning technology to model the time series characteristics and style information of the music rhythm. Compared with the prior art solutions that only rely on fixed action templates, it solves the deficiencies of single action style and inability to adapt to multiple music types, and significantly improves the diversity and expressiveness of the generated actions.
[0027] 3. The present invention optimizes the dance action time points through the action adjustment module in combination with the transmission probability and introduces a dynamic smoothing processing mechanism, achieving the technical effects of continuity and physical consistency of action generation. Compared with the solutions that simply rely on time point fine-tuning in the prior art, the present invention effectively avoids the problems of jumping or incoherence in the action time series, ensuring the authenticity and fluency of the generated actions.
[0028] 4. The present invention adopts an action display solution that maps the optimized dance actions to a virtual image or a 3D bone model, and combines various interpolation algorithms to generate smooth joint trajectories, achieving the technical effect of real-time visualization of dance actions. Compared with the problems of high action display delay and poor real-time performance in the prior art, the present invention solves the deficiency of out-of-sync between action generation and music playback, further enhancing the visual performance effect and user experience. Description of the Drawings
[0029] Figure 1 This is a schematic diagram of the system process of the present invention. Specific embodiments
[0030] Next, in conjunction with the accompanying drawings of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0031] Please refer to the attached Figure 1 , the embodiment of the present invention provides a real-time dance movement generation system based on AI music rhythm, including:
[0032] The input of the music rhythm extraction module is the audio signal received in real time, and this signal may be transmitted in the form of a streaming media or a pre-loaded music file. To ensure the accuracy and robustness of rhythm extraction, the audio signal first undergoes preprocessing steps, including denoising and framing operations.
[0033] In this embodiment, the audio signal is segmented into fixed-time frames of length N, and a window function (such as a Hanning window) is applied to each frame to reduce the influence of spectral leakage. Generally, the typical value range of the frame length N is from 1024 to 4096 sampling points, and the frame shift is 50% of the frame length. Through these settings, the module can balance the time resolution and the frequency resolution.
[0034] After the audio signal undergoes framing and windowing processing, the short-time Fourier transform (STFT) is used to extract spectral features. Specifically, the short-time Fourier transform formula is as follows:
[0035]
[0036] Where:
[0037] X(t,f) is the complex spectrum at time t and frequency f;
[0038] x[n] is the input audio signal in the time domain;
[0039] w[n] is the window function;
[0040] N is the frame length;
[0041] j is the imaginary unit.
[0042] In some embodiments, to improve the calculation efficiency, the fast Fourier transform (FFT) can be selected to implement the calculation of the above formula.
[0043] Specifically, the spectral data after Fourier transform is further processed to extract beat information. One possible implementation is to use a dynamic beat tracker to detect the periodic pulses in the audio signal. The dynamic beat tracker outputs a set of beat time points t i and the corresponding beat intensities s i . The beat intensity can be obtained by calculating the energy-weighted sum of the spectral amplitudes:
[0044]
[0045] where:
[0046] f min and f max are the lowest and highest frequency boundaries respectively, usually in the range of 20 Hz to 4000 Hz;
[0047] |X(t i ,f)| 2 represents the squared spectral amplitude, i.e., the energy of the signal at a specific frequency.
[0048] In some embodiments, the beat intensity s i can be further normalized to ensure the comparability of beat amplitudes for different music styles.
[0049] As an option, the module can also use wavelet transform instead of Fourier transform to extract the local time-frequency features of music. The formula for wavelet transform is:
[0050]
[0051] where:
[0052] W(t,s) is the wavelet coefficient at time t and scale s;
[0053] ψ is the mother wavelet function;
[0054] s is the scale parameter, which is used to control the time and frequency resolution.
[0055] In some embodiments, wavelet transform is more suitable for capturing the dynamic rhythm changes of non-regular music. The extracted beat time points t i and beat intensities s i are combined into a music rhythm distribution P(t), and its mathematical representation is as follows:
[0056]
[0057] where:
[0058] δ(t-t i) is the Dirac function, used to represent the position of the beat time point;
[0059] n is the total number of beats.
[0060] The music rhythm distribution P(t) will be transmitted to the dance movement generation module in real time, providing rhythm information for the generation of the initial dance movement. In the dynamic synchronization module, P(t) will also be used as an input for mapping calculation with the dance movement distribution Q(t′).
[0061] In a possible implementation, the module supports a sliding window mechanism to segment the input audio stream. Each segment is usually 2 seconds long, and there is a 50% overlap between windows. The design of the sliding window can ensure a balance between the real-time performance and data integrity of the system.
[0062] After receiving the music rhythm distribution P(t) output by the music rhythm extraction module, the dance movement generation module uses a pre-trained deep learning model to generate the initial dance movement distribution Q(t′). To ensure the accuracy of the generated movements, the module uses a deep neural network based on temporal modeling, such as the Transformer architecture or the recurrent neural network (RNN). Generally, these models have been trained using a large music-dance alignment dataset in the pre-training stage to learn the implicit mapping relationship between music features and dance movements.
[0063] Specifically, the music rhythm distribution P(t) includes the beat time point t i and the intensity s i , which are used as model inputs. The model encodes P(t) through a feature extraction layer, converting the time and intensity features into an implicit vector representation h(t):
[0064] h(t) = Encoder(P(t))
[0065] where:
[0066] h(t) is the implicit feature representation at time t;
[0067] Encoder is a feature extraction function, usually composed of a group of convolutional neural networks (CNNs) or multi-head self-attention mechanisms.
[0068] In some embodiments, the feature extraction layer can adopt a multi-layer Transformer network to model the time series characteristics of the music rhythm. Through the multi-head attention mechanism, the model can capture the global dependencies between different beats, making the generated movements smoother.
[0069] In the generation stage, the decoder part of the model generates the initial time point and intensity of the dance movement based on the encoded music features. The mathematical form of the movement distribution Q(t′) is:
[0070]
[0071] Wherein:
[0072] t j ′ is the time point of the j-th dance movement, generated by the decoder;
[0073] a j is the corresponding movement intensity, usually obtained by mapping the amplitude of the music features.
[0074] As an option, the model can also introduce physical constraints during the generation stage to ensure that the generated movements conform to the laws of human motion. For example, when generating the time point t j ′ of each movement, consider the interval Δt between the time points of the previous and subsequent movements:
[0075] t j ′ = t j-1 ′ + Δt
[0076] Wherein:
[0077] Δt is predicted by the model and satisfies the time interval distribution of human movements;
[0078] t j-1 ′ is the time point of the previous movement.
[0079] To further improve the physical consistency and style matching of the movements, the model adds a movement style control parameter z during the decoding stage. z represents the latent feature vector of a specific dance style, and generates stylized dance movements through the fusion with the music feature h(t):
[0080] h′(t) = f(h(t), z)
[0081] Wherein:
[0082] h′(t) is the fused feature vector;
[0083] f is the feature fusion function, usually implemented by a fully connected layer or a multi-head attention mechanism.
[0084] In some embodiments, to ensure the initial alignment of the movements generated by the model with the music rhythm, the module performs post-processing on the generated movement distribution Q(t′). The post-processing steps include the correction of the movement time points and the intensity normalization. Specifically, the module fine-tunes the generated movement time point t j ′ to reduce the initial deviation from the music beat time point t i :
[0085] t j ′ = t j ′ + α(ti -t j ′)
[0086] Wherein:
[0087] α is an adjustment coefficient used to control the amplitude of the time point correction;
[0088] t i and t j ′ are the nearest music beat time points.
[0089] During the intensity normalization process, the module adjusts the generated action intensity a j to make it conform to the overall distribution range of the music rhythm intensity:
[0090]
[0091] Wherein:
[0092] max(a j ) and max(s i ) are the maximum values of the generated action intensity and the music rhythm intensity respectively.
[0093] The generated initial dance action distribution Q(t′) will be output in a standardized format, including action time points, intensities, and related implicit features, facilitating subsequent processing by the dynamic synchronization module.
[0094] In a possible implementation, the module also supports user-defined style input. By adjusting the style parameter z, the user can select different dance styles (such as ballet, hip-hop, or modern dance), thereby generating more diverse actions.
[0095] In this embodiment, the specific implementation manner of the dynamic synchronization module is as follows:
[0096] The dynamic synchronization module first constructs the transmission relationship between the music rhythm distribution P(t) and the dance action distribution Q(t′). The music rhythm distribution P(t) can be expressed as:
[0097]
[0098] Wherein:
[0099] t i is the music beat time point;
[0100] s i is the corresponding beat intensity;
[0101] n is the total number of music beats.
[0102] The dance action distribution Q(t′) is expressed as:.
[0103]
[0104] Wherein:
[0105] t′ j is the time point of the dance movement;
[0106] a j is the corresponding movement intensity;
[0107] m is the total number of dance movements.
[0108] To achieve the synchronization of music rhythm and dance movements, the dynamic synchronization module defines the transmission density π(t, t′), which represents the mapping probability from the beat time point t in P(t) to the movement time point t′ in Q(t′).
[0109] Generally, the module optimizes the transmission density π(t, t′) through the optimal transport model to minimize the transmission cost between the music time point and the movement time point. The optimization objective function is:
[0110]
[0111] Wherein:
[0112] c(t, t′) = ||t - t′|| 2 is the transmission cost function, which is used to quantify the offset between the music time point t and the movement time point t′;
[0113] λ is the smoothness weight, which is used to adjust the continuity of the transmission density in the time dimension;
[0114] represents the time gradient of the transmission density.
[0115] As an option, the module adds the following constraint conditions during the optimization process:
[0116] Marginal distribution consistency:
[0117] ∫ T′ π(t, t′)dt′ = Pt, ∫ T π(t, t′)dt = Q(t′)
[0118] This constraint ensures that the marginal distributions of the transmission density are consistent with the music rhythm distribution and the dance movement distribution respectively.
[0119] Non-negativity constraint:
[0120] π(t, t′) ≥ 0
[0121] This constraint ensures that each term of the transmission density is non-negative.
[0122] In some embodiments, to improve the computational efficiency, the module supports a sliding window mechanism to segment the input rhythm distribution and action distribution. The length of each window is W, and there is a 50% overlap between windows. The boundary conditions of the sliding window are processed by incremental updates of the transmission density to avoid global recalculation.
[0123] In a possible implementation, the module also supports dynamically adjusting the value of the smoothness weight λ. For fast-paced music, λ can be appropriately reduced to enhance the responsiveness of the action to music changes; for slow-paced music, the value of λ can be increased to improve the smoothness and fluency of the action.
[0124] The action adjustment module receives the mapping relationship between the music rhythm and the dance action output by the dynamic synchronization module, and is used to adjust the time series and intensity of the initial dance action. Through optimized adjustment, this module enables the dance action to achieve a higher degree of matching with the music rhythm, while ensuring the fluency and physical consistency of action generation. Generally, the action adjustment module corrects the action time points through the transmission probability, and combines smoothing processing to achieve continuity optimization, and finally generates dance actions synchronized with the music.
[0125] First, the action adjustment module receives the transmission density output by the dynamic synchronization module, which describes the mapping probability between the music beat time points and the dance action time points. In addition, it also receives the initial dance action distribution, which includes multiple dance action time points and their corresponding intensities.
[0126] Generally, the action adjustment module corrects the initial dance action time points through the transmission density. Specifically, the module calculates the weighted average of each dance action time point and all music beat time points, and the weight is determined by the corresponding transmission probability. The corrected dance action time points are closer to the music beat time points, thus achieving the synchronization of the action and the music rhythm.
[0127] As an option, to enhance the fluency of the corrected action time series, the module performs smoothing processing on the adjusted time points. Time smoothing uses a weighted average method to make the relationship between the current time point and the previous time point more natural. By controlling the size of the smoothing factor, the needs of different styles of dance can be adapted. For example, for fast-paced dance, the smoothing factor can be reduced to enhance the flexibility of the action; for slow-paced dance, the smoothing factor can be increased to improve the smoothness of the action.
[0128] When adjusting the action intensity, the module redistributes the intensity of the dance action according to the intensity characteristics of the music beat. By superimposing the intensity distribution of the music beat on the intensity distribution of the dance action, the expressiveness of the action can be consistent with the rhythm changes of the music. For example, if the intensity of a certain music beat is high, the intensity of the corresponding dance action will be amplified, so as to better fit the emotional expression of the music.
[0129] In addition, the module also performs a physical consistency check on the adjustment results. For example, the time interval between consecutive actions is restricted within a certain range to avoid overly large or small time jumps between actions, thereby ensuring that the action generation conforms to the laws of human motion. Generally, the range of this time interval can be set according to the actual application scenario, usually smaller in fast-paced dances and larger in slow-paced dances.
[0130] In some embodiments, the module supports users to customize action style parameters for further adjusting the expression form of dance actions. For example, by introducing a style control factor, characteristics such as the undulation intensity and transition smoothness of the actions can be changed to generate more personalized dance actions.
[0131] After the above adjustments, the action adjustment module outputs an optimized distribution of dance actions, including the updated action time points and intensity values. The optimized result is transmitted to the action display module for generating animations of virtual avatars or 3D skeleton models. Through the action adjustment module, the time points and intensities of dance actions can be highly consistent with the music rhythm, while the continuity and physical consistency of the actions are ensured.
[0132] The action display module is a terminal component for mapping the adjusted dance action distribution into a virtual avatar or 3D skeleton model to achieve the visual expression of actions. This module directly receives the output of the action adjustment module, including the optimized dance action time points and action intensity data. Generally, the action display module needs to convert these time series data into dynamic joint trajectories or animation frame sequences of virtual characters and ensure the real-time synchronization of actions with the music rhythm. Through this module, users can intuitively see the matching effect of the generated dance actions and music.
[0133] In this embodiment, the specific implementation method of the action display module is as follows:
[0134] The action display module first parses the optimized data output by the action adjustment module, including the adjusted dance action time points and intensity values. These data are organized as a time series, denoted as t′ j , a′ j ), where (t′ j , a′ j ) is the occurrence time point of the dance action, and a′ j is the intensity of this action.
[0135] Generally, the module corresponds the action time points to the joint movements of the virtual character through a joint drive algorithm. Specifically, the optimized time point t′ j is mapped to the specific joint positions of the virtual character, and according to the action intensity value a′ jAdjust the movement range of joints. For example, if the action intensity is high at a certain time point, the joint movement range of the virtual character will increase accordingly to enhance the expressiveness of the action.
[0136] As an option, the module also supports multiple joint movement interpolation methods for generating smooth joint trajectories between time points. Specifically, the module can adopt the spline interpolation method to generate continuous joint movement trajectories based on discrete action time points. The basic form of the interpolation formula is as follows:
[0137]
[0138] Where:
[0139] x(t) is the position of the joint at time t;
[0140] c k is the interpolation coefficient, calculated from discrete time points and joint positions;
[0141] B k (t) is the spline basis function.
[0142] Through interpolation calculation, the module can generate continuous joint trajectories for the virtual character, making the action more natural.
[0143] Specifically, the module defines the initial position and movement range of joints according to the skeletal model of the virtual character. For example, a standard 3D skeletal model usually includes main joint points such as the head, torso, and limbs, and the position of each joint point is represented by three-dimensional coordinates (x, y, z). The module converts the time point t′ j and the corresponding intensity a′ j into three-dimensional movement data of the joints and applies it to the animation driving of the skeletal model.
[0144] In some embodiments, to enhance the visual performance effect of the action, the module adjusts the auxiliary parameters of the virtual character in combination with the action intensity value. For example, it dynamically changes the expression, limb posture, etc. of the virtual character according to the action intensity value, making the dance action more matching with the music emotion. In addition, the module can further highlight the visual impact of the dance action by adjusting the lighting or camera angle in the virtual scene.
[0145] In a possible implementation, the module supports multiple virtual character types, including anthropomorphic 3D models, cartoon-style characters, or characters with user-defined appearances. Users can select different character types according to specific needs, and the module will automatically map the dance actions to the corresponding character skeletal models. In addition, the module also supports the multi-character dance display function, and realizes the synchronous display of multi-person dance by allocating the optimized action data to multiple virtual characters.
[0146] To ensure the real-time display of actions, the module adopts a time-stepping mechanism to render the action sequence frame by frame. Within each time step, the module extracts the corresponding joint positions according to the current time point t and sends them to the 3D rendering engine for display. For example, when the current time t is between t′ j+1 and t″, the module calculates the joint positions through linear interpolation to generate continuous animation frames.
[0147] The module also supports dynamic synchronization with the music playback module. Generally, the module will detect the music playback time in real time and extract the corresponding action data according to the music timestamp for rendering. This synchronization mechanism can ensure that the action display is highly consistent with the music playback, avoiding action distortion or rhythm misalignment caused by time offset.
[0148] The action display module also supports additional personalized functions. For example, users can choose to add costumes, props, or scene elements to the virtual character, and these additional elements will be dynamically adjusted along with the character's dance movements, thus enhancing the overall display effect. In addition, the module also supports the action recording function, saving the generated dance action data in a standard animation file format (such as FBX or BVH) for subsequent editing or secondary application.
[0149] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A real-time dance movement generation system based on AI music rhythm, characterized in that: include: Music rhythm extraction module, used to extract the beat time and intensity characteristics of input music and generate music rhythm distribution; A dance movement generation module generates an initial dance movement distribution according to the extracted music rhythm distribution, and is used to generate an initial dance movement distribution based on a pre-trained deep learning model; A dynamic synchronization module, used to calculate the mapping relationship between music rhythm distribution and dance movement distribution through a dynamic optimal transmission method; A motion adjustment module is used to adjust the time series of the initial dance motion according to the mapping relationship and perform smoothing; The action display module is used to map the adjusted dance action sequence to the virtual image or 3D skeleton model, and play it in real time in synchronization with the music rhythm.
2. The real-time generation system of dance movements based on AI music rhythm according to claim 1 is characterized in that: The music rhythm distribution extracted by the music rhythm extraction module includes multiple beat time points and corresponding beat intensities, and each beat intensity is obtained from the music signal through short-time Fourier transform and dynamic beat tracker.
3. The real-time generation system of dance movements based on AI music rhythm according to claim 1 is characterized in that: The dance movement generation module generates an initial dance movement distribution using a pre-trained deep learning model. The dance movement distribution includes multiple movement occurrence time points and corresponding movement intensities. The generated movement time points are preliminarily aligned with the rhythm features of the input music.
4. The real-time generation system of dance movements based on AI music rhythm according to claim 1 is characterized in that: The dynamic synchronization module calculates the mapping relationship between the music rhythm distribution and the dance movement distribution through a dynamic optimal transmission method. The mapping relationship is achieved by minimizing a transmission cost function, which represents the offset square error between the music rhythm time point and the dance movement time point.
5. The real-time generation system of dance movements based on AI music rhythm according to claim 1 is characterized in that: The transmission density of the dynamic synchronization module structure meets the following constraints: The marginal sum of the transmission density in the music rhythm distribution and the marginal sum of the dance movement distribution are consistent; The transmission density is always non-negative.
6. The real-time generation system of dance movements based on AI music rhythm according to claim 1 is characterized in that: The action adjustment module adjusts the dance action time sequence by transmission density, and the adjusted time points meet the physical consistency of action generation through smoothing, and the smoothing is calculated and generated according to the time inertia factor of the action.
7. The real-time generation system of dance movements based on AI music rhythm according to claim 1 is characterized in that: The dynamic synchronization module iteratively solves the optimal transmission density through the Sinkhorn-Knopp algorithm. The algorithm takes the initial distribution parameters as input and finally obtains the mapping probability of the music rhythm time point to the dance movement time point through multiple rounds of normalization calculation.
8. The real-time generation system of dance movements based on AI music rhythm according to claim 1 is characterized in that: The action adjustment module superimposes the characteristic intensity of the music rhythm time point to the corresponding dance action time point through a mapping relationship.
9. The real-time generation system of dance movements based on AI music rhythm according to claim 1 is characterized in that: The system uses a sliding window mechanism to process the input music in segments, each segment of the music corresponds to an independently calculated music rhythm distribution and dance movement distribution, and the transmission density is quickly updated through incremental optimization between sliding windows.
10. The real-time generation system of dance movements based on AI music rhythm according to claim 1 is characterized in that: The action display module outputs dance action animation synchronized with the music rhythm by mapping the adjusted dance action time sequence to the joint trajectory or 3D skeleton model of the virtual image.
Citation Information
Cited By
Dance video generation method, device and equipment and readable storage medium
CN121037644A
SOR optimization system and method for tourism dance immersion experience
CN121300637A
Digital human dance motion generation method and system based on music feature matching
CN121353483A
A digital human dance action generation method and system based on music feature matching
CN121353483B