Artificial intelligence-based dance data generation method and device, equipment and medium

By acquiring music data and dance genre tags, processing music features using target professional expert modules and general expert modules, and combining them with a dance codebook prediction model, dance movements that conform to a specific genre style and have a reasonable basic pattern are generated. This solves the problem of dance movement generation not conforming to genre style in existing technologies, and improves generation quality and flexibility.

CN121924329APending Publication Date: 2026-04-24PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-19
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies lack explicit modeling of dance styles when generating dance movements, resulting in generated dance movements that deviate significantly from the typical characteristics of the target style, have stiff visual effects, fail to meet users' customization needs for specific dance styles, and have low flexibility and quality when handling complex musical features.

Method used

Using an AI-based approach, music data and dance genre labels are acquired. The music features are processed using a target expert module and a general expert module. Combined with a dance codebook prediction model, dance movements that conform to a specific genre style and have a reasonable basic pattern are generated.

Benefits of technology

The generated dance moves not only conform to the style requirements of a specific genre but also have a reasonable basic pattern, improving the quality and flexibility of dance move generation and meeting users' customization needs for specific styles of dance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121924329A_ABST
    Figure CN121924329A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to an artificial intelligence-based dance data generation method and device, equipment and a medium, and the method comprises the steps: carrying out the feature extraction of input music data, and obtaining a music feature; determining a target professional expert module based on the input dance genre label; performing feature processing on the music features and the dance genre labels based on a target professional expert module to obtain professional features, and performing feature processing on the music features based on a general expert module to obtain general features; performing feature fusion on the professional features and the general features to obtain fusion features; carrying out dance codebook prediction processing on the fusion features, the music features and dance genre labels based on a dance codebook prediction model to obtain a dance codebook sequence; and carrying out dance reconstruction on the dance codebook sequence to obtain a target dance action and outputting the target dance action. The method can be applied to dance data generation scenes related to insurance advertisement promotion in the field of financial science and technology, and the generation quality and the generation flexibility of dance actions are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and can be applied to the financial technology field, particularly to a method, apparatus, computer equipment, and storage medium for generating dance data based on artificial intelligence. Background Technology

[0002] In the field of artificial intelligence, music-driven 3D dance generation technology has received widespread attention in recent years as a key support for virtual reality, digital entertainment, creative content creation, and insurance advertising. Its core objective is to automatically generate matching 3D human motion sequences by analyzing music features, thereby reducing manual choreography costs and improving creative efficiency. However, existing technologies generally employ a single-stage generation method, directly mapping music features to human motion parameters. While this method can generate seemingly coherent dance movements, it has significant drawbacks: due to the lack of explicit modeling of dance styles, the generated dance movements often deviate severely from the typical characteristics of the target style, resulting in stiff visual effects, insufficient artistic expression, and difficulty in meeting users' customized needs for specific dance styles. Consequently, the flexibility and quality of dance movement generation are relatively low.

[0003] For example, in the context of insurance companies creating creative advertisements, if a virtual character is needed to perform a modern dance advertisement themed "Active Aging" to music, existing technology may fail to accurately match the fluid body language and musical rhythm of modern dance. This could lead to a disconnect between the character's movements and the advertising theme (such as stiff movements or stylistic mismatch), thereby weakening the advertisement's appeal and communicative effect. Furthermore, single-stage methods are less adaptable to complex musical features, and the flexibility and quality of generated movements further decline when dealing with music that blends multiple styles or has atypical rhythms.

[0004] Therefore, there is an urgent need for an intelligent method for generating dance movements that can deeply integrate musical characteristics and dance styles, in order to improve the quality and adaptability of generated dance movements and expand their commercial application value in fields such as insurance advertising, film and animation. Summary of the Invention

[0005] The purpose of this application is to propose a method, apparatus, computer device, and storage medium for generating dance data based on artificial intelligence, so as to solve the technical problems of low flexibility and low quality in the existing generation and processing of dance movements.

[0006] Firstly, an artificial intelligence-based method for generating dance data is provided, including: Obtain the input music data and dance genre tags; Based on a preset feature extraction strategy, feature extraction is performed on the music data to obtain corresponding music features; Based on the dance style tags, the corresponding target professional expert module is determined; Based on the target professional expert module, feature processing is performed on the music features and the dance style tags to obtain corresponding professional features, and based on the preset general expert module, feature processing is performed on the music features to obtain corresponding general features; The professional features and the general features are fused to obtain the corresponding fused features; Based on a preset dance codebook prediction model, the fusion features, the music features, and the dance genre tags are subjected to dance codebook prediction processing to obtain the corresponding dance codebook sequence. The dance codebook sequence is reconstructed to obtain the corresponding target dance movement, and the target dance movement is output.

[0007] Secondly, an artificial intelligence-based dance data generation device is provided, comprising: The acquisition module is used to acquire the input music data and dance genre tags; The extraction module is used to extract features from the music data based on a preset feature extraction strategy to obtain corresponding music features; The determination module is used to determine the corresponding target professional expert module based on the dance style label; The first processing module is used to perform feature processing on the music features and the dance genre tags based on the target professional expert module to obtain corresponding professional features, and to perform feature processing on the music features based on a preset general expert module to obtain corresponding general features. The fusion module is used to perform feature fusion on the professional features and the general features to obtain the corresponding fused features; The prediction module is used to perform dance codebook prediction processing on the fusion features, the music features and the dance genre tags based on a preset dance codebook prediction model to obtain the corresponding dance codebook sequence. The second processing module is used to perform dance reconstruction on the dance codebook sequence to obtain the corresponding target dance movement, and to output the target dance movement.

[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based dance data generation method.

[0009] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned artificial intelligence-based dance data generation method.

[0010] In the above-mentioned scheme implemented by the AI-based dance data generation method, device, computer equipment, and storage medium, the input music data and dance genre tags are first acquired; then, features are extracted from the music data based on a preset feature extraction strategy to obtain corresponding music features; and a corresponding target professional expert module is determined based on the dance genre tags; subsequently, the music features and dance genre tags are processed by the target professional expert module to obtain corresponding professional features, and the music features are processed by a preset general expert module to obtain corresponding general features; subsequently, the professional features and general features are fused to obtain corresponding fused features; further, the fused features, music features, and dance genre tags are processed by a preset dance codebook prediction model to obtain a corresponding dance codebook sequence; finally, the dance codebook sequence is reconstructed to obtain the corresponding target dance movement, and the target dance movement is output. Based on the above automated dance generation process, this application uses a combination of a target professional expert module, a general expert module, and a dance codebook prediction model to process the input music data and dance genre tags for dance generation. This enables the generated target dance movements to not only meet the style requirements of a specific genre but also have a reasonable basic pattern, ensuring the quality of the generated target dance movements and improving the flexibility of target dance movement generation. Attached Figure Description

[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the AI-based dance data generation method according to this application; Figure 3 This is a schematic diagram of the structure of an embodiment of the AI-based dance data generation device according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0016] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0017] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0019] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0020] It should be noted that the AI-based dance data generation method provided in this application is generally executed by a server / terminal device, and correspondingly, the AI-based dance data generation device is generally located in the server / terminal device.

[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0022] Continue to refer to Figure 2 This document illustrates a flowchart of an embodiment of the AI-based dance data generation method according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different requirements. The AI-based dance data generation method provided in this application can be applied to any scenario requiring dance generation, and therefore can be applied to products in these scenarios, such as dance generation scenarios in the financial insurance field. The AI-based dance data generation method includes the following steps: Step S201: Obtain the input music data and dance genre tags.

[0023] In this embodiment, the AI-based dance data generation method runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire input music data and dance genre tags via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future-developed wireless connection methods. The implementing entity of this application is specifically a data generation system, or a dance generation system, which can be simply referred to as the system. The aforementioned music data is pre-prepared music data used to generate the corresponding dance, for example, music data can be obtained from a music database, an online music platform, or pre-recorded audio files. The aforementioned dance genre tags are dance genre tags specified according to actual needs. This application can be applied to business scenarios in the financial insurance field where insurance companies use dance generation technology to create creative advertisements. The selection of music data and dance genre tags must closely revolve around the core value of the insurance product (such as protection, care, vitality, stability, etc.) and the preferences of the target audience. The following are two specific examples: Example 1: Health Insurance Advertisement – ​​Vitality and Protection. Objective: To convey the concept that "health protection makes life more free" and attract young people to health insurance products. Music Data: Style: Upbeat electronic pop, fast-paced (BPM 120-140), with positive melodies and drum beats. Content: Lyrics can include keywords such as "vitality," "protection," and "worry-free" (e.g., adapted from an existing song or custom lyrics), emphasizing the sense of security brought by health protection. Example Excerpt: A 20-second electronic music track. The intro uses synthesizer effects to create a technological feel, the verses incorporate upbeat guitar melodies, and the chorus emphasizes drum beats and vocal harmonies. Dance Genre Tags: Choose: Urban Dance or Jazz Funk. Reason: Urban Dance features fluid and rhythmic movements, reflecting the vitality of modern life; Jazz Funk combines the elegance of jazz with the rhythm of funk, suitable for expressing the "free and unrestrained" state brought by health insurance. Action design: The virtual character stretches, jumps, and spins in rhythm with the music, and ends with a confident pose (such as making an "OK" sign with both hands) to display the name of the insurance product.

[0024] Example 2: Pension Insurance Advertisement – ​​Stability and Care. Objective: To convey the concept of "sound planning for a more secure old age," attracting middle-aged consumers to pay attention to pension insurance products. Music Data: Style: Warm piano and string concerto, soothing tempo (BPM 60-80), gentle and emotional melody. Content: Instrumental music without lyrics, conveying the emotions of "companionship" and "peace of mind" through the rise and fall of the melody, or adding natural sound effects (such as birdsong and flowing water) to enhance the healing effect. Example Excerpt: A 30-second piano solo opens, gradually adding string harmonies, with gentle drumbeats simulating a heartbeat rhythm in the middle section, ending with a fading melody. Dance Genre Tags: Choose: Contemporary Dance or Lyrical Jazz. Reasons: Contemporary Dance features gentle and expressive movements, suitable for conveying the concept of "peaceful and quiet retirement"; Lyrical Jazz combines the grace of jazz with the narrative of modern dance, expressing "care" and "protection" through body language. Movement Design: The virtual character simulates a "companionship" scene with slow stretching, hugging, and lifting movements, ending with hands clasped in front of the chest displaying the name of the insurance product, symbolizing "peace of mind in entrusting your care."

[0025] Step S202: Based on a preset feature extraction strategy, feature extraction is performed on the music data to obtain the corresponding music features.

[0026] In this embodiment, the specific implementation process of extracting features from the music data based on the preset feature extraction strategy to obtain the corresponding music features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0027] Step S203: Determine the corresponding target professional expert module based on the dance style tag.

[0028] In this embodiment, the input dance genre tag is parsed to obtain the corresponding genre tag content, such as "Jazz". Then, by applying a hard routing mechanism, the corresponding target expert module is found based on the genre tag content. For example, for the "Jazz" tag, a dedicated expert module trained specifically for the Jazz genre is activated. This expert module has already learned the unique style and movement characteristics of Jazz dance during training. Simultaneously, a general expert module can be activated. This general expert module is shared by all genres and is responsible for modeling the common basic patterns of dance, such as rhythmic synchronization and biomechanical consistency.

[0029] Step S204: Based on the target professional expert module, feature processing is performed on the music features and the dance genre tags to obtain corresponding professional features, and based on the preset general expert module, feature processing is performed on the music features to obtain corresponding general features.

[0030] In this embodiment, the specific implementation process of obtaining corresponding professional features by performing feature processing on the music features and the dance genre tags based on the target professional expert module will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0031] The process of obtaining general features by processing music features based on a general expert module includes: 1) Capturing common local temporal patterns in dance through the Mamba module (selective state-space model) in the general expert module, such as the correspondence between music beats and joint movements (e.g., each beat corresponds to one leg lift), and sharing parameters: the Mamba parameters of the general expert are shared across all genres, ensuring generalization, thus outputting local basic features. 2) Modeling the global context using the Transformer self-attention layer (self-attention layer) in the general expert module, such as the biomechanical rationality of movements (e.g., the smoothness of center of gravity transfer) or the overall synchronization with the music. Constraint loss: Explicitly constraining the output to conform to the laws of human movement through auxiliary losses (e.g., joint angle loss, center of gravity loss). Finally, the global basic features, i.e., the aforementioned general features, are output.

[0032] Step S205: Perform feature fusion on the professional features and the general features to obtain the corresponding fused features.

[0033] In this embodiment, the feature fusion described above can be achieved using feature concatenation or fusion strategies. Feature concatenation involves concatenating the features processed by the target professional expert module and the general expert module. For example, the Jazz dance style feature vector extracted by the professional expert and the basic dance pattern feature vector extracted by the general expert can be concatenated in a specific order to form a comprehensive feature vector. Other fusion strategies include weighted fusion, which assigns different weights to the features of the target professional expert module and the general expert module based on their importance, and then performs a weighted sum to obtain the fused features.

[0034] Step S206: Based on the preset dance codebook prediction model, perform dance codebook prediction processing on the fusion features, the music features and the dance genre labels to obtain the corresponding dance codebook sequence.

[0035] In this embodiment, the specific implementation process of performing dance codebook prediction processing on the fusion features, the music features and the dance genre tags based on the preset dance codebook prediction model to obtain the corresponding dance codebook sequence will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0036] Step S207: Perform dance reconstruction on the dance codebook sequence to obtain the corresponding target dance movement, and output the target dance movement.

[0037] In this embodiment, the specific implementation process of reconstructing the dance codebook sequence to obtain the corresponding target dance movement and outputting the target dance movement will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0038] This application first acquires input music data and dance genre tags; then, it extracts features from the music data based on a preset feature extraction strategy to obtain corresponding music features; and determines the corresponding target professional expert module based on the dance genre tags; subsequently, it performs feature processing on the music features and dance genre tags based on the target professional expert module to obtain corresponding professional features, and performs feature processing on the music features based on a preset general expert module to obtain corresponding general features; next, it performs feature fusion on the professional features and general features to obtain corresponding fused features; further, it performs dance codebook prediction processing on the fused features, the music features, and the dance genre tags based on a preset dance codebook prediction model to obtain the corresponding dance codebook sequence; finally, it performs dance reconstruction on the dance codebook sequence to obtain the corresponding target dance movement, and outputs the target dance movement. Based on the above automated dance generation process, this application uses a combination of a target professional expert module, a general expert module, and a dance codebook prediction model to process the input music data and dance genre tags for dance generation. This enables the generated target dance movements to not only meet the style requirements of a specific genre but also have a reasonable basic pattern, ensuring the quality of the generated target dance movements and improving the flexibility of target dance movement generation.

[0039] Among the various implementation options, the core idea of ​​the dance generation system provided in this application is to explicitly decouple the consistency of dance choreography into "dance universality" and "genre specificity" for the first time. By modeling these two dimensions separately, high-quality 3D dances that are highly stylistically controllable, precisely synchronized with the music beat, and maintain the purity of genre style can be generated. The specific construction process of the system can be divided into two stages: the High Fidelity Dance Quantization (HFDQ) stage and the Genre-Aware Dance Generation (GADG) stage, which are detailed below: Phase 1: High-Fidelity Dance Quantization (HFDQ). The goal of this phase is to create a high-quality codebook of dance movement units (i.e., latent space representations). The specific implementation steps are as follows: 1. Motion Decomposition. Human motion is decomposed into upper body and lower body motion, with separate codebooks maintained for each. This is done to improve representational capabilities, as the motion patterns and characteristics of the upper and lower body differ, and separate codebooks can better capture their respective features.

[0040] 2. Finite Scalar Quantization (FSQ). Quantization process formula: The quantization process of FSQ is defined as follows: .in, These are the output characteristics of a dance encoder; It is a bounded function (such as sigmoid) used to map features to a specific range; It is an operation that rounds to the nearest integer; It is the stop gradient operator, used to prevent gradients from propagating to the earlier parts during backpropagation; These are the quantized features. The codebook collapse problem is solved by replacing the "argmin" lookup with a differentiable bounded rounding operation, achieving 100% codebook utilization (compared to 75% of VQ-VAE) and enhancing the diversity of the latent space.

[0041] 3. Motion Reconstruction and Kinematic-Dynamic Constraints. Loss Function Construction: When training HFDQ, not only are the SMPL parameters reconstructed, but kinematic and dynamic constraints are also introduced. The total loss function of HFDQ is... .in, It is the SMPL parameter ( The reconstruction loss is used to measure the difference between the reconstructed SMPL parameters and the original parameters; This is the reconstruction loss of 3D joints (J) derived from SMPL parameters through forward kinematics. This term is introduced to address the issue that SMPL parameter reconstruction ignores the hierarchical structure of human kinematics (e.g., root node errors propagate globally, while hand errors only have a local effect). Refinement of dynamic constraints: and All include dynamic constraints, namely, the reconstruction of velocity (first derivative) and acceleration (second derivative). The specific formula is as follows: ,in, Represents velocity or acceleration. , These are weighting coefficients. This constraint forces the model to learn spatiotemporally coherent and smooth motion, improving the accuracy of reconstruction and the spatiotemporal smoothness of generated actions.

[0042] Phase Two: Genre-Aware Dance Generation (GADG). The goal of this phase is to autoregressively predict the dance codebook sequence given music M and genre label g. The specific implementation steps are as follows: 1. Hybrid Expert (MoE) Architecture Setup. Specialized Experts Training: A separate expert module is trained for each dance genre (e.g., Jazz, Popping, Ballet). Each specialized expert focuses on capturing the subtle style and unique movements of a specific genre, such as locking moves in Popping or elegant postures in Ballet. Universal Expert Setup: A shared expert module is built for all genres. The universal expert is responsible for modeling the basic patterns common to all dances, such as rhythm synchronization and biomechanical consistency, ensuring that the generated dance conforms to human biomechanics and matches the musical rhythm in terms of basic movement principles. Routing Mechanism Determination: A "hard routing" mechanism is used based on the input genre label. Only the corresponding specialized expert is activated, while all inputs are processed by the universal expert. For example, when the input genre label is Popping, only the specialized expert corresponding to Popping is activated, and the universal expert also processes the input.

[0043] 2. Mamba - Transformer Hybrid Backbone Construction. Mamba Module Application: Each expert in the MoE (whether specialized or general) independently processes intramodal sequences (i.e., music, upper body, and lower body features) using a Mamba module (a selective state-space model). The Mamba module efficiently captures fine-grained local temporal dependencies, such as subtle changes in dance movements corresponding to a specific rhythmic point in the music. Transformer Self-Attention Layer Application: Cross-modal fusion is performed using Transformer's self-attention layer, capturing the global contextual relationships between music and dance. For example, understanding the overall style and emotional atmosphere of the music and mapping it to the overall style and emotional expression of the dance. Attention Mechanism Optimization: An attention mechanism is employed. Where Q, K, and V are the query, key, and value matrices, and M is the attention mask. A sliding window attention mask is specifically employed to match the training process on short sequences with the sliding window inference process on long sequences, thereby improving the model's generalization ability across sequences of different lengths.

[0044] 3. System Training and Optimization. Data Preparation: Collect large datasets containing various dance styles and music genres, such as the FineDance and AIST++ datasets. Preprocess the data, including music feature extraction and SMPL parameter representation of dance movements. Training Process: Input the prepared data into the constructed system, using music and style labels as input and dance codebook sequences as output, and train using an autoregressive approach. During training, continuously adjust model parameters and optimize model performance according to the set loss function (which can be determined based on the specific task and evaluation metrics, such as music-movement synchronicity loss, style purity loss, etc.). Model Evaluation and Tuning: Evaluate the dances generated by the model using objective metrics (such as relevant metrics on the FineDance and AIST++ datasets), and tune the model based on the evaluation results, such as adjusting the parameters of the expert module and the structure of the backbone network, until the model reaches a satisfactory performance level.

[0045] Through the training and implementation of the above two stages, the system can generate high-quality 3D dances that are highly controllable in style, precisely synchronized with the music beat, and maintain the purity of the genre style.

[0046] In some alternative implementations, step S202 includes the following steps: The music data is preprocessed to obtain the corresponding specified music data.

[0047] In this embodiment, the preprocessing includes operations such as audio sampling rate conversion and noise reduction. For example, the sampling rate of the music is uniformly converted to 44.1kHz to ensure consistency in subsequent model processing. Simultaneously, audio processing tools are used to remove background noise from the music, improving the quality of the music signal.

[0048] Obtain multiple preset feature extraction strategies.

[0049] In this embodiment, the feature extraction strategies described above can include at least two strategies: Strategy 1: Pre-trained audio encoder; using a pre-trained model (such as Hubert or Wav2Vec2) to extract music embedding vectors. Advantages: Directly utilizes features pre-trained from large-scale audio data, reducing reliance on dance data. Strategy 2: Hand-designed features + MLP. Extracts traditional audio features (such as MFCC, rhythm histogram, onset intensity). Projects these features onto the dimensions required for dance generation using a multilayer perceptron (MLP). Advantages: High interpretability and computational efficiency. Strategy 3: Lightweight temporal model. Uses 1D-CNN or TCN (Temporal Convolutional Network) to process audio waveforms and capture local rhythmic patterns. Advantages: Suitable for scenarios with high real-time requirements.

[0050] Based on preset processing requirements, a target feature extraction strategy is selected from all the aforementioned feature extraction strategies.

[0051] In this embodiment, the aforementioned processing requirements refer to the target processing requirements corresponding to the current dance data generation (such as reducing dependence on dance data / high computational efficiency / high real-time requirements). A strategy matching the aforementioned processing requirements can be selected from all feature extraction strategies to serve as the corresponding target feature extraction strategy.

[0052] Based on the target feature extraction strategy, feature extraction is performed on the specified music data to obtain the corresponding output features.

[0053] In this embodiment, feature processing of the specified music data can be performed according to the selected target feature extraction strategy to extract corresponding output features. These output features are the music features, which may include: basic rhythmic features: the strength of the beat, tempo (BPM), and beat grouping (e.g., 4 / 4 time, 3 / 4 time); intensity and energy features: the dynamic range of the music (e.g., volume, energy change curve), reflecting the amplitude and intensity of dance movements; melody and harmony features: melodic information such as pitch, intervals, and chord progressions, affecting the smoothness and emotional expression of dance movements; structural features: the division of musical sections (e.g., intro, verse, chorus), repetition patterns, and other structural information, guiding the choreography logic of dance movements; and temporal features: the temporal sequence characteristics of the music (e.g., the periodicity of rhythm, the fluctuation of melody), which must be aligned with the temporal pattern of the dance movements.

[0054] The output features are used as the music features.

[0055] This application preprocesses music data to obtain corresponding specified music data; then acquires multiple preset feature extraction strategies; and selects a target feature extraction strategy from all feature extraction strategies based on preset processing requirements; subsequently, it extracts features from the specified music data based on the target feature extraction strategy to obtain corresponding output features; and finally, it uses the output features as music features. Based on the above processing flow, this application obtains specified music data through preprocessing, then selects a target feature extraction strategy from all feature extraction strategies based on processing requirements, and then extracts features from the specified music data based on the selected target feature extraction strategy to obtain corresponding music features. This effectively improves the intelligence and adaptability of music feature extraction and is beneficial for providing an accurate data foundation for generating dance movements.

[0056] In some optional implementations of this embodiment, step S204, which involves performing feature processing on the music features and dance genre tags based on the target professional expert module to obtain corresponding professional features, includes the following steps: The music features and the dance genre tags are fused together to obtain the corresponding specified fusion features.

[0057] In this embodiment, the above-mentioned music features and dance genre tags can be fused by splicing or weighted summation to obtain the fused features, namely the specified fused features.

[0058] The specified fusion features are input into the target expert module; wherein, the target expert module includes a target selective state space model and a target self-attention layer.

[0059] In this embodiment, the aforementioned target selective state space model specifically refers to the target Mamba module within the target expert module, and the aforementioned target self-attention layer refers to the target Transformer self-attention layer within the target expert module.

[0060] The music features are processed based on the target-selective state-space model to obtain the corresponding local genre features.

[0061] In this embodiment, the aforementioned target selective state space model, namely the target Mamba module within the target professional expert module, captures local temporal dependencies through the state space model (SSM) and performs genre optimization: the parameters of Mamba (such as convolution kernels and state transition matrices) can be dynamically generated by dance genre labels (such as through HyperNetwork), making it focus on genre-specific patterns and ultimately outputting the corresponding local genre features.

[0062] Based on the target self-attention layer, the local genre features are mined to obtain the corresponding global genre features.

[0063] In this embodiment, the aforementioned target self-attention layer, namely the target Transformer self-attention layer within the target professional expert module, will mine the global context through the self-attention mechanism and guide the genre: introducing dance genre tags in the attention calculation (such as adding dance genre tags to the projection layer of Query / Key), making the attention biased towards genre-related patterns, and finally outputting the corresponding global genre features.

[0064] The global genre characteristics are used as the professional characteristics.

[0065] This application fuses music features with dance genre labels to obtain corresponding specified fused features. These fused features are then input into a target professional expert module, which includes a target selective state space model and a target self-attention layer. The music features are then processed based on the target selective state space model to obtain corresponding local genre features. Subsequently, the local genre features are mined using the target self-attention layer to obtain corresponding global genre features. Finally, the global genre features are used as professional features. Based on this benefit-driven processing flow, this application, through the use of the target professional expert module, focuses on processing the input specified fused features with a focus on the stylistic details of a specific genre. The target selective state space model captures local temporal dependencies, and the target self-attention layer mines global contextual information, jointly providing accurate data support for generating high-quality dance movements and effectively ensuring the accuracy of the generated professional features.

[0066] In some alternative implementations, step S206 includes the following steps: Call the preset dance codebook prediction model.

[0067] In this embodiment, the dance codebook prediction model described above can specifically employ a Transformer decoder, or it can also employ a Mamba module.

[0068] Based on the music features and the dance genre labels, the fused features are initialized and predicted using a dance codebook prediction model to obtain the corresponding initial prediction codebook.

[0069] In this embodiment, the music features can be encoded into music context using CNN or Transformer, and then the genre label and fusion features can be concatenated to obtain target features. The first frame of the music context and the target features can then be input into the dance codebook prediction model to output the first codebook probability distribution. Based on the first codebook probability distribution, the first codebook of the dance codebook sequence can be predicted, which is the initial prediction codebook.

[0070] Based on a preset autoregressive prediction method, the initial prediction codebook is used to perform stepwise prediction processing until the preset termination condition is met, and the generated target prediction codebook is obtained.

[0071] In this embodiment, an autoregressive prediction method is used to progressively predict subsequent codebooks based on previously predicted codebook sequence segments. When predicting each codebook, the dance codebook prediction model comprehensively considers musical features, genre labels, and previously predicted codebook information to ensure the continuity of dance movements and synchronization with the music. This process continues until a terminator (such as an EOS codebook) is predicted or the maximum length is reached. This ensures that the entire dance sequence remains consistent in time and space, matching the rhythm and emotion of the music. The final codebook obtained is then used as the target predicted codebook. The termination condition refers to reaching the maximum length or generating a terminator.

[0072] The target prediction codebook is used as the dance codebook sequence.

[0073] This application calls a preset dance codebook prediction model; then, based on the music features and the dance genre tags, it uses the dance codebook prediction model to initialize the fused features, obtaining a corresponding initial prediction codebook; subsequently, based on a preset autoregressive prediction method, it uses the initial prediction codebook for progressive prediction until a preset termination condition is met, obtaining a generated target prediction codebook; subsequently, the target prediction codebook is used as the dance codebook sequence. Based on the above processing flow, this application, by using an autoregressive prediction method based on a dance codebook prediction model, ensures the continuity of dance movements. By progressively predicting each codebook, the entire dance sequence remains consistent in time and space, effectively guaranteeing the accuracy and standardization of the generated dance codebook sequence.

[0074] In some alternative implementations, step S207 includes the following steps: Call the preset dance movement unit codebook.

[0075] In this embodiment, during the high-fidelity dance quantization stage, a high-quality dance movement unit codebook has been created, which contains a large number of original dance movement units and their corresponding latent space representations.

[0076] Find the dance movement unit that corresponds to the dance codebook sequence from the dance movement unit codebook.

[0077] In this embodiment, an index can be established in the dance movement unit codebook based on the predicted dance codebook sequence to find the matching dance movement unit.

[0078] The dance movement units are reconstructed based on a preset motion reconstruction strategy to obtain the corresponding dance movements.

[0079] In this embodiment, the dance movement unit is reconstructed based on the preset motion reconstruction strategy to obtain the specific implementation process of the corresponding dance movement. This application will describe this in more detail in subsequent specific embodiments, and will not elaborate further here.

[0080] The dance movement is taken as the target dance movement, and the target dance movement is output based on a preset output method.

[0081] In this embodiment, the specific implementation process of outputting the target dance movement based on the preset output method will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0082] This application utilizes a pre-defined dance movement unit codebook. It then retrieves the dance movement unit corresponding to the codebook sequence. Following this, it reconstructs the dance movement unit based on a pre-defined motion reconstruction strategy to obtain the corresponding dance movement. This dance movement is then used as the target dance movement and output according to a pre-defined output method. Based on this process, this application, by using a dance movement unit codebook, ensures accurate retrieval of the dance movement unit corresponding to the codebook sequence. The pre-defined motion reconstruction strategy then reconstructs the dance movement unit, and the reconstructed dance movement is output as the target dance movement. This effectively guarantees the quality of the output target dance movement, facilitating subsequent use based on the target dance movement to present a high-quality dance effect.

[0083] In some optional implementations of this embodiment, the process of reconstructing the dance movement unit based on a preset motion reconstruction strategy to obtain the corresponding dance movement includes the following steps: Based on the dance movement unit, the parameters of the preset parametric model are reconstructed to obtain the corresponding model parameters.

[0084] In this embodiment, the aforementioned parametric model is specifically an SMPL (Skinned Multi-Person Linear Model). SMPL is a parametric model used to represent the shape and posture of the human body. By reconstructing the SMPL parameters, the approximate shape and posture information of the human body can be obtained. Specifically, based on the found dance motion units, the SMPL parameters are first reconstructed to obtain the reconstructed model parameters, or SMPL parameter sequence. This SMPL parameter sequence includes multiple shape parameters and pose parameters. Shape parameters define the static shape of the human body (such as height and weight), which usually remains unchanged in dance movements (unless character deformation is involved). Pose parameters define the human body posture (joint rotation angles) in each frame and are the core of the movement.

[0085] Based on the principle of forward kinematics, the corresponding three-dimensional joints are derived from the model parameters.

[0086] In this embodiment, 3D joints, i.e., the aforementioned three-dimensional joints, are derived from the reconstructed SMPL parameters based on the principles of forward kinematics. These 3D joints represent the positions of various joints in the human body in three-dimensional space and are key information for describing dance movements. SMPL associates template mesh vertices with shape and pose parameters through linear blending skin (LBS). The 3D joint coordinates can be recursively derived by calculating the joint rotation matrix using forward kinematics and combining it with a pre-defined joint hierarchy tree (e.g., hip → knee → ankle).

[0087] The model parameters and the three-dimensional joints are optimized based on a preset dynamic constraint optimization strategy to obtain the corresponding target model parameters and target three-dimensional joints.

[0088] In this embodiment, the above optimization processes include: 1) Velocity / acceleration constraints: Calculate the first derivative (velocity) and second derivative (acceleration) of the joints. By optimizing and adjusting the posture parameters in the SMPL parameter sequence, the velocity / acceleration is made close to the range of real human motion (e.g., the foot acceleration does not exceed a certain range during running). 2) Physical contact constraints: Force the joints in the foot contact frames to be located near the ground (e.g., z-coordinate error < 1 cm). For non-contact frames, calculate reasonable joint torques through inverse dynamics to avoid the "floating" phenomenon of the suspended foot. 3) Spatiotemporal smoothing: Apply sliding window filtering (e.g., Gaussian filtering) to the posture parameters in the SMPL parameter sequence to reduce abrupt changes between adjacent frames. Output: Optimized SMPL parameters (i.e., target model parameters) and corresponding optimized 3D joints (target 3D joints), at which point the movements are coherent and conform to physical laws.

[0089] Based on the target model parameters and the target 3D joints, a corresponding 3D dance movement is generated.

[0090] In this embodiment, the optimized SMPL parameters and optimized 3D joints meet the core requirements of dance movements—style matching (selected via codebook) and physical feasibility (through dynamic constraints). They serve as an intermediate representation of the movements. The target model parameters and target 3D joints can be further rendered or exported to standard formats (such as BVH or FBX) to obtain visually appealing dance animations, i.e., 3D dance movements. Specifically, the target model parameters can be input into the SMPL model to generate a mesh sequence (visual movement), and surface materials (such as clothing textures) can be added. Then, based on the target 3D joints, the mesh sequence can be rendered into a video or exported to a common motion format (such as BVH) to generate a specific dance movement video.

[0091] The three-dimensional dance movements are used as the dance movements.

[0092] This application reconstructs the parameters of a pre-defined parametric model based on dance motion units to obtain the corresponding model parameters. Then, based on the principle of forward kinematics, the corresponding 3D joints are derived from the model parameters. Next, a pre-defined dynamic constraint optimization strategy is used to optimize the model parameters and 3D joints to obtain the corresponding target model parameters and target 3D joints. Subsequently, the corresponding 3D dance motion is generated based on the target model parameters and target 3D joints. Finally, the 3D dance motion is used as the dance motion. Based on the above processing flow, this application, by using a motion reconstruction strategy and combining SMPL parameter reconstruction, 3D joint derivation, and dynamic constraint application, ensures the quality of the generated dance motion from multiple aspects. Furthermore, by considering kinematic and dynamic constraints, the generated dance motion is spatiotemporally smooth and consistent with the laws of human movement, thus presenting a high-quality dance effect and effectively guaranteeing the quality of the generated dance motion.

[0093] In some optional implementations of this embodiment, outputting the target dance movement based on a preset output method includes the following steps: The target dance movement is subjected to motion smoothness optimization processing to obtain the corresponding first dance movement.

[0094] In this embodiment, the goal of the motion smoothness optimization is to eliminate abrupt changes in motion caused by codebook quantization or autoregressive prediction, ensuring the continuity of motion in time and space. Correspondingly, the motion smoothness optimization process includes: temporal filtering: applying sliding window filtering (such as Gaussian filtering or Savitzky-Golay filtering) to the reconstructed 3D joint sequence to smooth the first-order (velocity) and second-order (acceleration) derivatives and reduce high-frequency jitter. Dynamic constraint enhancement: explicitly constraining the magnitude of velocity and acceleration changes between adjacent frames during the optimization process (e.g., limiting joint angular velocity to no more than a physiological threshold).

[0095] The first dance move is optimized for physical rationality to obtain the corresponding second dance move.

[0096] In this embodiment, the goal of the aforementioned physical rationality optimization is to ensure that dance movements conform to human kinematics and dynamics (e.g., avoiding foot clipping and ground penetration). Correspondingly, the implementation process of physical rationality optimization includes: Contact constraints: detecting the contact state between the foot and the ground (using contact tags in the codebook or predicting contact probabilities), and forcing the foot joint position in contact frames to be fixed near the ground. Joint constraints: limiting joint angles to a physiological range (e.g., knee flexion angle 0-150 degrees), and performing projection correction on frames exceeding this range. Energy minimization: optimizing the joint torque sequence to make the overall movement energy consumption close to that of real human movement (e.g., minimizing the sum of squares of joint angular accelerations).

[0097] The second dance move is optimized for style consistency to obtain the corresponding third dance move.

[0098] In this embodiment, the goal of the style consistency optimization is to enhance the style matching between dance movements and input music genre labels (e.g., exaggerated movements in Jazz, rhythmic feel in Hip-hop). Correspondingly, the style consistency optimization process includes: rhythm alignment enhancement: analyzing the music beat (e.g., using a pre-trained beat detector) and aligning key movements (e.g., jumps, turns) to the strong beat positions. Motion amplitude scaling: adjusting the motion amplitude according to the genre label (e.g., amplifying arm swing amplitude in Jazz, reducing hip movement range in ballet), and scaling the joint trajectory through linear transformation.

[0099] Get the preset output method.

[0100] In this embodiment, the selection of the above output method is not specifically limited and can be determined according to actual business needs. For example, email sending, information sending, interface display, etc. can be used.

[0101] The third dance movement is output and processed based on the aforementioned output method.

[0102] In this embodiment, the generated third dance move can be output according to the selected output method.

[0103] This application optimizes the motion smoothness of a target dance movement to obtain a first dance movement; then optimizes the physical plausibility of the first dance movement to obtain a second dance movement; next, it optimizes the style consistency of the second dance movement to obtain a third dance movement; subsequently, it obtains a preset output method; and finally, it outputs the third dance movement based on the output method. Based on this processing flow, this application intelligently optimizes the quality of the target dance movement by optimizing its motion smoothness, physical plausibility, and style consistency, thus improving the quality of the generated third dance movement. Furthermore, it outputs the third dance movement based on the output method used, enhancing the intelligence of the dance movement output.

[0104] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.

[0105] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0106] Furthermore, this application, through the aforementioned technical solution, systematically solves the shortcomings of the prior art, bringing the following significant advantages: Excellent generation quality and music synchronization: Through genre decoupling and an efficient backbone network, this invention achieves state-of-the-art performance on multiple objective metrics on the two large datasets FineDance and AIST++.

[0107] Strong genre controllability: The MoE-based hard routing design effectively prevents interference and "mixing" between genres. Even in cases of cross-modal conflicts (such as generating a dance to Chinese music using the Popping genre), genre characteristics and music synchronization can be maintained.

[0108] High-fidelity dance quantization: FSQ technology successfully solved the codebook collapse problem of VQ-VAE, achieving 100% codebook utilization.

[0109] Spatiotemporal coherence: The kinematic-dynamic dual constraints (L_joint and velocity / acceleration loss) significantly improve the accuracy of reconstruction and the spatiotemporal smoothness of generated motion.

[0110] A high-efficiency and powerful backbone: The Mamba-Transformer hybrid architecture combines Mamba's efficiency in capturing local dependencies with Transformer's advantages in capturing global cross-modal contexts.

[0111] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0112] It should be emphasized that, in order to further ensure the privacy and security of the aforementioned target dance moves, the target dance moves can also be stored in a blockchain node.

[0113] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0114] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0115] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0116] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0117] Further reference Figure 3 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of an artificial intelligence-based dance data generation device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0118] like Figure 3 As shown, the AI-based dance data generation device 300 described in this embodiment includes: an acquisition module 301, an extraction module 302, a determination module 303, a first processing module 304, a fusion module 305, a prediction module 306, and a second processing module 307. Wherein: Module 301 is used to acquire input music data and dance genre tags; Extraction module 302 is used to extract features from the music data based on a preset feature extraction strategy to obtain corresponding music features; The determination module 303 is used to determine the corresponding target professional expert module based on the dance style label; The first processing module 304 is used to perform feature processing on the music features and the dance style tags based on the target professional expert module to obtain corresponding professional features, and to perform feature processing on the music features based on a preset general expert module to obtain corresponding general features. The fusion module 305 is used to perform feature fusion on the professional features and the general features to obtain corresponding fused features; Prediction module 306 is used to perform dance codebook prediction processing on the fusion features, the music features and the dance genre tags based on a preset dance codebook prediction model to obtain the corresponding dance codebook sequence; The second processing module 307 is used to perform dance reconstruction on the dance codebook sequence to obtain the corresponding target dance movement, and output the target dance movement.

[0119] In some optional implementations of this embodiment, the extraction module 302 includes: The preprocessing submodule is used to preprocess the music data to obtain the corresponding specified music data; The first acquisition submodule is used to acquire multiple preset feature extraction strategies; The filtering submodule is used to filter out the target feature extraction strategy from all the feature extraction strategies based on preset processing requirements; The extraction submodule is used to extract features from the specified music data based on the target feature extraction strategy to obtain the corresponding output features; The first determining submodule is used to use the output feature as the music feature.

[0120] In some optional implementations of this embodiment, the first processing module 304 includes: The fusion submodule is used to fuse the music features with the dance genre tags to obtain the corresponding specified fusion features; An input submodule is used to input the specified fusion features into the target expert module; wherein, the target expert module includes a target selective state space model and a target self-attention layer; The first processing submodule is used to process the music features based on the target selective state space model to obtain the corresponding local genre features; The second processing submodule is used to mine the local genre features based on the target self-attention layer to obtain the corresponding global genre features; The second determining submodule is used to use the global school characteristics as the professional characteristics.

[0121] In some optional implementations of this embodiment, the prediction module 306 includes: The first calling submodule is used to call the preset dance codebook prediction model; The first prediction submodule is used to perform initial prediction processing on the fused features based on the music features and the dance genre labels, using a dance codebook prediction model to obtain the corresponding initial prediction codebook. The second prediction submodule is used to perform stepwise prediction processing using the initial prediction codebook based on a preset autoregressive prediction method until a preset termination condition is met, and to obtain the generated target prediction codebook. The third determining submodule is used to use the target prediction codebook as the dance codebook sequence.

[0122] In some optional implementations of this embodiment, the second processing module 307 includes: The reconstruction unit is used to reconstruct the parameters of a preset parametric model based on the dance movement unit to obtain the corresponding model parameters. The derivation unit is used to derive the corresponding three-dimensional joint points from the model parameters based on the principle of forward kinematics. The first optimization unit is used to optimize the model parameters and the three-dimensional joints based on a preset dynamic constraint optimization strategy to obtain the corresponding target model parameters and target three-dimensional joints. A generation unit is used to generate corresponding three-dimensional dance movements based on the target model parameters and the target three-dimensional joints; A determining unit is used to identify the three-dimensional dance movements as the dance movements.

[0123] In some optional implementations of this embodiment, the reconstruction submodule includes: The second optimization unit is used to perform motion smoothness optimization processing on the target dance movement to obtain the corresponding first dance movement. The third optimization unit is used to perform physical rationality optimization on the first dance movement to obtain the corresponding second dance movement; The fourth optimization unit is used to perform style consistency optimization on the second dance movement to obtain the corresponding third dance movement; The acquisition unit is used to acquire the preset output mode; The output unit is used to process the third dance movement based on the output method.

[0124] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based dance data generation method in the aforementioned implementation method, and will not be repeated here.

[0125] In some optional implementations of this embodiment, the third processing submodule includes: The judgment module is used to determine whether the user-triggered business behavior operation on the target page has been received; The determination module is used to determine the URL of the redirect link page corresponding to the target page based on the business behavior operation if the condition is met. The acquisition module is used to acquire the position parameters corresponding to the resource positions in the target page based on the resource position embedding rules; The splicing module is used to splice the location parameters into the URL of the redirect link page.

[0126] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the artificial intelligence-based dance data generation method in the aforementioned implementation method, and will not be repeated here. To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed] for details. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0127] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0128] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0129] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for an artificial intelligence-based dance data generation method. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0130] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions for the AI-based dance data generation method.

[0131] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0132] Compared with the prior art, the embodiments of this application have the following beneficial effects: In this embodiment, the application uses a combination of a target professional expert module, a general expert module, and a dance codebook prediction model to process the input music data and dance genre tags for dance generation. This enables the generated target dance movements to not only meet the style requirements of a specific genre but also have a reasonable basic pattern, ensuring the quality of the generated target dance movements and improving the flexibility of the generated target dance movements.

[0133] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the artificial intelligence-based dance data generation method described above.

[0134] Compared with the prior art, the embodiments of this application have the following main advantages: In this embodiment, the application uses a combination of a target professional expert module, a general expert module, and a dance codebook prediction model to process the input music data and dance genre tags for dance generation. This enables the generated target dance movements to not only meet the style requirements of a specific genre but also have a reasonable basic pattern, ensuring the quality of the generated target dance movements and improving the flexibility of the generated target dance movements.

[0135] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0136] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A method for generating dance data based on artificial intelligence, characterized in that, Includes the following steps: Obtain the input music data and dance genre tags; Based on a preset feature extraction strategy, feature extraction is performed on the music data to obtain corresponding music features; Based on the dance style tags, the corresponding target professional expert module is determined; Based on the target professional expert module, feature processing is performed on the music features and the dance style tags to obtain corresponding professional features, and based on the preset general expert module, feature processing is performed on the music features to obtain corresponding general features; The professional features and the general features are fused to obtain the corresponding fused features; Based on a preset dance codebook prediction model, the fusion features, the music features, and the dance genre tags are subjected to dance codebook prediction processing to obtain the corresponding dance codebook sequence. The dance codebook sequence is reconstructed to obtain the corresponding target dance movement, and the target dance movement is output.

2. The method for generating dance data based on artificial intelligence according to claim 1, characterized in that, The step of extracting features from the music data based on a preset feature extraction strategy to obtain corresponding music features specifically includes: The music data is preprocessed to obtain the corresponding specified music data; Acquire multiple preset feature extraction strategies; Based on the preset processing requirements, a target feature extraction strategy is selected from all the aforementioned feature extraction strategies; Based on the target feature extraction strategy, feature extraction is performed on the specified music data to obtain the corresponding output features; The output features are used as the music features.

3. The method for generating dance data based on artificial intelligence according to claim 1, characterized in that, The step of obtaining corresponding professional features by performing feature processing on the music features and dance genre tags based on the target professional expert module specifically includes: The music features and the dance genre tags are fused together to obtain the corresponding specified fusion features; The specified fusion features are input into the target expert module; wherein, the target expert module includes a target selective state space model and a target self-attention layer; The music features are processed based on the target-selective state-space model to obtain the corresponding local genre features; Based on the target self-attention layer, the local genre features are mined to obtain the corresponding global genre features; The global genre characteristics are used as the professional characteristics.

4. The method for generating dance data based on artificial intelligence according to claim 1, characterized in that, The step of performing dance codebook prediction processing on the fused features, the music features, and the dance genre tags based on a preset dance codebook prediction model to obtain the corresponding dance codebook sequence specifically includes: Call the preset dance codebook prediction model; Based on the music features and the dance genre labels, the fused features are initialized and predicted using a dance codebook prediction model to obtain the corresponding initial prediction codebook. Based on a preset autoregressive prediction method, the initial prediction codebook is used to perform stepwise prediction processing until the preset termination condition is met, and the generated target prediction codebook is obtained. The target prediction codebook is used as the dance codebook sequence.

5. The method for generating dance data based on artificial intelligence according to claim 1, characterized in that, The step of reconstructing the dance codebook sequence to obtain the corresponding target dance movement specifically includes: Call the preset dance movement unit codebook; Find the dance movement unit that corresponds to the dance codebook sequence from the dance movement unit codebook; The dance movement unit is reconstructed based on a preset motion reconstruction strategy to obtain the corresponding dance movement. The dance movement is taken as the target dance movement, and the target dance movement is output based on a preset output method.

6. The method for generating dance data based on artificial intelligence according to claim 5, characterized in that, The step of reconstructing the dance movement unit based on a preset motion reconstruction strategy to obtain the corresponding dance movement specifically includes: Based on the dance movement units, the parameters of the preset parametric model are reconstructed to obtain the corresponding model parameters; Based on the principle of forward kinematics, the corresponding three-dimensional joint points are derived from the model parameters; The model parameters and the three-dimensional joints are optimized based on a preset dynamic constraint optimization strategy to obtain the corresponding target model parameters and target three-dimensional joints. Based on the target model parameters and the target 3D joints, generate corresponding 3D dance movements; The three-dimensional dance movements are used as the dance movements.

7. The method for generating dance data based on artificial intelligence according to claim 5, characterized in that, The step of outputting the target dance movement based on a preset output method specifically includes: The target dance movement is subjected to motion smoothness optimization processing to obtain the corresponding first dance movement; The first dance move is optimized for physical rationality to obtain the corresponding second dance move; The second dance move is optimized for style consistency to obtain the corresponding third dance move; Get the preset output mode; The third dance movement is output and processed based on the aforementioned output method.

8. A dance data generation device based on artificial intelligence, characterized in that, include: The acquisition module is used to acquire the input music data and dance genre tags; The extraction module is used to extract features from the music data based on a preset feature extraction strategy to obtain corresponding music features; The determination module is used to determine the corresponding target professional expert module based on the dance style label; The first processing module is used to perform feature processing on the music features and the dance genre tags based on the target professional expert module to obtain corresponding professional features, and to perform feature processing on the music features based on a preset general expert module to obtain corresponding general features. The fusion module is used to perform feature fusion on the professional features and the general features to obtain the corresponding fused features; The prediction module is used to perform dance codebook prediction processing on the fusion features, the music features and the dance genre tags based on a preset dance codebook prediction model to obtain the corresponding dance codebook sequence. The second processing module is used to perform dance reconstruction on the dance codebook sequence to obtain the corresponding target dance movement, and to output the target dance movement.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the artificial intelligence-based dance data generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the AI-based dance data generation method as described in any one of claims 1 to 7.