Dance motion generation method, virtual character control method, and electronic device

By extracting music features and predicting dance features, and using cascaded coding networks and diffusion models to generate dance movements corresponding to the music, the problem of incoordination in long-term dance sequences is solved, and high-quality dance effects are achieved.

CN116705064BActive Publication Date: 2026-07-14ALIBABA DAMO (HANGZHOU) TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA DAMO (HANGZHOU) TECH CO LTD
Filing Date
2023-04-04
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies, when generating long-term dance sequences, exhibit poor overall correlation between dance categories and inconsistent choreography styles, resulting in subpar dance performances for virtual digital humans.

Method used

By collecting target music information, extracting global and local music features, predicting global and local dance features, and generating dance movements based on these features, the cascaded coding network and diffusion model are used for feature decoding to generate dance movements corresponding to the music features.

Benefits of technology

It achieves local rhythm matching and global coordination of long-duration dance movements, generating natural, harmonious, and non-repetitive dance movement sequences, thus improving the presentation effect of dance movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116705064B_ABST
    Figure CN116705064B_ABST
Patent Text Reader

Abstract

The application discloses a dance action generation method, a virtual character control method, and an electronic device. The method comprises the following steps: collecting target music information; performing feature extraction on the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to represent the change rule of the target music information, and the local music features are used to represent the change rule of music elements in the target music information; based on the global music features and the local music features, global dance features and local dance features are predicted; and based on the global dance features and the local dance features, target dance actions corresponding to the target music information are generated. The application solves the technical problem of poor action effect in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of virtual scene processing, and more specifically, to a method for generating dance movements, a method for controlling virtual characters, and an electronic device. Background Technology

[0002] Currently, when generating corresponding dance movements based on music information, short dance sequences are generally generated. If long dance sequences are to be generated, problems such as overall incoordination, poor correlation of the overall dance category, and inconsistent choreography styles will occur, resulting in poor performance of virtual digital humans when displaying dance movements.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a method for generating dance movements, a method for controlling virtual characters, and an electronic device, in order to at least solve the technical problem of poor movement effects generated in related technologies.

[0005] According to one aspect of the embodiments of this application, a method for generating dance movements is provided, comprising: acquiring target music information; extracting features from the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of music elements in the target music information; predicting global dance features and local dance features based on the global music features and local music features; and generating a target dance movement corresponding to the target music information based on the global dance features and local dance features.

[0006] According to another aspect of the embodiments of this application, a method for generating dance movements is also provided, comprising: responding to an input command applied to an operation interface, displaying target music information on the operation interface; responding to a generation command applied to the operation interface, displaying a target dance movement corresponding to the target music information on the operation interface, wherein the target dance movement is generated based on global dance features and local dance features, the global dance features and local dance features are predicted based on global music features and local music features of the target music information, the global music features and local music features are obtained by feature extraction of the target music information, the global music features are used to characterize the variation law of the target music information, and the local music features are used to characterize the variation law of musical elements in the target music information.

[0007] According to another aspect of the embodiments of this application, a method for generating dance movements is also provided, comprising: obtaining target music information by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter is the target music information; performing feature extraction on the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to characterize the variation law of the target music information, and the local music features are used to characterize the variation law of music elements in the target music information; predicting global dance features and local dance features based on the global music features and local dance features; generating a target dance movement corresponding to the target music information based on the global dance features and local dance features; and outputting the target dance movement by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the target dance movement.

[0008] According to another aspect of the embodiments of this application, a method for controlling a virtual character is also provided, comprising: collecting target music information through a data acquisition device of a smart device; extracting features from the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of music elements in the target music information; predicting global motion features and local motion features based on the global music features and local music features; generating target limb movements corresponding to the target music information based on the global motion features and local motion features; and controlling a virtual character displayed on a smart device to perform the target limb movements.

[0009] According to another aspect of the embodiments of this application, a method for controlling a virtual character is also provided, comprising: displaying a virtual character on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; when the VR device or AR device collects target music information, extracting features from the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to characterize the variation law of the target music information, and the local music features are used to characterize the variation law of the music elements in the target music information; predicting global motion features and local motion features based on the global music features and local music features; generating target limb movements corresponding to the target music information based on the global motion features and local motion features; and driving the VR device or AR device to render and display the execution result of the virtual character, wherein the execution result is the result of controlling the virtual character to perform the target limb movements.

[0010] According to another aspect of the embodiments of this application, a method for controlling a virtual character is also provided, comprising: acquiring target music information by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter is the target music information; when the target music information is acquired by a VR device or an AR device, performing feature extraction on the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to characterize the change pattern of the target music information, and the local music features are used to characterize the change pattern of music elements in the target music information; predicting global motion features and local motion features based on the global music features and local music features; generating target limb movements corresponding to the target music information based on the global motion features and local motion features; controlling the virtual character to perform the target limb movements to obtain an execution result; and outputting the execution result by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the execution result.

[0011] In this embodiment, target music information can be collected; feature extraction can be performed on the target music information to obtain global music features and local music features, wherein the global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of the music elements in the target music information; based on the global music features and local music features, global dance features and local dance features are predicted; based on the global dance features and local dance features, target dance movements corresponding to the target music information are generated, thereby achieving the purpose of improving the display effect of the target dance movements; it is easy to note that the dance movements can be abstracted into two layers: local dance features and global dance features. When generating dance movements, the local music features and global music features of the music information are fused at the corresponding levels, so that the final generated target dance movements correspond to the music features both locally and globally, thereby obtaining target dance movements with higher display effects, thus solving the technical problem of poor generated movement effects in related technologies.

[0012] It is worth noting that the general description above and the detailed description that follow are merely for illustrative purposes and do not constitute a limitation on this application. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0014] Figure 1 This is a schematic diagram of the hardware environment of a virtual reality device for a method of generating dance movements according to an embodiment of this application;

[0015] Figure 2 This is a structural block diagram of the computational environment for a method of generating dance movements according to an embodiment of this application;

[0016] Figure 3 This is a flowchart of a method for generating dance movements according to Embodiment 1 of this application;

[0017] Figure 4a This is a schematic diagram of a training process for a concatenated coding network according to an embodiment of this application;

[0018] Figure 4b This is a schematic diagram of a dance movement sequence generation process according to an embodiment of this application;

[0019] Figure 5 This is a flowchart of a method for generating dance movements according to Embodiment 2 of this application;

[0020] Figure 6 This is a flowchart of a method for generating dance movements according to Embodiment 3 of this application;

[0021] Figure 7 This is a flowchart of a virtual character control method according to Embodiment 4 of this application;

[0022] Figure 8 This is a flowchart of a virtual character control method according to Embodiment 5 of this application;

[0023] Figure 9 This is a flowchart of a virtual character control method according to Embodiment 6 of this application;

[0024] Figure 10 This is a schematic diagram of a dance movement generation device according to Embodiment 7 of this application;

[0025] Figure 11 This is a schematic diagram of a dance movement generation device according to Embodiment 8 of this application;

[0026] Figure 12 This is a schematic diagram of a dance movement generation device according to Embodiment 9 of this application;

[0027] Figure 13 This is a schematic diagram of a virtual character control device according to Embodiment 10 of this application;

[0028] Figure 14 This is a schematic diagram of a control device for a virtual character according to Embodiment 11 of this application;

[0029] Figure 15 This is a schematic diagram of a control device for a virtual character according to Embodiment 12 of this application;

[0030] Figure 16 This is a structural block diagram of a computer terminal according to an embodiment of this application. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0034] Vector Quantised-Variational Auto Encoder (VQ-VAE): A generative model that encodes data into a discrete quantization latent space through an encoder, so that the quantized features conform to a conditional probability distribution, and then uses a decoder to restore the quantized features back to the original data.

[0035] Diffusion Model: A mathematical model that reconstructs the distribution of prior data step by step from random noise.

[0036] Music-generated dance: Generates corresponding dance movement sequences based on music type, melody, and rhythm.

[0037] Transformer: A network structure that uses a self-attention mechanism.

[0038] Virtual digital human technology: a form of human-computer interaction and service.

[0039] Example 1

[0040] According to an embodiment of this application, a method for generating dance movements is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0041] Figure 1 This is a schematic diagram of the hardware environment of a virtual reality device according to an embodiment of the present application of a method for generating dance movements. Figure 1 As shown, the virtual reality device 104 is connected to the terminal 106, and the terminal 106 is connected to the server 102 via a network. The virtual reality device 104 is not limited to: virtual reality helmets, virtual reality glasses, virtual reality all-in-one machines, etc. The terminal 104 is not limited to PCs, mobile phones, tablets, etc. The server 102 can be a server corresponding to a media file operator. The network includes, but is not limited to: wide area network, metropolitan area network, or local area network.

[0042] Optionally, the virtual reality device 104 in this embodiment includes a memory, a processor, and a transmission device. The memory stores an application program that can be used to perform the following: acquiring target music information; extracting features from the target music information to obtain global and local music features, wherein the global music features characterize the variation patterns of the target music information, and the local music features characterize the variation patterns of musical elements in the target music information; predicting global and local dance features based on the global and local music features; and generating target dance movements corresponding to the target music information based on the global and local dance features, thereby solving the technical problem of poor movement effects generated in related technologies.

[0043] The terminal in this embodiment can be used to display the target dance action corresponding to the target music information on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, and send the target dance action to the virtual reality device 104. After receiving the target dance action, the virtual reality device 104 displays it at the target projection position.

[0044] Optionally, the eye-tracking HMD (Head-Mounted Display) and eye-tracking module in the virtual reality device 104 of this embodiment function the same as in the embodiments described above. That is, the screen in the HMD is used to display real-time images, and the eye-tracking module in the HMD is used to acquire the real-time movement trajectory of the user's eyes. The terminal in this embodiment acquires the user's position and movement information in real three-dimensional space through the tracking system, and calculates the three-dimensional coordinates of the user's head in virtual three-dimensional space, as well as the user's field of vision orientation in virtual three-dimensional space.

[0045] Figure 1 The hardware structure block diagram shown can serve not only as an exemplary block diagram of the aforementioned AR / VR device (or mobile device), but also as an exemplary block diagram of the aforementioned server. In one optional embodiment, Figure 2 The use of the above is illustrated in a block diagram. Figure 1 The AR / VR device (or mobile device) shown is an example of a computing node in computing environment 201. Figure 2 This is a structural block diagram of the computational environment for a method of generating dance movements according to an embodiment of this application, such as... Figure 2 As shown, computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (represented as 210-1, 210-2, ..., in the diagram). Each computing node contains local processing and memory resources, and end user 202 can remotely run applications or store data within computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 within computing environment 201, representing services "A", "D", "E", and "H", respectively.

[0046] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).

[0047] The services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.

[0048] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, such as Figure 2 As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers within a Pod handle requests related to one or more corresponding functions of the service. Proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be equipped with similar Pods.

[0049] During operation, executing a user request from end user 202 may require calling one or more services in computing environment 201, and executing one or more functions of one service may require calling one or more functions of another service. For example... Figure 2 As shown, service "A" 220-1 receives user requests from terminal user 202 from ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to perform one or more functions.

[0050] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.

[0051] Under the aforementioned operating environment, this application provides the following: Figure 3 The method for generating dance movements shown is illustrated. It should be noted that the method for generating dance movements in this embodiment can be derived from... Figure 1The mobile terminal of the illustrated embodiment is executed. Figure 3 This is a flowchart of a method for generating dance movements according to Embodiment 1 of this application. Figure 3 As shown, the method may include the following steps:

[0052] Step S302: Collect target music information.

[0053] The target music information mentioned above includes, but is not limited to, music genre, melody, and rhythm.

[0054] The target music information mentioned above can be the music information needed to generate corresponding dance movements. This target music information can be given music information or music information collected by a sound acquisition device.

[0055] In one alternative embodiment, the target music information may be a complete piece of music, or a portion of a complete piece of music.

[0056] Step S304: Extract features from the target music information to obtain global and local music features of the target music information.

[0057] Among them, global music features are used to characterize the changing patterns of target music information, while local music features are used to characterize the changing patterns of music elements in target music information.

[0058] The musical elements mentioned above include, but are not limited to, the patterns of change in rhythm, dynamics, tempo, etc.

[0059] The aforementioned global music features are used to represent the variation patterns of the overall music segments in the target music information. Among them, the feature dimension of the global music features is higher than that of the local music features.

[0060] The aforementioned local musical features are used to represent the patterns of change in rhythm, dynamics, tempo, etc.

[0061] In one optional embodiment, feature extraction networks can be used to extract features from the target music information. During the encoding process, the time dimension is continuously aggregated and compressed, and the feature dimension is continuously increased. Higher-dimensional features have a larger time span and can represent more abstract meanings such as dance phrases and action paradigms, which can obtain the aforementioned global music features. Lower-dimensional features are more fine-grained short-term actions.

[0062] Furthermore, it is found that low-dimensional features generated in low-dimensional space contain local feature information, namely the aforementioned local musical features, while high-dimensional features generated in high-dimensional space contain global feature information, namely the aforementioned global musical features. Here, low-dimensional space is used to represent cases where the number of data features is relatively small; high-dimensional space is used to represent cases where the number of data features is relatively large. The number of musical features contained in low-dimensional space can be less than the number of musical features contained in high-dimensional space.

[0063] Step S306: Based on global music features and local music features, predict global dance features and local dance features.

[0064] The aforementioned global music features can be represented using high-level features (embedding) encoded by the encoder, while local music features can be represented using information such as beat, onset, and melody (MFCC) extracted from the music and audio processing analysis library (Librosa). Librosa can be used for music and audio processing and analysis.

[0065] The aforementioned overall characteristics of dance are reflected in dance categories, movement paradigms, and dance phrases.

[0066] The aforementioned local dance characteristics are manifested in the instantaneous speed changes and posture changes at key points.

[0067] In one alternative embodiment, dance movements can be abstracted into two layers: local dance features and global dance features. When generating dance movements, the local and global music features are fused at the corresponding levels, so that the final generated target dance movement corresponds to the music features both locally and globally.

[0068] Step S308: Based on global dance features and local dance features, generate the target dance movements corresponding to the target music information.

[0069] In one alternative embodiment, the local and global features of music and dance can be decoupled, and the local and global features of music and dance can be associated separately, thereby improving the correlation between the target dance movements and the target music information.

[0070] In practical applications, users can input or provide a certain type of music, and the above method can generate a harmonious and natural sequence of dance movements that match the music type and rhythm. This method can also generate long-term, globally unified dance movements.

[0071] Generating long-duration dance movement sequences generally requires consideration of global dance features, such as overall coordination, relevance to dance category, and consistency of choreography style. This application decouples the local and global features of music and dance, and then associates them separately. This ensures that the generated dance movements achieve good results in both local rhythm matching and global coordination, and can generate long-duration, natural, and non-repetitive dance movement sequences.

[0072] The method described in this application can be applied to virtual digital human scenarios. Virtual digital humans can serve as a form of human-computer interaction and service. Based on computer vision and graphics, virtual digital humans achieve near-human appearances and lifelike facial expressions and movements, possessing the ability to express emotions and communicate, thus providing innovative products for the digital industry. Among the technologies required for virtual digital humans, limb-driven technology is crucial. Limb-driven technology allows virtual digital humans to "move," and generating dance through music is a significant application of limb-driven technology, for example, in smart speakers with screens, smart vehicles, and smart karaoke screens.

[0073] In virtual digital human products, the virtual digital human can generate realistic and natural long-duration dance effects based on the type, melody, and rhythm of the music by having the user input or select various types of music in appropriate scenarios, using the dance movement generation method provided in this application. In smart speakers with screens, smart cars, and karaoke smart screens, the virtual digital human in the smart system can generate natural and realistic dance effects that conform to the type and rhythm of the music by receiving music input in real time.

[0074] Through the above steps, target music information can be collected; features can be extracted from the target music information to obtain global and local music features. Global music features characterize the variation patterns of the target music information, while local music features characterize the variation patterns of musical elements within the target music information. Based on these global and local music features, global and local dance features are predicted. Based on these global and local dance features, target dance movements corresponding to the target music information are generated, thus improving the display effect of the target dance movements. It is noteworthy that dance movements can be abstracted into two layers: local and global dance features. During dance movement generation, the local and global music features of the music information are fused at their respective levels, ensuring that the final generated target dance movement corresponds to the music features both locally and globally. This results in target dance movements with higher display effects, thus solving the technical problem of poorly generated movement effects in related technologies.

[0075] In the above embodiments of this application, global dance features and local dance features are predicted based on global music features and local music features, including: processing global music features using a first diffusion model to predict global dance features; and processing global dance features and local music features using a second diffusion model to predict local dance features.

[0076] The first diffusion model described above is used to predict global dance features. The second diffusion model described above is used to predict local dance features.

[0077] The first and second diffusion models mentioned above can be diffusion model networks. The diffusion model is divided into a diffusion stage and a de-diffusion stage. In the diffusion stage, noise can be continuously added to the original data to change the data from the original distribution to the desired distribution. For example, Gaussian noise can be continuously added to change the original data distribution to a normal distribution. In the de-diffusion stage, a neural network can be used to restore the data from the normal distribution to the original data distribution.

[0078] In one alternative embodiment, in the diffusion model, music features can be used as conditions for association to generate encoded information that conforms to the music features. For the first diffusion model, global music features can be added as conditions to the Transformer module of the first diffusion model to generate global dance features containing motion content; for the second diffusion model, local music features and global dance features can be added as conditions to the second diffusion model to generate local dance features that match local music features such as music rhythm.

[0079] In the above embodiments of this application, the global music features are processed using a first diffusion model to predict global dance features, including: inputting the global music features into the first diffusion model to obtain the first predicted features output by the first diffusion model; and discretizing the first predicted features to obtain global dance features.

[0080] In an optional embodiment, global music features can be input into a first diffusion model so that the global music features can be predicted by the first diffusion model to obtain high-dimensional first predicted features. The first predicted features can be encoded into a discrete quantization latent space to achieve discretization processing of the first predicted features, thereby obtaining global dance features.

[0081] In the above embodiments of this application, the second diffusion model is used to process global dance features and local music features to predict local dance features, including: inputting global dance features and local music features into the second diffusion model to obtain the second predicted features output by the second diffusion model; and discretizing the second predicted features to obtain local dance features.

[0082] In an optional embodiment, global dance features and local music features can be input into a second diffusion model so that the global dance features and local music features can be predicted by the second diffusion model to obtain low-dimensional second predicted features. The second predicted features can be encoded into a discrete quantization latent space to achieve discretization processing of the second predicted features, thereby obtaining local dance features.

[0083] In the above embodiments of this application, generating a target dance movement corresponding to the target music information based on global dance features and local dance features includes: using a decoder in a cascaded coding network to decode the global dance features and local dance features to generate the target dance movement.

[0084] In an optional embodiment, the quantized global and local dance features can be decoded by a decoder in a cascaded coding network to restore the global and local dance features to the original data, i.e., the target dance movement.

[0085] In another optional embodiment, after obtaining the global dance features and local dance features, the global dance features and local dance features can be displayed to the user, who can then judge whether the global dance features and local dance features are accurate. If they are accurate, the global dance features and local dance features can be directly decoded to generate the target dance movement; if they are inaccurate, the global dance features and local dance features can be adjusted, and the target dance movement can be generated based on the adjusted dance features.

[0086] In the above embodiments of this application, the cascaded coding network includes: multiple cascaded encoders and decoders. The cascaded coding network is trained based on preset dance movements and generated dance movements. The generated dance movements are generated by decoding high-dimensional dance features and low-dimensional dance features using the decoder. The low-dimensional dance features are obtained by extracting features from the preset dance movements using the encoder located at the lower level among the multiple encoders. The high-dimensional dance features are obtained by extracting features from the preset dance movements using the encoder located at the higher level among the multiple encoders.

[0087] The aforementioned preset dance moves can be pre-set dance moves, and there is no limitation on the preset dance moves here. Among them, preset dance moves can also be dance move sequences generated using a preset dance resource pool, combined with music and dance feature matching.

[0088] The aforementioned cascaded coding network can be a two-layer cascaded vector quantization variational autoencoder (VQ-VAE). By introducing a two-layer cascaded vector quantization variational autoencoder, the preset dance movements can be decoupled into low-dimensional and high-dimensional dance features in the feature dimension. While ensuring the effect, it also has the following applications: replacing the high-dimensional dance feature code with the same low-dimensional dance feature code can change the global category of the dance while maintaining the same dance rhythm; changing the low-dimensional dance feature code with the same high-dimensional dance feature code can change the specific movements of the dance while ensuring the consistency of the dance type, so that the generated dance movements can be diverse.

[0089] Figure 4a This is a schematic diagram of a cascaded coding network training process according to an embodiment of this application, such as... Figure 4a As shown, the aforementioned cascaded coding network can be a two-layer cascaded vector quantization variational autoencoder. During the training phase, a two-layer cascaded vector quantization variational autoencoder is first trained to decouple the preset dance movements into local and global dance features in the feature dimension, producing low-dimensional dance features encoded by the lower-layer features and high-dimensional dance features encoded by the higher-layer features. These low-dimensional and high-dimensional dance features can then be quantized using a cascaded decoder to generate the dance movements.

[0090] Figure 4b This is a schematic diagram of a dance movement sequence generation process according to an embodiment of this application, such as... Figure 4b As shown, pre-trained encoders and tools can be used to extract local and global music features, respectively. Then, high-level and low-level diffusion models with different training levels are used to predict high-dimensional dance features (top-code) and low-dimensional dance features (bottom-code), respectively. In the diffusion models at different levels, the music features of the corresponding level can be used as conditions to generate target dance movements that conform to the music features. In the high-level diffusion model, the high-dimensional dance features can be used as conditional codes to predict local music features, thereby obtaining local dance features.

[0091] In the above embodiments of this application, feature extraction is performed on the target music information to obtain global music features and local music features of the target music information, including: using a pre-trained encoder to extract features from the target music information to obtain global music features; and performing audio signal processing on the target music information to obtain local music features.

[0092] The aforementioned audio signals are used to represent the changing patterns of music in terms of beat, rhythm, dynamics, and tempo.

[0093] In one alternative embodiment, a pre-trained encoder can be used to extract the encoded features from the target music information to obtain global music features. Audio signal processing can be performed on the target music information to extract the rhythm, dynamics, tempo, and other variation patterns of the music in the audio signal to obtain local music features.

[0094] This application introduces a cascaded vector quantization variational autoencoder to encode dance movements into local and global features, which are then matched with the local and global features of the music, respectively. Sampling is performed using a diffusion model to obtain dance movement codes that match the music features, and a decoder then reconstructs the dance movement sequence. By decoupling local and global features, not only can local rhythm, global music type, melody, and other features be controlled separately, but the consistency of local rhythm and global dance can also be guaranteed, thus generating high-quality, long-duration dance sequences.

[0095] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0096] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0097] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0098] Example 2

[0099] According to an embodiment of this application, a method for generating dance movements is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0100] Figure 5 This is a flowchart of a method for generating dance movements according to Embodiment 2 of this application, as shown below. Figure 5 As shown, the method may include the following steps:

[0101] Step S502: In response to the input command applied to the operation interface, display the target music information on the operation interface.

[0102] Optionally, users can input commands on the operation interface by clicking, swiping, etc., and the target music information can be displayed on the operation interface, making it convenient for users to adjust or change the target music information.

[0103] The aforementioned user interface can be the user interface of a smart device, which includes, but is not limited to, mobile terminals, computer terminals, smart wearable devices, smart speakers, etc.

[0104] When the user interface is the smart speaker's interface, users can operate the smart speaker's interface in any one or more ways to generate input commands.

[0105] Step S504: In response to the generation command applied to the operation interface, the target dance movement corresponding to the target music information is displayed on the operation interface.

[0106] Among them, the target dance movement is generated based on global dance features and local dance features. The global dance features and local dance features are predicted based on the global music features and local music features of the target music information. The global music features and local music features are obtained by feature extraction of the target music information. The global music features are used to characterize the change pattern of the target music information, and the local music features are used to characterize the change pattern of the music elements in the target music information.

[0107] The aforementioned generation instructions can be generated by manipulating relevant controls or through other means when a user needs to generate target dance movements corresponding to target music information.

[0108] Through the above steps, in response to input commands applied to the operation interface, target music information is displayed on the operation interface; in response to generation commands applied to the operation interface, target dance movements corresponding to the target music information are displayed on the operation interface. The target dance movements are generated based on global and local dance features, which are predicted based on the global and local music features of the target music information. These global and local music features are obtained by feature extraction from the target music information. Global music features are used to characterize the variation patterns of the target music information, while local music features are used to characterize the variation patterns of musical elements within the target music information. This achieves the goal of improving the display effect of the target dance movements. It is noteworthy that the dance movements can be abstracted into two layers: local and global dance features. During dance movement generation, the local and global music features of the music information are fused at the corresponding levels, ensuring that the final generated target dance movements correspond to the music features both locally and globally. This results in target dance movements with higher display effects, thus solving the technical problem of poorly generated movement effects in related technologies.

[0109] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0110] Example 3

[0111] According to an embodiment of this application, a method for generating dance movements is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0112] Figure 6 This is a flowchart of a method for generating dance movements according to Embodiment 3 of this application, as shown below. Figure 6 As shown, the method may include the following steps:

[0113] Step S602: Obtain target music information by calling the first interface.

[0114] The first interface includes a first parameter, the value of which is the target music information.

[0115] The first interface in the above steps can be an interface for data interaction between the cloud server and the client. The client can pass the target music information into the interface function as the first parameter of the interface function to achieve the purpose of uploading the target music information to the cloud server.

[0116] Step S604: Extract features from the target music information to obtain global and local music features of the target music information.

[0117] Among them, global music features are used to characterize the changing patterns of target music information, while local music features are used to characterize the changing patterns of music elements in target music information.

[0118] Step S606: Based on global music features and local music features, predict global dance features and local dance features.

[0119] Step S608: Based on global dance features and local dance features, generate the target dance movements corresponding to the target music information.

[0120] Step S610: Output the target dance movement by calling the second interface.

[0121] The second interface includes a second parameter, the value of which is the target dance movement.

[0122] The aforementioned second interface can be an interface for data interaction between the cloud server and the client. The cloud server can pass the target dance movement into the interface function as the second parameter of the interface function, thereby achieving the purpose of sending the target dance movement to the client.

[0123] Through the above steps, target music information is obtained by calling a first interface, where the first interface includes a first parameter whose value is the target music information; feature extraction is performed on the target music information to obtain global music features and local music features, where global music features are used to characterize the variation patterns of the target music information, and local music features are used to characterize the variation patterns of musical elements in the target music information; based on the global and local music features, global and local dance features are predicted; based on the global and local dance features, target dance movements corresponding to the target music information are generated; and the target dance movements are output by calling a second interface, where the second interface includes a second parameter whose value is the target dance movements, thus achieving the goal of improving the display effect of the target dance movements. It is worth noting that the dance movements can be abstracted into two layers: local dance features and global dance features. When generating dance movements, the local and global music features of the music information are fused at the corresponding levels, so that the final generated target dance movements correspond to the music features both locally and globally, thereby obtaining target dance movements with higher display effects and solving the technical problem of poor generated movement effects in related technologies.

[0124] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0125] Example 4

[0126] According to an embodiment of this application, a method for controlling a virtual character is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0127] Figure 7 This is a flowchart of a virtual character control method according to Embodiment 4 of this application, as follows: Figure 7 As shown, the method may include the following steps:

[0128] Step S702: Collect target music information through the acquisition device of the smart device.

[0129] The aforementioned smart devices include, but are not limited to, mobile terminals, computer terminals, and smart wearable devices.

[0130] The aforementioned acquisition device can be a sound acquisition device.

[0131] Step S704: Extract features from the target music information to obtain global and local music features of the target music information.

[0132] Among them, global music features are used to characterize the changing patterns of target music information, while local music features are used to characterize the changing patterns of music elements in target music information.

[0133] Step S706: Based on global music features and local music features, predict global motion features and local motion features.

[0134] Step S708: Based on global motion features and local motion features, generate the target limb movements corresponding to the target music information.

[0135] Step S710: Control the virtual character displayed on the smart device to perform the target limb movement.

[0136] The virtual character images mentioned above can be preset images or can be changed according to needs. There are no restrictions on the images of the virtual characters here.

[0137] In one alternative embodiment, the smart device can be controlled to display a virtual character performing target limb movements on the smart device's display interface.

[0138] The target limb movements mentioned above can be dance movements, sports movements, etc., and no specific limitation is made to the target limb movements here.

[0139] The above methods can be applied to sports scenarios. The body movements can be sports movements, which can be rhythmic gymnastics, figure skating, synchronized swimming, etc. There are no restrictions here.

[0140] In one alternative embodiment, in rhythmic gymnastics, figure skating, and synchronized swimming scenarios, movements and music are generally required to be coordinated. Therefore, global sports movement features and local sports movement features can be predicted based on global and local music features. Sports movements corresponding to music information can be generated based on global and local sports movement features, thus making it applicable to sports-related scenarios.

[0141] Through the above steps, target music information is collected using the acquisition device of a smart device; features are extracted from the target music information to obtain global and local music features, where global music features characterize the variation patterns of the target music information, and local music features characterize the variation patterns of musical elements within the target music information; based on the global and local music features, global and local motion features are predicted; based on the global and local motion features, target limb movements corresponding to the target music information are generated; and the virtual character displayed on the smart device is controlled to perform the target limb movements, thereby improving the display effect of the target limb movements. It is noteworthy that limb movements can be abstracted into two layers: local and global motion features. During limb movement generation, the local and global music features of the music information are fused at their respective levels, ensuring that the final generated target limb movements correspond to the music features both locally and globally, resulting in target limb movements with higher display effects. This solves the technical problem of poorly generated movement effects in related technologies.

[0142] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0143] Example 5

[0144] According to the embodiments of this application, a control method for virtual characters in virtual reality scenarios such as virtual reality (VR) devices and augmented reality (AR) devices is also provided. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in a different order than that shown here.

[0145] Figure 8 This is a flowchart of a virtual character control method according to Embodiment 5 of this application. Figure 8 As shown, the method may include the following steps:

[0146] Step S802: Display the virtual character on the presentation screen of the virtual reality (VR) device or augmented reality (AR) device.

[0147] Step S804: When the VR device or AR device collects the target music information, feature extraction is performed on the target music information to obtain the global music features and local music features of the target music information.

[0148] Among them, global music features are used to characterize the variation pattern of target music information, while local music features are used to characterize the variation pattern of music elements in target music information.

[0149] Step S806: Based on global music features and local music features, predict global action features and local action features.

[0150] Step S808: Based on global motion features and local motion features, generate the target limb movements corresponding to the target music information.

[0151] Step S810: Drive the VR device or AR device to render and display the execution result of the virtual character, wherein the execution result is generated by controlling the virtual character to perform the target limb action.

[0152] Through the above steps, a virtual character is displayed on the screen of a virtual reality (VR) device or an augmented reality (AR) device. When the VR or AR device collects target music information, feature extraction is performed on the target music information to obtain global and local music features. Global music features characterize the variation patterns of the target music information, while local music features characterize the variation patterns of musical elements within the target music information. Based on the global and local music features, global and local motion features are predicted. Based on the global and local motion features, target limb movements corresponding to the target music information are generated. The VR or AR device is then driven to render and display the execution result of the virtual character. This execution result controls the virtual character to generate the target limb movements, thereby improving the display effect of the target limb movements. It is noteworthy that limb movements can be abstracted into two layers: local and global motion features. During limb movement generation, the local and global music features of the music information are fused at their respective levels, ensuring that the final generated target limb movements correspond to the music features both locally and globally. This results in target limb movements with higher display effects, thus solving the technical problem of poorly generated movement effects in related technologies.

[0153] Optionally, in this embodiment, the above-described virtual character control method can be applied to a hardware environment consisting of a server and a virtual reality device. The execution result of the virtual character is displayed on the presentation screen of the virtual reality (VR) device or augmented reality (AR) device. The server can be a server corresponding to a media file operator. The above-described network includes, but is not limited to, a wide area network (WAN), a metropolitan area network (MAN), or a local area network (LAN). The above-described virtual reality device is not limited to, for example, a virtual reality headset, virtual reality glasses, or a virtual reality all-in-one machine.

[0154] Optionally, the virtual reality device includes: a memory, a processor, and a transmission device. The memory stores an application that can be used to execute: displaying a virtual character on the screen of a virtual reality (VR) device or an augmented reality (AR) device; extracting features from target music information acquired by the VR or AR device to obtain global and local music features, wherein the global music features characterize the variation patterns of the target music information, and the local music features characterize the variation patterns of musical elements within the target music information; predicting global and local motion features based on the global and local music features; generating target limb movements corresponding to the target music information based on the global and local motion features; and driving the VR or AR device to render and display the execution result of the virtual character, wherein the execution result controls the virtual character to generate the target limb movements.

[0155] It should be noted that the above-described method for controlling virtual characters in VR or AR devices may include... Figure 8 The method of the illustrated embodiment is intended to achieve the purpose of driving VR devices or AR devices to display the execution results of virtual characters.

[0156] Optionally, the processor in this embodiment can invoke the application stored in the memory via the transmission device to perform the above steps. The transmission device can receive media files sent by the server via a network, and can also be used for data transmission between the processor and the memory.

[0157] Optionally, in a virtual reality device, there is a head-mounted display with eye tracking. The screen in the HMD is used to display the video footage. The eye tracking module in the HMD is used to acquire the real-time movement trajectory of the user's eyes. The tracking system is used to track the user's position and movement information in real three-dimensional space. The computing and processing unit is used to acquire the user's real-time position and movement information from the tracking system and calculate the three-dimensional coordinates of the user's head in the virtual three-dimensional space, as well as the user's field of vision orientation in the virtual three-dimensional space.

[0158] In this embodiment, the virtual reality device can be connected to a terminal, and the terminal and the server are connected via a network. The virtual reality device is not limited to virtual reality headsets, virtual reality glasses, virtual reality all-in-one machines, etc., and the terminal is not limited to PCs, mobile phones, tablets, etc. The server can be a server corresponding to a media file operator, and the network includes, but is not limited to, wide area networks, metropolitan area networks, or local area networks.

[0159] Example 6

[0160] According to an embodiment of this application, a method for controlling a virtual character is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0161] Figure 9 This is a flowchart of a virtual character control method according to Embodiment 6 of this application, as follows: Figure 9 As shown, the method may include the following steps:

[0162] Step S902: Obtain target music information by calling the first interface.

[0163] The first interface includes a first parameter, the value of which is the target music information.

[0164] The first interface in the above steps can be an interface for data interaction between the cloud server and the client. The client can pass the target music information into the interface function as the first parameter of the interface function to achieve the purpose of uploading the target music information to the cloud server.

[0165] Step S904: When the VR device or AR device collects the target music information, feature extraction is performed on the target music information to obtain the global music features and local music features of the target music information.

[0166] Among them, global music features are used to characterize the changing patterns of target music information, while local music features are used to characterize the changing patterns of music elements in target music information.

[0167] Step S906: Based on global music features and local music features, predict global action features and local action features.

[0168] Step S908: Based on global motion features and local motion features, generate the target limb movements corresponding to the target music information.

[0169] Step S910: Control the virtual character to perform the target limb action and obtain the execution result.

[0170] Step S912: Output the execution result by calling the second interface.

[0171] The second interface includes a second parameter, the value of which is the execution result.

[0172] The aforementioned second interface can be an interface for data interaction between the cloud server and the client. The cloud server can pass the execution result to the interface function as the second parameter of the interface function, thereby achieving the purpose of sending the execution result to the client.

[0173] Through the above steps, target music information is obtained by calling a first interface, where the first interface includes a first parameter whose value is the target music information. When the VR or AR device collects the target music information, feature extraction is performed on the target music information to obtain global and local music features. The global music features characterize the variation patterns of the target music information, while the local music features characterize the variation patterns of musical elements within the target music information. Based on the global and local music features, global and local motion features are predicted. Finally, based on the global and local motion features, target limb movements corresponding to the target music information are generated. The system controls a virtual character to perform target limb movements and obtains the execution result. The execution result is output by calling a second interface, which includes a second parameter whose value is the execution result. This improves the display effect of the target limb movements. It's noteworthy that limb movements can be abstracted into two layers: local movement features and global movement features. During limb movement generation, the local and global music features of the music information are fused at their respective levels, ensuring that the final generated target limb movement corresponds to the music features both locally and globally. This results in a target limb movement with a higher display effect, thus solving the technical problem of poorly generated movement effects in related technologies.

[0174] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0175] Example 7

[0176] According to an embodiment of this application, a dance movement generation apparatus for implementing the above-described dance movement generation method is also provided. Figure 10 This is a schematic diagram of a dance movement generation device according to Embodiment 7 of this application, as shown below. Figure 10 As shown, the device 1000 includes: a data acquisition module 1002, an extraction module 1004, a prediction module 1006, and a generation module 1008.

[0177] The system comprises the following modules: an acquisition module for acquiring target music information; an extraction module for extracting features from the target music information to obtain global and local music features, where global features characterize the variation patterns of the target music information and local features characterize the variation patterns of musical elements within the target music information; a prediction module for predicting global and local dance features based on the global and local music features; and a generation module for generating the target dance movements corresponding to the target music information based on the global and local dance features.

[0178] It should be noted that the acquisition module 1002, extraction module 1004, prediction module 1006, and generation module 1008 mentioned above correspond to steps S302 to S308 in Embodiment 1. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and run in the AR / VR device provided in Embodiment 1.

[0179] In the above embodiments of this application, the prediction module is further configured to process the global music features using a first diffusion model to predict global dance features, and to process the global dance features and local music features using a second diffusion model to predict local dance features.

[0180] In the above embodiments of this application, the prediction module is further configured to input global music features into the first diffusion model, obtain the first prediction features output by the first diffusion model, and discretize the first prediction features to obtain global dance features.

[0181] In the above embodiments of this application, the prediction module is further used to input global dance features and local music features into the second diffusion model, obtain the second prediction features output by the second diffusion model, and discretize the second prediction features to obtain local dance features.

[0182] In the above embodiments of this application, the generation module is further used to decode the global dance features and local dance features using the decoder in the cascaded coding network to generate the target dance movement.

[0183] In the above embodiments of this application, the cascaded coding network includes: multiple cascaded encoders and decoders. The cascaded coding network is trained based on preset dance movements and generated dance movements. The generated dance movements are generated by decoding high-dimensional dance features and low-dimensional dance features using the decoder. The low-dimensional dance features are obtained by extracting features from the preset dance movements using the encoder located at the lower level among the multiple encoders. The high-dimensional dance features are obtained by extracting features from the preset dance movements using the encoder located at the higher level among the multiple encoders.

[0184] In the above embodiments of this application, the extraction module is further used to extract features from the target music information using a pre-trained encoder to obtain global music features; and to perform audio signal processing on the target music information to obtain local music features.

[0185] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0186] Example 8

[0187] According to an embodiment of this application, a dance movement generation apparatus for implementing the above-described dance movement generation method is also provided. Figure 11 This is a schematic diagram of a dance movement generation device according to Embodiment 8 of this application, as shown below. Figure 11 As shown, the device 1100 includes: a first display module 1102 and a second display module 1104.

[0188] The first display module responds to input commands applied to the operation interface and displays target music information on the operation interface. The second display module responds to generation commands applied to the operation interface and displays the target dance movements corresponding to the target music information on the operation interface. The target dance movements are generated based on global and local dance features, which are predicted based on the global and local music features of the target music information. The global and local music features are obtained by feature extraction of the target music information. The global music features are used to characterize the variation patterns of the target music information, and the local music features are used to characterize the variation patterns of the musical elements in the target music information.

[0189] It should be noted that the first display module 1102 and the second display module 1104 mentioned above correspond to steps S502 to S504 in Embodiment 2. The two modules and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and run in the AR / VR device provided in Embodiment 1.

[0190] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0191] Example 9

[0192] According to an embodiment of this application, a dance movement generation apparatus for implementing the above-described dance movement generation method is also provided. Figure 12 This is a schematic diagram of a dance movement generation device according to Embodiment 9 of this application, as shown below. Figure 12 As shown, the device 12000 includes: an acquisition module 1202, an extraction module 1204, a prediction module 1206, a generation module 1208, and an output module 1210.

[0193] The system comprises the following modules: an acquisition module for acquiring target music information by calling a first interface, wherein the first interface includes a first parameter whose value is the target music information; an extraction module for extracting features from the target music information to obtain global and local music features, wherein the global music features characterize the variation patterns of the target music information, and the local music features characterize the variation patterns of musical elements within the target music information; a prediction module for predicting global and local dance features based on the global and local music features; a generation module for generating the target dance movement corresponding to the target music information based on the global and local dance features; and an output module for outputting the target dance movement by calling a second interface, wherein the second interface includes a second parameter whose value is the target dance movement.

[0194] It should be noted that the acquisition module 1202, extraction module 1204, prediction module 1206, generation module 1208, and output module 1210 mentioned above correspond to steps S602 to S610 in Embodiment 3. The five modules and their corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and run in the AR / VR device provided in Embodiment 1.

[0195] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0196] Example 10

[0197] According to an embodiment of this application, a control device for a virtual character used to implement the above-described virtual character control method is also provided. Figure 13 This is a schematic diagram of a virtual character control device according to Embodiment 10 of this application, as shown below. Figure 13 As shown, the device 1300 includes: a data acquisition module 1302, an extraction module 1304, a prediction module 1306, a generation module 1308, and a display module 1310.

[0198] The system comprises the following modules: a data acquisition module for acquiring target music information via a smart device; an extraction module for extracting features from the target music information to obtain global and local music features, where global features characterize the variation patterns of the target music information and local features characterize the variation patterns of musical elements within the target music information; a prediction module for predicting global and local motion features based on the global and local music features; a generation module for generating target limb movements corresponding to the target music information based on the global and local motion features; and a display module for controlling a virtual character displayed on the smart device to perform the target limb movements.

[0199] It should be noted that the acquisition module 1302, extraction module 1304, prediction module 1306, generation module 1308, and display module 1310 mentioned above correspond to steps S702 to S710 in Embodiment 4. The five modules and their corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and run in the AR / VR device provided in Embodiment 1.

[0200] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0201] Example 11

[0202] According to an embodiment of this application, a control device for a virtual character used to implement the above-described virtual character control method is also provided. Figure 14 This is a schematic diagram of a virtual character control device according to Embodiment 11 of this application, as shown below. Figure 14 As shown, the device 1400 includes: a display module 1402, an extraction module 1404, a prediction module 1406, a generation module 1408, and a driving module 1410.

[0203] The system comprises the following modules: a display module for showcasing virtual characters on the screen of a virtual reality (VR) device or an augmented reality (AR) device; an extraction module for extracting features from target music information collected by the VR or AR device, resulting in global and local music features, where global features characterize the variation patterns of the target music information and local features characterize the variation patterns of musical elements within the target music information; a prediction module for predicting global and local motion features based on the global and local music features; a generation module for generating target limb movements corresponding to the target music information based on the global and local motion features; and a driving module for driving the VR or AR device to render and display the execution results of the virtual character, where the execution results control the virtual character to generate the target limb movements.

[0204] It should be noted that the above-mentioned display module 1402, extraction module 1404, prediction module 1406, generation module 1408, and driving module 1410 correspond to steps S802 to S810 in Embodiment 5. The five modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above-mentioned modules or units can be hardware or software components stored in memory and processed by one or more processors. The above-mentioned modules can also be part of the device and run in the AR / VR device provided in Embodiment 1.

[0205] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0206] Example 12

[0207] According to an embodiment of this application, a control device for a virtual character used to implement the above-described virtual character control method is also provided. Figure 15 This is a schematic diagram of a virtual character control device according to Embodiment 12 of this application, as shown below. Figure 15 As shown, the device 1500 includes: an acquisition module 1502, an extraction module 1504, a prediction module 1506, a generation module 1508, a control module 1510, and an output module 1512.

[0208] The system comprises the following modules: an acquisition module for acquiring target music information by calling a first interface, wherein the first interface includes a first parameter whose value is the target music information; an extraction module for extracting features from the target music information when it is acquired by a VR or AR device, thereby obtaining global and local music features, wherein the global music features characterize the variation patterns of the target music information and the local music features characterize the variation patterns of musical elements within the target music information; a prediction module for predicting global and local motion features based on the global and local music features; a generation module for generating target limb movements corresponding to the target music information based on the global and local motion features; a control module for controlling a virtual character to perform the target limb movements and obtaining the execution result; and an output module for outputting the execution result by calling a second interface, wherein the second interface includes a second parameter whose value is the execution result.

[0209] It should be noted that the acquisition module 1502, extraction module 1504, prediction module 1506, generation module 1508, control module 1510, and output module 1512 mentioned above correspond to steps S902 to S912 in Embodiment 6. The six modules and their corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory and processed by one or more processors. The above modules can also be part of the device and run in the AR / VR device provided in Embodiment 1.

[0210] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0211] Example 13

[0212] Embodiments of this application may provide an electronic device, which can be any AR / VR device from a group of AR / VR devices. Optionally, in this embodiment, the aforementioned AR / VR device may also be replaced with a terminal device such as a mobile terminal.

[0213] Optionally, in this embodiment, the AR / VR device described above may be located in at least one of a plurality of network devices in a computer network.

[0214] In this embodiment, the AR / VR device described above can execute the following steps in the dance motion generation method: acquiring target music information; extracting features from the target music information to obtain global music features and local music features, wherein the global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of the music elements in the target music information; predicting global dance features and local dance features based on the global music features and local music features; and generating the target dance motion corresponding to the target music information based on the global dance features and local dance features.

[0215] Optionally, Figure 16 This is a structural block diagram of a computer terminal according to an embodiment of this application. Figure 16 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 102, memory 104, memory controller, and peripheral interfaces, wherein the peripheral interfaces are connected to a radio frequency module, an audio module, and a display.

[0216] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the dance movement generation method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned dance movement generation method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0217] The processor can access the information and application programs stored in the memory via a transmission device to perform the following steps: acquiring target music information; extracting features from the target music information to obtain global and local music features, wherein the global music features are used to characterize the variation patterns of the target music information, and the local music features are used to characterize the variation patterns of the musical elements in the target music information; predicting global and local dance features based on the global and local music features; and generating the target dance movements corresponding to the target music information based on the global and local dance features.

[0218] Optionally, the processor may also execute program code that performs the following steps: using a first diffusion model to process global music features and predict global dance features; using a second diffusion model to process global dance features and local music features and predict local dance features.

[0219] Optionally, the processor may also execute program code that performs the following steps: inputting global music features into the first diffusion model to obtain the first predicted features output by the first diffusion model; and discretizing the first predicted features to obtain global dance features.

[0220] Optionally, the processor may also execute program code that performs the following steps: inputting global dance features and local music features into the second diffusion model to obtain the second predicted features output by the second diffusion model; and discretizing the second predicted features to obtain local dance features.

[0221] Optionally, the processor may also execute program code that performs the following steps: using the decoder in the cascaded coding network to decode global and local dance features to generate the target dance movement.

[0222] Optionally, the processor may also execute program code with the following steps: The cascaded coding network includes multiple cascaded encoders and decoders. The cascaded coding network is trained based on preset dance movements and generated dance movements. The generated dance movements are generated by decoding high-dimensional dance features and low-dimensional dance features using the decoder. The low-dimensional dance features are obtained by extracting features from the preset dance movements using the encoder located at the lower level among the multiple encoders. The high-dimensional dance features are obtained by extracting features from the preset dance movements using the encoder located at the higher level among the multiple encoders.

[0223] Optionally, the processor may also execute program code that performs the following steps: extracting features from the target music information using a pre-trained encoder to obtain global music features; and performing audio signal processing on the target music information to obtain local music features.

[0224] The processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: responding to input commands applied to the operation interface, displaying target music information on the operation interface; responding to generation commands applied to the operation interface, displaying target dance movements corresponding to the target music information on the operation interface, wherein the target dance movements are generated based on global dance features and local dance features, which are predicted based on global and local music features of the target music information, and are obtained by feature extraction of the target music information. Global music features are used to characterize the variation patterns of the target music information, and local music features are used to characterize the variation patterns of musical elements in the target music information.

[0225] The processor can invoke information and application programs stored in the memory via a transmission device to perform the following steps: acquiring target music information by calling a first interface, wherein the first interface includes a first parameter, the value of which is the target music information; extracting features from the target music information to obtain global and local music features, wherein the global music features characterize the variation patterns of the target music information, and the local music features characterize the variation patterns of musical elements within the target music information; predicting global and local dance features based on the global and local music features; generating target dance movements corresponding to the target music information based on the global and local dance features; and outputting the target dance movements by calling a second interface, wherein the second interface includes a second parameter, the value of which is the target dance movement.

[0226] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: acquiring target music information through the acquisition device of the smart device; extracting features from the target music information to obtain global and local music features, wherein the global music features are used to characterize the variation patterns of the target music information, and the local music features are used to characterize the variation patterns of musical elements in the target music information; predicting global and local motion features based on the global and local music features; generating target limb movements corresponding to the target music information based on the global and local motion features; and controlling a virtual character displayed on the smart device to perform the target limb movements.

[0227] The processor can access information and applications stored in memory via a transmission device to perform the following steps: displaying a virtual character on the screen of a virtual reality (VR) device or an augmented reality (AR) device; extracting features from the target music information obtained by the VR or AR device to obtain global and local music features, wherein the global music features characterize the variation patterns of the target music information, and the local music features characterize the variation patterns of musical elements within the target music information; predicting global and local motion features based on the global and local music features; generating target limb movements corresponding to the target music information based on the global and local motion features; and driving the VR or AR device to render and display the execution result of the virtual character, wherein the execution result controls the virtual character to generate the target limb movements.

[0228] The processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: acquiring target music information by calling a first interface, wherein the first interface includes a first parameter, the value of which is the target music information; when the VR or AR device acquires the target music information, performing feature extraction on the target music information to obtain global and local music features, wherein the global music features characterize the variation patterns of the target music information, and the local music features characterize the variation patterns of musical elements within the target music information; predicting global and local motion features based on the global and local music features; generating target limb movements corresponding to the target music information based on the global and local motion features; controlling a virtual character to perform the target limb movements to obtain an execution result; and outputting the execution result by calling a second interface, wherein the second interface includes a second parameter, the value of which is the execution result.

[0229] Using the embodiments of this application, target music information can be collected; feature extraction can be performed on the target music information to obtain global music features and local music features of the target music information. The global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of the music elements in the target music information; based on the global music features and local music features, global dance features and local dance features are predicted; based on the global dance features and local dance features, target dance movements corresponding to the target music information are generated, thereby achieving the purpose of improving the display effect of the target dance movements. It is easy to note that the dance movements can be abstracted into two layers: local dance features and global dance features. When generating dance movements, the local music features and global music features of the music information are fused at the corresponding levels, so that the final generated target dance movements correspond to the music features both locally and globally, which can improve the display effect of the dance movements and thus solve the technical problem of poor generated movement effects in related technologies.

[0230] Those skilled in the art will understand that Figure 16 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, and other terminal devices. Figure 16 This does not limit the structure of the aforementioned electronic device. For example, computer terminal A may also include components that are more complex than those described above. Figure 16 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 16 The different configurations shown.

[0231] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0232] Example 14

[0233] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the aforementioned computer-readable storage medium can be used to store the program code executed by the dance motion generation method provided in Embodiment 1 above.

[0234] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in the AR / VR device terminal group in the AR / VR device network, or in any mobile terminal in the mobile terminal group.

[0235] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: acquiring target music information; extracting features from the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of music elements in the target music information; predicting global dance features and local dance features based on the global music features and local music features; and generating target dance movements corresponding to the target music information based on the global dance features and local dance features.

[0236] Optionally, the storage medium is further configured to store program code for performing the following steps: processing global music features using a first diffusion model to predict global dance features; and processing global dance features and local music features using a second diffusion model to predict local dance features.

[0237] Optionally, the storage medium is further configured to store program code for performing the following steps: inputting global music features into a first diffusion model to obtain a first predicted feature output by the first diffusion model; and discretizing the first predicted feature to obtain global dance features.

[0238] Optionally, the storage medium is further configured to store program code for performing the following steps: inputting global dance features and local music features into the second diffusion model to obtain the second predicted features output by the second diffusion model; and discretizing the second predicted features to obtain local dance features.

[0239] Optionally, the aforementioned storage medium is also configured to store program code for performing the following steps: decoding global and local dance features using a decoder in a cascaded coding network to generate a target dance movement.

[0240] Optionally, the aforementioned storage medium is further configured to store program code for performing the following steps: the cascaded coding network includes multiple cascaded encoders and decoders, the cascaded coding network is trained based on preset dance movements and generated dance movements, the generated dance movements are generated by decoding high-dimensional dance features and low-dimensional dance features using the decoder, the low-dimensional dance features are obtained by extracting features from the preset dance movements using the encoder located at the lower level among the multiple encoders, and the high-dimensional dance features are obtained by extracting features from the preset dance movements using the encoder located at the higher level among the multiple encoders.

[0241] Optionally, the storage medium is further configured to store program code for performing the following steps: extracting features from the target music information using a pre-trained encoder to obtain global music features; and performing audio signal processing on the target music information to obtain local music features.

[0242] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: in response to an input command applied to the operation interface, displaying target music information on the operation interface; in response to a generation command applied to the operation interface, displaying a target dance movement corresponding to the target music information on the operation interface, wherein the target dance movement is generated based on global dance features and local dance features, which are predicted based on global and local music features of the target music information, and are obtained by feature extraction of the target music information. The global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of musical elements in the target music information.

[0243] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining target music information by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter is the target music information; performing feature extraction on the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of music elements in the target music information; predicting global dance features and local dance features based on the global music features and local dance features; generating a target dance movement corresponding to the target music information based on the global dance features and local dance features; and outputting the target dance movement by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the target dance movement.

[0244] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: acquiring target music information through the acquisition device of the smart device; extracting features from the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of music elements in the target music information; predicting global motion features and local motion features based on the global music features and local music features; generating target limb movements corresponding to the target music information based on the global motion features and local motion features; and controlling the virtual character displayed on the smart device to perform the target limb movements.

[0245] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: displaying a virtual character on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device; when the VR device or AR device collects target music information, extracting features from the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of the music elements in the target music information; predicting global motion features and local motion features based on the global music features and local motion features; generating target limb movements corresponding to the target music information based on the global motion features and local motion features; and driving the VR device or AR device to render and display the execution result of the virtual character, wherein the execution result is the result of controlling the virtual character to perform the target limb movements.

[0246] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining target music information by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter is the target music information; when the VR device or AR device collects the target music information, performing feature extraction on the target music information to obtain global music features and local music features of the target music information, wherein the global music features are used to characterize the change pattern of the target music information, and the local music features are used to characterize the change pattern of music elements in the target music information; predicting global motion features and local motion features based on the global music features and local motion features; generating target limb movements corresponding to the target music information based on the global motion features and local motion features; controlling the virtual character to perform the target limb movements to obtain the execution result; and outputting the execution result by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the execution result.

[0247] Using the embodiments of this application, target music information can be collected; feature extraction can be performed on the target music information to obtain global music features and local music features of the target music information. The global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of the music elements in the target music information; based on the global music features and local music features, global dance features and local dance features are predicted; based on the global dance features and local dance features, target dance movements corresponding to the target music information are generated, thereby achieving the purpose of improving the display effect of the target dance movements. It is easy to note that the dance movements can be abstracted into two layers: local dance features and global dance features. When generating dance movements, the local music features and global music features of the music information are fused at the corresponding levels, so that the final generated target dance movements correspond to the music features both locally and globally, which can improve the display effect of the dance movements and thus solve the technical problem of poor generated movement effects in related technologies.

[0248] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0249] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0250] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0251] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0252] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0253] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0254] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for generating dance movements, characterized in that, include: Collect target music information; Feature extraction is performed on the target music information to obtain global music features and local music features of the target music information. The global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of music elements in the target music information. The global music features are high-dimensional music features generated in a high-dimensional space, and the local music features are low-dimensional music features generated in a low-dimensional space. The global music features are processed using a first diffusion model to predict global dance features; the global dance features and the local music features are processed using a second diffusion model to predict local dance features. Based on the global dance features and the local dance features, the target dance movements corresponding to the target music information are generated.

2. The method according to claim 1, characterized in that, The global music features are processed using a first diffusion model to predict the global dance features, including: The global music features are input into the first diffusion model to obtain the first predicted features output by the first diffusion model. The first predicted feature is discretized to obtain the global dance feature.

3. The method according to claim 1, characterized in that, The global dance features and the local music features are processed using a second diffusion model to predict the local dance features, including: The global dance features and the local music features are input into the second diffusion model to obtain the second predicted features output by the second diffusion model. The second predicted feature is discretized to obtain the local dance feature.

4. The method according to claim 1, characterized in that, Based on the global dance features and the local dance features, the target dance movements corresponding to the target music information are generated, including: The global dance features and the local dance features are decoded using a decoder in a cascaded coding network to generate the target dance movement.

5. The method according to claim 4, characterized in that, The cascaded coding network includes multiple cascaded encoders and the decoder. The cascaded coding network is trained based on preset dance movements and generated dance movements. The generated dance movements are generated by decoding high-dimensional and low-dimensional dance features using the decoder. The low-dimensional dance features are obtained by extracting features from the preset dance movements using the encoder at the lower level of the multiple encoders. The high-dimensional dance features are obtained by extracting features from the preset dance movements using the encoder at the higher level of the multiple encoders.

6. The method according to claim 1, characterized in that, Feature extraction is performed on the target music information to obtain global and local music features, including: The global music features are obtained by extracting features from the target music information using a pre-trained encoder. The target music information is subjected to audio signal processing to obtain the local music features.

7. A method for generating dance movements, characterized in that, include: In response to input commands applied to the user interface, the target music information is displayed on the user interface. In response to a generation command applied to the operation interface, a target dance movement corresponding to the target music information is displayed on the operation interface. The target dance movement is generated based on global and local dance features. The global and local dance features are obtained by processing the global music features using a first diffusion model, and the local dance features are obtained by processing the global and local music features using a second diffusion model. The global and local music features are obtained by feature extraction from the target music information. The global music features characterize the variation patterns of the target music information, and the local music features characterize the variation patterns of musical elements within the target music information. The global music features are high-dimensional music features generated in a high-dimensional space, and the local music features are low-dimensional music features generated in a low-dimensional space.

8. A method for generating dance movements, characterized in that, include: The target music information is obtained by calling the first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter is the target music information; Feature extraction is performed on the target music information to obtain global music features and local music features of the target music information. The global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of music elements in the target music information. The global music features are high-dimensional music features generated in a high-dimensional space, and the local music features are low-dimensional music features generated in a low-dimensional space. The global music features are processed using a first diffusion model to predict global dance features; the global dance features and the local music features are processed using a second diffusion model to predict local dance features. Based on the global dance features and the local dance features, the target dance movements corresponding to the target music information are generated; The target dance move is output by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the target dance move.

9. A method for controlling a virtual character, characterized in that, include: Collect target music information using the acquisition device of a smart device; Feature extraction is performed on the target music information to obtain global music features and local music features of the target music information. The global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of music elements in the target music information. The global music features are high-dimensional music features generated in a high-dimensional space, and the local music features are low-dimensional music features generated in a low-dimensional space. The global music features are processed using a first diffusion model to predict global motion features; the global motion features and the local music features are processed using a second diffusion model to predict local motion features. Based on the global motion features and the local motion features, the target limb movements corresponding to the target music information are generated; Control the virtual character displayed on the smart device to perform the target limb movement.

10. A method for controlling a virtual character, characterized in that, include: Displaying virtual characters on the screen of virtual reality (VR) devices or augmented reality (AR) devices; When the VR device or the AR device collects target music information, feature extraction is performed on the target music information to obtain global music features and local music features of the target music information. The global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of music elements in the target music information. The global music features are high-dimensional music features generated in a high-dimensional space, and the local music features are low-dimensional music features generated in a low-dimensional space. The global music features are processed using a first diffusion model to predict global motion features; the global motion features and the local music features are processed using a second diffusion model to predict local motion features. Based on the global motion features and the local motion features, the target limb movements corresponding to the target music information are generated; The VR device or AR device is driven to render and display the execution result of the virtual character, wherein the execution result is generated by controlling the virtual character to perform the target limb movement.

11. A method for controlling a virtual character, characterized in that, include: The target music information is obtained by calling the first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter is the target music information; When VR or AR devices collect target music information, feature extraction is performed on the target music information to obtain global music features and local music features of the target music information. The global music features are used to characterize the variation pattern of the target music information, and the local music features are used to characterize the variation pattern of music elements in the target music information. The global music features are high-dimensional music features generated in a high-dimensional space, and the local music features are low-dimensional music features generated in a low-dimensional space. The global music features are processed using a first diffusion model to predict global motion features; the global motion features and the local music features are processed using a second diffusion model to predict local motion features. Based on the global motion features and the local motion features, the target limb movements corresponding to the target music information are generated; Control the virtual character to perform the target limb action and obtain the execution result; The execution result is output by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the execution result.

12. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 11.