A method and system for driving a plurality of virtual digital humans

CN122653736APending Publication Date: 2026-08-28HANSONG NANJING TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610804182.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

然而,当前的虚拟数字人展示系统多采用预设动画或离线制作的方式,虽然能保证一定的精细度,但在实际应用中存在显著局限性:一方面,传统的虚拟数字人动作基于特定音频提前录制或手动编辑,难以对随机播放的音乐数据做出即时响应,实时交互性较差;另一方面,现有系统大多仅能识别音乐的音量或基本节拍,缺乏对音乐情感的深度解析,因而导致数字人的面部表情与音乐意境脱节,表现僵硬机械;此外,在处理节奏快速变化的旋律时,动作生成与音频信号之间的处理延迟常导致动作滞后或音画不同步,严重影响用户体验

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653736A_ABST
    Figure CN122653736A_ABST
Patent Text Reader

Abstract

The embodiment of the specification provides a driving method and system of multiple virtual digital persons. The method comprises: analyzing an audio signal, and extracting a multi-dimensional music feature of the audio signal; determining a target performance sequence corresponding to each virtual digital person based on the multi-dimensional music feature; and synchronously driving the multiple virtual digital persons to output at least one of an expression performance, a motion performance and a speech performance matched with the audio signal based on the target performance sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of virtual digital humans, and in particular to a driving method and system for multiple virtual digital humans. Background Technology

[0002] With the rapid development of the metaverse, virtual live streaming, and digital entertainment industries, virtual digital humans have been widely applied in various fields such as virtual idols, intelligent customer service, and digital tour guides, becoming an important medium connecting the virtual and the real. However, current virtual digital human display systems mostly adopt preset animations or offline production methods. Although these methods can ensure a certain level of precision, they have significant limitations in practical applications: On the one hand, traditional virtual digital human movements are based on pre-recorded or manually edited specific audio, making it difficult to respond instantly to randomly played music data, resulting in poor real-time interactivity; on the other hand, most existing systems can only recognize the volume or basic beat of the music, lacking in-depth analysis of the music's emotion, thus causing the digital human's facial expressions to be disconnected from the musical mood, resulting in stiff and mechanical performances; in addition, when processing melodies with rapidly changing rhythms, the processing delay between motion generation and audio signals often leads to motion lag or audio-visual asynchrony, seriously affecting the user experience.

[0003] Therefore, it is desirable to provide a driving method and system for multiple virtual digital humans that can accurately perceive audio signals and generate natural feedback in real time, so as to improve the expressiveness of multiple virtual digital humans and the user interaction experience. Summary of the Invention

[0004] This specification provides one or more embodiments of a method for driving multiple virtual digital humans. The method includes: analyzing an audio signal to extract multidimensional musical features from the audio signal; determining a target performance sequence for each virtual digital human based on the multidimensional musical features; and synchronously driving multiple virtual digital humans to output at least one of facial expressions, motor expressions, and speech expressions that match the audio signal based on the target performance sequence.

[0005] This specification provides a driving system for multiple virtual digital humans through one or more embodiments. The system includes: an analysis module configured to analyze an audio signal and extract multidimensional musical features from the audio signal; a determination module configured to determine a target performance sequence for each virtual digital human based on the multidimensional musical features; and a driving module configured to synchronously drive multiple virtual digital humans to output at least one of facial expressions, motor expressions, and speech expressions that match the audio signal based on the target performance sequence. Attached Figure Description

[0006] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0007] Figure 1 This is a schematic diagram illustrating an application scenario of a multi-virtual digital human driving system according to some embodiments of this specification; Figure 2 This is an exemplary block diagram of a multi-virtual digital human driving system according to some embodiments of this specification; Figure 3 This is an exemplary flowchart of a driving method for multiple virtual digital humans according to some embodiments of this specification; Figure 4 This is an exemplary schematic diagram illustrating the regeneration of the target performance sequence according to some embodiments of this specification; Figure 5 This is an exemplary schematic diagram illustrating the determination of the main group and the accompanying group according to some embodiments of this specification. Detailed Implementation

[0008] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0009] Figure 1 This is a schematic diagram illustrating an application scenario of a multi-virtual digital human driving system according to some embodiments of this specification.

[0010] In some embodiments, such as Figure 1 As shown, the application scenario 100 of the multi-virtual digital human driving system (hereinafter referred to as application scenario) may include a user terminal 110, a network 120, a processor 130, a memory 140, and an audio device 150.

[0011] In some embodiments, application scenario 100 may include scenarios such as virtual concerts, virtual live-streaming e-commerce, and emotional healing. For example, in a virtual concert scenario, the system can dynamically adjust the dance, expressions, and interactions of multiple virtual digital humans based on the played audio. For instance, the lead virtual digital singer might perform energetic movements during the climax of a rock song, while the backup virtual digital dancers switch to gentler body language during lyrical passages. In a virtual live-streaming e-commerce scenario, the system can coordinate the performance of multiple virtual digital humans based on product introductions and background music. For example, the lead virtual digital human might professionally explain product features and simulate trial actions, while the assistant virtual digital human might liven up the atmosphere with cheerful background music. In an emotional healing scenario, the system can drive the interaction of multiple virtual digital humans based on the user's voice or played audio. For example, a healing virtual digital human might comfort the user with a gentle tone when their voice is low, while a companion virtual digital human might guide relaxation movements to upbeat music.

[0012] User terminal 110 can be used to interact with users. Users can be one or more users, such as users directly using the multi-virtual digital human driving system, or other related users. In some embodiments, user terminal 110 may include a mobile phone 110-1, a portable computer 110-2, etc. User terminal 110 can be used to present multiple virtual digital humans 111.

[0013] The multiple virtual digital humans 111 can be multiple digital models generated by computer graphics technology, possessing anthropomorphic appearance and behavior. For example, the virtual digital humans 111 can be multiple virtual anchors, multiple virtual singers, etc. In some embodiments, the multiple virtual digital humans 111 can output at least one of motion performance, facial expression performance, and voice performance that matches the audio signal.

[0014] In some embodiments, the display device that presents multiple virtual digital humans 111 can be the display screen built into the user terminal 110, or it can be an independently set external display device (such as an external TV, projector, etc.).

[0015] In some embodiments, the user terminal 110 or the display device has a rendering engine. A rendering engine refers to a functional module or software component used to render the visual effects of multiple virtual digital humans in real time. For example, the rendering engine can employ a real-time graphics engine such as Unity or Unreal Engine, or be implemented based on graphics interfaces such as OpenGL or Vulkan, to support the real-time rendering and smooth display of complex visual effects of multiple virtual digital humans.

[0016] In some embodiments, the user terminal 110 further includes an audio acquisition device. The audio acquisition device is used to acquire the voice emitted by the user. For example, the audio acquisition device can be a microphone, a digital microphone array, etc.

[0017] Network 120 may include any suitable network capable of facilitating information and / or data exchange. In some embodiments, at least one component of application scenario 100 (e.g., user terminal 110, processor 130, memory 140, and audio device 150, etc.) may exchange information and / or data with at least one other component in application scenario 100 via network 120. For example, processor 130 may retrieve multidimensional music features, target performance sequences, performance driving parameters, etc., from memory 140 via network 120.

[0018] In some embodiments, network 120 can be any one or more of wired or wireless networks. For example, network 120 may include cable networks, fiber optic networks, telecommunications networks, cable connections, or any combination thereof. Network connections between components may employ one or more of the above methods. In some embodiments, the network may be a point-to-point, shared, centralized, or other topologies, or a combination of multiple topologies. In some embodiments, network 120 may include one or more network access points.

[0019] Processor 130 can process data and / or information obtained from other devices or system components. Based on this data, information, and / or processing results, processor 130 can execute program instructions to perform one or more functions described in this application. For example, processor 130 can acquire data such as audio signals from user terminal 110 via network 120 and execute related instructions such as extracting multi-dimensional musical features of the audio signals and determining the target performance sequence corresponding to each virtual digital human.

[0020] In some embodiments, processor 130 may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-chip processing device). By way of example only, processor 130 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), a microprocessor, or any combination thereof.

[0021] In some embodiments, the processor 130 may be integrated within the user terminal 110 or set up independently locally or in the cloud. In some embodiments, the processor may be the processor of a multi-virtual digital human driving system.

[0022] The memory 140 may store data, instructions, and / or any other information related to the driving system of the multiple virtual digital humans. In some embodiments, the memory 140 may store data and / or information (e.g., multidimensional music features, target performance sequences, performance driving parameters, etc.) acquired by the user terminal 110, the processor 130, and the audio device 150.

[0023] In some embodiments, memory 140 may store data and / or instructions used by processor 130 to execute or use in order to perform the exemplary methods described herein. For example, memory 140 may store performance-driven parameters corresponding to each virtual digital human generated by processor 130 based on a target performance sequence.

[0024] In some embodiments, memory 140 may include one or more storage units, each of which may be a separate device or part of another device. In some embodiments, memory 140 may be implemented on a cloud platform. In some embodiments, memory 140 may be part of user terminal 110, processor 130, and / or audio device 150.

[0025] Audio device 150 refers to a terminal device used for playing audio. For example, audio device 150 may include a combination of one or more devices with sound playback functions, such as speaker 150-1, headphones 150-2, and smart speaker 150-3.

[0026] In some embodiments, the audio device 150 may be an audio device embedded in the user terminal 110 (such as a speaker) or a stand-alone audio device (such as a smart speaker). When the audio device 150 is a stand-alone audio device, it can receive and play audio signals sent by the user terminal 110 via the network 120.

[0027] In some embodiments, in the application scenario of the multi-virtual digital human driving system, the processor 130 acquires the audio signal sent by the user terminal 110 through the network 120, analyzes the audio signal, and extracts multi-dimensional music features of the audio signal. Based on the multi-dimensional music features, the processor 130 determines the target performance sequence corresponding to each virtual digital human; based on the target performance sequence, it generates performance driving parameters corresponding to each virtual digital human, and sends the performance driving parameters to the user terminal 110 or a display device. After receiving the performance driving parameters, the user terminal 110 or the display device can control the multiple virtual digital humans 111 to output at least one of the facial expressions, action expressions, and voice expressions corresponding to the target performance sequence.

[0028] For further explanation of the above content, please refer to Figures 2 to 5 And its related descriptions.

[0029] Figure 2 This is an exemplary block diagram of a driving system for multiple virtual digital humans according to some embodiments of this specification.

[0030] In some embodiments, such as Figure 2 As shown, the multi-virtual digital human driving system 200 may include an analysis module 210, a determination module 220, and a driving module 230.

[0031] An analysis module is a module used to analyze audio signals.

[0032] In some embodiments, the analysis module is configured to analyze the audio signal and extract multidimensional musical features of the audio signal.

[0033] The determination module refers to the module used to determine the target performance sequence.

[0034] In some embodiments, the determining module is configured to: determine the target performance sequence corresponding to each virtual digital human based on multidimensional music features.

[0035] In some embodiments, the determining module is further configured to: determine at least one of the facial expression sequence and the voice sequence corresponding to each virtual digital human based on the emotional features in the multidimensional music features.

[0036] In some embodiments, the determining module is further configured to: generate one or more virtual digital persons belonging to the subject group and one or more virtual digital persons belonging to the accompaniment group; determine the target performance sequence corresponding to one or more virtual digital persons in the subject group based on the melody features and / or emotional features in the multidimensional music features; and determine the target performance sequence corresponding to one or more virtual digital persons in the accompaniment group based on the beat features in the multidimensional music features.

[0037] A driver module is a module used to drive multiple virtual digital humans to output data.

[0038] In some embodiments, the driving module is configured to: synchronously drive multiple virtual digital humans to output at least one of facial expressions, motion expressions, and voice expressions that match the audio signal, based on a target performance sequence.

[0039] In some embodiments, the driving module is further configured to: synchronously drive multiple virtual digital humans to output facial expressions, motion expressions, and voice expressions that match the audio signal based on the target performance sequence.

[0040] In some embodiments, the driving module is further configured to: monitor the consistency between the actual output performance of each virtual digital human and the audio signal; in response to the consistency meeting a preset condition, regenerate the target performance sequence, and based on the target performance sequence, synchronously drive multiple virtual digital humans to output at least one of facial expressions, motion performances, and voice performances that match the audio signal.

[0041] For further explanation of each of the above modules, please refer to [link / reference]. Figures 3-5 And its related descriptions.

[0042] It should be noted that the above description of the multi-virtual digital human driving system 200 and its modules is for ease of description only and should not be construed as limiting this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principles of this system, may arbitrarily combine the various modules or construct subsystems connected to other modules without departing from these principles. In some embodiments, Figure 1 The analysis module 210, determination module 220, and driving module 230 disclosed herein can be different modules within a single system, or a single module can implement the functions of two or more of the aforementioned modules. For example, the modules can share a single storage module, or each module can have its own separate storage module. Such variations are all within the scope of protection of this specification.

[0043] Figure 3 This is an exemplary flowchart illustrating a driving method for multiple virtual digital humans according to some embodiments of this specification. Figure 3 As shown, process 300 includes steps 310-330. In some embodiments, process 300 may be executed by a multi-virtual-digital-human driving system 200 or a processor 130.

[0044] Step 310: Analyze the audio signal and extract its multidimensional musical features.

[0045] In some embodiments, step 310 may be performed by the analysis module.

[0046] An audio signal is a data signal obtained by decoding an audio file. For example, an audio signal can be a music stream or the data signal corresponding to a performance accompaniment audio.

[0047] In some embodiments, the analysis module can acquire audio signals in various ways. For example, the analysis module can acquire audio signals from a user terminal. Another example is that the processor can download an audio file that the user wants to play from the cloud and parse the audio file to acquire the audio signal. For descriptions of audio acquisition devices and user terminals, please refer to [link to relevant documentation]. Figure 1 And its related descriptions.

[0048] Multidimensional musical features refer to a variety of characteristic information that can characterize the content and patterns of change in music. For example, multidimensional musical features include rhythmic features, melodic features, emotional features, musical structural features, and energy change features.

[0049] Beat characteristics refer to the speed characteristics that reflect the tempo of music, used to control the rhythmic synchronization of multiple virtual digital humans. For example, a beat characteristic can be the speed of the music's beat, expressed in BPM (beats per minute).

[0050] In some embodiments, the analysis module can perform frame-by-frame processing on the audio signal, calculate the short-time energy of each frame, convert the short-time energy of each frame into an energy envelope curve that varies with time, extract the average time interval between adjacent energy peaks using a peak interval statistical method, and convert the average time interval into the number of beats per unit time to obtain the beat characteristics.

[0051] Short-time energy can be used to detect the strength of each frame of audio signal. Frames with high short-time energy correspond to strong beats (such as drum beats), while frames with low short-time energy correspond to weak beats or silence. The peak interval statistics method involves detecting all local peak points on the energy envelope curve, filtering out weak peaks with energy below a preset peak threshold, and using strong peaks with energy greater than or equal to the preset peak threshold as beat candidate points. The time interval between adjacent strong peaks is calculated, and the average of all time intervals is obtained. A local peak point is a location where the energy is greater than the adjacent points before and after it. The preset peak threshold can be 30% of the maximum peak value.

[0052] Melodic features refer to the characteristic information reflecting the main melody's direction and the variation of pitch over time, used to drive the emotional expression of the virtual digital figures in the main body group. For example, melodic features include the main melody's direction and melodic variations. For more information on the main body group, see [link to relevant documentation]. Figure 5 And its related descriptions.

[0053] In some embodiments, the analysis module can obtain a time-spectrum graph by performing a short-time Fourier transform on the audio signal; extract the frequency component with the strongest energy at each time point as a candidate pitch; perform smoothing filtering on the candidate pitch sequence to remove abnormal jumps, and obtain a melody contour curve that changes over time, which serves as a melody feature. The candidate pitch sequence refers to the sequence composed of candidate pitches at each time point.

[0054] Emotional characteristics refer to the feature information that reflects the emotional tendency and intensity changes conveyed by music, and are used to determine the overall performance state and transitions of multiple virtual digital humans. In some embodiments, emotional characteristics may include emotional tendencies such as sadness, joy, excitement, relaxation, tension, and calmness. In some embodiments, emotional characteristics may be a probability vector, such as [probability of sadness, probability of joy, ...].

[0055] In some embodiments, the analysis module can determine the emotional characteristic based on beat characteristics and energy change rate by querying a first preset table. The first preset table may include a correspondence between multiple combinations of beat characteristics and energy change rates and emotional characteristics. For example, if the beat characteristic is greater than or equal to 120 BPM and the energy change rate is greater than or equal to 0.3, the emotional characteristic is rousing; if the beat characteristic is less than or equal to 80 BPM and the energy change rate is less than or equal to 0.15, the emotional characteristic is soothing. The first preset table can be set by a technician based on experience. Further explanation regarding energy change rate can be found in the description related to step 320.

[0056] In some embodiments, the analysis module can extract emotional features based on audio signals using a pre-defined emotion recognition model. The emotion recognition model can be a machine learning model. For example, the emotion recognition model may include any one or a combination of convolutional neural networks (CNNs) or other custom model structures.

[0057] In some embodiments, the input to the emotion recognition model may include audio signals, and the output may be emotion features.

[0058] In some embodiments, the emotion recognition model can be obtained by training a large number of first training samples with first training labels. In some embodiments, the first training samples may include sample audio signals, and the first training labels may be the emotion features corresponding to the sample audio signals. In some embodiments, the first training samples may be obtained based on historical data, and the first training labels may be manually labeled.

[0059] In some embodiments, the analysis module can be trained using various methods based on the first training samples and the first training labels. For example, training can be performed using gradient descent. As an example only, the analysis module can input multiple first training samples with the first training labels into the initial sentiment recognition model, construct a loss function using the first training labels and the results of the initial sentiment recognition model, and iteratively update the parameters of the initial sentiment recognition model based on the loss function. The model training is complete when the loss function of the initial sentiment recognition model satisfies a first preset condition, resulting in a trained sentiment recognition model. The first preset condition can be loss function convergence, the number of iterations reaching a threshold, etc.

[0060] Musical structural features refer to the structured information that reflects the overall division and temporal arrangement of music. For example, musical structural features include the division of music into sections such as intro, verse, chorus, interlude, and coda.

[0061] In some embodiments, the analysis module can obtain musical structural features by detecting periodic recurring peaks in the energy envelope curve of the audio signal. For example, a fixed duration interval corresponding to each periodically occurring peak in the energy envelope is marked as the chorus, the low-energy segment at the beginning of the energy envelope is marked as the intro, and the continuous low-energy segment at the end of the energy envelope is marked as the outro. In the remaining segments, the spectral similarity with the chorus is calculated, and segments with a spectral similarity greater than or equal to a preset similarity threshold are marked as the verse, while segments with a spectral similarity less than the preset similarity threshold are marked as interludes. A low-energy segment refers to an energy segment with an energy value lower than the median of the energy envelope. The preset similarity threshold can be set by a technician based on experience.

[0062] Energy variation characteristics refer to quantitative indicators that reflect the dynamic changes in music intensity over time. They are used to reflect the fluctuations, crescendos, diminuendos, and climaxes of music intensity, thereby dynamically adjusting the range of motion, spatial weights, and performance intensity of multiple virtual digital figures. For example, energy variation characteristics include the short-time energy, root mean square energy, energy envelope curve, and energy change rate curve of the audio signal.

[0063] In some embodiments, the analysis module can perform frame-by-frame processing on the audio signal, calculate the short-time energy or root mean square energy of each frame, perform smoothing filtering on the short-time energy or root mean square energy of each frame to obtain the energy envelope curve, and then perform point-by-point difference calculation on the energy envelope curve to obtain the energy change rate curve, thereby obtaining the energy change characteristics.

[0064] In some embodiments, the analysis module may also extract the maximum value and its corresponding time point in the energy envelope curve, the minimum energy value and its corresponding time point, and the peak value and its corresponding time point in the energy change rate curve as energy change characteristics.

[0065] In some embodiments, the analysis module can also extract multidimensional musical features of the audio signal through various other methods. For example, it can obtain multidimensional musical features of the audio signal by calling an open-source audio processing library or a preset audio parsing model.

[0066] Step 320: Based on multidimensional music features, determine the target performance sequence corresponding to each virtual digital human.

[0067] In some embodiments, step 320 may be performed by the determining module.

[0068] The target performance sequence refers to the sequence of performance content output by the virtual digital human within a preset future time period. The preset future time period refers to the duration of the audio signal.

[0069] In some embodiments, the target performance sequence includes at least one of an expression sequence, an action sequence, and a speech sequence. For example, for the lead singer virtual digital human in the main group, the target performance sequence 1 second before the chorus can be: at 0.0 seconds, the expression is a smile, the action is standing, and the speech is none; at 0.5 seconds, the expression transitions from a smile to excitement, the action is slowly spreading the arms, and the speech is the preparatory sound of inhalation; at 1.0 seconds, the expression is excitement, the action is spreading the arms to the horizontal, and the speech is the first syllable of the opening lyrics of the chorus.

[0070] An expression sequence refers to a sequence of changes in the facial expressions of a virtual digital human over a predetermined future time period. For example, an expression sequence could be: displaying a smiling expression in the first time period and switching to a focused expression in the second time period.

[0071] An action sequence refers to a sequence of limb movements or positional changes of a virtual digital human over a predetermined future period. For example, an action sequence could be: performing a lower limb stepping motion when the drumbeat appears, and performing an arm-expanding motion during the chorus.

[0072] A speech sequence refers to the sequence of voice output by a virtual digital human that changes over time within a preset future period. For example, a speech sequence could be: lowering the tone and volume of a preset speech segment along with a sigh or heavy breathing when the emotional characteristic is sadness; and raising the tone and volume of a preset speech segment along with a strong breath or short laugh when the emotional characteristic is excitement. A preset speech segment refers to the sound information that the audio signal needs to emit within a corresponding time period.

[0073] The determination module can determine the target performance sequence for each virtual digital human through multiple methods based on multi-dimensional music features.

[0074] In some embodiments, the determining module can determine the target performance sequence corresponding to each virtual digital human by matching with a first vector database based on multidimensional music features. For example, the determining module constructs a first target vector based on multidimensional music features. Based on the first target vector and the number of virtual digital humans, the determining module searches in a first vector sub-database corresponding to the number of virtual digital humans to obtain a first feature vector with the highest similarity to the first target vector, and uses the first label corresponding to the first feature vector with the highest similarity as the target performance sequence corresponding to each virtual digital human.

[0075] The first vector database may include multiple first vector sub-databases. Each first vector sub-database corresponds to the number of virtual digital humans. Each first vector sub-database includes multiple first feature vectors and their corresponding first labels. The first feature vectors are constructed based on multi-dimensional music features corresponding to historical audio signals.

[0076] In some embodiments, the determining module may use the set of target performance sequences that have the highest degree of fit with a preset evaluation index in multiple historical drives of the historical audio signals corresponding to the first feature vector as the first label corresponding to the first feature vector. Historical drive refers to the process of controlling and driving the behavior of multiple virtual digital humans based on historical target performance sequences determined by historical multidimensional music features. The highest degree of fit of the preset evaluation index can be defined as the minimum synchronization error between the performance of multiple virtual digital humans and the multidimensional music features. A set of target performance sequences refers to a collection of target performance sequences containing the same number as the number of virtual digital humans. For example, if the number of virtual digital humans is 2, the corresponding first label in the first vector sub-database is a set of target performance sequences containing 2 target performance sequences.

[0077] In some embodiments, the determining module may randomly assign a set of matched target performance sequences to a corresponding number of virtual digital humans.

[0078] In some embodiments, the determining module can determine the target performance sequence corresponding to each virtual digital human based on multidimensional music features and the number of virtual digital humans, using a preset machine learning model.

[0079] In some embodiments, the input to the machine learning model may include multidimensional musical features and the number of virtual digital humans, and the output may be a set of target performance sequences. The determining module may randomly assign the output set of target performance sequences to a corresponding number of virtual digital humans.

[0080] In some embodiments, the machine learning model can be trained using second training samples with second training labels. In some embodiments, the second training samples may include multidimensional musical features of the samples and the number of virtual digital humans. The second training label may be the set of target performance sequences that best matches a preset evaluation metric among multiple historical drives corresponding to the multidimensional musical features of the samples and the number of virtual digital humans. In some embodiments, the second training samples can be obtained from historical data, and the second training labels may be manually labeled.

[0081] The training process for machine learning models is similar to that for emotion recognition models; please refer to the training process for emotion recognition models.

[0082] In some embodiments, the action sequence includes a sequence of position changes for multiple virtual digital humans in a virtual scene. The determination module can generate spatial weights in the virtual scene based on musical structure features and energy change features in multidimensional musical features; based on the spatial weights, the spatial distribution state of the multiple virtual digital humans is adjusted to determine the position change sequence.

[0083] A virtual scene refers to a digital spatial environment where multiple virtual digital beings can engage in activities, display their talents, or perform. For example, a virtual scene can be a stage for a group of virtual digital beings to perform.

[0084] A position change sequence refers to the trajectory information of multiple virtual digital humans in a virtual scene as their position coordinates change over time. For example, a position change sequence can be the motion trajectory generated by connecting the spatial position coordinates of multiple virtual digital humans frame by frame.

[0085] Spatial weights refer to weighted parameters used to control the location distribution, clustering degree, and movement direction of multiple virtual digital beings in a virtual scene. For example, the determination module divides the virtual scene into multiple regions, each corresponding to a weight. Regions with high weights indicate that virtual digital beings tend to move towards or cluster in that region; regions with low weights indicate that virtual digital beings tend to leave or avoid staying in that region.

[0086] In some embodiments, the determining module can guide the virtual digital human to form different spatial distribution states by setting different weights for different regions. For example, when the weight of the central region of the virtual scene is significantly higher than that of the surrounding areas, a spatial distribution state of convergence towards the center is formed; when the weights of all regions are similar, a uniformly dispersed spatial distribution state is formed. The spatial weights are dynamic and can change in real time.

[0087] Spatial distribution refers to the relative positions, arrangement, or clustering / dispersion of multiple virtual digital figures in a virtual scene. For example, spatial distribution can include clustered layouts, dispersed layouts, linear layouts, and wraparound layouts.

[0088] The determination module can generate spatial weights in the virtual scene in various ways based on the musical structure features and energy change features in the multidimensional musical features.

[0089] In some embodiments, the determining module can determine the spatial distribution state and basic weight distribution corresponding to the current music segment by querying a second preset table based on the music structure characteristics; then, based on the energy change characteristics, obtain the energy change rate and change direction in the energy change characteristics, and determine the adjustment coefficient based on the energy change rate and change direction; finally, sort the weight elements in the basic weight distribution by size, multiply the first preset number (e.g., the first 50%) of weight elements by the adjustment coefficient, and divide the last preset number (e.g., the last 50%) of weight elements by the adjustment coefficient, thereby generating the spatial weights in the virtual scene.

[0090] The second preset table can include the correspondence between multiple musical structural features, spatial distribution states, and basic weight distributions. The second preset table can be set by technical personnel based on experience. For example, the intro corresponds to a dispersed layout, and the basic weight distribution can be [0.6, 0.2, 0.1, 0.2, 0.6], so that the virtual characters are evenly distributed around the edge of the stage; the verse corresponds to a linear layout, and the basic weight distribution can be [0.1, 0.6, 0.6, 0.6, 0.1], so that the virtual characters are concentrated on both sides of the center line of the stage; the chorus corresponds to a clustered layout, and the basic weight distribution can be [0.1, 0.2, 0.8, 0.2, 0.1], so that the virtual characters are concentrated in the central circular area; the interlude corresponds to a circular layout, and the basic weight distribution can be [0.7, 0.2, 0.1, 0.2, 0.7], so that the virtual characters are arranged in a ring around the stage. The areas corresponding to the basic weight distribution can be from left to right: left edge, left center, center, right center, and right edge.

[0091] The direction of energy change can include energy increase, energy decrease, and energy remaining constant. In some embodiments, the determining module can determine the direction of change by comparing the short-time energy of the current frame with the short-time energy of the previous frame. For example, if the short-time energy of the current frame is greater than that of the previous frame, the direction of change is determined to be energy increase; if the short-time energy of the current frame is less than that of the previous frame, the direction of change is determined to be energy decrease; and if the short-time energy of the current frame is equal to that of the previous frame, the direction of change is determined to be energy remaining constant.

[0092] The energy change rate refers to the magnitude of change in the short-time energy of the current frame compared to the short-time energy of the previous frame, used to reflect the drastic degree of change in music energy. In some embodiments, the determining module can calculate the absolute value of the difference between the short-time energy of the current frame and the short-time energy of the previous frame, and use the ratio of the absolute value of the difference to the short-time energy of the previous frame as the energy change rate.

[0093] As energy increases, the adjustment coefficient is positively correlated with the rate of energy change. For example, the determination module determines the adjustment coefficient based on the rate of energy change using formula (1): α=1+k×r (1) Where α is the adjustment coefficient, k is the preset adjustment intensity coefficient, and r is the energy change rate.

[0094] When energy decreases, the adjustment coefficient is negatively correlated with the rate of energy change. For example, the determination module determines the adjustment coefficient based on the rate of energy change using formula (2): α=1-k×r (2).

[0095] When the energy remains constant, the adjustment coefficient can be set to 1, indicating that no adjustment is performed.

[0096] When the adjustment coefficient is greater than 1 (i.e., energy increases), the high weights in the basic weight distribution become larger and the low weights become smaller, resulting in a strong clustering effect towards the center; when the adjustment coefficient is less than 1 (i.e., energy decreases), the weight differences in the basic weight distribution narrow and the overall distribution becomes more gradual, resulting in the population dispersing outward.

[0097] In some embodiments, the determining module can also generate spatial weights in the virtual scene based on the music structure features and energy change features in the multidimensional music features, using deep learning techniques. For example, by taking the music structure features and energy change features as input, a pre-trained spatiotemporal graph convolutional neural network model can directly output the spatial weights in the virtual scene.

[0098] The determination module can determine the sequence of position changes in various ways based on spatial weights.

[0099] In some embodiments, the determining module generates a spatial force field vector by comparing spatial weights with the actual distribution to adjust the spatial distribution state. The spatial weights define the proportion of the target number of people in each region, and the determining module statistically analyzes the actual distribution in real time. When the actual number of people in a region is greater than the target number, a spatial force field vector in a repulsive state is generated, driving the excess virtual digital people to spread outward; when the actual number of people in a region is less than the target number, a spatial force field vector in an attractive state is generated, driving the virtual digital people to converge towards the center of that region.

[0100] The target number of people refers to the number of people that should exist in each area, determined based on spatial weights and the total number of virtual digital humans. Actual distribution refers to the actual number of people in each area of ​​the virtual scene. The spatial force field vector is a dynamically adjusted vector determined based on spatial weights and the actual distribution, used to limit the movement of virtual digital humans between areas.

[0101] For example, a virtual scene is divided into three regions: left, center, and right. If the spatial weights are [0.2, 0.7, 0.1], and the actual distribution is [5 people, 2 people, 3 people], the determination module will identify that the actual number of people in the left and right regions exceeds the limit, generating an outward repulsive force in those regions. Conversely, the number of people in the center region will be insufficient, generating an inward attractive force. Each virtual digital person moves under the influence of this push-pull force field, and the determination module continuously adjusts the spatial force field vector until the actual distribution approaches the spatial weights.

[0102] In some embodiments, the determining module determines the sequence of position changes based on spatial weights and spatial force field vectors. For example, firstly, the determining module normalizes the spatial weights to obtain a probability distribution. Based on the probability distribution, a target region is randomly selected for each virtual digital human, and the target position is determined. For instance, if the spatial weight of region A is 0.2, the spatial weight of region B is 0.7, and the spatial weight of region C is 0.1, then 70% of the virtual digital humans tend to select region B (i.e., the target region). This results in a natural clustering effect in regions with high weights (such as region B), while regions with low weights (such as regions A and C) are sparsely distributed. The target region refers to the region that the virtual digital human needs to reach. The target position refers to the location that the virtual digital human needs to reach. In some embodiments, the target position can be a randomly generated location point within the target region or the center of the target region.

[0103] Subsequently, the determination module calculates the displacement vector from the current position to the target position, corrects it using the spatial force field vector, and obtains the direction and velocity of motion. For example, the displacement vector and the spatial force field vector are weighted and summed to obtain the direction of motion. For instance, the displacement vector and the spatial force field vector are multiplied by their respective weight coefficients and then summed to generate a new composite vector. This composite vector represents the direction of motion of the virtual digital person under the current spatial weights. This direction of motion simultaneously includes the intended direction of the displacement vector and the influence exerted by the spatial force field vector; the specific direction of motion is determined by the directions and weights of the displacement vector and the spatial force field vector. Based on the displacement vector and a preset maximum velocity constraint, the basic velocity is obtained. Then, the basic velocity is multiplied by the magnitude of the spatial force field vector to obtain the final velocity.

[0104] Finally, the determination module uses a kinematic integral algorithm to calculate the spatial position of each virtual digital human at each time point frame by frame, generating a position change sequence. When the movement trajectories of multiple virtual digital humans intersect, a collision avoidance algorithm is used to fine-tune the position change sequence to ensure that the virtual digital humans do not penetrate each other.

[0105] In some embodiments, the determining module can determine the position change sequence based on spatial weights using a pathfinding algorithm. For example, the reciprocal of the spatial weights can be used as the pathfinding cost to construct a navigation mesh cost map (NavMesh) for the virtual scene. The pathfinding algorithm can then plan the lowest-cost path from the current location to a high-weight target area for each virtual digital human, thereby generating the position change sequence.

[0106] Some embodiments in this specification introduce spatial weight sequences generated based on musical structure features and energy change features, thereby achieving synchronous adjustment of the spatial distribution of multiple virtual digital humans. This enables the position change sequences of the virtual digital humans to be linked in real time with musical segments and energy changes, overcoming the problem of traditional group formations being disconnected from musical segments. This creates a stage effect with highly synchronized audiovisual elements, significantly improving the coordination and immersion of the performance.

[0107] In some embodiments, the determining module may determine at least one of the facial expression sequence and the voice sequence corresponding to each virtual digital human based on the emotional features in the multidimensional music features.

[0108] In some embodiments, the determining module can determine the expression sequence corresponding to each virtual digital human based on the emotional features in the multidimensional music features. For example, the determining module determines the expression sequence based on the emotional features by searching a preset expression library. For instance, if the emotional feature is sadness, the retrieved expression sequence may include slightly drooping eyes, slightly pursed lips, and reduced blinking frequency. Another example is that the determining module inputs the emotional features into a pre-trained machine learning model, which then directly generates the corresponding expression sequence.

[0109] In some embodiments, the determining module can determine the speech sequence corresponding to each virtual digital human based on emotional features in multidimensional music features. For example, the determining module determines the speech sequence based on emotional features by searching a preset speech library. For instance, if the emotional feature is sadness, the retrieved speech sequence may include preset speech segments with lowered tone and volume, as well as sighs, heavy breathing sounds, etc. Another example is that the determining module inputs the emotional features into a pre-trained machine learning model, which automatically generates the corresponding speech sequence.

[0110] In some embodiments, the determining module can also adjust the facial expression sequence based on the speech sequence and emotional features. For example, the determining module can add subtle lip movements and other facial expressions corresponding to the speech sequence to the facial expression sequence. Alternatively, the determining module can input the emotional features and speech sequence into a pre-trained machine learning model, which can then directly generate the facial expression sequence.

[0111] In some embodiments of this specification, the problem of stiff facial expressions in virtual digital humans is solved by automatically mapping emotional features to facial expression sequences and speech sequences, achieving a high degree of consistency between emotions, facial micro-expressions, and speech rhythm. Simultaneously, the synchronized and smooth transition between facial expression sequences and speech sequences ensures audiovisual continuity, significantly enhancing audiovisual interactive feedback and the user's immersive experience.

[0112] In some embodiments, the determining module can generate one or more virtual digital humans belonging to a subject group and one or more virtual digital humans belonging to an accompaniment group; based on melody features and / or emotional features in multidimensional music features, it determines the target performance sequence corresponding to one or more virtual digital humans in the subject group; based on beat features in multidimensional music features, it determines the target performance sequence corresponding to one or more virtual digital humans in the accompaniment group. For further explanation of this section, please refer to [link / reference]. Figure 5 And its related descriptions.

[0113] Step 330: Based on the target performance sequence, simultaneously drive multiple virtual digital humans to output at least one of facial expressions, motion performances, and speech performances that match the audio signal.

[0114] In some embodiments, step 330 may be performed by the driver module.

[0115] Facial expressions refer to the visual images of multiple virtual digital humans' facial movements that are ultimately output on a display device. For example, facial expressions include smiling, focusing, or excitement.

[0116] Motion performance refers to the visual display of multiple virtual digital human body movements ultimately output on a display device. For example, motion performance includes body movements such as standing, stepping with the lower limbs, or extending the arms.

[0117] Voice performance refers to the audio feedback of multiple virtual digital humans ultimately output in an audio device. For example, voice performance includes raising the pitch, changing the volume, and other vocal states.

[0118] For more information on display and audio devices, please refer to [link / reference]. Figure 1 And its related descriptions.

[0119] The driving module can synchronously drive multiple virtual digital humans to output at least one of the following actions, expressions, and voices that match the audio signal, based on the target performance sequence and through various methods.

[0120] In some embodiments, the driving module can drive multiple virtual digital humans to output motion actions that match the audio signal by querying a driving resource dictionary based on the target performance sequence. The driving resource dictionary refers to a pre-built multi-dimensional mapping database used to transform the target performance sequence into underlying rendering parameters. For example, the driving resource dictionary may include an action mapping sub-table and a face mapping sub-table. The action mapping sub-table may include the correspondence between actions and skeletal rotation vectors. The face mapping sub-table may include the correspondence between facial expressions and speech and facial mesh deformation weights and lip movements.

[0121] For example, the driving module generates a skeletal rotation vector sequence that controls the body movement of each virtual digital human in the virtual scene by querying the action mapping sub-table based on the action sequence in the target performance sequence; then, based on the skeletal rotation vector sequence, it renders and outputs the corresponding action performance of each virtual digital human.

[0122] In some embodiments, the driving module can drive multiple virtual digital humans to output facial expressions that match the audio signal by querying a driving resource dictionary based on the target performance sequence. For example, based on the facial expression sequence and speech sequence in the target performance sequence, the driving module generates a facial mesh deformation weight sequence and a lip-sync sequence that control the facial movement of each virtual digital human in the virtual scene by querying a facial mapping sub-table; then, based on the facial mesh deformation weight sequence and lip-sync sequence, it renders, generates, and outputs the facial expression corresponding to each virtual digital human.

[0123] In some embodiments, the driving module can drive multiple virtual digital humans to output speech performances that match the audio signal based on a target performance sequence. For example, the driving module renders, generates, and outputs the speech performance corresponding to each virtual digital human based on the speech sequence in the target performance sequence.

[0124] In some embodiments, the target performance sequence includes an expression sequence, a motion sequence, and a speech sequence. The driving module can synchronously drive multiple virtual digital humans to output motion, expression, and speech performances that match the audio signal, based on the target performance sequence.

[0125] In some embodiments, the driving module can, based on the target performance sequence, query the driving resource dictionary to synchronously drive multiple virtual digital humans to output action performances, facial expressions, and voice performances that match the audio signal. For example, based on the action sequence in the target performance sequence, the driving module generates a skeletal rotation vector sequence that controls the body movement of each virtual digital human in the virtual scene by querying the action mapping sub-table; based on the facial expression sequence and voice sequence in the target performance sequence, it generates a facial mesh deformation weight sequence and lip-sync sequence that control the facial movement of each virtual digital human in the virtual scene by querying the facial mapping sub-table; then, based on the skeletal rotation vector sequence, facial mesh deformation weight sequence, and lip-sync sequence, it renders and outputs the action performances and facial expressions corresponding to each virtual digital human; simultaneously, based on the voice sequence in the target performance sequence, it renders and outputs the voice performance corresponding to each virtual digital human.

[0126] In some embodiments of this specification, multiple virtual digital humans are synchronously driven to output facial expressions, motion expressions, and voice expressions that match the audio signals through a unified target performance sequence. This achieves efficient and unified conversion of multimodal data, ensuring that the actions, lip movements, and audio signals of multiple virtual digital humans are strictly aligned in time. This effectively eliminates the sense of disharmony caused by audio-visual asynchrony and significantly improves the realism and smoothness of multi-virtual digital human interaction on the same screen.

[0127] In some embodiments, the driving module can monitor the consistency between the actual output performance of each virtual digital human and the audio signal; in response to the consistency meeting a preset condition, it regenerates the target performance sequence, and based on the target performance sequence, synchronously drives multiple virtual digital humans to output at least one of facial expressions, motion performances, and voice performances that match the audio signal. For further details on this section, please refer to [link / reference]. Figure 4 And its related descriptions.

[0128] In some embodiments of this specification, by extracting multidimensional musical features from audio signals, differentiated target performance sequences and performance driving parameters are generated for different virtual digital humans, enabling differentiated collaborative driving of multiple virtual digital humans for the same audio signal. Compared to traditional single synchronous driving, this method allows multiple virtual digital humans to present complementary or independent performances based on different dimensions of the music, significantly improving the layering and expressiveness of group performances and enhancing the viewing experience and realism of performances by multiple virtual digital humans.

[0129] It should be noted that the above description of process 300 is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art can make various modifications and changes to process 300 under the guidance of this specification. However, these modifications and changes remain within the scope of this specification.

[0130] Figure 4 This is an exemplary schematic diagram illustrating the generation of a target performance sequence according to some embodiments of this specification.

[0131] In some embodiments, such as Figure 4 As shown, the driving module can monitor the consistency 430 between the actual output performance 410 of each virtual digital human and the audio signal 420; in response to the consistency 430 meeting the preset condition 440, the target performance sequence 450 is regenerated, and based on the target performance sequence 450, the multiple virtual digital humans are synchronously driven to output at least one of the facial expressions 460, action expressions 470 and voice expressions 480 that match the audio signal 420.

[0132] For further explanation of audio signals, target performance sequences, facial expressions, motor expressions, and speech expressions, please refer to [link to relevant documentation]. Figure 3 And its related descriptions.

[0133] Actual output performance refers to the actual performance state of each virtual digital person at the current moment. For example, actual output performance can include facial expressions, physical actions, and vocal expressions.

[0134] In some embodiments, the driver module can acquire the actual output performance in various ways. For example, the driver module can collect the actual output performance of each virtual digital human in real time through the rendering engine's feedback interface. Another example is that the driver module can acquire the rendered screen in real time through an external application (such as screen recording software) and extract the actual output performance through visual recognition. For more information about the rendering engine, see [link to documentation]. Figure 1 And its related descriptions.

[0135] Consistency refers to the degree to which the actual output performance of a virtual digital human matches the audio signal in dimensions such as rhythm and emotion. In some embodiments, consistency can be represented numerically.

[0136] In some embodiments, the driver module can determine consistency based on the actual output performance and audio signal using a consistency judgment model.

[0137] A consistency judgment model is a model used to determine the consistency between the actual output performance and the audio signal. In some embodiments, the consistency judgment model can be a machine learning model. For example, the consistency judgment model may include any one or a combination of cross-modal attention models or other custom model structures.

[0138] In some embodiments, the input to the consistency judgment model may include time-aligned actual output performance and audio signal, and the output may be consistency. Time alignment means that the actual output performance and audio signal have the same duration and timestamp.

[0139] In some embodiments, the consistency judgment model can be obtained by training a third training sample with a third training label. In some embodiments, the third training sample may include time-aligned actual output performance of the sample and sample audio signal, and the third training label may be the actual consistency between the time-aligned actual output performance of the sample and sample audio signal. In some embodiments, the third training sample may be obtained from historical data, and the third training label may be manually labeled.

[0140] The training process of the consistency judgment model is similar to that of the emotion recognition model; please refer to the training process of the emotion recognition model.

[0141] Preset conditions refer to pre-defined conditions used to determine whether consistency requirements are met. For example, a preset condition could be that consistency is below a preset threshold. The preset threshold can be set by technical personnel based on experience.

[0142] In some embodiments, in response to situations where the consistency between the actual output performance of each virtual digital human and the audio signal does not meet preset conditions, the driving module can continue to synchronously drive multiple virtual digital humans to output at least one of facial expressions, motion expressions, and voice expressions that match the audio signal, based on the target performance sequence determined in step 320. For further details on this section, please refer to [link / reference]. Figure 3 And its related descriptions.

[0143] In some embodiments, in response to the condition that the consistency between the actual output performance of at least one virtual digital human and the audio signal meets a preset condition, the driving module can analyze subsequent audio signals starting from the current moment, re-extract multi-dimensional music features, and then regenerate the target performance sequence; then, based on the new target performance sequence, simultaneously drive multiple virtual digital humans to output at least one of facial expressions, motion expressions, and voice expressions that match the audio signal. For further explanation of this section, please refer to [link / reference]. Figure 3 And its related descriptions.

[0144] In some embodiments of this specification, adaptive error correction is achieved by real-time monitoring of the consistency between the actual output performance and the audio signal and introducing a regeneration mechanism triggered by preset conditions. This eliminates the cumulative error caused by clock drift or load, and can accurately identify abnormal situations such as music rhythm lag and emotional expression mismatch, quickly triggering recalculation to restore synchronization. This greatly enhances the robustness of the system under occasional stuttering and ensures the deep unity of audiovisual perception.

[0145] Figure 5 This is an exemplary schematic diagram illustrating the determination of the main group and the accompanying group according to some embodiments of this specification.

[0146] In some embodiments, such as Figure 5 As shown, the determining module can generate one or more virtual digital humans 510 belonging to the main group and one or more virtual digital humans 520 belonging to the accompanying group; based on the melody features 530 and / or emotional features 540 in the multidimensional music features, it determines the target performance sequence 550 corresponding to one or more virtual digital humans in the main group; based on the beat features 560 in the multidimensional music features, it determines the target performance sequence 570 corresponding to one or more virtual digital humans in the accompanying group.

[0147] For more information on multidimensional musical characteristics, melodic characteristics, emotional characteristics, rhythmic characteristics, and target performance sequences, please refer to [link to relevant documentation]. Figure 3 And its related descriptions.

[0148] A main group refers to a collection of virtual digital humans who undertake the primary performance tasks and are responsible for conveying the core emotions and melodic information of the music. Each virtual digital human in a main group is bound to a unique main role tag. For example, the main role tag is lead vocalist P.

[0149] In some embodiments, the determining module can determine the number of virtual digital humans in the subject group in a variety of ways. For example, the determining module can obtain a preset value as the number of virtual digital humans in the subject group.

[0150] In some embodiments, the number of virtual digital humans in the subject group can be positively correlated with the melodic complexity of the audio signal. For example, the greater the melodic complexity, the greater the number of virtual digital humans in the subject group.

[0151] Melodic complexity refers to the frequency and amplitude of pitch changes over time, reflecting the richness and expressiveness of music. In some embodiments, the determining module can statistically analyze the pitch and frequency of the audio signal, and calculate the entropy value of the pitch distribution using the Shannon entropy formula, which serves as the melody complexity. A higher entropy value indicates more unpredictable pitch changes and thus higher melody complexity.

[0152] A companion group refers to a collection of virtual digital humans that perform auxiliary tasks and cooperate with the main group to complete the overall performance. Each virtual digital human in a companion group is bound to a unique companion role tag. For example, the companion role tag is "Dancer A".

[0153] In some embodiments, the determining module can determine the number of virtual digital humans in the accompanying group in a variety of ways. For example, the determining module can obtain a preset value as the number of virtual digital humans in the accompanying group.

[0154] In some embodiments, the number of virtual digital humans in the accompaniment group can be positively correlated with the energy density of the audio signal. For example, the higher the energy density, the more virtual digital humans are in the accompaniment group.

[0155] Energy density refers to the average energy of an audio signal per unit time, reflecting the loudness, intensity, and impact of music. In some embodiments, the determining module can determine the energy density by calculating the short-time energy or root mean square (RMS) of the audio signal. For example, the determining module divides the audio signal into multiple time windows (e.g., 20ms-50ms), calculates the sum of squares of the amplitude of the audio signal within each time window, and divides the sum of squares of the amplitudes by the length of the time window to obtain the energy density of each time window. Based on the energy density of each time window, the determining module determines the number of virtual digital humans in the accompanying group during the corresponding time period.

[0156] In some embodiments, the determining module can determine the target performance sequence corresponding to one or more virtual digital humans in the subject group by looking up a table, based on melody features and / or emotional features in the multidimensional music features. For example, the determining module obtains a third preset table, which stores the mapping relationship between combinations of emotional features and melody features and target performance sequences. The determining module can look up the third preset table based on melody features and emotional features to obtain the target performance sequence corresponding to one or more virtual digital humans in the subject group. Taking the existence of multiple virtual digital humans in the subject group as an example, the target performance sequence corresponding to each virtual digital human in the subject group determined by the table lookup method can be the same. The third preset table can be set by technicians based on experience.

[0157] In some embodiments, the determining module can determine the target performance sequence corresponding to one or more virtual digital humans in the accompaniment group by looking up a table based on the beat features in the multidimensional music features. For example, the determining module obtains a fourth preset table, which stores the mapping relationship between beat features and target performance sequences. The determining module searches the fourth preset table based on the beat features to obtain the target performance sequence corresponding to one or more virtual digital humans in the accompaniment group. Taking the presence of multiple virtual digital humans in the accompaniment group as an example, the target performance sequence corresponding to each virtual digital human in the accompaniment group determined by the table lookup method can be the same. The fourth preset table can be set by a technician based on experience.

[0158] In some embodiments, if both the main group and the accompaniment group contain multiple virtual digital humans and the target performance sequence corresponding to each virtual digital human is different, the determining module can determine the target performance sequence corresponding to each virtual digital human in the main group and the accompaniment group based on multidimensional music features. For further explanation of this section, please refer to [link / reference]. Figure 3 And its related descriptions.

[0159] In some embodiments of this specification, multiple virtual digital humans are divided into a main group and an accompanying group, and driven differently based on melody features, emotional features, and rhythm features, respectively. This achieves efficient collaboration of multi-group performances, while enabling the number of virtual digital humans and the target performance sequence in the main group and the accompanying group to be dynamically adjusted in real time according to changes in melody complexity and energy density. This overcomes the rigidity of static grouping and greatly improves the flexibility of the virtual performance lineup and its compatibility with multi-dimensional musical features.

[0160] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.

[0161] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A method for driving multiple virtual digital humans, characterized in that, The method includes: Analyze the audio signal and extract its multidimensional musical features; Based on the multidimensional music features, a target performance sequence is determined for each virtual digital human, and the target performance sequence includes at least one of facial expression sequence, action sequence and voice sequence. Based on the target performance sequence, multiple virtual digital humans are synchronously driven to output at least one of facial expressions, motion expressions, and voice expressions that match the audio signal.

2. The method according to claim 1, characterized in that, The target performance sequence includes the facial expression sequence, the action sequence, and the speech sequence. The step of synchronously driving multiple virtual digital humans to output at least one of the facial expression, action, and speech performances matching the audio signal based on the target performance sequence includes: Based on the target performance sequence, the multiple virtual digital humans are synchronously driven to output action performances, facial expressions, and voice performances that match the audio signal.

3. The method according to claim 1, characterized in that, The determination of the target performance sequence for each virtual digital human based on the multidimensional music features includes: Based on the emotional features in the multidimensional music features, at least one of the facial expression sequence and the voice sequence corresponding to each virtual digital human is determined.

4. The method according to claim 1, characterized in that, The step of synchronously driving multiple virtual digital humans to output at least one of facial expressions, motor expressions, and voice expressions that match the audio signal based on the target performance sequence further includes: Monitor the consistency between the actual output performance of each virtual digital human and the audio signal; In response to the consistency meeting the preset conditions, the target performance sequence is regenerated, and based on the target performance sequence, the plurality of virtual digital humans are synchronously driven to output at least one of the facial expressions, the action expressions, and the voice expressions that match the audio signal.

5. The method according to claim 1, characterized in that, The determination of the target performance sequence for each virtual digital human based on the multidimensional music features includes: Generate one or more virtual digital persons belonging to the subject group and one or more virtual digital persons belonging to the companion group; Based on the melody features and / or emotional features in the multidimensional music features, determine the target performance sequence corresponding to the one or more virtual digital humans in the subject group; Based on the beat features in the multidimensional music features, the target performance sequence corresponding to the one or more virtual digital humans in the accompaniment group is determined.

6. A driving system for multiple virtual digital humans, characterized in that, The system includes: The analysis module is configured to analyze the audio signal and extract multidimensional musical features of the audio signal; The determining module is configured to determine a target performance sequence for each virtual digital human based on the multidimensional music features, wherein the target performance sequence includes at least one of an expression sequence, an action sequence, and a speech sequence; and The driving module is configured to synchronously drive multiple virtual digital humans to output at least one of facial expressions, motion expressions, and voice expressions that match the audio signal, based on the target performance sequence.

7. The system according to claim 6, characterized in that, The target performance sequence includes the facial expression sequence, the action sequence, and the speech sequence, and the driving module is further configured to: Based on the target performance sequence, the multiple virtual digital humans are synchronously driven to output action performances, facial expressions, and voice performances that match the audio signal.

8. The system according to claim 6, characterized in that, The determining module is further configured to: Based on the emotional features in the multidimensional music features, at least one of the facial expression sequence and the voice sequence corresponding to each virtual digital human is determined.

9. The system according to claim 6, characterized in that, The driver module is further configured to: Monitor the consistency between the actual output performance of each virtual digital human and the audio signal; In response to the consistency meeting the preset conditions, the target performance sequence is regenerated, and based on the target performance sequence, the plurality of virtual digital humans are synchronously driven to output at least one of the facial expressions, the action expressions, and the voice expressions that match the audio signal.

10. The system according to claim 6, characterized in that, The determining module is further configured to: Generate one or more virtual digital persons belonging to the subject group and one or more virtual digital persons belonging to the companion group; Based on the melody features and / or emotional features in the multidimensional music features, determine the target performance sequence corresponding to the one or more virtual digital humans in the subject group; Based on the beat features in the multidimensional music features, the target performance sequence corresponding to the one or more virtual digital humans in the accompaniment group is determined.