A virtual digital human driving method, system, device and medium

By analyzing audio signals to obtain musical features and generating visual performance data, the problem of synchronizing virtual digital humans with audio content has been solved, achieving efficient audiovisual interaction and enhanced immersion.

CN122492900APending Publication Date: 2026-07-31HANSONG NANJING TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANSONG NANJING TECH LTD
Filing Date
2026-05-12
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve precise synchronization between virtual digital humans and audio content, resulting in limited audiovisual interaction effects and immersion.

Method used

By analyzing audio signals, musical characteristics are obtained, and visual performance data of virtual digital humans is generated based on these characteristics, including virtual appearance, movement sequences and facial expressions. The virtual digital humans are then controlled to display visuals while the audio signals are playing.

Benefits of technology

It achieves precise matching and synchronization of the virtual digital human's image, movements, music rhythm, and emotions, improving the consistency and smoothness of the audiovisual presentation and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492900A_ABST
    Figure CN122492900A_ABST
Patent Text Reader

Abstract

This specification provides a virtual digital human driving method, system, device, and medium. The method includes: analyzing an audio signal to obtain the music features corresponding to the audio signal; generating visual performance data of the virtual digital human based on the music features, wherein the visual performance data includes at least one of virtual appearance, action sequence, and facial expression; and synchronously controlling the virtual digital human to perform visual display according to the visual performance data while the audio signal is playing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of virtual digital humans, and in particular to a virtual digital human driving method, system, device and medium. Background Technology

[0002] With the development of artificial intelligence, virtual reality, and augmented reality technologies, the demand for virtual digital humans in fields such as virtual entertainment, home furnishing, and advertising is growing. How to achieve deep integration and interaction between virtual digital humans and the music environment, providing users with a richer and more personalized immersive audiovisual experience, has become a key focus of current technological research and development. Currently, existing technologies struggle to establish an efficient dynamic link between audio analysis, virtual avatar generation, and real-time interaction, failing to achieve accurate synchronization between virtual digital humans and audio content, resulting in limited audiovisual interaction effects and immersion.

[0003] Therefore, there is a need to propose a virtual digital human driving method, system, device, and medium to achieve synchronization between the visual performance and audio content of the virtual digital human. Summary of the Invention

[0004] This specification provides one or more embodiments of a virtual digital human driving method, the method comprising: analyzing an audio signal to obtain music features corresponding to the audio signal; generating visual performance data of the virtual digital human based on the music features, the visual performance data including at least one of virtual appearance, action sequence and facial expression; and synchronously controlling the virtual digital human to perform visual display according to the visual performance data while the audio signal is playing.

[0005] This specification provides one or more embodiments of a virtual digital human driving system, the system including an analysis module, a processing module, and a display module; the analysis module is configured to analyze an audio signal to obtain the music features corresponding to the audio signal; the processing module is configured to generate visual performance data of the virtual digital human based on the music features, the visual performance data including at least one of virtual appearance, action sequence, and facial expression; the display module is configured to synchronously control the virtual digital human to perform visual display according to the visual performance data when the audio signal is played.

[0006] This specification provides a virtual digital human driving device according to one or more embodiments. The device includes at least one processor and at least one memory. The at least one memory is used to store computer instructions. The at least one processor is used to execute at least a portion of the computer instructions to implement the virtual digital human driving method described in the above embodiments.

[0007] This specification provides one or more embodiments of a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions from the storage medium, the computer executes the virtual digital human driving method described in the above embodiments. Attached Figure Description

[0008] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0009] Figure 1 These are schematic diagrams illustrating application scenarios of the virtual digital human driving system according to some embodiments of this specification; Figure 2 This is an exemplary block diagram of a virtual digital human driving system according to some embodiments of this specification; Figure 3 This is an exemplary flowchart of a virtual digital human driving method according to some embodiments of this specification; Figure 4 This is an exemplary flowchart illustrating the adjustment of multiple action subsequences according to some embodiments of this specification; Figure 5 This is an exemplary flowchart illustrating the adjustment of visual performance data according to some embodiments of this specification. Detailed Implementation

[0010] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the linguistic context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0011] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0012] Unless the context explicitly indicates an exception, words such as "a," "an," "a kind," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0013] Figure 1 This is a schematic diagram illustrating the application scenarios of a virtual digital human driving system according to some embodiments of this specification.

[0014] In some embodiments, the virtual digital human driving system can automatically correlate audio signals with the visual performance of the virtual digital human, and ensure that the visual display of the virtual digital human is synchronized with the audio playback. This can be widely applied in scenarios such as home life, short video creation, and emotional therapy. For example, in a home setting, users can play audio through devices such as tablets and speakers, and the virtual digital human can simultaneously display actions and facial expressions matching the audio on display devices (such as tablets and TVs), enhancing the user's audiovisual experience. In short video creation scenarios, the virtual digital human can present performance content that matches the background music, lowering the barrier to content creation. In emotional therapy scenarios, the virtual digital human can present soothing visual expressions that match calming music, assisting users in regulating their emotions.

[0015] like Figure 1 As shown, the application scenario of the virtual digital human driving system (hereinafter referred to as application scenario 100) may include user terminal 110, network 120, processor 130 and audio device 140.

[0016] User terminal 110 is a terminal device that interacts with a user. For example, user terminal 110 may include a mobile phone (such as mobile phone 110-1), a portable computer (such as portable computer 110-2), etc. User terminal 110 can be used to present a virtual digital human 111.

[0017] The virtual digital human 111 can be a digital model generated by computer graphics technology, possessing an anthropomorphic appearance and behavioral characteristics. For example, a virtual digital human can be a virtual anchor, a virtual singer, etc. In some embodiments, the virtual digital human 111 can be visually displayed on a display device according to visual performance data. See [link to documentation] for an explanation of visual performance data. Figure 3 And its related descriptions.

[0018] In some embodiments, the display device may be a display screen built into the user terminal 110, or it may be an independently set external display device (such as an external TV, projector, etc.).

[0019] In some embodiments, the user terminal 110 or the display device has a rendering engine. A rendering engine refers to a functional module or software component used to render the visual effects of the virtual digital human in real time. For example, the rendering engine can employ a real-time graphics engine such as Unity or Unreal Engine, or be implemented based on graphics interfaces such as OpenGL or Vulkan, to support the real-time rendering and smooth display of complex visual effects of the virtual digital human.

[0020] In some embodiments, the user terminal 110 further includes an image acquisition device. The image acquisition device is used to acquire image streams related to the user. For example, the image acquisition device can be a camera, a depth camera, or an infrared image acquisition device, etc.

[0021] In some embodiments, the user terminal 110 further includes an audio acquisition device. The audio acquisition device is used to acquire the voice emitted by the user. For example, the audio acquisition device can be a microphone, a digital microphone array, etc.

[0022] Network 120 may include any suitable network capable of facilitating information and / or data exchange. In some embodiments, at least one component of application scenario 100 (e.g., user terminal 110, processor 130, audio device 140, etc.) may exchange information and / or data with at least one other component in application scenario 100 via network 120. For example, processor 130 may obtain relevant information about user terminal 110 and audio device 140 via network 120.

[0023] In some embodiments, network 120 can be any one or more of wired or wireless networks. For example, network 120 may include cable networks, fiber optic networks, telecommunications networks, cable connections, or any combination thereof. Network connections between components may employ one or more of the above methods. Network 120 can have various topologies, such as point-to-point, shared, or centralized, or a combination of multiple topologies. Network 120 may include one or more network access points.

[0024] The processor 130 is used to process data, information, and / or processing results related to the application scenario 100 of the virtual digital human driving system, and to execute program instructions based on this data, information, and / or processing results to perform one or more functions described in this specification. For example, the processor 130 can obtain data such as audio signals from the user terminal 110 via the network 120, and execute related instructions such as obtaining the music features corresponding to the audio signals and generating visual performance data of the virtual digital human.

[0025] For more information on audio signals and musical characteristics, please refer to [link / reference]. Figure 3 Related descriptions.

[0026] In some embodiments, processor 130 may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-chip processing device). By way of example only, processor 130 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), a microprocessor, or any combination thereof.

[0027] In some embodiments, the processor 130 may be integrated within the user terminal 110 or disposed independently locally or in the cloud. The processor 130 or the user terminal 110 may integrate a storage device.

[0028] The storage device can store data, instructions, and / or any other information related to the virtual digital human driving system. The storage device may include mass storage, removable memory, or any combination thereof.

[0029] Audio device 140 refers to a terminal device used to play audio. For example, audio device 140 can be a speaker, loudspeaker, etc.

[0030] In some embodiments, the audio device 140 may be an audio device (such as a speaker) that is built into the user terminal 110 or a separately configured audio device (such as a separate speaker). When the audio device 140 is a separately configured audio device, it can receive audio signals sent by the user terminal 110 through the network 120 and perform playback.

[0031] In some embodiments, the application process of the virtual digital human driving system includes: the processor 130 acquiring the audio signal sent by the user terminal 110 through the network 120, analyzing the audio signal, and obtaining the corresponding music features. Based on the music features, the processor 130 generates visual performance data of the virtual digital human 111 and sends the visual performance data to the user terminal 110 or a display device. After receiving the visual performance data, the user terminal 110 or the display device can control the virtual digital human 111 to perform visual display according to the visual performance data. At the same time, the user terminal 110 or the processor 130 controls the audio device 140 to play the audio signal, thereby achieving time synchronization between the audio signal playback and the visual display of the virtual digital human 111.

[0032] For further explanation of the above content, please refer to [link / reference]. Figures 2 to 5 And its related descriptions.

[0033] Figure 2 This is an exemplary block diagram of a virtual digital human driving system according to some embodiments of this specification.

[0034] like Figure 2 As shown, the virtual digital human driving system 200 may include an analysis module 210, a processing module 220, and a display module 230. In some embodiments, the analysis module 210 and the processing module 220 may be integrated into the processor 130. The display module 230 may be... Figure 1 The display device described herein. For details regarding the processor 130 and the display device, see [link to documentation]. Figure 1 And its related descriptions.

[0035] In some embodiments, the analysis module 210 is configured to analyze the audio signal and obtain the music features corresponding to the audio signal.

[0036] In some embodiments, the processing module 220 is configured to generate visual performance data of the virtual digital human based on musical features.

[0037] In some embodiments, the display module 230 is configured to synchronously control the virtual digital human to perform visual display according to visual performance data while the audio signal is playing.

[0038] For further explanation of the above content, please refer to [link / reference]. Figures 3 to 5 And its related descriptions.

[0039] It should be noted that the above description of the virtual digital human driving system 200 and its modules is for convenience only and should not be construed as limiting this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principles of the system, may arbitrarily combine the various modules or construct subsystems connected to other modules without departing from these principles. In some embodiments, Figure 2 The analysis module 210, processing module 220, and display module 230 disclosed herein can be different modules within a single system, or a single module can implement the functions of two or more of the aforementioned modules. For example, the modules can share a single storage module, or each module can have its own separate storage module. Such variations are all within the scope of protection of this specification.

[0040] Figure 3 This is an exemplary flowchart of a virtual digital human driving method according to some embodiments of this specification. In some embodiments, process 300 may be executed by a processor (such as processor 130). Figure 3 As shown, process 300 includes the following steps: Step 310: Analyze the audio signal to obtain the corresponding music features.

[0041] Audio signals refer to data signals or electrical signals that carry sound information. For example, an audio signal can be the data signal or electrical signal corresponding to music that is being played or is about to be played.

[0042] In some embodiments, the processor may acquire audio signals in various ways. For example, the processor may acquire audio signals from a user terminal. Another example is that the processor may download an audio file that the user wants to play from the cloud and parse the audio file to acquire the audio signal. For descriptions of audio acquisition devices and user terminals, please refer to [link to relevant documentation]. Figure 1 Related descriptions.

[0043] Music features refer to structured sequence data used to describe music-related attributes. In some embodiments, music features may include at least one of the following: beat characteristics corresponding to each sub-segment of the audio signal, musical emotional attributes, musical emotional intensity, and musical style.

[0044] Sub-segments can be defined in several ways. For example, a sub-segment can be a time period obtained by the processor dividing the audio signal into equal segments based on a preset duration, such as 5 seconds or 10 seconds. Another example is a sub-segment that the processor divides the audio signal into segments based on the lyrics or paragraphs of a song.

[0045] Meter features refer to characteristics related to the beat of music. In some embodiments, meter features may include beat points (i.e., the moment corresponding to a beat) and beat frequency, etc. Beat frequency can be the number of times a beat occurs per unit time (e.g., 1 minute), used to characterize the speed or rate at which beats occur in music.

[0046] Musical emotional attributes refer to the types of emotions conveyed or implied by music. For example, musical emotional attributes can include excitement, joy, anger, sadness, etc.

[0047] Musical emotional intensity refers to data used to characterize the strength of the emotions conveyed or implied by music. In some embodiments, musical emotional intensity can be characterized by numerical values ​​or levels. For example, the processor can be set to a numerical range of 1-0, where a value closer to 1 indicates a stronger emotion, and a value closer to 0 indicates a weaker emotion.

[0048] Musical style refers to the genre or category to which a piece of music belongs. For example, musical styles can include pop, rock, classical, etc.

[0049] In some embodiments, the processor can analyze the audio signal by calling a pre-trained model or programming function library to obtain the corresponding audio features. For example, the processor can obtain beat features using functions such as Librosa.beat.beat_track. Another example is that the processor can obtain music style using pre-trained audio feature extractors (such as YAMNet) and classification layers. Yet another example is that the processor can obtain music emotional attributes using models such as MusicTagging Transformer. Furthermore, the processor can also obtain music emotional intensity using models such as PyTorch or TensorFlow. Besides the functions and models mentioned above, the processor can obtain the corresponding audio features of the audio signal using any other programming function library, deep learning network, pre-trained model, etc., that has audio signal processing capabilities; this specification does not specifically limit this.

[0050] In some embodiments, the processor can also call feature extraction functions from the programming function library to extract spectral features, and obtain music emotional attributes based on the spectral features using a classifier (such as support vector machine, random forest, etc.). For a description of spectral features, see [link to documentation]. Figure 4 And its related descriptions.

[0051] Step 320: Based on the musical characteristics, generate visual performance data for the virtual digital human.

[0052] For an explanation of virtual digital humans, please refer to [link / reference]. Figure 1 Related descriptions.

[0053] Visual performance data refers to the set of visual features that guide the visual display of a virtual digital human. In some embodiments, one sub-time period corresponds to one set of visual performance data, or multiple sub-time periods correspond to the same set of visual performance data.

[0054] In some embodiments, visual performance data may include virtual appearance, motion sequences, and facial expressions.

[0055] Virtual appearance refers to data related to the visual appearance of a virtual digital human. For example, virtual appearance can include the physical model of the virtual digital human (i.e., the virtual digital human's physique), clothing style, hairstyle, texture mapping, color scheme, etc.

[0056] A motion sequence refers to a series of continuous physical movements performed by a virtual digital human during audio playback. For example, a motion sequence can include the virtual digital human's limb positions, postures, and motion parameters. Postures can include limb extension, squatting, standing, and sitting postures.

[0057] Motion parameters refer to the parameters that control the movement performance of a virtual digital human. For example, motion parameters can include movement amplitude, movement speed, movement complexity, movement density, and space occupancy. Movement complexity is related to the number of bones involved in the movement; the more bones involved, the greater the movement complexity. Movement density can be the frequency of movement changes per unit of time. Space occupancy characterizes the area occupied by the movement in the virtual space in which the virtual digital human resides.

[0058] Virtual space refers to a digital three-dimensional environment in which virtual digital humans exist and move or interact.

[0059] In some embodiments, the processing device may determine the visual performance data corresponding to each sub-period by querying a preset parameter table based on music characteristics.

[0060] The preset parameter table records the correspondence between multiple sets of musical features and visual performance data. In some embodiments, the processor can preset and store the preset parameter table based on experience in matching musical features and visual performance data.

[0061] For example, matching experience can include matching musical characteristics such as sadness and low beat frequency with visual expression data that evokes sadness. Visual expression data that evokes sadness can be pre-set, such as including blue clothing, tearful facial expressions, sitting posture, and small, slow movements.

[0062] In some embodiments, the processor can also generate virtual appearances based on musical emotional attributes and musical style; generate action sequences based on beat features and musical emotional attributes; and generate facial expressions based on musical emotional attributes and musical emotional intensity.

[0063] In some embodiments, the processor can generate a virtual appearance by matching elements from a library of appearance elements based on musical emotional attributes and musical style. The library of appearance elements can be pre-set and includes elements such as the virtual digital human's physical model, clothing style, hairstyle, texture mapping, and color scheme, as well as corresponding musical emotional attribute tags or musical style tags for each element. For example, the processor first matches corresponding clothing styles, hairstyles, texture mappings, etc., based on musical style, and then matches corresponding color schemes, etc., based on musical emotional attributes, thereby generating the virtual appearance.

[0064] For example, classical music styles can be matched with formal dresses, long hair, etc., and the emotional attributes of "sad" music can be matched with blue or purple, etc.

[0065] In some embodiments, the processor can match a set of movements to a preset dance library based on the emotional attributes of the music, and determine the speed of the movements in the set based on beat characteristics, and use the set of movements with determined speeds as a sequence of movements. The preset dance library can be pre-set, including the correspondence between the emotional attributes of the music and the set of movements.

[0066] For example, when the music's emotional attribute is "exhilarating," the processor can query the preset dance library for the corresponding set of movements. The set of movements corresponding to "exhilarating" can be a set of movements with extended limbs and large movement parameters (such as amplitude, complexity, density, and space occupancy), such as street dance or rock headbanging. When the music's emotional attribute is "calm," the processor queries the preset dance library for the corresponding set of movements. The set of movements corresponding to "calm" can be a set of movements with a standing posture and small movement parameters (such as amplitude, complexity, density, and space occupancy), such as gentle swaying.

[0067] In some embodiments, the processor can time-scale the animation curve of the action set based on the beat frequency in the beat features to adjust the speed of each action in the action set. For example, if the beat frequency is 120 (i.e., one beat every 0.5 seconds), the processor times-scales the animation curve of the action set so that each action in the action set can be presented sequentially at time points of 0.5 seconds or integer multiples of 0.5 seconds. The animation curve is used to represent the correspondence between the change pattern of the virtual digital human transitioning from one action to another and time.

[0068] In some embodiments, the processor can also record the action sequence corresponding to each sub-time period as an action sub-sequence, and adjust multiple action sub-sequences based on adjustment parameters. For further details on this section, please refer to [link to documentation]. Figure 4 And its related descriptions.

[0069] In some embodiments, the processor can generate facial expressions based on musical emotional attributes and musical emotional intensity by querying a preset expression table. The preset expression table can be pre-set, including the correspondence between musical emotional attributes and deformation target points of the virtual digital human face, as well as the correspondence between musical emotional intensity and the weight coefficients of the deformation target points.

[0070] A deformation target point refers to the set of points that deform based on a standard facial model of a virtual digital human. The standard facial model can be a model of the virtual digital human's face in a calm state. Deformation target points can be pre-set; for example, technicians can set the corners of the mouth on the standard facial model to be upturned. The processor can store the set of points that cause the upturn as a deformation target point. The weight coefficient of the deformation target point can be used to determine the intensity of the deformation.

[0071] In some embodiments, the processor can determine the deformation target point corresponding to the music's emotional attribute by querying a preset expression table, and determine the weight coefficient corresponding to the deformation target point by querying the preset expression table based on the music's emotional intensity. For example, when the music's emotional attribute is "happy," the processor can query the preset expression table to determine deformation target points such as an upturned corner of the mouth. If the music's emotional intensity (e.g., 0.9) is higher than a first threshold, the processor can match the deformation target point in the preset expression table with the maximum weight parameter (e.g., 100%) to present a laughing facial expression. If the music's emotional intensity (e.g., 0.2) is lower than a second threshold, the processor can match the deformation target point in the preset expression table with the minimum weight parameter (e.g., 10%) to present a smiling facial expression. The first and second thresholds can be set by technicians based on historical experience, with the first threshold being greater than the second threshold.

[0072] In some embodiments, the processor can combine the virtual appearance, action sequence, and facial expression generated by the above method to obtain visual performance data.

[0073] In some embodiments of this specification, by extracting multi-dimensional features such as music beat, emotion, and style, the virtual appearance, motion time scaling, and facial deformation target point calculation of the virtual digital human are respectively mapped, thereby improving the fit between the visual performance of the virtual digital human and the music content, and realizing a refined and differentiated dynamic performance effect of the virtual digital human.

[0074] Step 330: While the audio signal is playing, the virtual digital human is synchronously controlled to perform visual display according to the visual performance data.

[0075] In some embodiments, the processor can send visual performance data for each sub-segment to the user terminal or display device. When the audio signal is playing, the processor can send rendering instructions to the user terminal or display device to synchronously control the virtual digital human on the user terminal or display device to perform visual display according to the visual performance data corresponding to the sub-segment, thus synchronizing the display time of the visual performance data with the playback time of the audio signal on the timeline. More information about display devices can be found in [link to relevant documentation]. Figure 1 Related descriptions.

[0076] User terminals or display devices can also autonomously and synchronously control the virtual digital human to perform visual displays according to the visual performance data corresponding to each sub-period.

[0077] Rendering commands are instructions used to control the rendering engine to render visual effects for a virtual digital human. For more information on rendering engines, see [link to documentation]. Figure 1 And its related descriptions.

[0078] In some embodiments, the processor may also determine a target time point based on music characteristics; determine the display time of visual performance data based on the target time point; and control the virtual digital human to perform visual display according to the visual performance data at the display time.

[0079] A target time point refers to a specific point in the audio stream that requires focused attention. In some embodiments, a target time point can be the start time of a sub-segment or a point within a sub-segment. A target time point can also be a point in time where the emotional attributes of the audio being played by the audio device change.

[0080] In some embodiments, the processor can determine the musical emotional attributes corresponding to each sub-segment based on musical features, and determine the time point at which the musical emotional attributes change (such as the start time of the sub-segment) as the target time point. The processor can also use the beat points in the beat features as the target time point.

[0081] The display moment refers to the theoretical moment when the virtual digital human begins to present visual performance data. A target time point corresponds to a display moment earlier than the target time point.

[0082] Understandably, due to the inherent delay between instruction sending and device execution, the virtual digital human's actual presentation of visual data occurs later than the processor's execution of rendering instructions. Therefore, if the processor sends rendering instructions at the target time, the actual display of the visual data will be later than the target time. However, if the processor sends rendering instructions at the display time (i.e., earlier than the target time), and after a certain delay, the actual display of the visual data will coincide with the target time on the timeline.

[0083] In some embodiments, the processor can determine the display time of the visual display data based on a target time point and a delay time. For example, for each target time point, the processor can subtract the target time point from the delay time, and the difference is the display time of the visual display data corresponding to the target time point (i.e., a time earlier than the target time point).

[0084] Latency can be the signal propagation delay caused by the performance of the communication link or the device itself. For example, latency can be the time it takes from when the processor sends rendering instructions to when the visual performance data of the virtual digital human is finally presented on the user terminal or display device.

[0085] In some embodiments, the processor may determine the current latency as the average of historical latency times.

[0086] In some embodiments, the processor controls the virtual digital human to perform visual display according to the visual performance data at the display time. This can be achieved by the processor issuing a rendering instruction to the rendering engine at the instant the system clock reaches the display time, so that after a delay, the moment when the virtual digital human actually presents the visual performance data can coincide with the target time point on the timeline.

[0087] In some embodiments of this specification, the target time point is determined by music characteristics, and the theoretical display time of the visual performance data is determined by combining system latency. Rendering of the virtual digital human is initiated when the system clock is about to reach the theoretical display time, achieving extremely low latency audio-visual synchronization response and ensuring the audio-visual consistency of the virtual digital human performance.

[0088] In some embodiments of this specification, visual performance data is generated based on music characteristics, and the display of the virtual digital human is controlled synchronously. This can automatically achieve precise matching and synchronization between the image, movements, and music rhythm and emotions of the virtual digital human, effectively improving the consistency and smoothness of the audiovisual presentation and enhancing the user experience.

[0089] In some embodiments, the processor can also generate spatial lighting parameters corresponding to the virtual space based on music features; and update the lighting effects of the virtual space based on the spatial lighting parameters when the audio signal is played.

[0090] For an explanation of virtual space, please refer to the relevant description in step 230.

[0091] Spatial lighting parameters refer to parameters used to control the rendering state of a 3D scene and lighting in a virtual space. In some embodiments, spatial lighting parameters may include the brightness and color of lights in the virtual space, the camera viewpoint, and the state of scene particles. The camera viewpoint can be the angle from which the virtual digital person and the virtual space are presented to the user.

[0092] Scene particles can be tiny visual units used to simulate objects in an environment or background (such as rain and snow particles to simulate rain and snow weather, fallen leaves and dust particles to simulate natural landscapes, and halos and smoke particles to simulate atmospheric effects). The state of scene particles can include their dynamic attributes (such as position and velocity) and visual parameters (such as color and texture).

[0093] In some embodiments, the processor can generate spatial lighting parameters corresponding to the virtual space based on music features. For example, the processor can construct a first target vector based on music features, match and obtain a first feature vector that meets the matching conditions with the first target vector from a first vector database, and use the label corresponding to the first feature vector as the spatial lighting parameters corresponding to the music features. The matching conditions may include the highest similarity between vectors. Vector similarity is negatively correlated with vector distance. Vector distance includes Euclidean distance, etc.

[0094] In some embodiments, the first vector database may be pre-configured based on historical data, including multiple first feature vectors constructed based on multiple historical music features. For each first feature vector, the processor may filter multiple historical playback processes corresponding to the first feature vector, and use the historical spatial lighting parameters used in the historical playback process with the highest user focus as the label of the first feature vector.

[0095] User focus refers to the degree of concentration a user has when viewing a virtual digital human. In some embodiments, the processor can acquire a user's facial image through the image acquisition device of the user terminal, extract features from the facial image using feature extraction algorithms or deep learning models, determine features such as the user's blinking frequency and head movement amplitude, and assign scores and perform weighted summation on the aforementioned features according to preset evaluation rules to obtain the user's focus.

[0096] In some embodiments, the preset evaluation rules can be that a lower blink frequency corresponds to a higher score, and a greater head movement amplitude corresponds to a lower score. The weights of features such as blink frequency and head movement amplitude can be preset based on historical experience. For a description of the image acquisition device, please refer to [link / reference]. Figure 1 And its related descriptions.

[0097] Lighting and shadow effects refer to the visual environment presented in virtual space. For example, lighting and shadow effects can include changes in the brightness of the environment, ambient lighting of specific colors, dynamic flickering of stage lights, and shadows cast by objects in the scene.

[0098] In some embodiments, when an audio signal is played, the processor can generate lighting control instructions based on spatial lighting parameters and send the lighting control instructions to the rendering engine to control the rendering engine to update the lighting effects of the virtual space according to the spatial lighting parameters.

[0099] In some embodiments of this specification, by generating spatial lighting parameters and synchronously updating the lighting effects of the virtual space, the environment of the virtual space can dynamically evolve with the music status, effectively enhancing the audiovisual integration and overall immersive experience of the overall visual presentation.

[0100] In some embodiments, the processor may also adjust the visual performance data based on user interaction data.

[0101] Interactive data refers to data generated based on user interactions, which can reflect the user's intent. In some embodiments, interactive data may include user voice, user gestures, user facial images, etc. One sub-time period corresponds to one piece of interactive data.

[0102] In some embodiments, the processor may acquire interactive data via an audio acquisition device and an image acquisition device, etc. For example, the processor may acquire the user's voice input signal via an audio acquisition device to obtain user speech, and may also acquire an image stream via an image acquisition device, recognizing and extracting the user's gestures and facial images based on the image stream. For a description of the audio acquisition device, please refer to [link to relevant documentation]. Figure 1 Related descriptions.

[0103] In some embodiments, for each sub-time period, the processor can adjust the visual performance data corresponding to that sub-time period based on the interaction data. For example, the processor can determine the user's display instructions based on the interaction data and adjust the visual performance data accordingly.

[0104] Explicit commands are instructions issued by the user to adjust specific parameters of visual performance data. For example, explicit commands may include changing the clothing style and hairstyle of the virtual avatar, or adjusting the posture of the virtual avatar.

[0105] For example, the processor can process the user's speech using technologies such as Natural Language Processing (NLP) and extract the display instructions contained therein. The processor can also match the display instructions corresponding to the user's gestures using a gesture command library.

[0106] In some embodiments, the gesture command library can be pre-set based on historical experience, including a variety of gestures and corresponding display instructions for each gesture. For example, if the user's gesture is to raise their hand, the corresponding display instruction can control the virtual digital human to change from a sitting posture to a standing posture, etc.

[0107] In some embodiments of this specification, a two-way interaction mechanism between the virtual digital human driving system and the user is constructed by adjusting visual performance data based on interactive data. This mechanism allows the virtual digital human to directly respond to external input signals, realizes real-time control and personalized customization of visual performance data, and improves the flexibility of human-computer interaction.

[0108] In some embodiments, the processor can also adjust the visual performance data based on the dominant mode. (For more information on this section, please refer to [link / reference]). Figure 5 And its related descriptions.

[0109] In some embodiments, the multiple action sub-sequences determined based on audio features do not take into account the different emotions and emotional expression methods expressed by the music in different sub-time periods, which leads to the problem of formulaic virtual digital human actions. Therefore, it is necessary to finely adjust the action sub-sequences in combination with the emotional features corresponding to different sub-time periods in order to improve the fit and expressiveness of the virtual digital human actions with emotional features.

[0110] Figure 4 This is an exemplary flowchart illustrating the adjustment of multiple action subsequences according to some embodiments of this specification. In some embodiments, process 400 may be executed by a processor (such as processor 130). Figure 4 As shown, process 400 includes the following steps: Step 410: Determine the emotional feature sequence corresponding to the audio signal based on the spectral characteristics of the audio signal.

[0111] Spectral features refer to characteristic data used to describe the distribution of sound frequencies. For example, spectral features may include audio spectrograms, chroma features, and spectral contrast. One sub-time period corresponds to one spectral feature. For more information on sub-time periods, please refer to [link to relevant documentation]. Figure 3 Related descriptions.

[0112] A sentiment feature sequence can be a data sequence composed of sentiment features corresponding to multiple sub-time periods.

[0113] Emotional features refer to feature data related to the emotions expressed by the audio played in a sub-segment. In some embodiments, emotional features may include audio structure, etc. Audio structure refers to the structural categories of audio in the music theory dimension, such as audio structure may include verse, chorus, intro, etc.

[0114] In some embodiments, the emotional features may further include musical emotional attributes corresponding to sub-time periods. For more information on musical emotional attributes, please refer to [link to relevant documentation]. Figure 3 Related descriptions.

[0115] In some embodiments, the processor can call feature extraction functions from a programming function library to extract spectral features. Based on these spectral features, the audio structure of each sub-segment can be determined using audio analysis functions. Feature extraction functions may include functions such as `librosa.stft`, `librosa.feature.chroma_stft`, and `librosa.feature.spectral_contrast`. Music analysis functions may include functions such as `librosa.segment.recurrence_matrix` and `librosa.segment.subsegment`. Specifically, `librosa.stft` can be used to extract audio spectrograms, `librosa.feature.chroma_stft` can be used to extract chroma features, and `librosa.feature.spectral_contrast` can be used to extract spectral contrast.

[0116] Step 420: Determine the adjustment parameters based on the sentiment feature sequence.

[0117] Adjustment parameters refer to the numerical values ​​used to adjust the motion parameters within a motion subsequence. For example, adjustment parameters may include the adjustment value for each parameter in the motion sequence. See [link to documentation] for an explanation of motion parameters. Figure 3 And its related descriptions.

[0118] In some embodiments, the processor can construct a target matching vector based on the emotional features corresponding to each sub-time period in the emotional feature sequence, match reference vector clusters in the style action library that satisfy preset conditions with the target matching vector, and determine the reference adjustment parameters corresponding to the reference vector clusters as the adjustment parameters of the action parameters corresponding to the sub-time periods. The preset conditions may include the target matching vector having the highest similarity to the center vector of the vector cluster.

[0119] In some embodiments, the processor can obtain a style action library through cluster analysis. For example, the processor constructs multiple cluster vectors based on multiple sets of historical sentiment features and historical action parameters, and then clusters these vectors using a clustering algorithm to obtain multiple vector clusters. The center vector of each vector cluster can be the mean vector of the multiple cluster vectors within the cluster, and the reference adjustment parameter corresponding to each vector cluster can be the mean of the labels of the multiple cluster vectors within the cluster. Here, a historical sub-stage corresponds to a set of historical sentiment features and historical action parameters.

[0120] Cluster vectors can be feature vectors constructed based on historical sentiment features and historical action parameters corresponding to the historical sub-stages in which users actually adjusted their action sub-sequences. The labels of the cluster vectors can be the adjustment parameters that users actually adjusted in the historical sub-stages of the action sub-sequence. Clustering algorithms include K-Means clustering, density-based clustering methods (DBSCAN), etc.

[0121] In some embodiments, the processor may also determine adjustment parameters based on a sequence of emotional features and performance stages.

[0122] A performance phase refers to a specific logical stage in the audio's structure. In some embodiments, a performance phase may include an intro phase, a build-up phase, and a final chorus phase. One sub-time period corresponds to one performance phase.

[0123] The beginning phase can be the time before the verse transitions to the chorus. The building-up phase can be the time before the climax of the chorus. The final chorus phase can be the time after the last chorus until the end or before the last verse.

[0124] In some embodiments, the processor can determine the performance stage corresponding to each sub-segment based on the audio structure corresponding to multiple sub-segments in the emotional feature sequence. For example, the processor can determine the order of appearance, start time, and end time of one or more verses and choruses based on multiple audio structures, thereby determining the performance stage corresponding to each sub-segment.

[0125] For example, the processor can identify the performance stage corresponding to the sub-segment before the chorus switches from the verse to the chorus as the starting stage; identify the performance stage corresponding to the sub-segment after the last chorus until the end or before the last verse appears as the final chorus stage; and identify the performance stage corresponding to the sub-segment after the chorus starts for a preset duration (such as 10 seconds) as the power-up stage.

[0126] In some embodiments, the processor may incorporate the performance stage corresponding to the sub-time period when constructing the target matching vector. Correspondingly, the processor may incorporate the historical performance stage corresponding to the historical sub-time period when constructing the clustering vector. Based on the new target matching vector, the processor may match reference vector clusters in the style action library that meet preset conditions with the target matching vector, and determine the reference adjustment parameters corresponding to the reference vector clusters as the adjustment parameters of the action parameters corresponding to the sub-time period.

[0127] In some embodiments of this specification, by combining emotional feature sequences and performance stages to determine adjustment parameters, the changing trends of the music structure are accurately matched and a motion adjustment mechanism is provided, thereby enabling the virtual digital human's physical performance to have a forward-looking nature based on the logic of music changes.

[0128] Step 430: Adjust multiple action subsequences based on the adjustment parameters.

[0129] In some embodiments, the processor may adjust multiple action parameters in an action subsequence based on adjustment parameters, such as adding the action parameters and adjustment values, or scaling the action parameters proportionally based on the adjustment values.

[0130] In some embodiments of this specification, emotional feature sequences are obtained based on music features and the spectral features of audio signals, and the parameters are adjusted by combining style action library mapping, thereby finely adjusting the action subsequences. This overcomes the formulaic problem of virtual digital human action generation and significantly improves the fit and expressiveness of virtual digital human actions with emotional features.

[0131] In some embodiments, in response to a switch in the audio structure, the processor can also smooth the adjustments of multiple motion subsequences based on a target smoothing algorithm. Smoothing can transform changes in motion parameters from abrupt changes to gradual transitions, preventing stuttering or stiffness in the virtual digital human's movements and achieving smooth motion transitions.

[0132] A target smoothing algorithm is an algorithm used to smoothly approximate the adjusted value of an action parameter from its current value. For example, target smoothing algorithms include linear interpolation algorithms, Bézier curve smoothing algorithms, spline interpolation algorithms, etc. The current value can be the action parameter determined in step 320. The adjusted value can be the action parameter after adjusting the action parameter determined in step 320 based on the adjustment parameter.

[0133] In some embodiments, the target smoothing algorithm may include smoothing time and smoothing curve.

[0134] Smoothing time refers to the duration required for a motion parameter to approximate its adjusted value from its current value. A smoothing curve characterizes the adjustment magnitude or value of the motion parameter at multiple points in time during the smoothing process. The horizontal axis represents multiple points in time, and the vertical axis represents the adjustment magnitude or value at a single point in time. In essence, the motion parameter can be smoothed by undergoing multiple adjustments from its current value to its adjusted value.

[0135] In some embodiments, the processor may determine the target smoothing algorithm based on musical features.

[0136] In some embodiments, the processor can construct a second target vector based on music features, match second feature vectors that meet matching conditions in a second vector database, and use the labels corresponding to the second feature vectors as the target smoothing algorithm. See [link to matching conditions] for details. Figure 3 And its related descriptions.

[0137] In some embodiments, the second vector database can be pre-constructed based on historical data, including multiple second feature vectors constructed based on multiple historical music features. The processor can filter multiple historical playback processes corresponding to the historical music features and use the historical smoothing algorithm used in the historical playback process with the highest user feedback value as the label of the second feature vector.

[0138] User feedback values ​​characterize a user's satisfaction with adjustments to a sequence of actions. The processor can extract features from the user's facial image using feature extraction algorithms or deep learning models, obtaining various facial expression features such as the amplitude of a smile and frown after adjustments to historical action sequences. These features are then assigned scores and weighted and summed according to pre-defined feedback evaluation rules to obtain the user feedback value. The weight of each facial expression feature can be pre-set based on historical experience.

[0139] In some embodiments, the feedback value evaluation rules can be preset. For example, the larger the smile, the higher the score corresponding to the smile; the larger the frown, the higher the score corresponding to the frown (the score corresponding to the frown is negative), etc.

[0140] In some embodiments of this specification, the target smoothing algorithm is dynamically determined based on the music characteristics, so that the transition method of the smoothing process can adapt to the music characteristics, thereby ensuring that the visual quality of the action switching is highly consistent with the current music style and significantly improving the visual expressiveness.

[0141] In some embodiments, the processor can use a target smoothing algorithm to calculate the smoothing time and smoothing curve, and update the action parameters hourly based on the adjustment values ​​at multiple times on the smoothing curve within the smoothing time, thereby completing the smoothing process of adjusting the action subsequence.

[0142] In some embodiments of this specification, a target smoothing algorithm is used to smooth the adjustment of the action subsequence during audio structure switching, so that the action parameters gradually transition during the switching, effectively avoiding the abrupt changes and mechanical feeling of the digital human's actions during music structure switching, and improving the continuity.

[0143] In some embodiments, since the expected values ​​of the visual performance data determined by the processor and the visual performance data reflected by the user's interaction data may differ, it is necessary to determine whether the final visual performance data is more inclined towards the visual performance data determined by the processor or the expected value of the visual performance data reflected by the user's interaction data.

[0144] Figure 5 This is an exemplary flowchart illustrating the adjustment of visual performance data according to some embodiments of this specification.

[0145] Step 510: Based on the interaction data, determine the user's emotional attributes, emotional intensity, and interaction clarity.

[0146] For more information on interactive data, musical emotional attributes, and visual performance data, please refer to [link / reference needed]. Figure 3 Related descriptions.

[0147] User emotional attributes refer to the types of emotions that users convey or imply. For example, user emotional attributes can include sadness, anger, happiness, excitement, etc.

[0148] User emotional intensity refers to the strength of the emotions conveyed or implied by a user. In some embodiments, user emotional intensity can be represented by numerical values ​​or levels, for example, user emotional intensity can be represented by a numerical range of 0-1, with values ​​closer to 1 (such as 0.9) indicating stronger emotions and values ​​closer to 0 (such as 0.2) indicating weaker emotions, etc.

[0149] Interaction clarity refers to an indicator used to measure the accuracy of interaction data. In some embodiments, interaction clarity can be represented by numerical values ​​or levels. For example, interaction clarity can be represented by a numerical range of 0-1, where the closer to 1 (e.g., 0.9) the more accurate the interaction data, and the closer to 0 (e.g., 0.2) the less accurate the interaction data.

[0150] In some embodiments, the processor can determine a user's emotional attributes, emotional intensity, and interaction clarity based on interaction data in various ways.

[0151] For example, the processor can extract facial features from user facial images in the interaction data using feature extraction models or algorithms such as the MediaPipe Face Mesh model. Based on the extracted facial features, it can perform emotion recognition using models such as DeepFace, and output the corresponding user emotion attributes and intensity. The processor can obtain the mean confidence score output by models such as DeepFace during emotion recognition and use the mean confidence score as the interaction clarity.

[0152] For example, the processor can also analyze user voice data in the interaction data using voice emotion recognition models such as Emotion2Vec+Large, output the corresponding user emotion attributes and user emotion intensity, and use the average confidence score output by the voice emotion recognition model as the interaction clarity.

[0153] like Figure 5 As shown, the processor can determine whether the user's emotional attribute differs from the music's emotional attribute. If the user's emotional attribute differs from the music's emotional attribute, the processor executes steps 520, 530, and 540. If the user's emotional attribute is the same as the music's emotional attribute, the processor executes step 550 (without adjusting the visual performance data).

[0154] Step 520: Determine the processor dominance coefficient based on the intensity of musical emotion, the structural importance of audio structure, the intensity of user emotion, and the clarity of interaction.

[0155] For an explanation of the intensity of musical emotion, please refer to [link / reference]. Figure 3 And its related descriptions.

[0156] Structural importance refers to the degree of importance of audio structure within the overall audio. In some embodiments, the structural importance of audio structures is preset, such as 0.3 for the verse and 0.7 for the chorus. More information on audio structure can be found in [link to relevant documentation]. Figure 4 Related descriptions.

[0157] The processor dominance coefficient is used to determine whether the decision-making power of the virtual digital human's actions is more biased towards the processor. The larger the processor dominance coefficient, the more the decision-making power of the virtual digital human's actions is biased towards the processor.

[0158] In some embodiments, the processor can determine the processor dominance coefficient in various ways based on musical emotional intensity, the structural importance of the audio structure, the user emotional intensity, and the clarity of the interaction. For example, the processor multiplies the musical emotional intensity by the structural importance of the audio structure to obtain a musical emotional score, and multiplies the user emotional intensity by the clarity of the interaction to obtain a user emotional score. The processor dominance coefficient can be the ratio of the musical emotional score to the total score (the sum of the musical emotional score and the user emotional score).

[0159] For example, the processor can pre-build a lookup table based on historical data, which contains a mapping relationship between multiple combinations of music emotional intensity, structural importance, user emotional intensity, and interaction clarity and the processor's dominance coefficient. By querying the lookup table, the processor dominance coefficient corresponding to the current combination of music emotional intensity, structural importance, user emotional intensity, and interaction clarity can be determined.

[0160] In some embodiments, the processor may also perform smoothing processing on the music sentiment intensity sequence, structural importance sequence, user sentiment intensity sequence, and interaction clarity sequence within a historical period based on a target smoothing algorithm, to determine the smoothed music sentiment score and the smoothed user sentiment score; and determine the processor dominance coefficient based on the smoothed music sentiment score and the smoothed user sentiment score.

[0161] The target smoothing algorithms mentioned here can include exponential moving average algorithms, simple moving average algorithms, etc. For more information on target smoothing algorithms, please refer to [link to relevant documentation]. Figure 4 And its related descriptions.

[0162] A historical period refers to a continuous time interval preceding the current moment. For example, a historical period could be 30 seconds before the current moment.

[0163] The music emotional intensity sequence, structural importance sequence, user emotional intensity sequence, and interaction clarity sequence are sequences formed by arranging the music emotional intensity, structural importance, user emotional intensity, and interaction clarity in chronological order for each sub-period within a historical period.

[0164] Smoothed music sentiment score refers to the music sentiment score for the current sub-period after smoothing the music sentiment intensity sequence and structural importance sequence. Smoothed user sentiment score refers to the user sentiment score for the current sub-period after smoothing the user sentiment intensity sequence and interaction clarity sequence.

[0165] In some embodiments, the processor can utilize a target smoothing algorithm to smooth the music sentiment intensity sequence, structural importance sequence, user sentiment intensity sequence, and interaction clarity sequence within a historical time period, obtaining new music sentiment intensity, structural importance, user sentiment intensity, and interaction clarity corresponding to the current sub-time period. Based on this data, the processor determines the music sentiment score and user sentiment score corresponding to the current sub-time period, i.e., the smoothed music sentiment score and the smoothed user sentiment score. The explanation regarding the determination of the music sentiment score and user sentiment score is provided above and in its related description.

[0166] In some embodiments, the processor can also determine the music emotion score and user emotion score corresponding to each sub-period within the historical period based on the music emotion intensity sequence, structural importance sequence, user emotion intensity sequence, and interaction clarity sequence within the historical period, and use a target smoothing algorithm to smooth the music emotion score and user emotion score corresponding to each sub-period within the historical period to obtain the new music emotion score and user emotion score corresponding to the current sub-period, i.e., smoothed music emotion score and smoothed user emotion score.

[0167] In some embodiments, the processor may use the ratio of the smoothed music sentiment score to the average total score (the sum of the smoothed music sentiment score and the smoothed user sentiment score) as the processor dominant coefficient.

[0168] In some embodiments, if there is no interaction data in the historical period, the above sequence is not smoothed, and the processor uses the processor dominance coefficient determined based on the music emotional intensity, structural importance, user emotional intensity and interaction clarity corresponding to the current sub-period.

[0169] In some embodiments of this specification, a target smoothing algorithm is used to smooth the relevant data sequence to obtain a smoothing score and calculate the processor dominance coefficient, which avoids frequent changes in the behavior of the digital virtual human caused by instantaneous input fluctuations or transient changes in music, and improves the stability of behavioral decisions.

[0170] Step 530: Determine the dominant mode of the virtual digital human based on the processor dominance coefficient. A dominant mode refers to the decision-making mode that determines the control of a virtual digital human's actions. In some embodiments, the dominant mode may include a user-dominated mode, a processor-dominated mode, and a combined mode of both. A combined mode refers to the user and the processor jointly controlling the visual performance of the virtual digital human.

[0171] In some embodiments, the processor can compare a processor dominance coefficient with a coordination interval to determine the dominance mode. For example, if the processor dominance coefficient is within the coordination interval, the processor can determine the dominance mode as a two-way coordination mode. If the processor dominance coefficient is less than the lower limit of the coordination interval, the processor determines the dominance mode as a user-dominated mode. If the processor dominance coefficient is greater than the upper limit of the coordination interval, the processor determines the dominance mode as an audio-dominated mode. The coordination interval can be preset based on historical experience, such as 0.3-0.7.

[0172] In some embodiments, in response to a dominant mode that is a collaborative mode, the processor maps user emotion-related data (including user emotion attributes and user emotion intensity) and music emotion-related data (including music emotion attributes and music emotion intensity) into a two-dimensional emotion space, obtaining corresponding user emotion vectors and audio emotion vectors. The processor uses the processor's dominant coefficient as the weight of the audio emotion vector and the difference between 1 and the processor's dominant coefficient as the weight of the user emotion vector. The two emotion vectors are then weighted and summed to obtain a fused emotion vector. The X-axis of the two-dimensional emotion space represents pleasure (from negative (sadness or anger, etc.) to positive (happiness or excitement, etc.)), and the Y-axis represents excitement (from low emotion intensity to high emotion intensity).

[0173] The fused emotion vector, obtained by weighting and fusing the user emotion vector and the audio emotion vector, includes feature values ​​representing the fused emotion attributes and intensity. In some embodiments, the processor can redetermine the virtual appearance, action sequence, and facial expressions in the visual performance data based on the emotion attributes and intensity in the fused emotion vector, combined with the music style and beat features in the music features, through the method in step 320.

[0174] In some embodiments, in response to the dominant mode being user-dominated mode, the processor via Figure 3 The method for adjusting visual performance data based on interactive data adjusts the visual performance data. In response to the dominant mode being audio-dominated, the processor ignores the interactive data and maintains the visual performance data determined in step 320 unchanged.

[0175] In some embodiments of this specification, by comprehensively calculating the processor dominance coefficient and determining the dominance mode based on multi-dimensional features, a dynamic allocation mechanism for control is realized when there is a conflict between the user and the music's emotional state. This effectively solves the problem of chaotic and rigid virtual human behavior logic caused by multimodal input command conflicts.

[0176] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.

[0177] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0178] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and are considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A virtual digital human driving method, characterized in that, The method includes: Analyze the audio signal to obtain the musical features corresponding to the audio signal; Based on the musical characteristics, visual performance data of the virtual digital human is generated, and the visual performance data includes at least one of virtual appearance, action sequence and facial expression; While the audio signal is playing, the virtual digital human is synchronously controlled to perform visual display according to the visual performance data.

2. The method as described in claim 1, characterized in that, The musical features include at least one of rhythmic features, musical emotional attributes, musical emotional intensity, and musical style; The visual performance data of the virtual digital human generated based on the music features includes at least one of the following: The virtual appearance is generated based on the musical emotional attributes and the musical style; The action sequence is generated based on the beat features and the musical emotional attributes; The facial expression is generated based on the musical emotional attributes and the intensity of the musical emotion.

3. The method as described in claim 1, characterized in that, The action sequence includes multiple action sub-sequences, and the method further includes: Based on the spectral characteristics of the audio signal, determine the emotional feature sequence corresponding to the audio signal; Based on the emotional feature sequence, the adjustment parameters are determined; Based on the adjustment parameters, the multiple action sub-sequences are adjusted.

4. The method as described in claim 3, characterized in that, The method further includes: In response to a change in audio structure, the adjustment of the multiple action subsequences is smoothed based on a target smoothing algorithm.

5. The method as described in claim 1, characterized in that, The method further includes: The visual performance data is adjusted based on user interaction data.

6. The method as described in claim 1, characterized in that, The step of synchronously controlling the virtual digital human to perform visual display according to the visual performance data while the audio signal is playing includes: Based on the aforementioned musical characteristics, the target time point is determined; Based on the target time point, determine the display time of the visual performance data; The virtual digital human is controlled to perform a visual display according to the visual performance data at the display time.

7. A virtual digital human driving system, characterized in that, The system includes an analysis module, a processing module, and a display module; The analysis module is configured to analyze the audio signal and obtain the music features corresponding to the audio signal; The processing module is configured to generate visual performance data of the virtual digital human based on the music features, the visual performance data including at least one of virtual appearance, action sequence and facial expression; The display module is configured to synchronously control the virtual digital human to perform visual display according to the visual performance data when the audio signal is played.

8. The system as described in claim 7, characterized in that, The musical features include at least one of rhythmic features, musical emotional attributes, musical emotional intensity, and musical style, and the processing module is further configured to: The virtual appearance is generated based on the musical emotional attributes and the musical style; The action sequence is generated based on the beat features and the musical emotional attributes; The facial expression is generated based on the musical emotional attributes and the intensity of the musical emotion.

9. A virtual digital human driving device, characterized in that, The device includes at least one processor and at least one memory; The at least one memory is used to store computer instructions; The at least one processor is configured to execute at least a portion of the computer instructions to implement the method of any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions. When the computer reads the computer instructions from the storage medium, the computer executes the method as described in any one of claims 1 to 6.