A voice-driven expression generation method, computer device and robot
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU YUNMU ZHIZAO TECH CO LTD
- Filing Date
- 2026-07-09
- Publication Date
- 2026-08-07
AI Technical Summary
[0008]为解决现有语音驱动表情生成中,宏观语义动作与微观生理动作因物理属性和时序特性不同而难以通过同一模型稳定生成,且简单叠加微观动作容易导致动作抖动、冲突或节律脱节的问题,本申请提供一种语音驱动表情生成方法、计算机设备及语音驱动表情生成机器人
第一,本申请将语音驱动表情生成过程中的宏观语义动作生成和微观生理动作生成进行解耦,使表情生成神经网络主要承担语音语义、语音韵律与嘴型、眉眼、脸颊等宏观面部动作之间的映射任务,而由独立的微观生理动作生成逻辑生成用于表征呼吸、微颤、微眼动或自然眨眼等微观生理动作的扰动作用。由此,可以避免神经网络在同一输出空间中同时学习语义相关性较强的宏观动作和随机性、弱语义相关性较强的微观动作,降低模型学习难度,减少微观生理动作被削弱为噪声或被生成为规律性伪影的情况,提高宏观语义表情和微观生理动作各自的生成稳定性。
Smart Images

Figure CN122531409A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent robot technology, and in particular to a voice-driven facial expression generation method, computer equipment, and robot. Background Technology
[0002] With the development of technologies such as virtual digital humans, embody intelligent robots, bionic companion robots, guide robots, and human-robot collaborative robots, voice-driven facial expression generation technology is widely used to generate facial expressions synchronized with speech based on speech content and state. This type of technology typically requires generating motion parameters for parts such as the mouth shape, jaw, eyebrows, eyes, cheeks, eyelids, and nose based on the input speech, so that the subject exhibits a more natural expressive state during vocalization, dialogue, or interaction.
[0003] In one existing implementation, voice-driven facial expression generation can be based on rules. For example, mouth movements can be generated based on the mapping relationship between phonemes and visual pixels, or blinking and eye movements can be generated based on preset time intervals and preset movement curves. This type of method has the advantages of simple implementation and controllable movements. However, since rules are usually difficult to fully express the complex influence of speech semantics, speech rhythm, and emotional state on facial movements, it is easy to lead to stiff mouth shape changes, insufficient expression levels, and difficulty in producing natural changes based on the differences of the same phoneme in different contexts and emotional states.
[0004] In another existing implementation, speech-driven facial expression generation can be based on neural networks. For example, speech features can be input into a neural network, which then predicts facial key points, blendshape parameters, expression weights, head pose, or control variables of robotic facial actuators. This type of method can learn the nonlinear mapping relationship between speech content, speech rhythm, and lip shape, jaw opening and closing, and macro-expressions, thus exhibiting good performance in lip-syncing and macro-expression generation.
[0005] However, facial movements include not only macroscopic movements strongly driven by speech and semantics, but also a large number of high-frequency, low-amplitude, weakly semantically related, or physiologically rhythmic microscopic movements, such as natural blinking, eyelid tremors, micro-eye movements, breathing fluctuations, and the subtle twitching of the nose, face, or jaw caused by breathing. These microscopic physiological movements differ significantly from macroscopic semantic movements such as lip-syncing and emotional eyebrow changes in terms of amplitude, frequency range, triggering method, physical continuity, and temporal regularity. If macroscopic semantic movements and microscopic physiological movements are treated as the same type of output in the same model for prediction, neural networks are prone to treating the more random and semantically weakly related microscopic movements as noise and weakening them, or generating overly regular micro-movement artifacts that lack natural variation, leading to problems such as local stillness of the subject being driven, lifeless eyes, and unnatural facial micro-movements.
[0006] Furthermore, directly superimposing preset micro-motions onto the final position, weights, or control values of the neural network output can easily create new problems. For example, micro-motions may be attached to facial movements like external shaking, lacking the inertia, elasticity, and damping continuity that facial tissues or mechanical actuators should have. Or, for example, when macro-movements require closing the eyes, half-closing the eyes, opening the mouth wide, or maintaining a specific emotional state, independently triggered natural blinking, micro-eye movements, or breathing micro-movements may conflict with the macro-movements, causing twitching, abrupt changes, or rhythmic discontinuities in facial expressions.
[0007] Therefore, how to enable neural networks to focus on generating macroscopic facial movements driven by speech semantics during the speech-driven expression generation process, while simultaneously generating and integrating microscopic physiological movements in a way that coordinates with the macroscopic movements, so that microscopic physiological movements no longer serve as simple positional shifts or weight superpositions, but participate in facial movement generation in a way that has physical continuity, has become a technical problem that needs to be solved. Summary of the Invention
[0008] To address the challenges in existing voice-driven facial expression generation, where macroscopic semantic actions and microscopic physiological actions differ in physical properties and temporal characteristics, making it difficult to stably generate them using the same model, and where simply superimposing microscopic actions can easily lead to jitter, conflict, or rhythmic discontinuity, this application provides a voice-driven facial expression generation method, a computer device, and a voice-driven facial expression generation robot.
[0009] This application provides a voice-driven facial expression generation method in a first aspect, comprising: acquiring a voice signal for driving an object to be driven, and extracting features from the voice signal to obtain voice features for characterizing voice semantics and / or prosodic states; inputting the voice features into an facial expression generation neural network to generate macroscopic semantic action parameters, the macroscopic semantic action parameters being used to characterize macroscopic facial movements driven by voice semantics; generating microscopic perturbation forces according to microscopic physiological action generation logic independent of the facial expression generation neural network, the microscopic perturbation forces being used to characterize the perturbation effect of microscopic physiological actions on facial movements; determining a target equilibrium state of a virtual mass-spring-damped system according to the macroscopic semantic action parameters, and using the microscopic perturbation forces as an external force acting on the virtual mass-spring-damped system. Force input, the target equilibrium state includes the target rest point and / or the spring rest length; the macroscopic semantic state is determined based on the speech features, the macroscopic semantic action parameters and / or the intermediate features of the expression generation neural network, and the generation parameters of the microscopic physiological action generation logic and / or the stiffness parameters and / or damping parameters of the virtual mass-spring-damped system are adjusted based on the macroscopic semantic state; the fused facial action control data is calculated based on the target equilibrium state, the microscopic disturbance force, the stiffness parameters and / or the damping parameters; the facial actuator of the object to be driven is controlled according to the facial action control data, or the facial key points, Blendshape parameters or facial animation model of the object to be driven are updated to generate facial expression actions corresponding to the speech signal.
[0010] Furthermore, the micro-perturbation force includes at least one of the following: respiratory fluctuation perturbation force, micro-tremor perturbation force, micro-eye movement perturbation force, and natural blinking perturbation force.
[0011] Furthermore, the breathing fluctuation disturbance force is generated according to the breathing rhythm, which is adjusted according to the speech activity detection results, speech pause status, speech speed and / or volume, so that the breathing amplitude decreases and the breathing frequency increases during the speaking state, and the breathing amplitude increases and the breathing frequency decreases during the pause state.
[0012] Furthermore, the muscle tension coefficient is determined based on the macroscopic semantic state, and the stiffness parameter and / or the damping parameter is adjusted based on the muscle tension coefficient.
[0013] Furthermore, the fused facial motion control data is obtained by solving the following dynamic relationship or its equivalent form: ,in, Indicates virtual quality. Indicates the damping parameter. Represents the stiffness parameter. This indicates the facial movement state after fusion. This represents the target equilibrium state determined by macroscopic semantic action parameters, or the macroscopic target action state corresponding to the target equilibrium state. This refers to the microscopic perturbation force.
[0014] Furthermore, the micro-perturbation force is generated based on the noise force field, periodic driving force, randomly triggered perturbation force, and / or micro-action primitive template.
[0015] Furthermore, before obtaining the fused facial motion control data, a bionic motion scheduling state machine is used to handle conflicts between the macroscopic semantic motion parameters corresponding to the macroscopic semantic motion parameters and the microscopic physiological motions corresponding to the microscopic perturbation forces. The conflict handling includes: when the macroscopic semantic motion meets a preset preemption condition, blocking the triggering of microscopic physiological motions in the corresponding facial region; when the macroscopic semantic motion meets a preset superposition condition, correcting the motion baseline of the microscopic physiological motions based on the current macroscopic motion state; and when the macroscopic semantic motion meets a preset coordination condition, triggering microscopic physiological motions that coordinate with the macroscopic semantic motions.
[0016] Furthermore, the speech start time of the speech signal is obtained based on speech activity detection, and the inspiratory peak in the breathing rhythm is adjusted to a preset time window before the speech start time, so that the breathing fluctuation disturbance force acts on the virtual mass-spring-damping system before the speech start time.
[0017] In a second aspect, this application provides a computer device including a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, implements the voice-driven facial expression generation method described in any of the preceding claims.
[0018] This application provides a voice-driven facial expression generation robot in a third aspect, comprising: a robot body; a facial actuator disposed on the robot body for driving the robot's face to generate facial expressions; a voice acquisition device or a voice receiving device for acquiring voice signals; a processor and a memory, wherein the memory stores a computer program, and the processor is connected to the voice acquisition device or voice receiving device and the facial actuator respectively; wherein, when executed by the processor, the computer program implements the voice-driven facial expression generation method described above, and controls the facial actuator to move according to the facial motion control data obtained by the voice-driven facial expression generation method, so as to generate robot facial expression actions corresponding to the voice signals.
[0019] The above technical solution has the following advantages compared to the existing technology: First, this application decouples the generation of macroscopic semantic actions and microscopic physiological actions in the speech-driven facial expression generation process. The facial expression generation neural network primarily undertakes the mapping task between speech semantics, speech prosody, and macroscopic facial actions such as mouth shape, eyebrows, eyes, and cheeks. Meanwhile, an independent microscopic physiological action generation logic generates perturbations representing microscopic physiological actions such as breathing, micro-tremors, micro-eye movements, or natural blinking. This avoids the neural network simultaneously learning semantically highly correlated macroscopic actions and random, weakly semantically correlated microscopic actions in the same output space, reducing the model's learning difficulty, minimizing the possibility of microscopic physiological actions being weakened into noise or generated as regular artifacts, and improving the generation stability of both macroscopic semantic expressions and microscopic physiological actions.
[0020] Second, this application transforms macroscopic semantic actions into a target equilibrium state within a virtual mass-spring-damped system, and converts microscopic physiological actions into disturbances acting on the system. This eliminates the need for macroscopic actions to be combined with microscopic actions through simple positional shifts, weight shifts, or the superposition of final control quantities. Instead, they are fused within the same virtual dynamic system through elastic, damping, and inertial responses. This reduces mechanical jitter, abrupt changes, and patch-like micro-motions caused by directly superimposing micro-motions, resulting in fused facial movements with better physical continuity and naturalness.
[0021] Third, this application also utilizes speech features, macroscopic action parameters, or intermediate features of neural networks to determine the macroscopic semantic state, and adjusts the generation intensity, rhythm parameters, or physical fusion parameters of microscopic physiological actions based on this macroscopic semantic state. This allows microscopic physiological actions to change with variations in speaking state, pauses, speech rate, volume, emotional intensity, or muscle tension, thereby avoiding independent playback of microscopic actions at fixed frequencies and amplitudes and improving the coordination between macroscopic semantic expression and microscopic physiological rhythm. Through the above processing, this application can improve the lifelikeness, physical realism, and behavioral coordination of facial micro-movements while ensuring the synchronization of speech-driven facial expressions, reducing unnatural facial expressions caused by local stillness, stiff micro-movements, or mismatch between macroscopic and microscopic movements.
[0022] The above technical solution also has the following advantages: By further classifying micro-perturbation forces into different types such as respiratory fluctuations, micro-tremors, micro-eye movements, and natural blinking, a more suitable generation method can be adopted for the frequency, amplitude, and triggering rules of different micro-physiological movements, thereby improving the sense of hierarchy and biomimetic expressiveness of micro-movements; by adjusting the breathing rhythm according to the state of speech activity, pause, speech rate, or volume, the breathing micro-movements can be adapted to the speaking state, so that the object to be driven presents different breathing fluctuation effects in continuous speaking, short pauses, or relaxed states; by adjusting stiffness or damping using muscle tension, facial movements can present different elasticity and damping responses in different states such as intense expression, calm expression, or fatigued expression, thereby enhancing the subtle changes in facial expressions; by using noise force fields, periodic driving forces, random triggering perturbation forces, or micro-movement primitive templates to generate micro-perturbation forces, micro-tremors and breathing can be taken into account. This application addresses the varying needs of actions like blinking and micro-eye movements in terms of randomness, periodicity, and template-based generation, improving the adaptability of micro-disturbance generation methods. By setting up a bionic motion scheduling state machine, it can preempt, superimpose, or coordinate actions that conflict with micro-movements such as natural blinking and micro-eye movements, reducing motion twitches and unreasonable superpositions. By establishing a correlation between the start of speech and the phase of the breathing rhythm, it can make breathing fluctuations play a preparatory role before the onset of speech, thereby enhancing details such as inhalation before speaking and slight facial twitching, reducing the problem of the mouth shape being out of sync with the breathing rhythm. By enabling facial motion control data to be used to control the robot's facial actuators, or to update facial key points, blendshape parameters, or facial animation models, this application is also compatible with various application scenarios such as bionic robots, virtual digital humans, 3D avatars, and interactive digital characters. Attached Figure Description
[0023] Figure 1 A flowchart illustrating a voice-driven facial expression generation method provided in an embodiment of this application; Figure 2 This application provides a schematic diagram of the architecture of a voice-driven facial expression generation system. Figure 3 A schematic diagram illustrating the decoupling and fusion of macroscopic semantic actions and microscopic physiological actions provided in an embodiment of this application; Figure 4 A schematic diagram illustrating the separation and heterogeneous characterization of macroscopic motion data and microscopic physiological motion data provided in the embodiments of this application; Figure 5 This is a schematic diagram illustrating the generation and synthesis of microscopic perturbation forces in an embodiment of this application. Figure 6 A schematic diagram of a virtual mass-spring-damping system provided in an embodiment of this application; Figure 7This is a schematic diagram of a biomimetic motion scheduling state machine provided in an embodiment of this application; Figure 8 This is a schematic diagram of the speech initiation time and respiratory rhythm phase anchoring provided in the embodiments of this application; Figure 9 This is a schematic diagram of the structure of a voice-driven facial expression generation robot provided in an embodiment of this application.
[0024] Explanation of reference numerals in the attached figures: 100. Speech Acquisition Unit; 110. Speech Feature Extraction Unit; 120. Facial Expression Generation Neural Network; 130. Microscopic Physiological Action Generation Unit; 140. State Modulation Unit; 150. Physical Fusion Unit; 160. Drive Output Unit; 170. Object to be Driven; 200. Macro-Micro Data Separation Unit; 210. Low-Pass Filtering Unit; 220. Residual Extraction Unit; 230. Extremely Low Frequency Extraction Unit; 240. Eyelid Event Segmentation Unit; 250. Macroscopic Data Storage Unit; 260. Microscopic Data Storage Unit; 300. Respiratory Disturbance Dynamics Generation Unit; 310 320 Noise disturbance force generation unit; 330 Random trigger disturbance force generation unit; 340 Template disturbance force generation unit; 400 Microscopic disturbance force synthesis unit; 410 Target end; 420 Spring; 430 Damper; 500 Virtual mass block; 600 Bionic motion scheduling state machine; 600 Voice start time; 610 Preset time window; 620 Inhalation peak value; 700 Robot body; 710 Voice acquisition device or voice receiving device; 720 Processor; 730 Memory; 740 Facial actuator; 750 Robot face. Detailed Implementation
[0025] The embodiments of this application will now be described with reference to the accompanying drawings. It should be understood that the following embodiments are used to illustrate the technical solutions of this application, and not to limit the scope of protection of this application. Where there is no conflict, the technical features in the embodiments of this application can be combined with each other. For those skilled in the art, various substitutions, modifications, or equivalent adjustments can be made to the following embodiments without departing from the technical concept of this application.
[0026] In the description of this application, unless otherwise expressly defined, terms such as "comprising," "including," and "having" should be understood as open-ended expressions, used to indicate the presence of the described features, steps, units, modules, components, or combinations thereof, but not excluding the possibility of the presence or addition of other features, steps, units, modules, components, or combinations thereof. "At least one" means one or more; "multiple" means two or more. "And / or" is used to indicate any one, any multiple, or all combinations of the objects listed before and after it, for example, "A and / or B" can indicate only A, only B, or both A and B. For ordinal numbers such as "first" and "second" used in this application, they are only used to distinguish objects of the same or similar types and should not be construed as indicating an order of importance, chronological order, spatial order, or quantity limitation. The terms "connection," "coupling," or "linked" used in this application, unless otherwise expressly defined, can be a direct connection or an indirect connection implemented through intermediate components, communication links, signal lines, buses, networks, or other structures; for connections between data processing modules or functional units, they can also be understood as the transmission relationship of data, control commands, or parameters. The "unit," "module," and "logic" described in this application can be implemented by hardware, software, firmware, or a combination thereof. For example, they can be implemented by a processor executing a program in memory, or by a dedicated circuit, controller, computing module, or multiple functional components working together. The term "preset" in this application can refer to pre-setting, calibrating, training, storing, or configuring before executing a corresponding step, or it can refer to updating or selecting based on the type of object to be driven, facial expression style, voice state, or control requirements. Unless the execution order of steps is explicitly defined, the step numbers and flow directions in the accompanying drawings in the embodiments of this application are mainly for illustrative purposes and should not be construed as an absolute restriction on the execution order of steps; some steps can be executed in parallel, alternately, or with adjusted execution order without affecting the data dependencies and processing logic.
[0027] In this embodiment, the object to be driven can be a virtual digital human, a 3D avatar, an interactive digital character, an embodied intelligent robot, a bionic companion robot, a guide robot, or other objects capable of generating facial expressions and movements based on control data. The facial movements of the object to be driven can be achieved through facial actuators, or through facial key points, Blendshape parameters, a skeletal controller, mesh vertices, or a facial animation model.
[0028] In this application embodiment, macroscopic semantic actions refer to facial movements strongly driven by phonological semantics, phonological rhythm, or emotional state, such as changes in mouth shape, jaw opening and closing, stretching of the corners of the mouth, raising of eyebrows, frowning, cheek lifting, emotional eye closing, expressions of surprise, or expressions of anger. Microscopic physiological actions refer to subtle movements whose frequency, amplitude, triggering method, or temporal pattern differs from macroscopic semantic actions, such as natural blinking, eyelid tremors, micro-eye movements, micro-saccades, breathing fluctuations, and the slight twitching of the nose, facial skin, or jaw caused by breathing.
[0029] In this embodiment, the micro-perturbation force refers to the external force input used to characterize the perturbation effect of micro-physiological movements on facial movements. The micro-perturbation force is not directly superimposed on the macro-movement as facial position offset, blendshape weight offset, or final actuator control quantity, but rather as an external force input acting on the virtual mass-spring-damped system, allowing micro-physiological movements to participate in facial movement generation through elastic, damping, and inertial responses. The target equilibrium state refers to the target state of the virtual dynamic system determined by macro-semantic action parameters, which can be the target rest point, spring rest length, target posture, target displacement, target angle, target blendshape weight, or an equivalent macro-movement state.
[0030] Figure 1 This is a flowchart illustrating a voice-driven facial expression generation method provided in this application. Figure 1 As shown, the method in this embodiment may include the following steps.
[0031] S101, acquire the voice signal used to drive the object to be driven.
[0032] In one embodiment, the speech signal can be obtained from a microphone, microphone array, robot speech acquisition device, remote communication interface, audio file reading interface, or real-time speech streaming interface. The speech signal can be real-time speech from when a user interacts with a robot or virtual digital human, or it can be a pre-recorded speech segment. The speech signal can undergo preprocessing such as resampling, noise reduction, endpoint detection, and loudness normalization to facilitate subsequent speech feature extraction.
[0033] S102, extract features from the speech signal to obtain speech features used to characterize speech semantics and / or prosodic states.
[0034] In one embodiment, speech features may include at least one of acoustic features, semantic features, and prosodic features. Acoustic features may include Mel spectrum, MFCC, energy envelope, fundamental frequency, volume, zero-crossing rate, or spectrogram features; semantic features may include speech recognition text, phoneme sequences, phoneme duration, word-level or sentence-level semantic vectors; prosodic features may include speech rate, stress, pauses, speech activity detection results, speech start time, speech end time, emotion label, or emotion intensity. Speech features may be aligned with facial motion control frames at a preset frame rate, for example, to 25 frames / second, 30 frames / second, 50 frames / second, or other control frame rates.
[0035] S103 inputs speech features into the facial expression generation neural network to generate macroscopic semantic action parameters.
[0036] In one embodiment, the facial expression generation neural network can be a temporal neural network, a Transformer network, a convolutional temporal network, a recurrent neural network, a diffusion model, an encoder-decoder network, or other networks capable of generating facial expression action parameters based on speech features. The input to the facial expression generation neural network can be a sequence of continuous frame speech features, and the output can be a sequence of macroscopic semantic action parameters.
[0037] Macro-semantic action parameters can be used to describe macro-movements of the mouth shape, jaw, corners of the mouth, eyebrows, eye sockets, cheeks, nostrils, or other facial areas. For example, in a virtual digital human scenario, macro-semantic action parameters can be several macro-blendshape target weights; in a facial keypoint scenario, macro-semantic action parameters can be keypoint target positions; in a robot scenario, macro-semantic action parameters can be intermediate representations of the facial actuator target pose, target displacement, or target angle.
[0038] In this embodiment, the facial expression generation neural network is mainly used to generate macroscopic facial movements related to speech semantics and prosody. Microscopic physiological movements such as natural blinking, breathing fluctuations, micro-eye movements, and micro-tremors are not directly output as the final control quantity by the facial expression generation neural network, but are instead generated by independent microscopic physiological movement generation logic to produce corresponding microscopic perturbations. This reduces the difficulty for the neural network to simultaneously learn macroscopic semantic movements and random microscopic physiological movements.
[0039] S104 generates micro-perturbation forces based on micro-physiological action generation logic independent of the facial expression generation neural network.
[0040] In one embodiment, the micro-physiological action generation logic can be composed of one or more of rule-based generation logic, statistical generation logic, physical generation logic, noise generation logic, and action template generation logic. The micro-physiological action generation logic can receive speech activity state, pause state, speech rate, volume, emotional state, or macro-semantic action state as modulation input, but its generation process is separated from the macro-action prediction task of the facial expression generation neural network.
[0041] Microscopic perturbations can include at least one of the following: respiratory fluctuation perturbation, micro-tremor perturbation, micro-eye movement perturbation, and natural blinking perturbation. Microscopic perturbations can be represented as vectors of the same dimension as the facial motion state vector, or as local vectors acting only on a portion of the facial area or a portion of the actuators. For example, respiratory fluctuation perturbation can act on the nasal alae, cheeks, jaw, or chest cavity linkage area; natural blinking perturbation can act on the eyelid-related dimension; micro-eye movement perturbation can act on the dimension related to the minute movements of the eyeball or eyelids; and micro-tremor perturbation can act on areas such as the cheeks, corners of the mouth, and eye sockets.
[0042] S105, determine the target equilibrium state of the virtual mass-spring-damped system based on macroscopic semantic action parameters.
[0043] In one embodiment, macroscopic semantic action parameters can be mapped to a target equilibrium state. For Blendshape animation models, It can represent the macroscopic blendshape target weights; for facial keypoint models, It can represent the location of key target points; for robotic facial actuators, It can represent the target angle of a servo motor, the target tension of a cable, the equivalent displacement corresponding to the target current of a shape memory alloy, the target deformation of a soft-drive component, or the target state of other actuators. In this application, the target equilibrium state can be understood as the macroscopic action state tended by the virtual mass-spring-damping system when it is not subjected to microscopic disturbance forces.
[0044] S106, adjust the generation parameters of the microscopic physiological action generation logic and / or the stiffness and damping parameters of the virtual mass-spring-damped system according to the macroscopic semantic state.
[0045] In one embodiment, the macro-semantic state can be determined based on speech features, macro-semantic action parameters, and / or intermediate features of the facial expression generation neural network. The macro-semantic state may include speaking state, pause state, speech rate state, volume state, emotional intensity state, muscle tension state, fatigue state, or high arousal state, etc.
[0046] Micromotion generation parameters refer to the parameters used to control the micro-perturbation force output by the logic for generating micro-physiological movements. For example, the micromotion generation parameters corresponding to respiratory fluctuations may include respiratory amplitude, respiratory rate, respiratory phase, inspiratory duration, and expiratory duration; the micromotion generation parameters corresponding to micro-tremors may include noise amplitude, noise frequency, and smoothing coefficient; the micromotion generation parameters corresponding to natural blinking may include blink trigger probability, blink interval, blinking speed, blinking speed, and movement template amplitude; and the micromotion generation parameters corresponding to micro-eye movements may include micro-saccade occurrence probability, angle range, and duration.
[0047] Stiffness parameters and damping parameters These are the physically integrated parameters in a virtual mass-spring-damped system. Stiffness parameters. Damping parameters can affect the strength of the regression of facial movements to the target equilibrium state. It can influence the velocity decay and response smoothness during changes in facial movement states. This is achieved by adjusting the micro-motion generation parameters and stiffness parameters based on the macro-semantic state. and / or damping parameters This allows for the coordination of microscopic physiological movements with the current state of speech expression.
[0048] S107, based on the target equilibrium state, micro-disturbance force, stiffness parameter and / or damping parameter, calculates the fused facial motion control data.
[0049] In one embodiment, the target equilibrium state can be... and microscopic disturbances Input a virtual mass-spring-damped system. The system outputs the facial motion state. It can be used as fused facial motion control data, or it can be processed by limiting, mapping, smoothing, and actuator constraint to form facial motion control data.
[0050] S108 controls the facial actuator of the object to be driven based on facial motion control data, or updates the facial key points, Blendshape parameters or facial animation model of the object to be driven, in order to generate facial expression movements corresponding to the speech signal.
[0051] In one embodiment, for a robot, facial motion control data can be converted into control commands such as servo angle, motor rotation angle, cable tension, shape memory alloy drive current, pneumatic actuator pressure, or software actuator displacement; for a virtual digital human, facial motion control data can be converted into Blendshape weights, skeletal controller parameters, facial key point positions, or mesh vertex displacements.
[0052] Figure 2This is a schematic diagram of the architecture of a voice-driven facial expression generation system provided in this application. Figure 2 As shown, the system may include a voice acquisition unit 100, a voice feature extraction unit 110, an expression generation neural network 120, a microphysiological action generation unit 130, a state modulation unit 140, a physical fusion unit 150, a drive output unit 160, and an object to be driven 170.
[0053] The speech acquisition unit 100 is used to acquire speech signals. The speech feature extraction unit 110 is used to extract speech features from the speech signals. The facial expression generation neural network 120 is used to generate macroscopic semantic action parameters based on the speech features. The microscopic physiological action generation unit 130 is used to generate microscopic perturbation forces based on microscopic physiological action generation logic. The state modulation unit 140 is used to determine the macro-semantic state based on speech features, macro-semantic action parameters, and / or intermediate features of the facial expression generation neural network 120, and outputs the micro-motion generation parameter adjustment amount and... Adjustment amount. The physical fusion unit 150 is used to receive the target equilibrium state and micro-disturbance force corresponding to the macro-semantic action parameters. as well as The adjustment amount is calculated to obtain facial motion control data. The drive output unit 160 is used to output the facial motion control data to the object to be driven 170. The object to be driven 170 may include a robot facial actuator, or it may include facial key points, Blendshape parameters, or a facial animation model.
[0054] Figure 2 In this model, the facial expression generation neural network 120 and the microscopic physiological action generation unit 130 operate in parallel processing paths. The facial expression generation neural network 120 generates macroscopic semantic action parameters, while the microscopic physiological action generation unit 130 generates microscopic perturbation forces. The two are not directly added together, but are jointly input into the physical fusion unit 150, which generates facial motion control data according to the virtual mass-spring-damping response relationship.
[0055] Figure 3 This is a schematic diagram illustrating the decoupling and fusion of macroscopic semantic actions and microscopic physiological actions provided in this application. (See diagram below.) Figure 3 As shown, the macroscopic layer includes speech features, facial expression generation neural network 120, macroscopic semantic action parameters, and target equilibrium state. . Figure 3 The micro-layer includes micro-physiological action generation unit 130 and micro-physiological action types such as breathing, micro-tremor, micro-eye movement, and natural blinking.
[0056] In the macroscopic layer, the facial expression generation neural network 120 outputs macroscopic semantic action parameters based on speech features, which are further converted into a target equilibrium state. Target equilibrium state It is not the final superimposed facial movement state, but the macroscopic target state in the virtual mass-spring-damping system.
[0057] In one specific embodiment, the output of the facial expression generation neural network 120 may not be directly used as the final facial movement position or the final actuator control quantity, but rather as a means to determine the target equilibrium state. The intermediate control quantities. For example, the expression generation neural network 120 can output the target equilibrium position of facial key points, the target equilibrium weights of the blendshape, the target pose of the robot's facial actuator, or output macroscopic control quantities corresponding to the rest length of the spring. For a virtual mass-spring-damped system, this target equilibrium state This is equivalent to the equilibrium position or equilibrium posture that the system tends to when it is not subjected to microscopic perturbations; when subjected to microscopic perturbations... When in operation, the actual facial movement state output by the system. It will be in a state of equilibrium around the target. It generates dynamic responses that conform to elastic, damped, and inertial constraints. Thus, the facial expression generation neural network 120 can focus on predicting macroscopic semantic action targets without directly outputting the final action state that includes microscopic physiological perturbations.
[0058] At the microscopic level, the microscopic physiological action generation unit 130 generates microscopic perturbation forces based on microscopic physiological actions such as breathing, micro-tremors, micro-eye movements, and natural blinking. This microscopic disturbance It is not the final position offset or the final weight offset, but the external force input acting on the physical fusion unit 150.
[0059] The state modulation unit 140 can output micro-motion generation parameter adjustment amounts to the micro-physiological action generation unit 130 based on the macro-semantic state, and output K / C adjustment amounts to the physical fusion unit 150. For example, when the speech state is continuous speaking, the state modulation unit 140 can reduce the breathing amplitude and increase the breathing frequency; when the emotional state is intense or tense, the state modulation unit 140 can increase the stiffness parameter K of some facial areas, making the facial expressions and movements present a tighter response.
[0060] Physical fusion unit 150 receives target equilibrium state Microscopic disturbances And the state-modulated K / C parameters, in the same physical fusion unit for macroscopic targets and Microscopic perturbations are decoupled and fused to output facial motion control data. Thus, microscopic movements act on the macroscopic motion process through physical responses, rather than being simply superimposed onto the final movement.
[0061] Figure 4 This diagram illustrates the separation and heterogeneous representation of macroscopic motion data and microscopic physiological motion data provided in this application. (See diagram for example.) Figure 4 As shown, in one optional embodiment, in order to train the facial expression generation neural network 120, or to construct the parameters of the micro-physiological action generation unit 130, the facial action data can first be separated into macro-action data and micro-physiological action data.
[0062] Specifically, facial movement sequences can be... Input macro-micro data separation unit 200. Macro-micro data separation unit 200 can obtain macro trajectory through low-pass filter unit 210. The low-pass filter unit 210 can employ a zero-phase low-pass filter, a Butterworth low-pass filter, a moving average filter, or other low-pass filters to preserve macroscopic movement trends such as mouth shape, eyebrows, eyes, and cheeks.
[0063] In one specific embodiment, the low-pass filter unit 210 can use a zero-phase low-pass filter to filter the facial action sequence. Processing is performed to reduce the impact of filter phase delay on lip-sync or facial expression timing. For example, a zero-phase low-pass filter with a cutoff frequency of approximately 4Hz can be used to extract the macroscopic trajectory. This allows low-frequency or low-to-mid-frequency macroscopic movements such as mouth opening and closing, jaw movement, and changes in facial expression to be preserved. The residual extraction unit 220 can then extract data based on the original facial movement sequence. With macro trajectory The difference between them yields the high-frequency micro-motion residual. ,Right now This high-frequency micro-motion residual can be used to characterize microscopic physiological movements such as eyelid tremors, micro-eye movements, and cheek tremors. The ultra-low frequency extraction unit 230 can use an ultra-low frequency filtering method below the macroscopic facial expression frequency range to extract respiratory baseline drift. For example, it can extract low-frequency components below 0.5Hz or corresponding to the respiratory cycle to obtain the slow fluctuations in the coordinated areas of the nose, face, jaw, or chest cavity caused by breathing. The cutoff frequencies mentioned above are only examples and can be adjusted according to the sampling frame rate, the type of object to be driven, the characteristics of the motion capture data, or the expression style.
[0064] Macro trajectory It can enter the macroscopic data storage unit 250, which can store temporal dynamic data such as position, velocity and acceleration.
[0065] The macro-micro data separation unit 200 can also extract high-frequency micro-motion residuals through the residual extraction unit 220. For example, the original facial motion sequence can be extracted. The macroscopic trajectory obtained by subtracting the low-pass filter The high-frequency micro-motion residual is obtained. The high-frequency micro-motion residual can be used to describe microscopic movements such as eyelid tremors, subtle facial tremors, or micro-eye movements, and is then stored in the microscopic data storage unit 260.
[0066] The macro-micro data separation unit 200 can also extract respiratory baseline drift through the ultra-low frequency extraction unit 230. Respiratory baseline drift corresponds to low-frequency changes in the coordinated areas of the nose, face, jaw, or chest cavity caused by respiration and is then stored in the micro data storage unit 260. The micro data storage unit 260 can store heterogeneous data such as statistical parameters, noise parameters, respiratory parameters, and motion templates.
[0067] For eyelid movements, since the eyelids may participate in both macroscopic emotional expression and microscopic natural blinking, the eyelid movement sequence can be input into the eyelid event segmentation unit 240. The eyelid event segmentation unit 240 can classify eyelid movements into macroscopic emotional eyelid closing and microscopic natural blinking based on factors such as movement duration, closing amplitude, and whether they are accompanied by significant eyebrow or mouth movements. For example, eyelid closing with a longer duration and accompanied by emotional eyebrow and eye movements can be classified as macroscopic emotional eyelid closing, while eyelid closing with a shorter duration and random triggering characteristics can be classified as microscopic natural blinking.
[0068] In one specific embodiment, the eyelid event segmentation unit 240 can classify eyelid movements into macro-events and micro-events by combining the duration of eyelid closure, the amplitude of eyelid closure, the shape of the closure curve, and whether it is accompanied by other macro-expression movements. For example, when the duration of an eyelid closure event is less than 0.4s and is not accompanied by large movements of the eyebrows, corners of the mouth, or jaw, the eyelid closure event can be judged as a micro-natural blink; when the duration of an eyelid closure event is greater than 0.5s, or when the eyelid closure event is accompanied by lowered eyebrows, tightened mouth, changes in the corners of the mouth, or other emotional macro-movements, the eyelid closure event can be judged as a macro-emotional eye closure. For eyelid events with durations in the middle range, further judgment can be made by combining the emotional state of the voice, the semantic context, the eyelid closure speed, or the closure amplitude. In this way, the problem of mixing natural blinks and emotional eye closures in the model can be reduced, allowing the expression generation neural network to mainly learn semantically related macro-eyelid movements, while natural blinks are treated as micro-physiological movements and processed by the micro-physiological movement generation logic.
[0069] Macro-level emotional closing of the eyes can enter macro-level data storage unit 250, while micro-level natural blinking can enter micro-level data storage unit 260.
[0070] In a specific example, the macroscopic data storage unit 250 can represent macroscopic actions as a temporal dynamic tensor containing position, velocity, and acceleration; the microscopic data storage unit 260 can represent microscopic actions as statistical parameters, noise parameters, respiratory parameters, and action templates. For example, a natural blink can correspond to the blink interval distribution and eyelid closure template; a micro-tremor can correspond to noise amplitude and noise frequency; breathing can correspond to respiratory frequency, respiratory amplitude, and respiratory phase; and micro-eye movements can correspond to micro-saccade probability and angular distribution. This heterogeneous representation method allows macroscopic semantic actions and microscopic physiological actions to adopt representations more suited to their physical properties and temporal characteristics.
[0071] In one specific embodiment, the statistical parameters stored in the microscopic data storage unit 260 may include the trigger interval distribution parameters of natural blinking, the probability of microsaccades, the distribution parameters of microsaccade angles, and the distribution parameters of microtremor amplitude. For example, the trigger interval of natural blinking may be described using an exponential distribution or a Poisson process. The mean of the exponential distribution may be set to 3s to 5s, and may also be adjusted according to macroscopic semantic states such as tension, fatigue, or relaxation. Noise parameters may include the amplitude, frequency, phase, and smoothing coefficient of Perlin noise or smoothed random noise, used to generate eyelid microtremors, cheek microtremors, or micro-eye movement perturbations. Respiratory parameters may include the basic respiratory rate, respiratory amplitude, respiratory phase, inspiratory duration, expiratory duration, and the mapping relationship between the speech activity state and the respiratory phase. Action templates may include natural blinking templates, microsaccade templates, respiratory traction templates, or cheek micro-movement templates; wherein, the natural blinking template may use an asymmetric Bézier curve to describe the process of rapid eyelid closure, brief pause, and relatively slow opening. By storing microphysiological actions as statistical parameters, noise parameters, respiratory parameters, and action templates, instead of uniformly storing them as the final position sequence output by the neural network, the logic for generating microphysiological actions can be better adapted to the randomness, periodicity, and templated characteristics of microphysiological actions.
[0072] Figure 5 This is a schematic diagram illustrating the generation and synthesis of microscopic perturbations provided in this application. Figure 5 As shown, the microscopic physiological action generation unit 130 may include a respiratory disturbance dynamics generation unit 300, a noise disturbance dynamics generation unit 310, a random trigger disturbance dynamics generation unit 320, a template disturbance dynamics generation unit 330, and a microscopic disturbance dynamics synthesis unit 340.
[0073] The respiratory disturbance generation unit 300 can generate respiratory disturbance according to the respiratory rhythm. The respiratory rhythm can be determined by the baseline respiratory rate, respiratory amplitude, respiratory phase, inspiratory duration, and expiratory duration. In one embodiment, the respiratory disturbance force can take the form of a periodic driving force, such as a sinusoidal, quasi-sinusoidal, or segmented inspiratory-expiratory curve. The respiratory disturbance force can act on the state dimensions corresponding to the linked areas of the nose, jaw, cheek, or chest cavity.
[0074] The noise disturbance generation unit 310 can generate noise disturbance forces based on micro-vibration or micro-eye movement noise. The noise can be Perlin noise, smoothed random noise, or other continuous low-amplitude random noise. Noise disturbance dynamics. It can be used to simulate subtle physiological tremors in the cheeks, eye sockets, corners of the mouth, or near the eyelids.
[0075] The random triggering disturbance generation unit 320 can generate random triggering disturbances based on natural blinking or micro-jump visual triggering events. Triggering events can be generated based on Poisson processes, exponential distributions, preset trigger probabilities, speech stress, speech energy mutations, or macroscopic action coordination conditions. For example, natural blinking can be randomly triggered based on the average blink interval; microsaccades can be randomly triggered based on preset micro-eye movement probabilities.
[0076] Template disturbance generation unit 330 can generate template disturbance based on micro-motion primitive templates. Microscopic motion primitive templates can include natural blinking templates, asymmetric eyelid closing-opening curve templates, micro-eye movement templates, respiratory traction templates, or cheek micro-movement templates. Templates can be converted into perturbation forms through amplitude scaling, time scaling, or baseline correction.
[0077] The micro-perturbation dynamic synthesis unit 340 can... , , and Synthesis was performed to obtain microscopic perturbation forces. The composition method can be summation, weighted summation, region selection, amplitude-limited superposition, or composition after state machine scheduling. For example, it can be represented as: .
[0078] The state modulation unit 140 can output micro-motion generation parameter adjustment amounts to the breathing disturbance force generation unit 300, the noise disturbance force generation unit 310, the random trigger disturbance force generation unit 320, the template disturbance force generation unit 330, and / or the micro-disturbance force synthesis unit 340. The micro-motion generation parameter adjustment amounts can be used to adjust the amplitude, frequency, trigger probability, template intensity, template duration, or disturbance force application area.
[0079] Figure 6A schematic diagram of the virtual mass-spring-damping system provided in this application. Figure 6 As shown, the system may include a target end 400, a spring 410, a damper 420, and a virtual mass block 430. The target end 400 corresponds to the target equilibrium state. Spring 410 corresponding stiffness parameters Damping parameters corresponding to damper 420 Virtual mass block 430 corresponds to virtual mass and facial movements Microscopic disturbances The response status of virtual mass block 430 is applied to virtual mass block 430. Used to output fused facial motion control data.
[0080] In one embodiment, the fused facial motion control data can be obtained by solving the following dynamic relationship or its equivalent form: ; in, Indicates virtual quality. Indicates the damping parameter. Represents the stiffness parameter. This indicates the facial movement state after fusion. It represents the target equilibrium state determined by macroscopic semantic action parameters, or the macroscopic target action state corresponding to the target equilibrium state. This represents the microscopic perturbation force.
[0081] In one embodiment, , and All can be Dimensional vector. It can correspond to the dimensions of facial key points, Blendshape parameters, number of robot facial actuators, or other facial control states. C and K can be scalars, diagonal matrices, region parameter matrices, or parameters set by dimension. For the mouth region, a larger macroscopic motion following weight can be set to make mouth shape synchronization more stable; for the nose and cheek regions, a more obvious respiratory disturbance response can be set; for the eyelid region, stiffness and damping parameters can be set to adapt to natural blinking and eyelid micro-tremors.
[0082] In a discrete control implementation, the frame period can be controlled by facial motion. Perform numerical solutions. For example, based on the current state... Current speed Target state and microscopic disturbances Calculate acceleration And update speed and status: ; ; .
[0083] Updated This can be used as the fused facial motion state of the current control frame. To ensure robot execution safety or animation parameter stability, it can also be... , Alternatively, facial motion control data may be subjected to amplitude limiting, speed constraint, or actuator stroke constraint processing.
[0084] In one embodiment, the state modulation unit 140 determines the muscle tension coefficient based on the macro-semantic state. Muscle tension coefficient It can be determined based on the intermediate features of the neural network 120, such as emotional intensity, voice volume, speech rate, macroscopic movement amplitude, macroscopic movement speed, or facial expression. It can be normalized to the interval [0,1], where the larger values are... It indicates a more intense, tense, or highly aroused state of expression.
[0085] In a specific example, the stiffness parameter K and / or damping parameter C can be adjusted according to the muscle tension coefficient, for example: ; .
[0086] Where K0 is the basic stiffness parameter and C0 is the basic damping parameter. and This is an adjustment coefficient. In another embodiment, the stiffness parameter K can be increased and the damping parameter C in some areas can be decreased in the high-awakening state to make the action response faster and tighter. Different adjustment methods can be selected according to the material of the object to be driven, the performance of the actuator, or the animation style.
[0087] In the breathing rhythm modulation embodiment, the breathing fluctuation disturbance force can be adjusted according to the speech activity detection result, speech pause status, speech rate and / or volume. When the speech activity detection result indicates a continuous speaking state, the breathing amplitude can be reduced and the breathing frequency can be increased, making the breathing appear short or slightly fluctuating; when the speech activity detection result indicates a pause state, the breathing amplitude can be increased and the breathing frequency can be reduced, making the object to be driven appear relaxed or in a breathing state. For example, in the speaking state, the breathing amplitude can be 0.4 to 0.8 times the resting amplitude, and the breathing frequency can be 1.1 to 1.8 times the resting frequency; in the pause state, the breathing amplitude can be 1.0 to 1.5 times the resting amplitude, and the breathing frequency can be 0.6 to 1.0 times the resting frequency. The above values are only examples and can be adjusted according to the object type, facial expression style, or actuator response capability.
[0088] Figure 7 This is a schematic diagram of the biomimetic motion scheduling state machine provided in this application. Figure 7 As shown, in one embodiment, before the fused facial motion control data is calculated, conflict resolution between macroscopic semantic actions and microscopic physiological actions can be performed based on the bionic motion scheduling state machine 500.
[0089] The biomimetic motion scheduling state machine 500 can receive macroscopic semantic motion states and microscopic physiological motion trigger events. Macroscopic semantic motion states can include the degree of eyelid closure, the degree of jaw opening, the degree of mouth corner stretching, the degree of eyebrow raising, or emotional state, etc. Microscopic physiological motion trigger events can include natural blinking trigger events, microsaccade trigger events, respiratory disturbance peak events, or microtremor enhancement events, etc.
[0090] When the preemption condition is met, the bionic motion scheduling state machine 500 can block micro-motion triggers in the corresponding facial area. For example, when the macro-semantic motion state indicates that the degree of eyelid closure is greater than a first threshold, and the duration of the eye-closing action is greater than a preset time, it can be considered that the current state is emotional eye-closing or large eye-closing. At this time, if a natural blinking trigger event occurs, the natural blinking trigger event is blocked to avoid superimposing another blinking action in the already closed eye state. The first threshold can be 70%, 80%, or other values of the degree of eyelid closure.
[0091] When the superposition condition is met, the biomimetic motion scheduling state machine 500 can modify the motion baseline of the microscopic physiological action based on the current macroscopic motion state before superimposing it. For example, when the macroscopic semantic motion state indicates that the eyelid is in a half-closed state, the baseline of the natural blink template can be adjusted to the current half-closed state, and then a small-amplitude eyelid micro-movement can be superimposed on this baseline, instead of performing a complete blinking action starting from the fully open baseline. The half-closed state can be determined by the degree of eyelid closure being between 20% and 60%.
[0092] When the coordination conditions are met, the biomimetic motion scheduling state machine 500 can trigger microscopic physiological actions that coordinate with macroscopic semantic actions. For example, when the mandibular opening degree exceeds a second threshold, a tiny blink or a slight twitch of the cheek can be triggered to simulate the natural linkage when speaking with a wide-open mouth, expressing surprise, or making a forceful sound. The second threshold can be 60%, 70%, or other values of the maximum mandibular opening range. If the preemption condition, superposition condition, and coordination condition are not met, the microscopic disturbance force can be normally input into the physical fusion unit 150 for fusion.
[0093] Figure 8 This is a schematic diagram illustrating the phase anchoring of speech initiation time and respiratory rhythm provided in this application. Figure 8 As shown, in one embodiment, the speech start time 600 of the speech signal can be obtained based on speech activity detection. The speech start time 600 can be the moment when the VAD state switches from a non-speech state to a speech state, or it can be the start time when the speech energy exceeds a preset threshold and continues for a preset time.
[0094] After obtaining the speech start time 600, a preset time window 610 can be set before the speech start time 600, and the inspiratory peak 620 in the breathing rhythm can be adjusted to fall within the preset time window 610. For example, the preset time window 610 can be located 100ms to 300ms before the speech start time 600, and preferably the inspiratory peak 620 can be located approximately 200ms before the speech start time 600. The above time values are merely examples and do not limit the scope of protection of this application.
[0095] In one specific embodiment, the breathing phase can be corrected based on the speech start time 600. Let the breathing phase be... The phase corresponding to the inspiratory peak is It can be based on the voice start time Tstart and the preset advance duration. Determine the target peak time And adjust the breathing phase to make Preset advance time The time frame can be between 100ms and 300ms, preferably around 200ms. If the inspiratory peak corresponding to the current respiratory phase is not within the preset time window 610, the inspiratory peak 620 can be made to fall within the preset time window 610 by phase shifting, local time scaling, or selecting a target peak in adjacent respiratory cycles. The phase-corrected respiratory rhythm is used to generate respiratory fluctuation disturbance force, which acts on the virtual mass-spring-damped system before the onset of speech.
[0096] After the peak inhalation 620 is adjusted to within the preset time window 610, the respiratory fluctuation disturbance force can act on the virtual mass-spring-damping system before the speech initiation time 600, causing a slight response in facial movements before the speech begins. For example, the nostrils may open slightly, the jaw may tend to open slightly, and the skin or cheeks may twitch slightly. This creates preparatory facial micro-movements before speaking, making the speech initiation action more coordinated with the breathing rhythm in time.
[0097] In one embodiment, the preset time window 610 can be dynamically adjusted according to speech rate, voice intensity, emotional state, or the type of object to be driven. For example, when the speech rate is fast, the preset time window can be shortened; when the tone is strong or when emphasis is needed, the peak inhalation 620 can be advanced appropriately; when the robot actuator responds slowly, the duration of the breathing fluctuation disturbance can also be advanced according to the actuator response delay.
[0098] like Figure 9 As shown, Figure 9 This application provides a structural schematic diagram of a voice-driven facial expression generation robot. The robot may include a robot body 700, a voice acquisition device or voice receiving device 710, a processor 720, a memory 730, a facial actuator 740, and a robot face 750.
[0099] The voice acquisition device or voice receiver 710 can be installed on the robot body 700, or it can be connected to the robot body 700 via a communication interface. The voice acquisition device or voice receiver 710 is used to acquire voice signals and send the voice signals to the processor 720. The memory 730 stores a computer program, and when the processor 720 executes the computer program, it can implement the voice-driven facial expression generation method in any of the foregoing embodiments.
[0100] The processor 720 can generate facial motion control data based on voice signals and send the facial motion control data to the facial actuator 740. The facial actuator 740 may include servos, cables, shape memory alloys, soft actuators, micro motors, pneumatic actuators, or other mechanisms capable of driving the robot's face 750 to produce facial expressions. The facial actuator 740 can drive the robot's face 750 to produce mouth shape changes, jaw opening and closing, eyelid opening and closing, nasal wing micro-movements, cheek micro-movements, or other facial expressions based on the facial motion control data.
[0101] In one embodiment, the facial motion control data output by the processor 720 can first be converted into actuator control commands. For example, when the facial actuator 740 includes a servo motor, the facial motion state can be... Mapped to servo angle; when the facial actuator 740 includes a cable-driven structure, it can map the facial motion state. Mapped to cable tension or cable contraction; when the facial actuator 740 includes a shape memory alloy, it can map facial movement states. Mapped to drive current or heating time; when the facial actuator 740 includes a soft actuator, the facial motion state can be... The mapping can be expressed as pressure, flow rate, or deformation. The mapping method can be proportional mapping, lookup table mapping, calibration function mapping, or inverse kinematics mapping.
[0102] In one embodiment, the processor 720 can be deployed inside the robot body 700, or it can be located in a server, edge computing device, or control terminal that communicates with the robot. The memory 730 can be a read-only memory, random access memory, flash memory, solid-state drive, or other storage media. When the computer program is executed by the processor 720, it can perform functions such as speech feature extraction, macroscopic semantic action parameter generation, microscopic perturbation force generation, state modulation, physical fusion, and drive output.
[0103] In the above embodiments, although robots and virtual digital humans were used as examples, the method of this application can also be applied to augmented reality characters, virtual anchors, game characters, virtual avatars for online meetings, emotional interaction terminals, or other objects that need to generate facial expressions and actions based on speech. As long as it adopts the method of using macroscopic semantic actions driven by speech as the target equilibrium state and microscopic physiological actions as the perturbation force and integrating them through a virtual mass-spring-damping system, it can fall within the scope of the technical concept of this application.
[0104] The above embodiments are only for illustrating the technical concept and features of this application, and are intended to enable those skilled in the art to understand the content of this application and implement it accordingly. They should not be used to limit the scope of protection of this application. It is obvious to those skilled in the art that this application is not limited to the details of the above exemplary embodiments, and that this application can be implemented in other specific forms without departing from the spirit or basic characteristics of this application. Therefore, all embodiments should be considered exemplary and non-limiting, and all changes within the meaning and scope of equivalent elements of the technical features of this application are included within this application.
Claims
1. A voice-driven facial expression generation method, characterized in that, include: Acquire the speech signal used to drive the object to be driven, and extract features from the speech signal to obtain speech features that characterize speech semantics and / or prosodic states; The speech features are input into the facial expression generation neural network to generate macroscopic semantic action parameters, which are used to characterize macroscopic facial movements driven by speech semantics. Microscopic perturbation forces are generated based on microscopic physiological action generation logic independent of the facial expression generation neural network. These microscopic perturbation forces are used to characterize the perturbation effect of microscopic physiological actions on facial movements. The target equilibrium state of the virtual mass-spring-damped system is determined based on the macroscopic semantic action parameters, and the microscopic disturbance force is used as the external force input acting on the virtual mass-spring-damped system. The target equilibrium state includes the target rest point and / or the spring rest length. The macro-semantic state is determined based on the speech features, the macro-semantic action parameters, and / or the intermediate features of the facial expression generation neural network, and the generation parameters of the micro-physiological action generation logic and / or the stiffness parameters and / or damping parameters of the virtual mass-spring-damping system are adjusted based on the macro-semantic state. Based on the target equilibrium state, the micro-disturbance force, the stiffness parameter and / or the damping parameter, the fused facial motion control data is calculated. The facial actuator of the object to be driven is controlled according to the facial motion control data, or the facial key points, Blendshape parameters or facial animation model of the object to be driven are updated to generate facial expression movements corresponding to the speech signal.
2. The voice-driven facial expression generation method according to claim 1, characterized in that, The micro-perturbation forces include at least one of the following: respiratory fluctuation perturbation force, micro-tremor perturbation force, micro-eye movement perturbation force, and natural blinking perturbation force.
3. The voice-driven facial expression generation method according to claim 2, characterized in that, The breathing fluctuation disturbance force is generated according to the breathing rhythm, which is adjusted according to the speech activity detection results, speech pause status, speech speed and / or volume, so that the breathing amplitude decreases and the breathing frequency increases during the speaking state, and the breathing amplitude increases and the breathing frequency decreases during the pause state.
4. The voice-driven facial expression generation method according to claim 1, characterized in that, The muscle tension coefficient is determined based on the macroscopic semantic state, and the stiffness parameter and / or the damping parameter is adjusted based on the muscle tension coefficient.
5. The voice-driven facial expression generation method according to claim 1, characterized in that, The fused facial motion control data is obtained by solving the following dynamic relationship or its equivalent form: , in, Indicates virtual quality. Indicates the damping parameter. Represents the stiffness parameter. This indicates the facial movement state after fusion. This represents the target equilibrium state determined by macroscopic semantic action parameters, or the macroscopic target action state corresponding to the target equilibrium state. This refers to the microscopic perturbation force.
6. The voice-driven facial expression generation method according to claim 1, characterized in that, The micro-disturbance force is generated based on the noise force field, periodic driving force, randomly triggered disturbance force and / or micro-action primitive template.
7. The voice-driven facial expression generation method according to claim 1, characterized in that, Before obtaining the fused facial motion control data, conflict resolution is performed on the macro-semantic motion corresponding to the macro-semantic motion parameters and the micro-physiological motion corresponding to the micro-perturbation force based on the biomimetic motion scheduling state machine. The conflict resolution includes: When the macroscopic semantic action meets the preset preemption condition, the triggering of microscopic physiological actions in the corresponding facial area is blocked. When the macroscopic semantic action meets the preset superposition conditions, the action baseline of the microscopic physiological action is corrected based on the current macroscopic action state. When the macroscopic semantic action meets the preset coordination conditions, a microscopic physiological action coordinated with the macroscopic semantic action is triggered.
8. The voice-driven facial expression generation method according to claim 3, characterized in that, The speech start time of the speech signal is obtained based on speech activity detection, and the inspiratory peak in the breathing rhythm is adjusted to a preset time window before the speech start time, so that the breathing fluctuation disturbance force acts on the virtual mass-spring-damping system before the speech start time.
9. A computer device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, implements the voice-driven facial expression generation method according to any one of claims 1 to 8.
10. A voice-driven facial expression generation robot, characterized in that, include: The robot itself; A facial actuator, located on the robot body, is used to drive the robot's face to produce facial expressions. A voice acquisition device or voice receiving device is used to acquire voice signals; A processor and a memory, wherein the memory stores a computer program, and the processor is connected to the voice acquisition device or voice receiving device and the facial actuator, respectively. When the computer program is executed by the processor, it implements the voice-driven facial expression generation method according to any one of claims 1 to 8, and controls the facial actuator to perform actions according to the facial motion control data obtained by the voice-driven facial expression generation method, so as to generate robot facial expression actions corresponding to the voice signal.