Digital human eye follow-up method based on sound source localization
By using a microphone array and sound source localization algorithm to drive the digital human to adjust its gaze direction, the problem of the digital human not being able to follow movement was solved, and real gaze following was achieved in multiple conversational parties or in scenarios not directly in front of the user, thus improving the interactive experience.
Patent Information
- Application Number
- CN202510748975.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-11-04
AI Technical Summary
Existing digital human interaction systems cannot achieve eye movement in real-life human-to-human dialogue scenarios, nor can they adjust the direction of the gaze according to the position of the other party, resulting in an inability to realistically simulate eye contact when there are multiple parties or when the conversation is not directly in front of the other party.
By introducing a microphone array and a sound source localization algorithm, the voice location of the speaker is obtained. Combined with deep learning technology, the digital human adjusts its gaze direction. By using the sound source localization algorithm and the digital human generation model to introduce the speaker's direction information during the training phase, the digital human's gaze follows the speaker's movements.
It enables digital humans to realistically simulate eye-following in multiple conversational scenarios or in non-directly-facing scenes, enhancing the sense of presence and realism of the interaction. It is applicable to various lighting environments and reduces costs, making it suitable for scenarios where cameras cannot be used, such as in classified units.
Smart Images

Figure CN120892007A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of digital people and computer programs, and relates to a digital person gaze following method based on sound source positioning. BACKGROUND
[0002] With the continuous development and maturity of artificial intelligence technology, digital people technology has been widely applied to various scenes, such as customer service, advertising, teaching, tourism, medical treatment and various scenes. In particular, the access of large language model (LLM, Large Language Module) technology enables the digital people system to realize the dialogue ability close to human beings. The digital clone or virtual person created by digital people technology can well communicate with people. In various interactive scenes, the mainstream digital people interactive system adopts the technical framework as shown in the figure, which first extracts the speech information of the participants in the dialogue through the ASR (Automatic Speech Recognition) module, converts it into text, then sends the converted text into the LLM model to generate the dialogue content of the digital people, sends the generated dialogue content into the TTS (Text to Speech) to generate the speech of the digital people, finally combines the generated speech with the pre-trained digital people generation model to synthesize the digital people video and plays it, completing the dialogue loop. Figure 1
[0003] In this framework, the audio-driven face generation technology is the core supporting technology for the natural interaction of digital people with people, and the goal is to generate real and natural head movement gestures through speech combined with action models, which are synchronized with pronunciation and lip shape, and matched with tone and expression. According to the technical route adopted, audio-driven face generation mainly includes three types of methods based on generative adversarial networks (Generative Adversarial Networks), diffusion models (Diffusion Model) and neural radiation fields (Neural Radiance Fields).
[0004] The generative adversarial network includes a generator for generating data and a discriminator for judging whether the generated data is real, and the two are in confrontation in the training process. Wav2Lip proposed by K R Prajwal et al. uses a pre-trained lip-sync expert discriminator to guide the generator to generate more accurate lip movements; DINet proposed by Zhimeng Zhang et al. aligns the feature map of the reference image with the driving audio through spatial deformation, and then fuses the features of the source image through repair, thereby better preserving the texture details. The core idea of the diffusion model-based method is to first add noise to the data step by step, and then remove the noise through a neural network model. AniPortraitwav2vec model proposed by Huawei Wei et al. extracts audio features, and then uses a fully connected layer and a Transformer decoder to convert these features into a 3D face mesh and head pose, respectively, and finally obtains a 2D face landmark through perspective projection; DreamActo-M1 model proposed by Yuxuan Luo et al. realizes robust control of facial expressions and body movements by introducing mixed control signals of 3D head spheres and 3D body skeletons. The neural radiance field obtains a spatial code according to the audio information, and then models it using a neural network model. HAvatar model proposed by XiaoChen Zhao et al. is a lightweight 3D avatar modeling method that integrates the expressive power of NeRF and the prior information of the parameterized template, enabling the synthesis of high-resolution, realistic, and perspective-consistent dynamic head appearances, and allowing fine-grained control of head poses and facial expressions. DFA-NeRF framework proposed by Shunyu Yao et al. decouples facial attributes into head pose, blinking features, and lip features, combines audio features, and performs high-quality image synthesis through neural radiance field; the above audio-driven face generation has achieved good results in terms of lip shape, expression, and movement generation, but the results generated by these methods are limited by the driving source. Trying to maximize the use of the information of the output audio to construct digital human head and body movements from audio is completely sufficient in non-human interaction scenarios, such as virtual news anchors, virtual advertising sales, various demonstration video characters, etc., but there are still many defects in scenarios that require interaction with people. In real human-to-human conversation scenarios, not only the content of the speech, the tone of the voice, the facial expression, and the body language can convey information, but also the most important eye contact. Usually, both parties in a conversation constantly watch each other during the conversation, and the digital human cannot simulate this point only through the speech information prepared by the digital human.
[0005] The current mainstream digital human interaction system cannot achieve eye tracking for the conversation party, so it can only default the position of the conversation party to be directly in front of the digital human playback device, and the face and gaze of the digital human are also directly in front.
[0006] Such a setting cannot simulate the real scene for the following cases:
[0007] Scene 1. Digital human and single dialogue party dialogue, when the position of the dialogue party is not in front of the digital human playing device, the digital human still gazes at the front;
[0008] Scene 2. Digital human and multiple dialogue parties dialogue, different dialogue parties are at different positions, and the digital human cannot change the position of the gaze when different dialogue parties participate in the dialogue;
[0009] Scene 3. When the digital human is having a dialogue, a loud sound or similar sound that can attract people's attention suddenly comes from somewhere, and the digital human cannot turn the gaze to the direction of the sound like a real person.
[0010] The reason for the above problems is that in the current digital human interaction system, the spatial orientation of the dialogue party relative to the digital human is not considered and is not taken as the driving data of the digital human, so when the digital human video is synthesized, the direction of the gaze of the digital human cannot be adjusted accordingly. SUMMARY
[0011] The present application aims at the problems of the prior art and provides a digital human gaze following method based on sound source positioning.
[0012] A digital human gaze following method based on sound source positioning, comprising the following steps: by introducing a microphone array, the position of the dialogue party or the sound in the environment that can attract people's attention is obtained through a sound source positioning algorithm, and the digital human is driven to make a corresponding response.
[0013] The advantage of the present application is that the inference of the target sound source direction by means of the microphone array, and then the digital person's gaze following to the dialogue party, can make the dialogue party and the user feel a more realistic experience in the interaction process. By means of the microphone array, the gaze switching between multiple dialogue parties can be carried out through voice. When multiple dialogue parties do not speak at the same time, the reaction mode is the same as that of a single dialogue party; when multiple dialogue parties speak at the same time, the identity of the multiple dialogue parties can be standardized through voiceprint features, and the reply can be carried out one by one according to a certain order, and the line of sight is turned to the direction of the corresponding dialogue party. The position of the dialogue party is confirmed by the sound source positioning method, which is not affected by the light condition of the dialogue environment. Through sound source positioning, various types of sounds (human voice or non-human voice) in the dialogue environment can be fed back in terms of expression, action or even conversation content, and the digital person is given more emotional expression ability. The microphone array is used as the key positioning information receiving device, which has a lower cost than the camera scheme, and is suitable for some scenes where the camera cannot be used, such as environments where shooting is prohibited in secret units. The technical scheme is less changed for the current mainstream digital person generation model scheme, and only the dialogue party direction information needs to be added to the data label in the training stage, so that the gaze following effect can be realized. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, below will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor. As shown in the drawings:
[0015] Figure 1 It is a digital person interaction system framework of the prior art.
[0016] Figure 2 It is a digital person interaction system framework of the present application introducing sound source positioning.
[0017] Figure 3 It is a general principle diagram of deep learning sound source positioning. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0019] Embodiment 1: as Figure 2As shown, a digital human gaze following method based on sound source positioning, by introducing a microphone array, through a sound source positioning algorithm, to obtain the position of the sound that can attract the attention of the interlocutor or the environment, and drive the digital human to make corresponding reactions.
[0020] A digital human gaze following method based on sound source positioning, containing the following steps: adding a microphone matrix on the digital human playing device, the microphone matrix transmits the received sound into the SSL (Sound Source Location) sound positioning module, the sound positioning module calculates the position of the sound source relative to the digital human eye position, and transmits this parameter to the digital human synthesis module for digital human video generation.
[0021] Also containing the following steps: sound source positioning step, introducing digital human motion model step of sound source position.
[0022] The sound source positioning step also contains the following steps:
[0023] First, determine the microphone array. In order to get the azimuth and elevation angle of the sound source, at least 3 microphones are needed, considering the amount of redundancy to be saved, it is more appropriate to configure 4 microphones. For larger devices, 6 or more microphones can be configured, and for smaller devices, only two can be equipped, then only the azimuth angle information can be obtained and the elevation angle cannot be obtained.
[0024] In this scheme, only the case of 4 or more microphones is discussed.
[0025] Second, determine the direction of the digital human's gaze. The reference point is the position of the digital human's lips, and the target point is the position of the interlocutor's lips. The position of the target point relative to the reference point is equal to the position of the interlocutor's lips relative to the digital human's lips.
[0026] The position of the digital human's lips can be calculated based on the position of the microphone array, combined with the position and size of the digital human on the display device; the position of the target point is the position of the sound source.
[0027] Therefore, the DOA (Direction of Arrival) of the sound source obtained in this scheme is the direction of the digital human's gaze.
[0028] If the slight difference caused by the different mouth-eye distances of the interlocutor is considered, the age, gender and height of the interlocutor can also be inferred from the voice, and then the relationship between the mouth-eye distance, gender, age and height is corrected.
[0029] Third, select a sound source positioning method based on deep learning, one is because it has better positioning effect, in addition, the sound source positioning and biological information extraction (gender, age, height, mouth-eye distance) can be trained at the same time to get a reliable model.
[0030] Sound source localization (SSL) refers to determining the position of one or more target points (sound sources) relative to a reference point (usually a microphone) by analyzing received acoustic signals. In practical applications, only the direction of origin (DOA), including azimuth and pitch angles, is usually estimated, while the distance between the reference point and the sound source is ignored.
[0031] Traditional sound source localization methods based on signal / channel models and signal processing techniques, despite years of development, still face many challenges in scenarios with high noise, reverberation, and multiple sound sources. In contrast, deep learning-based sound source localization methods have developed rapidly over the past decade and have achieved higher localization accuracy than traditional methods.
[0032] General principles of sound source localization based on deep learning ( Figure 3 The function is to extract features from the multi-channel input audio signals received by the microphone array and then feed these features into a deep neural network to obtain the DOA parameters.
[0033] Furthermore, the choice of microphone array is also crucial for sound source localization. Typically, the number of microphones should be greater than the localization dimension plus one, and adding redundant microphones further improves the localization effect. However, the number of microphones is also limited by device size and cost, and increasing it also leads to increased computational complexity.
[0034] In order to enable the digital human's gaze to follow the dialogue partner, this application aims to obtain the position of the dialogue partner's eyes relative to the digital human's eyes, and therefore needs to make corrections based on the sound source localization.
[0035] The steps of introducing a digital human motion model based on sound source location include establishing a digital human generation model. The digital human motion model is a part of the digital human generation model. During the training of the digital human generation model, the training data is decoupled into lip movements, facial expressions, head movements, eye movements, etc. Among them, head movements and eye movements are closely related to the digital human's gaze direction. By introducing DOA labels into the training set, the digital human motion model can generate corresponding head movements and gaze directions based on the DOA.
[0036] Digital Human Generation Model: Based on the existing digital human generation model, a digital human motion model with DOA correction is introduced. When voice information and DOA information are input into the digital human generation model at the same time, the head movements and eye movements generated by the digital human motion model are fused with facial expressions and lip movements to generate a digital human video.
[0037] Audio-driven digital human generation models generate videos of digital humans speaking by using input speech information.
[0038] In the present application, the horizontal angle and the pitch angle parameters of the dialogue party are introduced into the digital human generation model, wherein the horizontal angle can help the digital human head and line of sight to adjust in the horizontal direction, and the pitch angle can help the digital human head and line of sight to adjust in the vertical direction.
[0039] On the training end, on the basis of the current audio-driven digital human generation model, a large amount of labeled training data is used to train the model, and the label information contains the direction information (i.e. horizontal angle and pitch angle) of the gaze fixation; on the inference end, when the voice and the direction information of the gaze fixation are sent into the generation model together, the digital human speaking video that fixes on the dialogue party can be obtained.
[0040] The audio-driven digital human generation model synthesizes the digital human speaking video synchronized with the voice content by inputting voice information.
[0041] On the basis of the existing audio-driven framework, the present application introduces the spatial orientation parameters of the target dialogue party as the control condition, and the specific implementation is as follows:
[0042] Fourth, the orientation parameter definition: the spatial position of the dialogue party relative to the digital human is described by the azimuth angle and the pitch angle, the head horizontal turning and the vertical inclination are controlled respectively, the line of sight direction is adjusted through parameter mapping, and the digital human presents the natural gaze fixation behavior.
[0043] Fifth, the training method: a ternary training data set containing voice-video-orientation label is constructed, the digital human generation model is trained, and the consistency of the lip synchronization, head movement and line of sight direction is optimized.
[0044] Sixth, the inference process: the voice signal to be output and the orientation information are input into the digital human generation model, and the digital human is driven to synthesize the speaking video that fixes on the target direction.
[0045] The present application solves the problem that the traditional model digital human cannot realize the following movement, and can improve the sense of presence in the scene of interacting with the dialogue party. The present application is applicable to the generative adversarial network, the diffusion model and the neural radiation field.
[0046] As shown in Figure 2 The hardware aspect of the present application adds a microphone matrix on the digital human playing device, the microphone matrix transmits the received sound into the SSL sound positioning module, the sound positioning module calculates the orientation of the sound source relative to the digital human eye position, and transmits this parameter into the digital human synthesis module to generate the digital human video.
[0047] In Figure 2In the SSL module, the azimuth and the pitch angle of the interlocutor relative to the digital human are inferred according to the voice received through multiple channels, and the azimuth and the pitch angle are sent to the Avatar digital human generation module. After receiving the voice information, the ASR module extracts the text information therein and sends it to the LLM large language model to generate a reply text information which is sent to the TTS module to convert the text information into the voice of the digital human, and then the voice of the digital human is sent to the Avatar module. The voice information of the digital human, the azimuth and the pitch angle of the interlocutor, together drive the Avatar module to generate a digital human video with the gaze directed at the interlocutor. Finally, the digital human video is played through the display of the digital human playing device, and the digital human voice is played through the loudspeaker of the digital human playing device.
[0048] The above merely describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for digital human eye-tracking based on sound source localization, characterized in that, It includes the following steps: by introducing a microphone array and using a sound source localization algorithm, the location of attention-grabbing sounds from the conversation partner or environment is obtained, and the digital human is driven to make corresponding responses.
2. The digital human eye-tracking method based on sound source localization according to claim 1, characterized in that, It includes the following steps: adding a microphone matrix to the digital human playback device, the microphone matrix transmitting the received sound to the SSL (Sound Source Location) sound localization module, the sound localization module calculating the orientation of the sound source relative to the position of the digital human's eyes, and transmitting this parameter to the digital human synthesis module to generate digital human video, and also includes the following steps: sound source localization step, and introducing the digital human motion model of the sound source position step.
3. The digital human eye-tracking method based on sound source localization according to claim 2, characterized in that, The sound source localization process also includes the following steps: First, determine the microphone array. To obtain both the azimuth and elevation angles of the sound source simultaneously, at least three microphones are needed. Secondly, determine the direction of the digital human's gaze. The reference point is the position of the digital human's lips, and the target point is the position of the speaker's lips. The position of the target point relative to the reference point is approximately equal to the position of the speaker's lips relative to the digital human's lips. The position of the digital human's lips is calculated based on the microphone array position and the position and size of the digital human on the display device; the position of the target point is the sound source position, and the obtained sound source DOA (Direction of Arrival), including the azimuth angle and pitch angle, is the direction of the digital human's gaze.
4. The digital human gaze tracking method based on sound source localization according to claim 3, characterized in that, To allow for redundancy, four microphones are configured, but larger devices can be configured with six or more.
5. The digital human gaze tracking method based on sound source localization according to claim 3, characterized in that, Considering the subtle differences caused by varying mouth-eye distances, we can infer the age, gender, and height of the speakers from their speech, and then correct these differences by examining the relationship between mouth-eye distance and gender, age, and height.
6. The digital human eye-tracking method based on sound source localization according to claim 2, characterized in that, The steps of introducing a digital human motion model based on sound source location include establishing a digital human generation model. Digital Human Motion Model: The digital human motion model is a part of the digital human generation model. During the training process of the digital human generation model, the training data is decoupled into lip movements, facial expressions, head movements, and eye movements. Among them, head movements and eye movements are closely related to the digital human's gaze direction. By introducing DOA (Directory of Adaptive) labels into the training set, the digital human motion model can generate corresponding head movements and gaze directions based on the DOA. Digital Human Generation Model: Based on existing digital human generation models, a DOA-corrected digital human motion model is introduced. When voice information and DOA information are simultaneously input into the digital human generation model, head movements, eye movements, facial expressions, and lip movements generated by the digital human motion model are fused together to generate a digital human video. Audio-driven digital human generation models generate videos of digital humans speaking by using input speech information. This is achieved by introducing horizontal and pitch angle parameters into the model. The horizontal angle helps adjust the digital human's head and gaze horizontally, while the pitch angle helps adjust the head and gaze vertically. On the training side, based on the current audio-driven digital human generation model, a large amount of labeled training data is used to train the model. The label information includes gaze direction information (i.e., horizontal angle and pitch angle). On the inference side, when the speech and gaze direction information are fed into the generation model together, a video of a digital human speaking while looking at the other party can be obtained. Audio-driven digital human generation models synthesize digital human speaking videos synchronized with the input speech information. The steps for introducing the spatial orientation parameters of the target interlocutor as control conditions are as follows: First, the orientation parameters are defined: azimuth and pitch angles are used to describe the spatial position of the speaker relative to the digital human, controlling the horizontal head turn and vertical tilt angle respectively. Through parameter mapping, the gaze direction is adjusted, enabling the digital human to exhibit natural eye contact behavior. Second, training method: A three-element training dataset containing speech-video-location labels is constructed to train the digital human generation model, and the consistency of lip-sync, head movement and gaze direction is jointly optimized. Third, the reasoning process: The voice signal and location information to be output are input into the digital human generation model, which drives the digital human to synthesize a speaking video that looks in the direction of the target.
7. The digital human gaze tracking method based on sound source localization according to claim 1, characterized in that, A microphone matrix is added to the digital human playback device. The microphone matrix transmits the received sound to the SSL sound localization module. The sound localization module calculates the orientation of the sound source relative to the position of the digital human's eyes and transmits this parameter to the digital human synthesis module to generate digital human video. When the conversation partner (the person interacting with the digital human) speaks, their voice is transmitted to the microphone array. The voice received by the microphone array is transmitted to the SSL module and the ASR module respectively. In the SSL module, the azimuth and pitch angles of the conversation partner relative to the digital human are inferred from the voice received through multiple channels, and the azimuth and pitch angles are sent to the Avatar digital human generation module. After receiving voice information, the ASR module extracts the text information and sends it to the LLM large language model to generate a reply text information, which is then sent to the TTS module to convert the text information into the digital human's voice. The digital human's voice is then sent to the Avatar module. The digital human's voice information, the azimuth angle and pitch angle of the conversation partner together drive the Avatar module to generate a digital human video with the eyes facing the conversation partner. Finally, the digital human video is played on the display of the digital human playback device, and the digital human voice is played through the loudspeaker of the digital human playback device.
Citation Information
Cited By
Eye animation display method and device based on image rendering
CN122023611A