AI digital human device and method based on voiceprint discrimination and capable of performing face-to-face communication with multiple persons
By combining microphone arrays, infrared laser ranging sensors, facial tracking technology and voiceprint recognition technology, the problem of being unable to accurately locate and identify speakers in a multi-person environment is solved, and high-precision speaker recognition and positioning is achieved, providing a natural and smooth multi-person dialogue and interaction experience.
Patent Information
- Application Number
- CN202510630345.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-17
AI Technical Summary
In the prior art, in multi-person voice interaction and intelligent facial tracking, it is impossible to accurately identify the identities of different speakers, locate their locations, and achieve personalized interactions based on the conversation content.
By combining microphone arrays, infrared laser ranging sensors, facial tracking technology, and voiceprint recognition, multiple speakers are accurately positioned and identified. The beamforming algorithm and RNNoise noise reduction module are used to improve speech clarity, and voiceprint feature extraction and comparison are used using Mel frequency cepspectral coefficient (MFCC) and ECAPA-TDNN deep learning model.
It realizes accurate identification and positioning of each speaker in a multi-person environment, solves the problem of identification errors and information confusion, provides a natural and smooth multi-person dialogue and interaction experience, and ensures efficient communication between digital people and multiple speakers.
Smart Images

Figure CN120162601A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of AI digital humans, and particularly to an AI digital human device and method based on voiceprint discrimination and face-to-face communication with multiple people. Background Art
[0002] With the rapid development of artificial intelligence technology, the applications of virtual characters and human-computer interaction devices are becoming more and more extensive. In many fields, how to enable virtual digital humans to interact with users in a more natural and realistic way has become an urgent technical problem to be solved.
[0003] In the existing technologies, although some virtual human interaction devices can locate users through voice recognition and image processing technologies, most devices can only provide limited interaction experiences.
[0004] Therefore, how to enable virtual digital humans to communicate with users more naturally and smoothly, and enhance the immersion and realism, has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides an AI digital human device and method based on voiceprint discrimination and face-to-face communication with multiple people, aiming to further enhance the immersion and realism of the interaction between users and virtual digital humans.
[0006] In a first aspect, an AI digital human method based on voiceprint discrimination and face-to-face communication with multiple people, the method includes:
[0007] A display screen for presenting an AI animated digital human;
[0008] A plurality of microphone units distributed based on the display screen, for independently collecting the voice information of users near the display screen respectively;
[0009] A plurality of ranging sensors distributed based on the display screen, for measuring the distances between them and the users respectively;
[0010] A plurality of cameras distributed based on the display screen, for collecting the face information of the users;
[0011] A sound source localization module for determining the sound source position according to the voice information collected by the plurality of microphone units and in combination with the distances between the plurality of ranging sensors and the users;
[0012] The visual tracking module is used to determine, based on the position of the sound source, the camera among the multiple cameras that is closest to the current sound source position, control this camera to capture an image of the user's face, and detect face feature information based on the user's face image; perform subsequent face matching and tracking based on the face feature information, and control the AI animated digital human to move within the display screen following the actual position change of the user, so that the AI animated digital human and the user maintain a face-to-face effect;
[0013] The voiceprint tracking module is used to obtain the sound signal at the position of the sound source through the beamforming algorithm, perform noise processing on the received sound signal through the noise reduction module to obtain the voice information of the target speaker, extract high-dimensional feature vectors from the voice information of the target speaker using a deep learning model, construct a voiceprint sample of the speaker based on these features, determine the target speaker by comparing with the voice information of the target speaker pre-stored in the voiceprint library, and then control the AI animated digital human to maintain a face-to-face effect with the target speaker through the visual tracking module.
[0014] In the above solution, optionally, the voiceprint library is used to store and compare the voiceprint features of speakers, and adopts a hash table storage method;
[0015] Store the voiceprint features of different speakers through a hash table, and compare them according to the voiceprint similarity to identify the speaker.
[0016] In the above solution, optionally, the display screen is an arc-shaped cylinder;
[0017] The cross-section of the display screen is semi-circular.
[0018] In the above solution, optionally, the multiple microphone units correspond one-to-one with the ranging sensors, and the microphone unit and its corresponding ranging sensor are integrated into one body;
[0019] The multiple microphone units are respectively installed on the upper border, lower border, left border and right border of the display screen;
[0020] There are multiple microphone units on the upper border of the display screen and on the lower border of the display screen, and the numbers are equal; there are multiple microphone units on the left border of the display screen and on the right border of the display screen, and the numbers are equal;
[0021] The multiple cameras are installed on the upper border of the display screen.
[0022] In the above solution, optionally, all the microphone units on the upper border of the display screen are divided into several groups with equal numbers, and the multiple cameras are respectively distributed at intervals between adjacent groups of microphone units.
[0023] In the above solution, optionally, for subsequent face matching and tracking based on the face feature information, it includes further determining the eye position coordinates of the user according to the characteristics of the facial features, and controlling the AI animated digital human to face the user directly by simulating the ray selection algorithm of the mouse for the three-dimensional virtual simulation model object.
[0024] In a second aspect, a control method for an AI digital human device based on voiceprint discrimination and face-to-face communication with multiple people includes the following steps:
[0025] Step 1: Receive the voices from multiple speakers through a microphone array, and measure the distance between each speaker and the device through an infrared laser ranging sensor;
[0026] Step 2: Use a sound localization algorithm combined with the infrared ranging data to determine the spatial position of each speaker, and perform facial tracking through a micro camera to accurately locate the eye position of the speaker;
[0027] Step 3: Enhance the sound signal from the speaker through a beamforming algorithm, and use an RNNoise noise reduction module to process the received voice signal for noise to ensure clear voice;
[0028] Step 4: Extract features from the received voice signal, and use the Mel Frequency Cepstral Coefficient technology to generate the voice feature vector of the speaker;
[0029] Step 5: Use a deep learning model to extract high-dimensional voiceprint features from the voice signal to generate a personalized voiceprint sample of the speaker;
[0030] Step 6: Compare the voiceprint features of the speaker with the data in the stored voiceprint library, and judge the identity of the speaker through cosine similarity;
[0031] Step 7: When the speaker is recognized, the digital human interacts according to the context of the conversation and adjusts the orientation of the AI animated digital human so that the digital human always faces the current speaker.
[0032] In the above solution, optionally, the voiceprint comparison in Step 6 is judged by cosine similarity. When the similarity is greater than 0.85, the identity of the speaker is confirmed.
[0033] In the above solution, optionally, the voice feature extraction in Step 4 uses the Mel Frequency Cepstral Coefficient technology and generates a 39-dimensional feature vector, including the first-order and second-order differences of the voice signal.
[0034] Compared with the prior art, the present application has at least the following beneficial effects:
[0035] This application is based on further analysis and research of the problems in the prior art, and recognizes that in multi-person voice interaction and intelligent face tracking, the prior art cannot accurately identify the identities of different speakers, locate their positions, and achieve personalized interactions according to the conversation content. By combining multiple technologies such as microphone arrays, infrared laser ranging sensors, face tracking technology, and voiceprint recognition, it is possible to accurately locate and identify multiple speakers, solving the problems of incorrect identification and information confusion in multi-person interaction in the background technology. First, the device receives sound signals from different directions through the microphone array, accurately measures the distance between the speaker and the device through the infrared ranging sensor, and at the same time uses a micro camera for face tracking to ensure that the position and eye position of each speaker can be accurately captured. Secondly, the combination of the beamforming algorithm and the RNNoise noise reduction module enables the device to accurately extract the voice of the target speaker from a complex environment and remove background noise, improving speech clarity. Through the efficient extraction of speech features by the Mel Frequency Cepstral Coefficients (MFCC) and the ECAPA-TDNN deep learning model, the system can achieve high-precision voiceprint recognition. Combining with the hash table storage method, it quickly compares the voiceprint features of the speakers, further enhancing the recognition accuracy. The digital human can dynamically adjust its display direction according to the real-time position of each speaker, always facing the current speaker, ensuring the correct connection of the context during multi-person conversations. Generally speaking, the present invention not only solves the problem of the inability to efficiently identify and locate speakers in a multi-person environment in the prior art, but also provides a natural and smooth multi-person conversation interaction experience, ensuring efficient communication between the digital human and multiple speakers. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 FIG. is a diagram of an AI digital human device based on voiceprint discrimination and face-to-face communication with multiple people provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0038] In one embodiment, as Figure 1 shown, an AI digital human device based on voiceprint discrimination and face-to-face communication with multiple people is provided, including:
[0039] A display screen for presenting an AI animated digital human;
[0040] A plurality of microphone units distributed based on the display screen for independently collecting the sound information of users near the display screen;
[0041] A plurality of ranging sensors distributed based on the display screen are used to measure the distances between them and the user respectively;
[0042] A plurality of cameras distributed based on the display screen are used to collect the face information of the user;
[0043] A sound source localization module is used to determine the sound source position according to the sound information collected by the plurality of microphone units and in combination with the distances between the plurality of ranging sensors and the user;
[0044] A visual tracking module is used to determine, according to the sound source position, the camera among the plurality of cameras that is closest to the current sound source position, control the camera to capture the user's face image, and detect face feature information based on the user's face image; perform subsequent face matching and tracking based on the face feature information, and control the AI animated digital human to move within the display screen following the actual position change of the user, so that the AI animated digital human and the user maintain a face-to-face effect;
[0045] A voiceprint tracking module is used to obtain the sound signal at the sound source position through a beamforming algorithm, perform noise processing on the received sound signal through a noise reduction module to obtain the target speaker's voice information, extract high-dimensional feature vectors from the target speaker's voice information using a deep learning model, construct a voiceprint sample of the speaker based on these features, determine the target speaker by comparing with the target speaker's voice information pre-stored in the voiceprint library, and then control the AI animated digital human to maintain a face-to-face effect with the target speaker through the visual tracking module.
[0046] In this embodiment, the voiceprint library is used to store and compare the voiceprint features of speakers, and adopts a hash table storage method;
[0047] Store the voiceprint features of different speakers through a hash table and perform comparison according to the voiceprint similarity to identify the speaker.
[0048] In this embodiment, the display screen is an arc-shaped cylinder;
[0049] The cross-section of the display screen is semi-circular.
[0050] In this embodiment, the plurality of microphone units correspond one-to-one with the ranging sensors, and the microphone unit and its corresponding ranging sensor are integrated into one body;
[0051] The plurality of microphone units are respectively installed on the upper frame, lower frame, left frame and right frame of the display screen;
[0052] There are multiple microphone units located on the upper border of the display screen and multiple microphone units located on the lower border of the display screen, and the numbers are equal; there are multiple microphone units located on the left border of the display screen and multiple microphone units located on the right border of the display screen, and the numbers are equal;
[0053] The multiple cameras are installed on the upper border of the display screen.
[0054] In this embodiment, all the microphone units located on the upper border of the display screen are divided into several groups with equal numbers, and the multiple cameras are respectively distributed at intervals between adjacent groups of microphone units.
[0055] In this embodiment, the subsequent face matching and tracking based on the face feature information includes further determining the eye position coordinates of the user according to the facial features of the face, and controlling the AI animated digital person to face the user directly through the ray selection algorithm that simulates the mouse for the three-dimensional virtual simulation model object.
[0056] In this embodiment, the microphone arrays installed on the upper, lower, left, and right borders of the screen have equal spacing between each microphone. Each microphone is equipped with an infrared laser range finder. The infrared laser range finder can measure the distance between the speaker and the device in real time, and combined with the sound signals received by the microphone array, it provides data support for further sound localization and determination of the speaker's position. This configuration can not only improve the directional sound pickup accuracy of the microphone array, but also accurately measure the distance of the speaker, making the sound localization more accurate.
[0057] After receiving the sound signals from the microphone array, the system determines the spatial position of the speaker through the sound source localization algorithm combined with the infrared laser range data. At the same time, the miniature camera installed in the device is responsible for real-time tracking of the speaker's facial position, and further optimizes the speaker localization through facial recognition technology. With the dual technical support of sound localization and facial tracking, this device can accurately obtain the position of each speaker, including the position of their eyes, so as to achieve accurate face-to-face interaction.
[0058] In this embodiment, through the beamforming algorithm, the system can enhance the sound signals from the speaker, ensuring effective separation and enhancement of the sound signals from different directions. Through precise beamforming, environmental noise can be effectively suppressed and the speech signals of the target speaker can be enhanced. The supporting RNNoise noise reduction module can process the sound signals received by the microphone array, remove background noise, and improve speech clarity, thus avoiding speech recognition errors caused by noise interference.
[0059] After completing the acquisition and processing of the sound signal, the system extracts the features of the speech signal through the Mel Frequency Cepstral Coefficient (MFCC) technology, generating a set of high-dimensional speech feature vectors. These feature vectors are used to construct the voiceprint features of each speaker, and then high-dimensional voiceprint features are extracted through a deep learning model (such as ECAPA-TDNN). The voiceprint features of each speaker are stored in the voiceprint library, and a hash table is used for efficient storage and management. Each time a new speaker passes through the device, the system will compare their voiceprint features with the data in the voiceprint library and use cosine similarity (greater than 0.85) to determine whether it is a known speaker.
[0060] This device displays a virtual 3D digital human through a flexible LED screen. According to the position of the speaker, the digital human dynamically adjusts the display direction of its virtual image to always face the current speaker. The system can continuously update the facing direction of the digital human according to the real-time position of each speaker to ensure that during a multi-person conversation, the digital human can always effectively interact with the current speaker.
[0061] This embodiment combines a microphone array, an infrared laser ranging sensor, and face tracking technology to solve the problem that traditional devices cannot accurately locate and identify each speaker in a multi-person environment. The voiceprint recognition technology can accurately identify the identity of each speaker and compare it with their voiceprint when multiple people are speaking simultaneously, avoiding misidentification and information confusion.
[0062] Through the beamforming algorithm and the RNNoise noise reduction module, the system can effectively enhance the speech signal of the target speaker and remove background noise, improving the clarity of the speech. Even in a noisy environment, the system can still maintain a high speech recognition accuracy, ensuring that the digital human's response to each speaker during multi-party interaction is clear and accurate.
[0063] The digital human can dynamically adjust its display direction according to the position of each speaker and always face the current speaker for interaction. This makes the communication process of the digital human more natural and can achieve an interactive experience similar to face-to-face communication. In this way, when the digital human switches between multiple speakers, it always maintains an efficient and error-free conversation fluency.
[0064] Through precise voiceprint comparison and dynamic position adjustment, the present invention can ensure that during a multi-person conversation, the digital human can always respond correspondingly according to the identity of each speaker. Through context management, the digital human can continuously interact reasonably with each speaker and adjust the conversation content according to the speaker's context to keep the conversation consistent and coherent.
[0065] The AI digital human device provided in this embodiment effectively solves the problems of recognition accuracy and information connection during multi-person interaction in the prior art through technological innovations such as precise speaker localization, voiceprint recognition, voice clarity enhancement, and dynamic display control. It can not only accurately identify and locate each speaker, but also maintain a smooth and natural interaction experience, providing a more intelligent and efficient solution for multi-person communication.
[0066] The AI digital human device in this embodiment can accurately locate and identify multiple speakers by combining multiple technologies such as a microphone array, an infrared laser ranging sensor, facial tracking technology, and voiceprint recognition, solving the problems of recognition errors and information confusion during multi-person interaction in the background art. First, the device receives sound signals from different directions through the microphone array, accurately measures the distance between the speaker and the device through the infrared ranging sensor, and at the same time uses a micro camera for facial tracking to ensure that the position and eye position of each speaker can be accurately captured. Second, the combination of the beamforming algorithm and the RNNoise noise reduction module enables the device to accurately extract the voice of the target speaker from a complex environment and remove background noise, improving voice clarity. Through the efficient extraction of voice features by the Mel Frequency Cepstral Coefficients (MFCC) and the ECAPA-TDNN deep learning model, the system can achieve high-precision voiceprint recognition. Combined with the hash table storage method, it can quickly compare the voiceprint features of the speaker, further enhancing the recognition accuracy. The digital human can dynamically adjust its display direction according to the real-time position of each speaker, always facing the current speaker, ensuring the correct connection of context during multi-person conversations. Generally speaking, the present invention not only solves the problem of inefficiently identifying and locating speakers in a multi-person environment in the prior art, but also provides a natural and smooth multi-person conversation interaction experience, ensuring efficient communication between the digital human and multiple speakers.
[0067] In one embodiment, a control method for an AI digital human device that discriminates voiceprints and communicates face-to-face with multiple people is provided, which is characterized by including the following steps:
[0068] Step 1: Receive the voices of multiple speakers through the microphone array and measure the distance between each speaker and the device through the infrared laser ranging sensor;
[0069] Step 2: Use the sound localization algorithm combined with the infrared ranging data to determine the spatial position of each speaker, and perform facial tracking through the micro camera to accurately locate the eye position of the speaker;
[0070] Step 3: Enhance the sound signal from the speaker through the beamforming algorithm and use the RNNoise noise reduction module to process the received voice signal for noise to ensure clear voice;
[0071] Step 4: Extract features from the received voice signal, and use Mel Frequency Cepstral Coefficient technology to generate the voice feature vector of the speaker;
[0072] Step 5: Use a deep learning model to extract high-dimensional voiceprint features from the voice signal and generate a personalized voiceprint sample of the speaker;
[0073] Step 6: Compare the voiceprint features of the speaker with the data in the stored voiceprint library, and judge the identity of the speaker through cosine similarity;
[0074] Step 7: When the speaker is recognized, the digital human interacts according to the context of the conversation and adjusts the orientation of the AI animated digital human so that the digital human always faces the current speaker.
[0075] In this embodiment, the device first receives the voices from multiple speakers through a microphone array, and measures the distance between each speaker and the device through an infrared laser ranging sensor. The system accurately determines the position of each speaker through a sound source localization algorithm combined with the infrared ranging data. In this process, the infrared laser ranging sensor can achieve high-precision distance measurement. By combining with the data of the microphone array, the system can locate the spatial coordinates of each speaker, providing support for subsequent face tracking and sound localization.
[0076] Next, the system uses a micro camera to perform real-time tracking of the faces of the speakers. The face tracking technology accurately determines the eye positions of the speakers by collecting images and using face recognition algorithms. The sound source localization and face tracking information are combined to ensure that the eye positions of each speaker are accurately identified and tracked, thereby realizing the adjustment of the digital human to face the speaker.
[0077] This control method can effectively solve the problem of inaccurate speaker localization when multiple people speak simultaneously. Through the combination of an infrared laser ranging sensor, a microphone array, and face tracking technology, the system can accurately obtain the spatial positions of each speaker and realize the orientation adjustment of the digital human. This technical solution solves the dilemma that speakers cannot be accurately located during multi-person interaction in the background technology, enabling the digital human to quickly respond and accurately face each speaker during multi-person conversations, maintaining a natural and smooth interaction. This solution greatly improves the naturalness and realism of the interaction and avoids the limitations of traditional devices that cannot adapt to multi-party conversation scenarios.
[0078] In this embodiment, the voiceprint comparison in Step 6 uses cosine similarity for judgment. When the similarity is greater than 0.85, the identity of the speaker is confirmed.
[0079] In this embodiment, the speaker localization in Step 2 is achieved through the combination of sound localization and face tracking technologies with infrared ranging to ensure the accurate identification of the speaker's position.
[0080] In this embodiment, the voice feature extraction in step 4 adopts the Mel Frequency Cepstral Coefficient (MFCC) technology and generates a 39-dimensional feature vector, including the first-order and second-order differences of the voice signal.
[0081] In one embodiment, as Figure 1 shown, an AI animated digital human device that uses voiceprint to distinguish multiple people and achieve face-to-face communication with multiple people is provided; the hollow circles are microphone devices with infrared ranging functions; the solid circles are miniature cameras installed vertically on the screen facing outward; the semi-circular part is a flexible LED screen device, and the figure inside is a virtual simulation three-dimensional digital human.
[0082] Array microphones are installed on the upper, lower, left, and right borders of the screen. The distance between each microphone is equal. An infrared laser ranging sensor is installed at the position of each microphone to enable it to have a ranging function.
[0083] When someone approaches, the eye position of the questioner is located by combining infrared ranging, sound source localization by listening, and face tracking, so as to achieve face-to-face conversation.
[0084] When the microphone receives sound, noise reduction processing is performed through the Beamforming algorithm in cooperation with the RNNoise noise reduction module.
[0085] When multiple people ask questions, first, the voice features are extracted: the Mel Frequency Cepstral Coefficient (MFCC) is used. MFCC (Mel Frequency Cepstral Coefficient): Extract the spectral features of the voice signal and simulate the auditory characteristics of the human ear. Usually, 12 - 13-dimensional MFCC coefficients are extracted, plus the first-order and second-order differences, to form a 39-dimensional feature vector.
[0086] A pre-trained deep learning model (ECAPA-TDNN) is used to extract high-dimensional feature vectors, and a voiceprint sample is defined based on the feature vectors.
[0087] After identifying and initializing a voiceprint, the voiceprint feature value of this person is stored in a voiceprint library.
[0088] Remember the voiceprints of different people in front of the screen, form memories through continuous sampling, and input them into the memory library to form a voiceprint library.
[0089] The voiceprint library is stored using a hash table storage method.
[0090] When multiple people discuss with the digital human, the digital human will compare the newly sampled voiceprint with the voiceprints in the original voiceprint library. The comparison method: by comparing the threshold, when the cosine similarity > 0.85, it is confirmed as the same voiceprint (cosine similarity: calculate the cosine angle between two voiceprint feature vectors, and the value closer to 1 indicates a higher similarity).
[0091] After the digital human determines the voiceprint of a person, it will communicate according to the Q&A in the context.
[0092] The digital human can locate the eye position of the questioner by combining infrared ranging, sound source localization by listening, and face tracking. When there are multiple people in a conversation, it will face the questioner directly.
[0093] The digital human can continuously adjust its body position with multiple people to ensure the correctness of the context of the multi-person conversation and ensure face-to-face conversation with multiple people.
[0094] In this embodiment, the digital human can find the person for context communication according to the voiceprint, ensure the connection and uniqueness of the communication information, and adjust its body position back and forth among multiple people through the positioning system to ensure that there are no mistakes in the multi-person discussion.
[0095] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
Claims
1. An AI digital human device based on voiceprint recognition and face-to-face communication with multiple people, characterized in that: include: Display screen, used to present AI animated digital humans; Multiple microphone units distributed based on the display screen are used to independently collect voice information of users near the display screen; A plurality of distance measuring sensors distributed based on the display screen, used for respectively measuring the distance between the sensors and the user; A plurality of cameras distributed based on the display screen are used to collect facial information of the user; A sound source localization module, used to determine the location of the sound source based on the sound information collected by the multiple microphone units and the distance between the multiple ranging sensors and the user; A visual tracking module, configured to determine a camera among the multiple cameras that is closest to the current sound source position according to the sound source position, control the camera to capture a user's face image, and obtain facial feature information based on the user's face image detection; Perform subsequent face matching and tracking based on the face feature information, and control the AI animated digital human to move within the display screen following the actual position change of the user, so that the AI animated digital human and the user maintain a face-to-face effect; The voiceprint tracking module is used to obtain the sound signal of the sound source position through the beamforming algorithm, and to process the received sound signal through the noise reduction module to obtain the sound information of the target speaker, and to extract high-dimensional feature vectors from the sound information of the target speaker using a deep learning model, and to construct the voiceprint sample of the speaker based on these features, and to determine the target speaker by comparing the sound information of the target speaker pre-stored in the voiceprint library, and then to control the AI animated digital human to maintain a face-to-face effect with the target speaker through the visual tracking module.
2. The AI digital human device according to claim 1, characterized in that: The voiceprint library is used to store and compare the voiceprint features of the speaker, and adopts a hash table storage method; The voiceprint features of different speakers are stored in a hash table and compared based on the voiceprint similarity to identify the speaker.
3. The device according to claim 1, characterized in that The display screen is a curved cylindrical surface; The cross section of the display screen is semicircular.
4. The device according to claim 1, characterized in that The plurality of microphone units correspond to the distance measuring sensors one by one, and the microphone units and the corresponding distance measuring sensors are integrated into one body; The plurality of microphone units are respectively mounted on the upper frame, the lower frame, the left frame and the right frame of the display screen; There are multiple microphone units located on the upper frame of the display screen and the microphone units located on the lower frame of the display screen, and the number of the microphone units is equal; There are multiple microphone units located on the left frame of the display screen and the microphone units located on the right frame of the display screen, and the number of the microphone units is equal; The multiple cameras are installed on the upper frame of the display screen.
5. The device according to claim 4, characterized in that All microphone units located on the upper frame of the display screen are divided into a number of groups of equal number, and the multiple cameras are distributed between adjacent groups of microphone units at intervals.
6. The device according to claim 1, characterized in that The subsequent face matching and tracking based on the face feature information includes further determining the user's eye position coordinates according to the facial features, and controlling the AI animated digital human to keep facing the user by simulating the ray selection algorithm of the mouse for the three-dimensional virtual simulation model object.
7. A control method for an AI digital human device based on voiceprint recognition and face-to-face communication with multiple people, characterized in that: The following steps are involved: Step 1: Receive voices from multiple speakers through a microphone array and measure the distance between each speaker and the device through an infrared laser ranging sensor; Step 2: Use the sound localization algorithm combined with infrared ranging data to determine the spatial position of each speaker, and use a micro camera to perform facial tracking to accurately locate the speaker's eye position; Step 3: Enhance the sound signal from the speaker through the beamforming algorithm, and use the RNNoise noise reduction module to process the received voice signal to ensure clear speech; Step 4: Extract features from the received speech signal and generate the speaker's speech feature vector using Mel frequency cepstral coefficient technology; Step 5: Use a deep learning model to extract high-dimensional voiceprint features from the speech signal and generate a personalized voiceprint sample for the speaker; Step 6: Compare the speaker's voiceprint features with the data in the stored voiceprint library, and determine the speaker's identity through cosine similarity; Step 7: After identifying the speaker, the digital human interacts according to the context of the conversation and adjusts the direction of the AI animated digital human so that the digital human always faces the current speaker.
8. The control method according to claim 7, characterized in that: The voiceprint comparison in step 6 is judged by cosine similarity. When the similarity is greater than 0.85, the identity of the speaker is confirmed.
9. The control method according to claim 7, characterized in that: The speech feature extraction in step 4 adopts the Mel-frequency cepstral coefficient technology and generates a 39-dimensional feature vector including the first-order and second-order differences of the speech signal.
Citation Information
Patent Citations
Sound source positioning method, apparatus and device, and storage medium
CN112578338A
Virtual digital human sight line following system and method based on sound and vehicle
CN116643713A
Method and device for identifying and tracking spokesman, electronic equipment and medium
CN118470592A