Multi-seat speaker detection and speaker separation using biometrics

By integrating biometric data like voice characteristics and facial images with spatial acoustic cues, the system accurately identifies the speaker's position and enhances their audio input, addressing the limitations of spatial acoustic cues in voice-based assistants.

WO2026019987A1PCT designated stage Publication Date: 2026-01-22CERENCE OPERATING CO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/038023
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-17
Filing Date
2025-07-17
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Voice-based assistants in vehicles struggle to accurately detect which passenger is speaking, especially when the desired speaker is outside their normal acoustic zone, leading to degraded command detection and speech enhancement due to limitations in spatial acoustic cues.

Method used

Utilizing user biometrics, such as voice characteristics and facial images, in combination with spatial acoustic cues to enhance speech detection and processing, including the use of voiceprints and visual information to determine the speaker's position and enhance their audio input.

Benefits of technology

Improves the accuracy of speech recognition and wake-up word detection by identifying the speaker's position, even when they are temporarily outside their typical zone, reducing false triggers and enhancing the quality of audio input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025038023_22012026_PF_FP_ABST
    Figure US2025038023_22012026_PF_FP_ABST
Patent Text Reader

Abstract

A method for vehicle occupant input includes acquiring a vehicle occupant data from a plurality of occupants of a vehicle, where the vehicle occupant data includes an acoustic spatial data and biometric data including at least voice characteristic data. The method further includes associating one or more positions of a plurality of positions in the vehicle with an occupant information. The occupant information associated with a position includes the biometric data for an occupant at the position based on the vehicle occupant data acquired and processing a first audio input from a first occupant of the vehicle using the voice characteristic data associated with a first position of the first occupant in the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

MULTI- SEAT SPEAKER DETECTION AND SPEAKER SEPARATION USING BIOMETRICSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. provisional application Serial No. 63 / 672,448 filed July 17, 2024, the disclosure of which is hereby incorporated in its entirety by reference herein.TECHNICAL FIELD

[0002] The present disclosure is generally directed to speaker detection and signal separation using biometrics.BACKGROUND]0003[ For cars with multiple microphones (e.g., mounted for each seat or a microphone array), voice-based assistants supporting multiple seats / passengers in a car have been successfully deployed. Acoustic cues may be used for spatial localization and / or spatial filtering to enhance performance of such assistants. For instance, acoustic signals from a desired speaker may be enhanced while suppressing noise and signals from other speakers. Also, when a command is detected to have come from a particular occupant, a subsequent dialog may be conducted with that particular occupant.SUMMARY

[0004] Use of spatial acoustic cues may have limitations. When a desired speaker is outside their normal acoustic zone, for instance by leaning while speaking, command detection and speech enhancement may be degraded.

[0005] In a general aspect, approaches described herein make use of user biometric, alone or in combination with spatial acoustic cues, to detect and process speech input from particularoccupants of a vehicle. In particular, voice characteristics may be used for such purposes, and facial images or video may further contribute.(0006] In one aspect, in general, a method for vehicle occupant input includes acquiring vehicle occupant data from a plurality of occupants of a vehicle. The vehicle occupant data includes acoustic spatial data, and biometric data including at least voice characteristic data. One or more positions of a number of different positions in the vehicle are associated with occupant information. The occupant information associated with a particular position includes biometric data for an occupant of that position based on the vehicle occupant data acquired. A first audio input from a first occupant of the vehicle is processed using voice characteristic data associated with a first position of said first occupant in the vehicle.]0007| In another aspect, in general, a method for vehicle occupant input includes acquiring vehicle occupant data from a plurality of occupants of a vehicle. This vehicle occupant data includes acoustic spatial data, and biometric data including at least voice characteristic data. One or more positions of a number of different positions in the vehicle are associated with occupant information. The occupant information associated with a particular position includes biometric data for an occupant of that position based on the acquired vehicle occupant data. A first audio input from a first occupant of the vehicle is then processed. This processing includes determining a first position associated with the first occupant based at least in part on the voice characteristic data and the occupant information, and processing the first audio data according the determined first position.(0008] Additional features and advantages are evident from the following description of one or more embodiments, and from the appended claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] FIG. 1A is a schematic side view of a vehicle with occupants in accordance with the present disclosure;

[0010] FIG. IB is a schematic top view of the vehicle of FIG. 1A in accordance with the present disclosure; and

[0011] FIG. 2 is a flowchart of an operating example of an infotainment system of the vehicle in accordance with the present disclosure.DETAILED DESCRIPTION|0012| As required, detailed embodiments of the present invention are disclosed herein; however, it is to be understood that the disclosed embodiments are merely exemplary of the invention that may be embodied in various and alternative forms. The figures are not necessarily to scale; some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ the present invention.

[0013] A voice-based assistant employed in a vehicle may not be able to accurately detect which passenger is speaking and / or uttering a command and may not enhance the audio of the speaker’s speech.

[0014] Referring to FIG. 1, a vehicle 120 includes an infotainment system 130 that provides a multi-modal user interface for interacting with one or more occupants 110 of the vehicle 120. In some implementations, the infotainment system 130 employs voice interaction as an input mode for the occupant (also referred to as “user”). This infotainment system 130 receives input from the interior cabin of the vehicle 120 via microphones 140, and via one or more cameras 150 as well as via other input modalities such as via touch sensors (colloquially “buttons”). The system 130 provides output to the occupants 110 via audio speakers 160 and visual displays 170.|0015| In some aspects, there are generally a number of defined positions for occupants 110 of the vehicle 120, such as “driver’s seat”, “front passenger seat”, “left rear passenger seat”, etc. Inputs from the occupants 110 may be interpreted by the system 130 according to the position of the occupant 110. For instance, certain commands may be exclusive to the driver position, whileother commands may be permitted by users at any position in the vehicle 120. Similarly, interpretation of a command may be based on the position of the speaker, for example, to process a command such as “open my window” or “make it warmer for me.”10016 There may be multiple microphones 140 in the vehicle 120 enabling some ability to determine the position of the user 110 that is speaking. For example, there may be different microphones associated with different available positions of the occupant 110 or zones in the vehicle 120. In addition to or in lieu of assigned microphones, there may be an array of microphones 140 that together determine directionality (e.g., direction of arrival or source location) of an audio signal emitted by a user. Directionality may be determined using the audio data from the set of available microphone signals by evaluating the signal levels between the microphones (e.g., by power ratios) or the relative time delays of the signal between microphones 140. The processing for determining directionality is often done in the frequency domain where the spectra of the microphones 140 are evaluated in terms of magnitude and phase differences. Such directionality information may be used to determine the position of a speaking occupant 110, and thereby affect the operation of the infotainment system 130.

[0017] There may be times that such zone or directionality information could be inadequate to determine the position of the occupant 110 that is speaking. For example, when there are strong interfering sounds or background noises in the vehicle 120, such as a smartphone playing from a backseat passenger, other passengers in the car are speaking at the same time, or the vehicle 120 is driving at high speed. A false detection of the position of the speaking occupant 110 may occur, if the driver is leaning toward the middle of the vehicle 120 as they adjust some controls, the system 130 may not be able to reliably distinguish between the driver and the passenger. Another aspect related to the determination of the position of the speaker is the process of avoiding false triggers of the system 130. For instance, if the system 130 is configured to use a “wakeup” word as a prefix to commands, the system 130 may preferentially trigger with such a word having a source at the driver’s position rather than from elsewhere. Again, if the driver is leaning into the center area, the system 130 may disregard their utterance of the wakeup word, or associate it with a wrong user / position. The succeeding dialog steps may result in wrong actions, e.g., for acommand “open my window” the wrong window may be opened. In the adverse acoustic situations described above, it might happen that multiple speech recognizers recognize

[0018] Voice characteristics of users may be helpful in determining which user is speaking and for enhancing the audio input to preferentially retain a particular user’s voice and suppress non-speech sounds and the speech of other users in the vehicle 120. An aspect of determining which user is speaking may involve “wakeup word arbitration,” which is the process of determining not only whether a wakeup word has been uttered by someone, but more importantly whether it has been uttered by a particular user.

[0019] More generally, a combination of directionality / zone information of the voice source and characteristics of the voice can be used to make a decision as to whether the occupant 110 seated at a particular position is speaking. One approach is to classify the input according to position independently based on acoustic directionality and on voice characteristics, and then “averaging” the classifications in some manner (e.g., if the classifications provide scores or probabilities for the various positions, then perhaps multiplying those probabilities or scores). More generally, a joint processing of directionality and voice characteristics may result in a more accurate classification.

[0020] One way of knowing the voice characteristics of the users is to explicitly enroll each potential user of a vehicle, for example, by having them speak a number of utterances into an enrollment application (which could be hosted in the infotainment system, or at a server computer), and processing those utterances to retain a representation of their voice. Such a representation may be referred to as a “voiceprint” (i.e., a voice fingerprint), and / or as a vector “embedding” (e.g., a numerical array, for instance with 256 or more elements) of their voice. To the extent that the infotainment system is informed as to which user is in which position in the vehicle 120, the audio processing of the system 130 may be able to react to a driver’s utterance even if they have moved from an acoustic input zone associated with the driver position in the vehicle 120, and may be able to enhance the acquired audio to preferably pass through the driver’s voice rather than noise of interfering voices.

[0021] But explicit voice enrollment, and explicit registering of users according to their seating in the vehicle 120 may not be practical or desirable from a usability point of view. Therefore, there is a need to be able to use other technical means as an alternative.[0022 A first means is to make use of acquired speech in the vehicle 120 while the occupants 110 are in their present positions. Audio may be acquired as the users speak, whether in conversation or in interaction with the infotainment system 130, and the system 130 builds a voiceprint associated with each position. For example, the system 130 may recognize that there are three occupants (e.g., a driver, a front passenger, and a left rear passenger), and build a voiceprint for those three otherwise unknown occupants 110. The voiceprint of the driver can then be used to enhance their speech, and to detect commands (e.g., wakeup words) even when they are temporarily outside the zone associated with their normal position (e.g., leaning into the middle of the vehicle).

[0023] A variant of the first means retains some history between trips, recognizing that many users may make repeated trips in a particular vehicle 120, although they might occupy different positions on different trips. For example, a husband and wife may use the vehicle 120 individually, and when they drive together, they may alternate driving and being a passenger. The system 130 may build long-term voice profiles of the users, and when initial audio is acquired on a particular trip, rather than having to build a voiceprint for a user from scratch, the system 130 may be able to identify the repeated occupant and quickly adopt an existing voiceprint. The advantages of such a variant may be in reducing the amount of acquired audio to build a voiceprint for a particular user (and their occupant position) and in increasing the accuracy or utility of the voiceprint for a given amount of acquired audio.

[0024] A second means, which may be used in combination with the first means, is to use technical means to determine which user is at which occupant position. Cameras that may acquire facial images of the occupant 110 is one such technical means. The system 130 may acquire such an image of an occupant and discriminate between different individuals so that it can determine whether a previously seen user is occupying a position. For instance, the system 130 may simply keep an enumeration of individuals (e.g., user-1, user-2, ...) and data suitable for determiningwhen a new facial image matches one of those previously seen users. As an example, facial images are acquired and processed to form vector embeddings and prototypical such embeddings (e.g., determined by a clustering or averaging approach) are retaining in association with an identifier (e.g., user-n).[0025| When atrip begins, the system 130 then identifies the occupants 110 as new or previously seen. For a new user, the system 130 may build a voiceprint when it acquires speech from that user, and store the voiceprint in association with that user’s facial prototype. For a user that is identified as have been previously seen, and to the extent that the voiceprint for that user is already stored, the system 130 may adopt that voiceprint without acquiring any audio from that user on a particular trip. In the scenario in which a husband and wife share a car, the system 130 is able to quickly identify who is driving, and to adopt that user’s voiceprint based on the facial image.

[0026] Image data is not necessarily limited to use for identification and assignment of voiceprints to occupant positions. For example, acquired images or video may be used to estimate the speaking direction, and that direction may be used to improve directional acquisition of audio. In some instances, the choice of microphone sensitivity beam patterns may be based not only on the acoustic source location (i.e., where is the speaker’s head) but also in the direction of acoustic emissions (i.e., in what direction are they speaking). As another example of use of video information, lip movement may be used to enhance the detection of when the user is speaking - it is unlikely that a driver has uttered a wakeup word if his or her lips did not move. More detailed information may also be gleaned from the particular pattern of lip movement, and wakeup word detection or full utterance transcription may make use of the evolution of lip movement and acoustic spectral characteristics. Such a combination may be particularly effective in noisy environments.

[0027] There are other technical means of determining which user is occupying which position in the vehicle 120. For example, locations of smartphones or other devices (e.g., watches) in the vehicle 120 may be used assuming that repeated occupants are carrying their phone on separate trips. Yet other biometric information might be used, such as the weights of occupants in the vehicle 120 with weight sensors in the seats.

[0028] In some implementations, the vehicle cabin is associated with distinct spatial zones, with each zone associated with a particular vehicle position (e.g., driver’s seat). In operation, speech recognition and wakeup word detection is performed in parallel for each of the zones. To the extent that each occupant is within a zone for their position, it is expected that their voice will be strongest in the audio acquisition for that zone (e.g., using a microphone or an array beam pattern for that zone). In such a case, then the corresponding recognition and wakeup word process detects the occupant speaking and processes the audio and contained language accordingly. Due to remaining crosstalk (that might even be present after speech signal enhancement and speaker separation) multiple recognizers may recognize speech for the same utterance that is spoken by a user from his zone in the car. It is the task of an “arbitration” mechanism needs to decide for which zone the recognition result is accepted. To the extent that the occupant 110 is detected in a zone that is not associated with their position, for example, because they are leaning or directing their speech in an atypical direction, the system 130 may detect speech from another zone, but nevertheless be able to determine the position of the occupant 110 speaking based on a match of their voice characteristics.

[0029] In some examples, a relatively small or limited number of voice profdes maybe known for actual or candidate occupants of the vehicle 120. For instance, if occupants 110 have been identified and associated with occupant positions using cameras in the vehicle 120, a database of voice profiles may be accessed to obtain the voice profiles to be considered when processing audio input. Such restriction to voice profiles of occupants 110 may increase accuracy (e.g., increase detection rate and / or reduce false acceptance rate) as compared to considering a larger number of voice profiles that may include profiles of people not in the vehicle 120.

[0030] In some examples, the process of determining the position of the speaker may be related to a process for speech enhancement. For example, a number of speech enhancement techniques benefit from knowing a voice profile of a target speaker. Essentially, with such a known voice profile for a desired speaker, voice and other acoustic signals that do not match that profile may be more easily suppressed or cancelled. Therefore, upon acquisition of a speech input, the joint process of determining the speaker and their position along with enhancing the speech of thatspeaker may provide a high-quality audio input along with the speaker and / or position information that is used to process the content of that audio input.(0031] Referring to FIG. 2, an operating example 200 of an embodiment can be understood with reference to a number of functional modules, and data processing and transformation steps implemented by those modules. In a non-limiting example, the operating example including the modules are performed by the infotainment system 130. Microphones or other audio sources (e.g., microphones 140) provide raw audio acquired in the vehicle cabin, at operation 282. A speech signal enhancement (SSE) module 210 uses the raw audio to preform speaker segregation of speaker signals at operation 284 in the raw audio. An automatic speech recognition (ASR) module 220 detects a wake-up word that is spoken from any of the positions / seats in the vehicle 120 at operation 286. A voice biometry module 230 further processes the speaker signals to detect who has spoken the wake-up word at operation 288. This detection is informed by a relevant subset of user profiles, which includes voice profiles of multiple users. A camera or other video source (e.g., camera 150) acquires images of people in the vehicle 120, at operation 292, and a face biometry module 260 detects the people, at operation 294 and links the people to profiles from a set of profiles 240. A set of relevant profiles are determined based on the detected people to form the relevant subset of profiles 250, and each of the profiles is associated with a seat / position in the vehicle 120 based on visual information at operation 296. Both the detection of who has spoken the wake-up word and the seat / position of that user are used to determine the seat from which the wake-up word was spoken at operation 298.(0032] The process 200 and other aspects of the present disclosure may provide improved voice processing, for instance improved speech recognition and / or wake-up word detection by using the voice characteristics. The present disclosure may also provide the possibility of processing audio input according to the position of the speaker, even if that speaker has temporarily moved outside a typical spatial zone associated with the speaker’s determined position. For instance, processing of a driver’s input may be specific to the driver, even if the driver has leaned outside the typical driver’s zone when speaking.

[0033] Use of visual information may have further benefits. In the case of multi-seat wake-up word detection, edge cases where passengers are moving their heads from the nominal positions, can be identified in advance. For passengers leaning into the middle of the car, speech signal enhancement (SSE) can be informed, and SSE can adjust its internal models for seat detection and adaptation control. With visual information, lip movements can also be detected. With this additional input, SSE can benefit in voice activity detection, internal adaptation control, and SSE- based zone detection.

[0034] A number of embodiments of the invention have been described. Nevertheless, it is to be understood that the foregoing description is intended to illustrate and not to limit the scope of the invention, which is defined by the scope of the following claims. Accordingly, other embodiments are also within the scope of the following claims. For example, various modifications may be made without departing from the scope of the invention. Additionally, some of the steps described above may be order independent, and thus can be performed in an order different from that described.

[0035] Aspects of the present embodiments such as, but not limited to, the infotainment system 130, the SSE module 210, the ASR module, the voice biometry 230, the face biometry 260, and / or the process 200, may be embodied as a system, method, or computer program product. Accordingly, the aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a “module” or “system.” Furthermore, the aspects of the present disclosure may take the form of a computer program product embodied in one or more computer readable medium(s) having computer readable program code embodied thereon.

[0036] Any combination of one or more computer readable medium(s) may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: an electrical connectionhaving one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (erasable programmable read-only memory (EPROM) or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0037] Aspects of the present disclosure are described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general -purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / acts specified in the flowchart and / or block diagram block or blocks. Such processors may be, without limitation, general purpose processors, special-purpose processors, application-specific processors, or field- programmable.

[0038] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagramsand / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0039] The description of the disclosure is merely exemplary in nature and, thus, variations that do not depart from the substance of the disclosure are intended to be within the scope of the disclosure. Such variations are not to be regarded as a departure from the spirit and scope of the disclosure.

Claims

WHAT IS CLAIMED IS:

1. A method for vehicle occupant input, comprising: acquiring a vehicle occupant data from a plurality of occupants of a vehicle, the vehicle occupant data comprising an acoustic spatial data and biometric data including at least voice characteristic data; associating one or more positions of a plurality of positions in the vehicle with an occupant information, wherein the occupant information associated with a position comprises the biometric data for an occupant at the position based on the vehicle occupant data acquired; and processing a first audio input from a first occupant of the vehicle using the voice characteristic data associated with a first position of the first occupant in the vehicle.

2. The method of claim 1, wherein: acquiring the vehicle occupant data comprises acquiring audio signals in association with the acoustic spatial data, and associating the one or more positions with the occupant information includes processing the audio signals determined to have a source at the first position in the vehicle to determine the voice characteristic data associated with the first position.

3. The method of claim 1, acquiring the vehicle occupant data comprises acquiring facial characteristic data, the biometric data comprising the facial characteristic data.

4. The method of claim 3, wherein acquiring facial characteristic data comprises acquiring a facial image of the first occupant at the first position, determining the facial characteristic data from the facial image, and associating the voice characteristic data for the first occupant with the facial characteristic data.

5. The method of claim 2, wherein processing the first audio input comprises at least one of: enhancing the audio signals acquired, suppressing audio from other occupants provided in the audio signals,suppressing non-speech sounds in the audio signals, or determining a spoken command, using the voice characteristic data associated with the first position.

6. The method of claim 1, wherein the processing of the first audio input includes interpreting a spoken command using prior inputs spoken by an occupant of the first position in the vehicle.

7. The method of claim 6, wherein the first position in the vehicle is a driver position in said vehicle.

8. A method for vehicle occupant input, comprising: acquiring a vehicle occupant data from a plurality of occupants of a vehicle, the vehicle occupant data comprising an acoustic spatial data and biometric data including at least a voice characteristic data; associating one or more positions of a plurality of positions in the vehicle with an occupant information, wherein the occupant information associated with a position comprises biometric data for an occupant of the position based on the vehicle occupant data acquired; processing a first audio data from a first occupant of the vehicle, including determining a first position associated with the first occupant based at least in part on the voice characteristic data and the occupant information, and processing the first audio data according to the first position determined.

9. The method of claim 8, wherein processing the first audio data comprises detecting a wake-up word in the first audio data.

10. The method of claim 9, wherein processing the first audio data comprises determining whether to act on a command in the first audio data based on the first position determined.

11. The method of claim 8, further comprising processing a first audio input using the voice characteristic data associated with the first position.

12. A machine readable medium comprising instructions stored thereon, said instructions when executed by a processing system cause said system to perform the method of claim 8.

13. A vehicle infotainment system for use in a vehicle, the system comprising a processor configured to perform the method of claim 8.

14. The vehicle infotainment system of claim 13, further comprising a non- transitory memory for storing the biometric data for a plurality of potential occupants of the vehicle.

15. The vehicle infotainment system of claim 14, further comprising microphones for acquiring audio input from an occupant of the vehicle or one or more cameras for acquiring facial images of and the occupant of the vehicle.

Citation Information

Patent Citations

  • Vehicle based determination of occupant audio and visual input

    US20140214424A1

  • Speaker-specific speech filtering for multiple users

    US20240212689A1

  • Voice assistant optimization dependent on vehicle occupancy

    WO2023122283A1