Speaker-specific speech filtering for multiple users

By enabling speaker-specific voice input filters for each user, the confusion and interruption problems of voice assistant system in multi-person environments are solved, and independent voice processing and efficient user experience are achieved.

CN120457484APending Publication Date: 2025-08-08QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084802.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-11-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In a multi-person environment, it is difficult for the prior art to effectively separate and process voice input from multiple users, resulting in confusion of voice assistant systems and poor user experience, waste of resources and interjection problems.

Method used

By selectively enabling speaker-specific voice input filters, a personalized voice output signal is generated based on the voice signature data of each user, filtering out voice interference from other users, and realizing independent voice processing.

Benefits of technology

Improves user experience, reduces resource waste, allows multiple users to have independent voice assistant conversations at the same time, avoids interruptions and confusion, and improves the efficiency and accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120457484A_ABST
    Figure CN120457484A_ABST
Patent Text Reader

Abstract

An apparatus includes one or more processors configured to detect speech of a first user and a second user, and obtain first speech signature data associated with the first user and second speech signature data associated with the second user. The one or more processors are configured to selectively enable a first speaker-specific voice input filter based on the first voice signature data to generate a first voice output signal corresponding to the voice of the first user. The one or more processors are further configured to selectively enable a second speaker-specific voice input filter based on the second voice signature data to generate a second voice output signal corresponding to the voice of the second user.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] I. Cross-reference to Related Applications

[0002] This application claims priority to commonly owned U.S. non-provisional patent application No. 18 / 069,649, filed on December 21, 2022, the contents of which are expressly incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure generally relates to filtering audio data to process the speech of multiple users.

[0004] III. Related Technology

[0005] Technological advances have resulted in smaller and more powerful computing devices. Many of these devices can communicate voice and data packets over wired or wireless networks. Furthermore, many such devices incorporate additional functionality, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Furthermore, such devices can process executable instructions, including software applications, such as web browser applications, which can be used to access the Internet.

[0006] Many of these devices incorporate functionality for interacting with users via voice commands. For example, a computing device may include a voice assistant application and one or more microphones to generate audio data based on detected sounds. In this example, the voice assistant application is configured to perform various operations in response to the user's voice, such as transmitting commands to other devices, retrieving information, and the like.

[0007] While voice assistant applications enable hands-free interaction with computing devices, controlling computing devices using voice is not without its complications. For example, when a computing device is in a noisy environment, it can be difficult to separate speech from background noise. As another example, when multiple people are present, speech from multiple people may be detected, resulting in mixed input to the computing device and a less-than-satisfactory user experience. Summary of the Invention

[0008] According to one embodiment of the present disclosure, a device includes one or more processors configured to detect speech of a first user and a second user, and obtain first speech signature data associated with the first user and second speech signature data associated with the second user. The one or more processors are configured to selectively enable a first speaker-specific speech input filter based on the first speech signature data to generate a first speech output signal corresponding to the speech of the first user. The one or more processors are further configured to selectively enable a second speaker-specific speech input filter based on the second speech signature data to generate a second speech output signal corresponding to the speech of the second user.

[0009] According to another specific implementation of the present disclosure, a method includes: detecting speech of a first user and a second user at one or more processors; and obtaining, at the one or more processors, first speech signature data associated with the first user and second speech signature data associated with the second user. The method includes: selectively enabling, at the one or more processors, a first speaker-specific speech input filter based on the first speech signature data to generate a first speech output signal corresponding to the speech of the first user. The method also includes: selectively enabling, at the one or more processors, a second speaker-specific speech input filter based on the second speech signature data to generate a second speech output signal corresponding to the speech of the second user.

[0010] According to another embodiment of the present disclosure, a non-transitory computer-readable medium stores instructions that are executable by one or more processors to cause the one or more processors to detect speech of a first user and a second user, and obtain first speech signature data associated with the first user and second speech signature data associated with the second user. The one or more processors are executable to selectively enable a first speaker-specific speech input filter based on the first speech signature data, thereby generating a first speech output signal corresponding to the speech of the first user. The one or more processors are also executable to selectively enable a second speaker-specific speech input filter based on the second speech signature data, thereby generating a second speech output signal corresponding to the speech of the second user.

[0011] According to another embodiment of the present disclosure, a device includes components for detecting speech of a first user and a second user. The device includes components for obtaining first speech signature data associated with the first user and second speech signature data associated with the second user. The device includes components for selectively enabling a first speaker-specific speech input filter based on the first speech signature data to generate a first speech output signal corresponding to the speech of the first user. The device also includes components for selectively enabling a second speaker-specific speech input filter based on the second speech signature data to generate a second speech output signal corresponding to the speech of the second user.

[0012] Other aspects, advantages, and features of the present disclosure will become apparent upon review of the entire application, including the Brief Description of the Drawings, Detailed Description, and Claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a block diagram of certain illustrative aspects of a system operable to perform speaker-specific speech filtering for multiple users according to some examples of the present disclosure.

[0014] Figure 2 is a diagram of a first example of a vehicle operable to perform speaker-specific speech filtering for multiple users according to some examples of the present disclosure.

[0015] Figure 3A is a diagram of illustrative aspects of operations associated with speaker-specific speech filtering for multiple users, according to some examples of the present disclosure.

[0016] Figure 3B is a diagram of illustrative aspects of operations associated with speaker-specific speech filtering for multiple users, according to some examples of the present disclosure.

[0017] Figure 3C is a diagram of illustrative aspects of operations associated with speaker-specific speech filtering for multiple users, according to some examples of the present disclosure.

[0018] Figure 4 is a diagram of illustrative aspects of operations associated with speaker-specific speech filtering for multiple users, according to some examples of the present disclosure.

[0019] Figure 5 is a diagram of illustrative aspects of operations associated with speaker-specific speech filtering for multiple users, according to some examples of the present disclosure.

[0020] Figure 6 is a diagram of illustrative aspects of operations associated with speaker-specific speech filtering for multiple users, according to some examples of the present disclosure.

[0021] Figure 7 is a diagram of a voice-controlled speaker system operable to perform speaker-specific speech filtering for multiple users according to some examples of the present disclosure.

[0022] Figure 8 An example of an integrated circuit operable to perform speaker-specific speech filtering for multiple users according to some examples of the present disclosure is illustrated.

[0023] Figure 9 is a diagram of a mobile device operable to perform speaker-specific voice filtering for multiple users according to some examples of the present disclosure.

[0024] Figure 10 is a diagram of a wearable electronic device operable to perform speaker-specific speech filtering for multiple users according to some examples of the present disclosure.

[0025] Figure 11 is a diagram of a camera operable to perform speaker-specific speech filtering for multiple users according to some examples of the present disclosure.

[0026] Figure 12 is a diagram of a head-mounted device (such as a virtual reality, mixed reality, or augmented reality head-mounted device) that is operable to perform speaker-specific speech filtering for multiple users according to some examples of the present disclosure.

[0027] Figure 13 is a diagram of a second example of a vehicle operable to perform speaker-specific speech filtering for multiple users according to some examples of the present disclosure.

[0028] Figure 14 According to some examples of the present disclosure Figure 1 A diagram of illustrative aspects of the operation of components of a system.

[0029] Figure 15 According to some examples of the present disclosure, Figure 1 FIG. 1 is a diagram illustrating a specific implementation of a method for speaker-specific speech filtering for multiple users performed by a device.

[0030] Figure 16 According to some examples of the present disclosure, Figure 1 FIG. 1 is a diagram illustrating a specific implementation of a method for speaker-specific speech filtering for multiple users performed by a device.

[0031] Figure 17 According to some examples of the present disclosure, Figure 1 FIG. 1 is a diagram illustrating a specific implementation of a method for speaker-specific speech filtering for multiple users performed by a device.

[0032] Figure 18 is a block diagram of a particular illustrative example of a device operable to perform speaker-specific speech filtering for multiple users in accordance with some examples of the present disclosure. DETAILED DESCRIPTION

[0033] According to certain aspects disclosed herein, speaker-specific voice input filters are selectively used to generate voice input for one or more voice assistants from multiple users. For example, in some implementations, each speaker-specific voice input filter is activated in response to detecting the voice of a corresponding user from a plurality of users (such as a wake-up word in an utterance). In such implementations, each speaker-specific voice input filter, when enabled, is configured to process the received audio data to enhance the voice of the specific user associated with the speaker-specific voice input filter. Enhancing the voice of a specific user may include, for example, reducing background noise in the audio data, removing the voices of one or more other people from the audio data, and the like.

[0034] Traditionally, voice assistants enable hands-free interaction with computing devices; however, when multiple people are present, the operation of the voice assistant may be interrupted or confused by the voices from multiple people. As an example, a first person may initiate an interaction with the voice assistant by saying a wake-up word and a command in sequence. In this example, if a second person speaks while the first person is speaking to the voice assistant, the first person's voice and the second person's voice may overlap, making it impossible for the voice assistant to correctly interpret the command from the first person. This confusion results in an unsatisfactory user experience and waste (because the voice assistant processes audio data without generating the requested results). For illustration, this confusion may result in inaccurate speech recognition, resulting in an inappropriate response from the voice assistant.

[0035] Another example may be referred to as barging in. In the barging case, a first person may initiate an interaction with the voice assistant by speaking a wake word followed by a first command, in sequence. In this example, a second person may interrupt the interaction between the first person and the voice assistant by speaking the wake word (possibly followed by a second command) before the voice assistant completes the action associated with the first command. When the second person barges in, the voice assistant may stop performing the action associated with the first command to process the input from the second person (e.g., the second command). Barging in results in a less than satisfactory user experience and is wasteful in a similar manner to obfuscation because the voice assistant processes the audio data associated with the first command without generating the requested result.

[0036] Because of these issues, systems that provide conventional voice assistant services to multiple people (such as in a car) limit voice assistant interactions to one person at a time, even if the system can support multiple voice assistants. For example, when a car occupant interacts with a particular voice assistant by saying its first wake-up word (e.g., "Hey, Assistant"), all subsequently said wake-up words for the particular voice assistant and other supported voice assistants are disabled while the particular voice assistant is in listening mode. The user experience for the car occupants would be improved if they could interact with the voice assistants simultaneously rather than one at a time.

[0037] According to certain aspects, selectively enabling speaker-specific voice input filters enables an improved user experience and more efficient use of resources (e.g., power, processing time, bandwidth, etc.). For example, the speaker-specific voice input filter can be enabled in response to detecting a wake-up word in an utterance from a first person. In this example, the speaker-specific voice input filter is configured to provide filtered audio data corresponding to speech from the first person to the voice assistant based on voice signature data associated with the first person. The speaker-specific voice input filter is configured to remove speech from other people from the filtered audio data provided to the voice assistant. As a result, the first person can conduct the voice assistant session uninterrupted, thereby improving resource utilization and improving the user experience.

[0038] Another benefit of selectively enabling speaker-specific voice input filters for multiple users is that multiple virtual assistant sessions can be conducted simultaneously because each speaker-specific voice input filter is configured to remove speech from other users. For example, the speech of each user participating in a virtual assistant session is removed from the speech of each other user provided to the corresponding virtual assistant session of the other user. Thus, even when the users are in close proximity to each other, such as when the users are occupants of a car, airplane, or other vehicle, each of the multiple users can simultaneously participate in a different corresponding voice assistant session without interference between the multiple voice assistant sessions.

[0039] In the context of a car or other vehicle, a voice assistant service provided by the vehicle may allow multiple passengers to conduct multiple conversations concurrently. According to some aspects, when a first passenger in a vehicle cabin invokes a voice assistant, passengers in other cabins may also invoke the voice assistant while the first passenger's voice assistant conversation is ongoing. For example, passenger identities and regional information regarding the passenger's location within the vehicle may be used to isolate and distinguish the voices of multiple passengers to reduce or eliminate interference between multiple concurrent voice assistant conversations.

[0040] According to some aspects, once the vehicle is in motion, one or more other modalities and controller area network (CAN) bus information (such as seat weight sensor information) can be used to track the number of seated passengers. By monitoring speech in the vehicle cabin, the identity of each seated passenger can be established and their identity "locked" relative to their location in the cabin, regardless of the vehicle's voice activation or operating conditions. Based on the locked identity of the passenger in each zone, speaker-dependent speech enhancement is provided in that zone to generate identity-aware regional "voice bubbles." Based on each passenger's identity and regional information, other passengers can be enabled to invoke the assistant in parallel or interrupt an existing assistant session. The regional voice and CAN bus weight sensors in the vehicle cabin can be continuously monitored to update passenger identity information.

[0041] Certain aspects of the present disclosure are described below with reference to the accompanying drawings. In this specification, common features are designated by common reference numerals. As used herein, various terms are used only for the purpose of describing a particular implementation and are not intended to limit the implementation. For example, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. In addition, some features described herein are singular in some implementations and plural in other implementations. For the purpose of illustration, Figure 1 Depicts a system comprising one or more processors ( Figure 1 190), indicating that in some implementations, the device 102 includes a single processor 190, while in other implementations, the device 102 includes multiple processors 190. For ease of reference herein, such features are generally introduced as "one or more" features and are subsequently referred to in the singular or optionally in the plural (as generally indicated by "(s)") unless aspects relating to multiples of a feature are being described.

[0042] In some of the drawings, multiple instances of a particular type of feature are used. Although the features are physically and / or logically different, the same reference numeral is used for each feature, and the different instances are distinguished by adding a letter to the reference numeral. When features as a group or type are referred to herein (e.g., when no reference is made to a particular one of the features), the reference numeral is used without a distinguishing letter. However, when a particular feature of a plurality of features of the same type is referred to herein, the reference numeral is used with a distinguishing letter. For example, reference numerals are used for the features of the same type. Figure 2 , a plurality of microphones are illustrated and are associated with reference numerals 104A to 104F. When referring to a specific one of these microphones, such as microphone 104A, a distinguishing letter "A" is used. However, when referring to any of these microphones or to the microphones as a group, reference numeral 104 is used without a distinguishing letter.

[0043] As used herein, the term "comprise" may be used interchangeably with "include". Additionally, the term "wherein" may be used interchangeably with "wherein". As used herein, "exemplary" indicates an example, a specific implementation and / or an aspect and should not be interpreted as limiting or indicating a preference or preferred specific implementation. As used herein, ordinal terms (e.g., "first," "second," "third," etc.) used to modify an element (such as a structure, component, operation, etc.) do not themselves indicate any priority or order of the element relative to another element, but simply distinguish the element from another element with the same name (but using an ordinal term). As used herein, the term "set" refers to one or more specific elements of a specific element, while the term "plurality" refers to a plurality (e.g., two or more) of specific elements.

[0044] As used herein, "coupling" may include "communicatively coupled," "electrically coupled," or "physically coupled," and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., a wired network, a wireless network, or a combination thereof), etc. As illustrative, non-limiting examples, two electrically coupled devices (or components) may be included in the same device or in different devices, and may be connected via electronics, one or more connectors, or inductive coupling. In some implementations, two devices (or components) that are communicatively coupled (such as electrically connected) may transmit and receive signals (e.g., digital signals or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, "direct coupling" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without an intermediate component.

[0045] In this disclosure, terms such as "determine," "calculate," "estimate," "shift," "adjust," etc. may be used to describe how to perform one or more operations. It should be noted that such terms should not be interpreted as limiting, and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, "generate," "calculate," "estimate," "use," "select," "access," and "determine" may be used interchangeably. For example, "generating," "calculating," "estimating," or "determining" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or may refer to using, selecting, or accessing a parameter (or signal) that has already been generated (such as by another component or device).

[0046] Figure 1A specific implementation of a system 100 is illustrated that is operable to perform speaker-specific voice filtering for multiple users to selectively filter audio data provided to one or more voice assistant applications. The system 100 includes a device 102 that includes one or more processors 190 and a memory 142. The device 102 is coupled to or includes: one or more microphones 104 coupled via an input interface 114 of the processor 190, and one or more audio transducers 164 (e.g., speakers) coupled via an output interface 160 of the processor 190.

[0047] exist Figure 1 , a microphone 104 is positioned in an acoustic environment to receive sounds 106. Sounds 106 may include, for example, speech 108 from one or more people 180, ambient sounds 112, or both. Microphone 104 is configured to provide a signal to input interface 114 to generate audio data 116 representing sounds 106. Audio data 116 is provided to processor 190 for processing, as further described below.

[0048] exist Figure 1 In the illustrated example, the processor 190 includes an audio analyzer 140. The audio analyzer 140 includes an audio preprocessor 118 and a multi-stage speech processor that includes a first-stage speech processor 124 and a second-stage speech processor 154. In a particular implementation, the first-stage speech processor 124 is configured to perform wake-up word detection, and the second-stage speech processor 154 is configured to perform more resource-intensive speech processing, such as speech-to-text conversion, natural language processing, and related operations. In order to conserve resources (e.g., power, processor time, etc.) associated with the resource-intensive speech processing performed at the second-stage speech processor 154, the first-stage speech processor 124 is configured to provide audio data 150 to the second-stage speech processor 154 after the first-stage speech processor 124 detects the wake-up word 110 in the speech 108 from the person 180. In some implementations, the secondary voice processor 154 remains in a low-power or standby state until the primary voice processor 124 signals the secondary voice processor 154 to wake up or enter a high-power state to process the audio data 150. In some such implementations, the primary voice processor 124 operates in an always-on mode, such that the primary voice processor 124 is always listening for the wake word 110. However, in other such implementations, the primary voice processor 124 is configured to be activated by some additional action, such as a button press.

[0049] A technical benefit of such a multi-stage speech processor is that the most resource-intensive operations associated with speech processing can be offloaded to the second-stage speech processor 154, which can be active only when a voice assistant session is ongoing after the wake word 110 is detected, thereby saving power, processor time, and other computing resources associated with the operation of the second-stage speech processor 154. In specific implementations where power, processor time, and other computing resources are relatively abundant, such as when Figure 2 When implemented in a passenger vehicle as described in , the first stage speech processor 124 , the second stage speech processor 154 , or both may remain active and may be combined into a single processor stage.

[0050] Although the second level speech processor 154 Figure 1 1 as being included in the device 102, but in some implementations, the second-level speech processor 154 is remote from the device 102. For example, the second-level speech processor 154 may be located at a remote voice assistant server. In such implementations, after the first-level speech processor 124 detects the wake word 110, the device 102 sends the audio data 150 to the second-level speech processor 154 via one or more networks. A technical benefit of this arrangement is that because the audio data 150 transmitted to the second-level speech processor 154 only represents a subset of the audio data 116 generated by the microphone 104, communication resources associated with sending the audio data to the second-level speech processor 154 are conserved. Additionally, by not transmitting all of the audio data 116 to the remote voice assistant server, power, processor time, and other computing resources associated with the operation of the second-level speech processor 154 at the remote voice assistant server are conserved.

[0051] exist Figure 1, the audio preprocessor 118 includes a plurality of speech input filters 120 that are configurable to operate as speaker-specific speech input filters. In this context, a "speaker-specific speech input filter" refers to a filter that is configured to enhance the speech of one or more specified persons. For example, a speaker-specific speech input filter associated with person 180A may be operable to enhance the speech of utterance 108A from person 180A. To illustrate, enhancing the speech of person 180A may include attenuating portions (or components) of the audio data 116 that do not correspond to speech from person 180A, such as portions of the audio data 116 representing ambient sounds 112, portions of the audio data 116 representing utterance 108B from person 180B, or both. Similarly, the speaker-specific speech input filter associated with person 180B may be operable to enhance speech from person 180B's speech 108B, which may include attenuating portions (or components) of the audio data 116 representing ambient sounds 112, portions of the audio data 116 representing person 180A's speech 108A, or both.

[0052] exist Figure 1 In the illustrated implementation, speech input filter 120A is configured as a speaker-specific speech input filter to receive audio data 116 and generate speech output signal 152A, wherein portions or components of audio data 116 that do not correspond to speech from person 180A are attenuated or removed. Similarly, speech input filter 120B is configured as a speaker-specific speech input filter to receive audio data 116 and generate speech output signal 152B, wherein portions or components of audio data 116 that do not correspond to speech from person 180B are attenuated or removed. Speech input filter 120 may include one or more additional input filters (not shown) that are not configured as speaker-specific speech input filters and that may apply general signal filtering (e.g., echo cancellation, noise suppression, etc.) to audio data 116 to generate output signals. Speech input filter 120 may also include one or more additional filters (not shown) configured as speaker-specific speech input filters for other users. (As used herein, a "user" of device 102 is a person who has initiated a voice interaction with device 102.) Specifically, although the operation of device 102 is generally described in the context of providing speaker-specific voice input filtering for person 180A and person 180B, device 102 may be operable to provide speaker-specific voice input filtering for any number of users. The output signal generated by voice input filter 120 is provided to first-stage voice processor 124 as filtered audio data 122. Filtered audio data 122 may include multi-channel data. For example, filtered audio data 122 may include a different channel for the output of each active voice input filter 120.

[0053] In certain implementations, processor 190 is configured to selectively enable voice input filter 120 to operate as a speaker-specific voice input filter, such as based on detecting wake word 110. For example, in response to detecting wake word 110A in utterance 108A from person 180A, processor 190 retrieves voice signature data 134A associated with person 180A, and voice input filter 120A uses voice signature data 134A to generate a voice output signal 152A corresponding to the voice of person 180A based on audio data 116. As a simplified example, voice input filter 120A compares input audio data (e.g., audio data 116) with voice signature data 134A to generate voice output signal 152A that attenuates (e.g., removes) portions or components of the input audio data that do not correspond to the voice of person 180A. Similarly, in response to detecting the wake word 110B in the speech 108B from the person 180B, the processor 190 retrieves the voice signature data 134B associated with the person 180B, and the voice input filter 120B uses the voice signature data 134B to generate a voice output signal 152B corresponding to the voice of the person 180B based on the audio data 116. In some implementations, the voice input filter 120 includes one or more trained models, such as those described in reference to FIG. Figures 4 to 6 As further described, the speech signature data 134 includes one or more speaker embeddings that are provided as input to the speech input filter 120 along with the audio data 116 to customize the speech input filter 120 to operate as a speaker-specific speech input filter.

[0054] In a particular implementation, the audio analyzer 140 includes a speaker detector 128 operable to determine a speaker identifier 130 for each person 180 whose voice is detected or whose wake word 110 is detected to be spoken. Figure 1 , the audio preprocessor 118 is configured to provide filtered audio data 122 to the first-level speech processor 124. In this example, before detecting the wake word 110 (e.g., when no voice assistant session is in progress), the audio preprocessor 118 may perform non-speaker specific filtering operations, such as noise suppression, echo cancellation, etc. In this example, the first-level speech processor 124 includes a wake word detector 126 and a speaker detector 128. The wake word detector 126 is configured to detect one or more wake words, such as the wake word 110A in the utterance 108A from the person 180A and the wake word 110B in the utterance 108B from the person 180B. As further described below, different wake words 110 can be used to initiate a session with different voice assistant applications 156.

[0055] In response to detecting the wake word 110, the wake word detector 126 causes the speaker detector 128 to determine an identifier (e.g., speaker identifier 130) of a person 180 associated with the utterance 108 in which the wake word 110 was detected. In certain implementations, the speaker detector 128 is operable to generate voice signature data based on the utterance 108 and compare the voice signature data with voice signature data 134 in the memory 142. The voice signature data 134 in the memory 142 may be included in the registration data 136 associated with the set of registered users associated with the device 102. In other implementations, the device 102 uses sensor data (e.g., image data of the user's face or other biometric data) to identify the person 180 via comparison with corresponding user identification data associated with the voice signature data 134, instead of or in addition to using the generated voice signature data. The speaker detector 128 provides the speaker identifier 130 for each detected user to the audio preprocessor 118, and the audio preprocessor 118 retrieves the configuration data 132 based on each speaker identifier 130. The configuration data 132 may include, for example, voice signature data 134 for each person 180 associated with the utterance 108 in which the wake word 110 was detected.

[0056] In some implementations, the configuration data 132 includes other information in addition to the voice signature data 134 of the person 180 associated with the utterance 108 in which the wake word 110 was detected. For example, the configuration data 132 may include voice signature data 134 associated with multiple people 180 (such as a child and the child's parent) who may be allowed to jointly participate in a voice assistant session at the device 102. In such implementations, the configuration data 132 enables one of the voice input filters 120 to generate a voice output signal 152 based on the speech of two or more specific people.

[0057] Therefore, in Figure 1In the illustrated example, after identifying a particular person 180 (or after detecting the wake word 110 in the utterance 108 from the particular person 180), one or more of the voice input filters 120 is configured to operate as a speaker-specific voice input filter associated with the identified particular person 180. The portion of the audio data 116 following the wake word 110 is processed by the speaker-specific voice input filter such that the audio data 150 provided to the second-stage voice processor 154 includes the speech of the particular person 180 and omits or attenuates other portions of the audio data 116. For example, when persons 180A and 180B are speaking simultaneously, a first channel of audio data 150 provided to the second-stage voice processor 154 includes a voice output signal 152A for person 180A, and a second channel of audio data 150 provided to the second-stage voice processor 154 includes a voice output signal 152B for person 180B.

[0058] The second level speech processor 154 includes one or more voice assistant applications 156 that are configured to perform voice assistant operations in response to commands detected within the speech output signal 152. For example, the voice assistant operation may include accessing information from the memory 142 or from another memory, such as the memory of a remote server device. To illustrate, the speech output signal 152 may include an inquiry about local weather conditions, and in response to the inquiry, the voice assistant application 156 may determine the location of the device 102 and transmit a query to a weather database based on the location of the device 102. As another example, the voice assistant operation may include instructions for controlling other devices (e.g., smart home devices), outputting media content, or other similar instructions. When appropriate, the voice assistant application 156 may generate a voice assistant response 170, and the processor 190 may transmit the output audio signal 162 to the audio transducer 164 to output the voice assistant response 170. Although Figure 1 The example illustrates the voice assistant response 170 being provided via the audio transducer 164 , but in other implementations, the voice assistant response 170 can be provided via a display device or another output device coupled to the output interface 160 .

[0059] In some implementations, the audio analyzer 140 is configured to provide the speech output signal 152A as an input to a first voice assistant instance 158A and to provide the speech output signal 152B as an input to a second voice assistant instance 158B that is different from the first voice assistant instance 158A. For example, in some implementations, the second-level speech processor 154 is configured to activate the first voice assistant instance 158A based on detecting the first wake word 110 in the speech output signal 152A and to activate the second voice assistant instance 158B based on detecting the second wake word 110 in the speech output signal 152B. In examples where the device 102 supports multiple voice assistant applications 156, the first-level speech processor 124 provides the second-level speech processor 154 with an indication of the wake word 110A spoken by the person 180A, an indication of which voice assistant application 156 corresponds to the wake word 110A, or both. Similarly, the first stage speech processor 124 provides an indication of the wake word 110B spoken by the person 180B or an indication of which voice assistant application 156 corresponds to the wake word 110B to the second stage speech processor 154 .

[0060] In some examples where wake word 110A is the same as wake word 110B, voice assistant instances 158A and 158B are instances of the same voice assistant application 156 to provide independent voice assistant sessions in parallel to person 180A and person 180B. For illustration, first voice assistant instance 158A corresponds to a first instance of first voice assistant application 156, and second voice assistant instance 158B corresponds to a second instance of first voice assistant application 156. In other examples where wake word 110A is different from wake word 110B, voice assistant instances 158A and 158B are instances of two different voice assistant applications 156 to provide independent voice assistant sessions in parallel to person 180A and person 180B. For illustration, the first voice assistant instance 158A corresponds to a first voice assistant application 156 (e.g., a voice assistant application local to the processor 190), and the second voice assistant instance 158B corresponds to a second voice assistant application 156 that is different from the first voice assistant application 156 (e.g., a third-party voice assistant application installed on the device 102).

[0061] The use of a speaker-specific voice input filter at voice input filter 120A to generate voice output signal 151A substantially prevents the voice of person 180B from interfering with the voice assistant conversation of person 180A with first voice assistant instance 158A. Similarly, the use of a speaker-specific voice input filter at voice input filter 120B to generate voice output signal 152B substantially prevents the voice of person 180A from interfering with the voice assistant conversation of person 180B with second voice assistant instance 158B.

[0062] A technical benefit of filtering audio data 116 to remove or attenuate portions of audio data 116 other than the speech of the particular person 180 who spoke the wake word 110 is that such audio filtering prevents (or reduces the likelihood of) other people from interjecting into the voice assistant session. For example, when person 180A speaks wake word 110A, device 102 launches first voice assistant instance 158A, initiates a voice assistant session associated with person 180A, and configures voice input filter 120A to attenuate portions of audio data 116 other than the speech of person 180A. In this example, another person 180B cannot interject into the voice assistant session because the portion of audio data 116 associated with person 180B's utterance 108B is not provided to second-stage voice processor 154 in the same channel of audio data 150 as the speech output signal 152A for person 180A's session with first voice assistant instance 158A. When the utterance 108B of person 180B is irrelevant to the voice assistant session associated with person 180A, reducing interruptions can improve the user experience associated with the voice assistant application 156 and can conserve resources of the second-stage speech processor 154. In addition, irrelevant speech can cause the first voice assistant instance 158A to misunderstand the speech of person 180A associated with the voice assistant session, thereby causing person 180A to have to repeat the speech and the voice assistant application 156 to have to repeat operations to analyze the speech. In addition, irrelevant speech can reduce the accuracy of speech recognition operations performed by the first voice assistant instance 158A.

[0063] In some cases, the speech of person 180A and the speech of person 180B overlap in time. In such cases, the first speaker-specific speech input filter (speech input filter 120A) suppresses the speech of person 180B during the generation of speech output signal 152A, and the second speaker-specific speech input filter (speech input filter 120A) suppresses the speech of person 180A during the generation of speech output signal 152B. Thus, each person 180A and 180B is prevented from interrupting the voice assistant session of the other person 180A or 180B, thereby enhancing the user experience by enabling concurrent voice assistant sessions to proceed without interfering with each other.

[0064] In some implementations, the interrupting voice may be allowed when it is relevant to the ongoing voice assistant session. Figure 6 As further described, when the audio data 116 includes “interrupted speech” (e.g., speech not associated with the person 180 who spoke the wake word 110 to initiate the voice assistant session), the interrupted speech is processed to determine a relevance score, and only the interrupted speech associated with a relevance score that satisfies the relevance criteria is provided to the voice assistant application 156.

[0065] As an example of the operation of system 100, microphone 104 detects sound 106 including speech 108A of person 180A and provides audio data 116 to processor 190. Before identifying person 180A and detecting wake word 110A, audio preprocessor 118 performs non-speaker-specific audio preprocessing operations, such as echo cancellation, noise reduction, etc. Additionally, in some implementations, second-stage speech processor 154 remains in a low-power state until wake word 110A is detected. In some such implementations, first-stage speech processor 124 operates in an always-on mode, and second-stage speech processor 154 operates in a standby mode or low-power mode until activated by first-stage speech processor 124. Audio preprocessor 118 provides filtered audio data 122 (without speaker-specific speech output signal 152) to first-stage speech processor 124, which executes wake word detector 126 to process filtered audio data 122 to detect wake word 110A and speaker detector 128 to identify person 180A.

[0066] The wake word detector 126 detects the wake word 110A, and the speaker detector 128 determines a speaker identifier 130 associated with the person 180A based on the voice signature data of the filtered audio data 122, biometric or other sensor data, or a combination thereof. In some implementations, the speaker detector 128 provides the speaker identifier 130 to the audio preprocessor 118, and the audio preprocessor 118 obtains voice signature data 134A associated with the person 180A. In other implementations, the speaker detector 128 provides the voice signature data 134A as the speaker identifier 130 to the audio preprocessor 118. The voice signature data 134A and optional other configuration data 132 are provided to the speech input filter 120A to enable the speech input filter 120A to operate as a speaker-specific speech input filter 120A associated with the first person 180A and generate a speaker-specific speech output signal 152A.

[0067] Additionally, based on detecting the wake word 110A, the wake word detector 126 activates the second-stage speech processor 154 and causes a speech output signal 152A to be provided to the second-stage speech processor 154. The speech output signal 152A includes a portion of the audio data 116 after being processed by the speaker-specific speech input filter 120A. For example, the speech output signal 152A may include the entire utterance 108A containing the wake word 110A based on the processing of the audio data 116 by the speaker-specific speech input filter 120A. For example, the audio analyzer 140 may store the audio data 116 in a buffer and, in response to detecting the wake word 110A and identifying the person 180A, cause the audio data 116 stored in the buffer to be processed by the speaker-specific speech input filter 120A. In this illustrative example, the portion of the audio data 116 received before the speech input filter 120A is configured as speaker-specific may still be filtered using the speaker-specific speech input filter 120A before being provided to the second-stage speech processor 154.

[0068] According to some implementations, in addition to detecting the wake word 110A, the second-level speech processor 154 initiates the first voice assistant instance 158A based on an indication of the wake word 110A, the particular voice assistant application 156 associated with the wake word 110A, or both, from the first-level speech processor 124. While the voice assistant session between the person 180A and the first voice assistant instance 158A is ongoing, the second-level speech processor 154 continues to route the channel of audio data 150 corresponding to the speech output signal 152A to the first voice assistant instance 158A.

[0069] In some implementations, after speaker-specific voice input filter 120A is enabled, utterance 108B of person 180B is included in audio data 116 while person 180A continues to speak during the voice assistant session. Audio data 116 is filtered by both speaker-specific voice input filter 120A and voice input filter 120B. The output of speaker-specific voice input filter 120A can be received at first-stage voice processor 124 (e.g., as a first channel of filtered audio data 122) and routed to second-stage voice processor 154 as voice output signal 152A. Additionally, the output of voice input filter 120B can be concurrently provided to first-stage voice processor 124 (e.g., as a second channel of filtered audio data 122) for use in wake-up word detection and speaker detection processing.

[0070] In response to the wake word detector 126 detecting the wake word 110B in the output of the speech input filter 120B and the speaker detector 128 identifying the person 180B as the speaker of the wake word 110B, the audio preprocessor 118 obtains speech signature data 134B associated with the person 180B in a manner similar to that described above. The speech signature data 134B and optional other configuration data 132 are provided to the speech input filter 120B to enable the speech input filter 120B to operate as a speaker-specific speech input filter 120B associated with the person 180B and generate a speech output signal 152B. The speech output signal 152 is transmitted to the first stage speech processor 124 (e.g., as the second channel of the filtered audio data 122) and is routed to the second stage speech processor 154 as the second channel of the audio data 150. Additionally, the audio pre-processor 118 may specify another speech input filter 120 (not shown) to continue performing non-speaker specific filtering (generating a third channel of filtered audio data 122) so that wake word processing and speaker detection processing can continue at the first stage speech processor 124 to detect any wake word 110 that may be spoken by another person 180 (not shown).

[0071] According to some implementations, in addition to detecting the wake word 110B, the second-level speech processor 154 initiates a second voice assistant instance 158B, such as based on an indication of the wake word 110B, a particular voice assistant application 156 associated with the wake word 110B, or both, from the first-level speech processor 124. While the voice assistant session between the person 180B and the second voice assistant instance 158B is ongoing, the second-level speech processor 154 continues to route the channel of audio data 150 corresponding to the speech output signal 152B to the second voice assistant instance 158B.

[0072] In certain implementations, each voice assistant session continues until a termination condition for that session is met. For example, a termination condition with a particular person 180 may be met when a certain duration of the voice assistant session has elapsed, when a voice assistant operation that does not require a response or further interaction with the particular person 180 is performed, or when the particular person 180 commands termination of the voice assistant session.

[0073] In some implementations, the configuration data 132 provided to the audio preprocessor 118 to configure the voice input filter 120 is based on voice signature data 134 associated with multiple people. In such implementations, the configuration data 132 enables the voice input filter 120 to operate as a speaker-specific voice input filter 120 associated with multiple people. For example, when the configuration data 132 provided to a single voice input filter 120 is based on voice signature data 134A associated with person 180A and voice signature data 134B associated with person 180B, the voice input filter 120 can be configured to operate as a speaker-specific voice input filter 120 associated with both person 180A and person 180B. An example of an implementation that can use voice signature data 134 based on the voices of multiple people includes a scenario where person 180A is a child and person 180B is a parent. In this scenario, based on the configuration data 132, the parent can have permissions that allow the parent to interject into any voice assistant conversation initiated by the child.

[0074] In certain implementations, the voice signature data 134 associated with a particular person 180 includes a speaker embedding. For example, during a registration operation, the microphone 104 may capture the speech of the person 180, and the speaker detector 128 (or another component of the device 102) may generate a speaker embedding. The speaker embedding may be stored at the memory 142 along with other data, such as a speaker identifier for the particular person 180, as the registration data 136. Figure 1 In the illustrated example, the enrollment data 136 includes three sets of voice signature data 134, including voice signature data 134A, voice signature data 134B, and voice signature data 134N. However, in other implementations, the enrollment data 136 includes more than three sets of voice signature data 134 or fewer than three sets of voice signature data 134. The enrollment data 136 optionally also includes information specifying the sets of voice signature data 134 to be used together, such as in the above-described example in which a parent's voice signature data 134 is provided to the audio preprocessor 118 along with the child's voice signature data 134.

[0075] In some implementations, once a particular person 180 is identified, the device 102 records data indicating the location of the particular person 180, and the speaker detector 128 can use the location data to identify the particular person 180 as the source of future utterances. For example, the microphone 104 can correspond to a microphone array, and the audio preprocessor 118 can obtain the location data of the particular person 180 via one or more location or source separation techniques (such as time of arrival, angle of arrival, multilateration, etc.). In some implementations, the device 102 assigns each detected person to a particular area of a plurality of logical areas based on the location of the person, and can perform beamforming or other techniques to attenuate speech originating from people in other areas, such as reference areas. Figure 2 Further described.

[0076] Figure 2 is a diagram of an example of a vehicle 250 operable to perform speaker-specific speech filtering for multiple users according to some examples of the present disclosure. Figure 2 In the embodiment, the system 100 or a portion thereof is integrated into a vehicle 250, Figure 2 In the example of FIG. 2 , the vehicle is illustrated as a car including a plurality of seats 252A-252E. Although the vehicle 250 is Figure 2 2 is illustrated as a car, but in other implementations, the vehicle 250 is a bus, train, airplane, boat, or another type of vehicle configured to transport one or more passengers (which may optionally include a vehicle operator).

[0077] The vehicle 250 includes the audio analyzer 140 and one or more audio sources 202. The audio analyzer 140 and the audio sources 202 are coupled via a codec 204 to the microphone 104, the audio transducer 164, or both. Figure 2 The vehicle 250 also includes one or more vehicle systems 270 , some or all of which may be coupled to the audio analyzer 140 to enable the voice assistant application 156 to control various operations of the vehicle systems 270 .

[0078] exist Figure 2 In FIG, the vehicle 250 includes a plurality of microphones 104A-104F. For example, in Figure 2 In FIG, each microphone 104 is positioned adjacent to a corresponding one of seats 252A-252E. Figure 2 In the example of , the positioning of the microphone 104 relative to the seat 252 enables the audio analyzer 140 to distinguish the audio area 254 of the vehicle 250. Figure 2, there is a one-to-one relationship between audio zones 254 and seats 252. In some other implementations, one or more of audio zones 254 include more than one seat 252. For illustration, seats 252C-252E can be associated with a single "back seat" audio zone.

[0079] although Figure 2 The vehicle 250 is illustrated as including multiple microphones 104A-104F arranged to detect sounds within the vehicle 250 and optionally enable the audio analyzer 140 to distinguish which audio region 254 includes the source of the sound, but in other implementations, the vehicle 250 includes only a single microphone 104. In other implementations, the vehicle 250 includes multiple microphones 104, and the audio analyzer 140 does not distinguish between the audio regions 254.

[0080] exist Figure 2 In the embodiment, the audio analyzer 140 includes an audio preprocessor 118, a first-stage speech processor 124 and a second-stage speech processor 154, each of which is as shown in FIG. Figure 1 Follow the instructions in Figure 2 In the particular example illustrated, the audio pre-processor 118 includes a speech input filter 120 that is configurable to operate as a speaker-specific speech input filter to selectively filter audio data for speech processing.

[0081] Figure 2 The audio preprocessor 118 in the embodiment of the present invention further includes an echo cancellation and noise suppression (ECNS) unit 206 and an adaptive interference canceller (AIC) 208. The ECNS unit 206 and the AIC 208 are operable to filter the audio data from the microphone 104 independently of the voice input filter 120. For example, the ECNS unit 206, the AIC 208, or both may perform non-speaker specific audio filtering operations. For purposes of illustration, the ECNS unit 206 is operable to perform echo cancellation operations, noise suppression operations (e.g., adaptive noise filtering), or both. The AIC 208 is configured to distinguish between audio regions 254 and, optionally, to limit the audio data provided to the individual voice input filters 120 to audio from specific corresponding one or more audio regions in the audio regions 254. To illustrate, when a user is detected in a particular zone (e.g., person 180 occupying one of seats 252), the AIC 208 may generate an audio signal for the particular zone (illustrated as zone audio signal 260) that attenuates or removes audio from sources outside of the particular zone.

[0082] The audio analyzer 140 is configured to selectively enable the individual voice input filters 120 to operate as speaker-specific voice input filters 120 based on detecting the locations of the users within the vehicle 250. To illustrate, when a first user and a second user (e.g., person 180A and person 180B, respectively) are in the vehicle 250, the audio analyzer 140 is configured to selectively enable the first speaker-specific voice input filter 120A based on the first user's first seating location within the vehicle 250, and to selectively enable the second speaker-specific voice input filter 120B based on the second user's second seating location within the vehicle 250.

[0083] For illustration, the audio analyzer 140 is configured to detect that a first user is in a first seating position and a second user is in a second seating position based on sensor data from one or more sensors of the vehicle 250. As an example, the sensor data may correspond to the audio data 116 received via the microphone 104 and used to identify the seating position of each speech source (e.g., each user speaking) detected in the vehicle 250 based on the operation of the AIC 208 and the speaker detector 128, and to identify the identity of each detected user via comparison of voice signatures as described above. Alternatively or additionally, the sensor data may correspond to data generated by one or more cameras, seat weight sensors, other sensors that can be used to locate the seating position of occupants in the vehicle 250, or a combination thereof.

[0084] In some implementations, selectively enabling the speaker-specific speech input filter 120 is performed on a per-zone basis and includes generating different per-zone audio signals. For illustration, the audio analyzer 140 (e.g., AIC 208) processes the audio data 116 received from the microphone 104 to generate a first zone audio signal 260A. The first zone audio signal 260A includes sounds originating from a first zone (e.g., zone 254A including the first user's seating position) of the plurality of logical zones 254 of the vehicle 250 and at least partially attenuates sounds originating from outside the first zone. The audio analyzer 140 also generates a second zone audio signal 260B that includes sounds originating from a second zone (e.g., zone 254B including the second user's seating position) and at least partially attenuates sounds originating from outside the second zone.

[0085] The audio analyzer 140 enables the selected speech input filter 120 to be used as a speaker-specific speech input filter for the specific regional audio signal 260 associated with the detected user, thereby generating identity-aware regional speech bubbles for each identified user. For example, audio source separation applied in conjunction with the regions 254 separates the speech by virtue of each user's location, and speaker-specific speech enhancement in each region 254 produces additional isolation of each user's speech. For example, if a first user in a first region 254A leans into a second region 254B occupied by a second user and speaks, regional source separation alone may not be able to filter the first user's speech from the second user's speech in the second region 254B; however, the first user's speech is filtered by the speaker-dependent speech input filtering applied to the audio in the second region 254B.

[0086] In one example, as part of a first filtering operation on the first region audio signal 260A, the first speaker-specific voice input filter 120A is enabled to enhance the voice of the first user, attenuate sounds other than the first user's voice, or both, thereby generating a first voice output signal 152A. Similarly, as part of a second filtering operation on the second region audio signal 260B, the second speaker-specific voice input filter 120B is enabled to enhance the voice of the second user, attenuate sounds other than the second user's voice, or both, thereby generating a second voice output signal.

[0087] In some implementations, when a particular user is detected in a particular area but no voice signature data 134 is available for the user, such as when the particular user is a visitor in a vehicle 250, the audio analyzer 140 processes the regional audio signal 260 for the particular area using the (non-speaker-specific) voice input filter 120. The audio analyzer 140 may also process the particular user's speech to generate the user's voice signature data 134. Although filtering using an initial version of the user's voice signature data 134 may be relatively ineffective at distinguishing the particular user's speech from the speech of others based on the relatively small number of utterances processed by the device 102, as more speech of the particular user becomes available for processing, one or more updated versions of the voice signature data 134 may be generated, thereby increasing the effectiveness of the voice signature data 134 and enabling the use of the voice input filter 120 as a speaker-specific voice input filter 120. Thus, the particular user's voice signature data 134 may be added to the registration data 136 and used to identify the user and enable speaker-specific voice filtering even when the particular user does not participate in the registration operation.

[0088] During operation, one or more of microphones 104 can detect sounds within vehicle 250 and provide audio data representing the sounds to audio analyzer 140. In the example of person 180A seated in zone 254A, when a voice assistant session for zone 254A is not in progress, ECNS unit 206, AIC 208, or both process the audio data to generate filtered audio data (e.g., filtered audio data 122) that attenuates sounds from sources outside of zone 254A and provide the filtered audio data to first-stage speech processor 124 as zone audio signal 260.

[0089] In some implementations, the filtered audio data for region 254A is processed by speaker detector 128 to identify person 180A as the user whose voice is included in the filtered audio data based on voice signature comparison. In other implementations, the user is not detected until wake word detector 126 detects a wake word (e.g., Figure 1 After the wake word 110 is received (e.g., the wake word 110 is received), the speaker detector 128 operates to identify the person 180A. In response to identifying the person 180A in the area, the speech input filter 120A is activated as the speaker-specific speech input filter for the area 254A. The speaker-specific speech input filter 120A processes the area audio signal 260A from the AIC 208 and generates a speech output signal 152A for the speech of the person 180A in the area 254A.

[0090] In addition, the wake-up word detector 126 processes the filtered audio data for region 254A (if person 180A has not been identified) or the voice output signal 152A for region 254A (if person 180A has been identified). In response to detecting the wake-up word, if the second-level voice processor 154 is not already active, the wake-up word detector 126 activates the second-level voice processor 154 to initiate a voice assistant session associated with region 254A. The first-level voice processor 124 provides the voice output signal 152A to the second-level voice processor 154 and may also provide an indication of the wake-up word spoken by person 180A or an indication of which voice assistant application 156 is associated with the wake-up word. The second-level voice processor 154 initiates a first voice assistant instance 158A of the voice assistant application 156 associated with the wake-up word and routes the voice output signal 152A associated with region 254A to the first voice assistant instance 158A while the voice assistant session between person 180A and the first voice assistant instance 158A is ongoing.

[0091] Based on the speech content represented in the audio data from person 180A in area 254A, first voice assistant instance 158A can control the operation of audio source 202, control the operation of vehicle systems 270, or perform other operations, such as retrieving information from a remote data source.

[0092] The response from the first voice assistant instance 158A (e.g., voice assistant response 170) can be played to the occupants of the vehicle 250 via the audio transducer 164. Figure 2 In the illustrated example, the audio transducer 164 is positioned near or within a particular audio area in the audio region 254 , which enables a separate instance of the voice assistant application 156 to provide a response to a particular occupant (e.g., the occupant who initiated the voice assistant session) or to multiple occupants of the vehicle 250 .

[0093] The example described above regarding the operation of detecting the speech of the occupant in zone 254A may also be repeated for each zone 254 in which an audio source (e.g., an occupant) is detected. Thus, the system 100 enables multiple occupants of the vehicle 250 to simultaneously participate in a voice assistant conversation using a dedicated speaker-specific voice input filter 120 and a corresponding dedicated voice assistant instance 158 for each occupied zone 254.

[0094] The selective operation of the voice input filter 120 as a speaker-specific voice input filter enables the voice assistant application 156 to perform more accurate speech recognition because noise and irrelevant speech are removed from the audio data provided to the voice assistant application 156. Additionally, the selective operation of the voice input filter 120 as a speaker-specific voice input filter and the interference cancellation performed by the AIC 208 limit the ability of other occupants of the vehicle 250 to interrupt the voice assistant session. For example, if the driver of the vehicle 250 initiates a voice assistant session to request driving directions, the voice assistant session may be associated only with the driver (or with one or more other persons, as described above), such that other occupants of the vehicle 250 cannot interrupt the voice assistant session.

[0095] Figures 3A to 3C Illustrate aspects of operations associated with speaker-specific speech filtering for multiple users according to some examples of the present disclosure. Figure 3A , illustrating a first example 300. In the first example 300, the configuration data 132 for configuring the speech input filter 120 to operate as a speaker-specific speech input filter 310 includes first speech signature data 306. The first speech signature data 306 includes, for example, a first person (such as Figure 1 Speaker embeddings associated with person 180A).

[0096] In a first example 300, audio data 116 provided as input to a speaker-specific voice input filter 310 includes ambient sounds 112 and speech 304. The speaker-specific voice input filter 310 is operable to generate audio data 150 (e.g., a speech output signal 152) as output based on the audio data 116. In the first example 300, the audio data 150 includes speech 304 and excludes or attenuates the ambient sounds 112. For example, the speaker-specific voice input filter 310 is configured to compare the audio data 116 with first voice signature data 306 to generate the audio data 150. The audio data 150 attenuates portions of the audio data 116 that do not correspond to speech 304 from the person associated with the first voice signature data 306.

[0097] exist Figure 3A In the illustrated first example 300, audio data 150 representing speech 304 is provided to a voice assistant application 156 as part of a voice assistant session. Additionally, a portion of the audio data 116 representing ambient sounds 112 is attenuated or omitted from the audio data 150 provided to the voice assistant application 156. A technical benefit of filtering the audio data 116 to attenuate or omit the ambient sounds 112 from the audio data 150 is that such filtering enables the voice assistant application 156 to more accurately identify speech in the audio data 150, which reduces the error rate of the voice assistant application 156 and improves the user experience.

[0098] refer to Figure 3B , illustrating a second example 320. In the second example 320, the configuration data 132 for configuring the speech input filter 120 includes Figure 3A For example, the first voice signature data 306 includes the first person (such as Figure 1 Speaker embeddings associated with person 180A).

[0099] In a second example 320, the audio data 116 provided as input to the speaker-specific speech input filter 310 includes multiple voices 322, such as Figure 1 The speaker-specific speech input filter 310 is operable to generate audio data 150 as output based on the audio data 116. In the second example 320, the audio data 150 includes the speech 324 of a single person, such as the speech of person 180A. In this example, the speech of one or more other persons, such as the speech of person 180B, is omitted or attenuated in the audio data 150. For example, the audio data 150 attenuates portions of the audio data 116 that do not correspond to speech from the person associated with the first voice signature data 306.

[0100] exist Figure 3B In the illustrated second example 320, audio data 150 representing a single person's voice 324 (e.g., the voice of the person who initiated the voice assistant session) is provided to the voice assistant application 156 as part of the voice assistant session. In addition, a portion of the audio data 116 representing the voices of other persons (e.g., the voices of the person who did not initiate the voice assistant session) is attenuated or omitted from the audio data 150 provided to the voice assistant application 156. A technical benefit of filtering the audio data 116 to attenuate or omit the voices of the person who did not initiate the particular voice assistant session is that such filtering limits the ability of such other persons to interrupt the voice assistant session.

[0101] although Figure 3B The ambient sounds 112 are not specifically illustrated in the audio data 116 provided to the speaker-specific speech input filter 310, but in some implementations, the audio data 116 in the second example 320 also includes the ambient sounds 112. In such implementations, the speaker-specific speech input filter 310 performs both speaker separation (e.g., to distinguish between single-speaker speech 324 and multiple-speaker speech 322) and noise reduction (e.g., to remove or attenuate the ambient sounds 112).

[0102] refer to Figure 3C , illustrating a third example 340. In the third example 340, the configuration data 132 for configuring the voice input filter 120 includes first voice signature data 306 and second voice signature data 342. For example, the first voice signature data 306 includes a first person (such as Figure 1 The second speech signature data 342 includes a speaker embedding associated with a second person such as Figure 1 180B) associated speaker embeddings.

[0103] In the third example 340, the audio data 116 provided as input to the speaker-specific voice input filter 310 includes ambient sounds 112 and speech 344. Speech 344 may include the speech of a first person, the speech of a second person, the speech of one or more other persons, or any combination thereof. The speaker-specific voice input filter 310 is operable to generate audio data 150 as output based on the audio data 116. In the third example 340, the audio data 150 includes speech 346. Speech 346 includes the speech of the first person (if present in the audio data 116), the speech of the second person (if present in the audio data 116), or both. Furthermore, in the audio data 150, the ambient sounds 112 and the speech of the other persons are attenuated (e.g., attenuated or removed). That is, portions of the audio data 116 that do not correspond to the speech of the first person associated with the first voice signature data 306 or the speech of the second person associated with the second voice signature data 342 are attenuated in the audio data 150.

[0104] exist Figure 3C In the illustrated third example 340, audio data 150 representing speech 346 is provided to the voice assistant application 156 as part of a voice assistant session. Furthermore, a portion of the audio data 116 representing ambient sound 112 or other people's speech is attenuated or omitted from the audio data 150 provided to the voice assistant application 156. A technical benefit of filtering the audio data 116 to attenuate or omit the speech of some individuals (e.g., individuals not associated with the first voice signature data 306 or the second voice signature data 342) while still allowing multiple people's speech (e.g., speech from individuals associated with the first voice signature data 306 or the second voice signature data 342) to pass to the voice assistant application 156 is that such filtering limits the ability for a particular user to interject. For example, multiple members of a family may be allowed to interject on each other's voice assistant sessions while preventing others from interjecting on a voice assistant session initiated by a member of the family.

[0105] Figure 4 A specific example of the voice input filter 120 is illustrated. Figure 4 In the illustrated example, the speech input filter 120 includes or corresponds to one or more speech enhancement models 440. The speech enhancement models 440 include one or more machine learning models that are configured and trained to perform speech enhancement operations such as denoising, speaker separation, etc. Figure 4In the illustrated example, the speech enhancement model 440 includes a dimensionality reduction network 410, a combiner 416, and a dimensionality expansion network 418. The dimensionality reduction network 410 includes a plurality of layers (e.g., neural network layers) arranged to perform convolution, pooling, concatenation, etc. to generate a latent space representation 412 based on the audio data 116. In one example, the audio data 116 is input to the dimensionality reduction network 410 as a series of input feature vectors, where each input feature vector in the series represents one or more audio data samples (e.g., a frame or another portion) of the audio data 116, and the dimensionality reduction network 410 generates a latent space representation 412 associated with each input feature vector. The input feature vectors may include, for example, values representing spectral features (e.g., a complex spectrum, an amplitude spectrum, a mel spectrum, a bark spectrum, etc.) of a time-windowed portion of the audio data 116, cepstral features (e.g., mel-frequency cepstral coefficients, bark-frequency cepstral coefficients, etc.) of the time-windowed portion of the audio data 116, or other data representing the time-windowed portion of the audio data 116.

[0106] The combiner 416 is configured to combine the speaker embedding 414 and the latent space representation 412 to generate a combined vector 417 as an input to the dimension expansion network 418. In one example, the combiner 416 includes a concatenator configured to concatenate the speaker embedding 414 to the latent space representation 412 of each input feature vector to generate the combined vector 417.

[0107] The dimension-expanding network 418 includes one or more recurrent layers (e.g., one or more gated recurrent unit (GRU) layers) and multiple additional layers (e.g., neural network layers) arranged to perform convolution, pooling, concatenation, etc. to generate audio data 150 based on the combined vector 417.

[0108] Optionally, the speech enhancement model 440 may further include one or more skip connections 419. Each skip connection 419 connects the output of one of the layers of the dimensionality reduction network 410 to the input of a corresponding one of the layers of the dimensionality expansion network 418.

[0109] During operation, audio data 116 (or a feature vector representing audio data 116) is provided as input to speech enhancement model 440. Audio data 116 may include speech 402, ambient sounds 112, or both. Speech 402 may include the speech of a single person or the speech of multiple people.

[0110] Based on the architecture and training of the dimensionality reduction network 410, the dimensionality reduction network 410 processes each feature vector of the audio data 116 through a series of convolution operations, pooling operations, activation layers, recursive layers, other data manipulation operations, or any combination thereof to generate a latent space representation 412 of the feature vector of the audio data 116. Figure 4In the illustrated example, the generation of the latent space representation 412 of the feature vector is performed independently of the speech signature data 134. Thus, the same operations are performed regardless of who initiates the voice assistant session.

[0111] Speaker embeddings 414 are speaker-specific and are selected based on the specific person (or persons) whose speech is to be enhanced. Each latent space representation 412 is combined with the speaker embedding 414 to generate a corresponding combined vector 417, and the combined vector 417 is provided as input to the extended network 418. As described above, the extended network 418 includes at least one recursive layer, such as a GRU layer, so that each output vector of the audio data 150 depends on a series (e.g., more than one) of combined vectors 417. In some implementations, the extended network 418 is configured (and trained) to generate enhanced speech 420 for a specific person as the audio data 150. In such implementations, the specific person whose speech is enhanced is the person whose speech is represented by the speaker embedding 414. In some implementations, the extended network 418 is configured (and trained) to generate enhanced speech 420 for more than one specific person as the audio data 150. In such implementations, the specific person whose speech is enhanced is the person associated with the speaker embedding 414.

[0112] The dimensionality expansion network 418 can be considered a generative network that is configured and trained to recreate that portion of the input audio data stream (e.g., audio data 116) that resembles the speech of a particular person (e.g., the person associated with the speaker embedding 414). Thus, the speech enhancement model 440 can use a set of machine learning operations to perform both noise reduction and speaker separation to generate enhanced speech 420.

[0113] Figure 5 Another specific example of the voice input filter 120 is illustrated. Figure 5 In the illustrated example, the speech input filter 120 includes or corresponds to one or more speech enhancement models 440. Figure 4 As shown, the speech enhancement model 440 includes one or more machine learning models that are configured (and trained) to perform speech enhancement operations such as denoising, speaker separation, etc. Figure 5 In the illustrated example, the speech enhancement model 440 includes a dimensionality reduction network 410 coupled to a switch 502. The switch 502 may include, for example, a logic switch configured to select which of a plurality of subsequent processing paths to execute. Figure 4 The method operates as described to generate a latent space representation 412 associated with each input feature vector of the audio data 116 .

[0114] exist Figure 5In the illustrated example, switch 502 is coupled to a first processing path including combiner 504 and multi-dimensional network 508, and switch 502 is also coupled to a second processing path including combiner 512 and multi-person multi-dimensional network 518. In this example, the first processing path is configured (and trained) to perform operations associated with enhancing the speech of a single person, and the second processing path is configured (and trained) to perform operations associated with enhancing the speech of multiple people. Thus, switch 502 is configured to Figure 1 The first processing path is selected when the configuration data 132 for the speaker includes a single speaker embedding 506 or otherwise indicates that the speech of a single identified speaker is to be enhanced to generate a single person's enhanced speech 510. In contrast, the switch 502 is configured to Figure 1 The second processing path is selected when the configuration data 132 includes multiple speaker embeddings (such as the first speaker embedding 514 and the second speaker embedding 516) or otherwise indicates that the speech of multiple identified speakers is to be enhanced to generate the enhanced speech 520 for multiple persons.

[0115] The combiner 504 is configured to combine the speaker embedding 506 and the latent space representation 412 to generate a combined vector as an input to the dimension expansion network 508. The dimension expansion network 508 is configured to process the combined vector as shown in FIG. Figure 4 As described, to generate enhanced speech 510 for a single person.

[0116] The combiner 512 is configured to combine two or more speaker embeddings (e.g., the first speaker embedding 514 and the second speaker embedding 516) and the latent space representation 412 to generate a combined vector as an input to the multi-person extended-dimensional network 518. The multi-person extended-dimensional network 518 is configured to process the combined vector as shown in FIG. Figure 4 As described above, the enhanced speech of multiple persons is generated 520. Although the first processing path and the second processing path perform similar operations, Figure 5 Different processing paths are used in the illustrated example because the combined vectors generated by combiners 504 and 512 have different dimensions. Therefore, the extended-dimensional network 508 and the multi-person extended-dimensional network 518 have different architectures to accommodate the combined vectors of different dimensions.

[0117] Alternatively, in some implementations, Figure 5 Different processing paths are used in the example to account for the different operations performed by combiners 504, 512. For example, combiner 512 can be configured to combine speaker embeddings 514, 516 in an element-by-element manner such that the combined vectors generated by combiners 504, 512 have the same dimensions. To illustrate, combiner 512 can add or average the value of each element of first speaker embedding 514 with the value of the corresponding element of second speaker embedding 516.

[0118] Figure 6 Another specific example of the voice input filter 120 is illustrated. Figure 6 In the illustrated example, the speech input filter 120 includes or corresponds to one or more speech enhancement models 440. Figure 4 and Figure 5 As shown, the speech enhancement model 440 includes one or more machine learning models that are configured (and trained) to perform speech enhancement operations such as denoising, speaker separation, etc. Figure 6 In the illustrated example, the speech enhancement model 440 includes a dimensionality reduction network 410, which is as shown in FIG. Figure 4 The method operates as described to generate a latent space representation 412 associated with each input feature vector of the audio data 116 .

[0119] exist Figure 6 In the illustrated example, the dimensionality reduction network 410 is coupled to a first processing path including the combiner 602 and the dimensionality expansion network 606, and to a second processing path including the combiner 610 and the dimensionality expansion network 614. In this example, the first processing path is configured (and trained) to perform operations associated with enhancing the speech of a first person (e.g., the person initiating a particular voice assistant session), and the second processing path is configured (and trained) to perform operations associated with enhancing the speech of one or more second persons (e.g., based on the speech of the first person). Figure 1 The configuration data 132 is authorized to operate in association with the voice of a person (a person interjecting into a voice assistant session) in certain circumstances.

[0120] The combiner 602 is configured to combine the speaker embedding 604 (e.g., the speaker embedding associated with the person who spoke the wake word 110 to initiate the voice assistant session) and the latent space representation 412 to generate a combined vector as an input to the dimension expansion network 606. The dimension expansion network 606 is configured to process the combined vector as shown in FIG. Figure 4 As described, the first person's enhanced voice 608 is generated. Since the first person is the one who initiated the voice assistant session, the first person's enhanced voice 608 is provided to the voice assistant application 156 for processing.

[0121] The combiner 610 is configured to combine the speaker embedding 612 (e.g., the speaker embedding associated with the second person who did not speak the wake word 110 to initiate the voice assistant session) and the latent space representation 412 to generate a combined vector as an input to the dimension expansion network 614. The dimension expansion network 614 is configured to process the combined vector, as shown in FIG. Figure 4 (or in the case where the speaker embedding 612 corresponds to multiple persons (collectively referred to as “second persons”) Figure 5) to generate the enhanced speech of the second person 616. Note that at any given time, the latent space representation 412 may include the speech of the first person, the speech of the second person, neither, or both. Thus, in some implementations, each latent space representation 412 may be processed via both the first processing path and the second processing path.

[0122] The second person has conditional access to the voice assistant session. Therefore, the second person's enhanced voice 616 is further analyzed to determine whether the conditions for providing the second person's voice 616 to the voice assistant application 156 are met. Figure 6 In the illustrated example, the enhanced speech 616 of the second person is provided to a natural language processing (NLP) engine 620. Additionally, contextual data 622 associated with the enhanced speech 608 of the first person is provided to the NLP engine 620. The contextual data 622 may include, for example, the enhanced speech 608 of the first person, data summarizing the enhanced speech 608 of the first person (e.g., keywords from the enhanced speech 608 of the first person), results generated by the voice assistant application 156 in response to the enhanced speech 608 of the first person, other data indicating the content of the enhanced speech 608 of the first person, or any combination thereof.

[0123] The NLP engine 620 is configured to determine whether the second person's speech (as represented in the second person's enhanced speech 616) is contextually relevant to the voice assistant request, command, query, or other content of the first person's speech as indicated by the context data 622. As an example, the NLP engine 620 may perform context-aware semantic embedding of the context data 622, the second person's enhanced speech 616, or both to determine a value of a relevance metric associated with the second person's enhanced speech 616. In this example, the context-aware semantic embedding may be used to map the second person's enhanced speech 616 to a feature space in which semantic similarity may be estimated based on the distance between two points (e.g., cosine distance, Euclidean distance, etc.), and the relevance metric may correspond to the value of the distance metric. If the relevance metric meets a threshold, the content of the second person's enhanced speech 616 may be considered relevant to the voice assistant session.

[0124] If the content of the second person's enhanced speech 616 is deemed relevant to the voice assistant session, the NLP engine 620 provides the second person's relevant speech 624 to the voice assistant application 156. Otherwise, if the content of the second person's enhanced speech 616 is deemed not relevant to the voice assistant session, the second person's enhanced speech 616 is discarded or ignored.

[0125] Figure 7is an implementation in which the system 100 is integrated into a wireless speaker and voice activated device 700. The wireless speaker and voice activated device 700 may have wireless network connectivity and be configured to perform voice assistant operations. Figure 7 In FIG. 7 , the audio analyzer 140 , the audio source 202 , and the codec 204 are included in the wireless speaker and voice activated device 700 . The wireless speaker and voice activated device 700 also includes the audio transducer 164 and the microphone 104 .

[0126] During operation, one or more of the microphones 104 can detect sounds near the wireless speaker and voice activation device 700, such as in the room in which the wireless speaker and voice activation device 700 are located. The microphone 104 provides audio data representing the sounds to the audio analyzer 140. When no voice assistant session is in progress, the ECNS unit 206, the AIC 208, or both processes the audio data to generate filtered audio data (e.g., filtered audio data 122) and provides the filtered audio data to the wake word detector 126. If the wake word detector 126 detects a wake word (e.g., Figure 1 10), the wake word detector 126 signals the speaker detector 128 to identify the person who said the wake word. In addition, the wake word detector 126 activates the second-stage speech processor 154 to initiate the voice assistant session. The speaker detector 128 provides an identifier of the person who said the wake word (e.g., speaker identifier 130) to the audio preprocessor 118, and the audio preprocessor 118 obtains configuration data (e.g., configuration data 132) to activate the voice input filter 120 as a speaker-specific voice input filter. The above process can be repeated for each different person who said the wake word, thereby enabling multiple concurrent voice assistant sessions to be performed for multiple users. In some specific implementations, the wake word detector 126 can also provide information to the AIC 208 to indicate the direction, location, or audio area from which each detected wake word originated, and the AIC 208 can perform beamforming or other directional audio processing to filter the audio data provided to the voice input filter 120 based on the direction, location, or audio area from which each person's speech originated.

[0127] The speaker-specific voice input filter is used to filter the audio data and provide the filtered audio data to the corresponding instance of the voice assistant application 156, as shown in FIG. Figures 1 to 6Based on the speech content represented in the filtered audio data, the voice assistant application 156 performs one or more voice assistant operations, such as transmitting a command to a smart home device, playing media, or performing other operations, such as retrieving information from a remote data source. A response from the voice assistant application 156 (e.g., voice assistant response 170) may be played via the audio transducer 164.

[0128] The selective operation of the voice input filter 120 as a speaker-specific voice input filter enables the voice assistant application 156 to perform more accurate speech recognition because noise and irrelevant speech are removed from the audio data provided to each instance of the voice assistant application 156. Additionally, the selective operation of the voice input filter 120 as a speaker-specific voice input filter limits the ability of multiple people in a room with respective voice assistant sessions with the wireless speaker and voice-activated device 700 to interject into each other's voice assistant sessions.

[0129] Figure 8 The device 102 is depicted as an implementation 800 of an integrated circuit 802 that includes one or more processors 190, including one or more components of the audio analyzer 140. The integrated circuit 802 also includes input circuitry 804 (such as one or more bus interfaces) to enable the receipt of audio data 116 for processing. The integrated circuit 802 also includes output circuitry 806 (such as a bus interface) to enable the transmission of output data 808 from the integrated circuit 802. For example, the output data 808 may include Figure 1 As another example, the output data 808 may include commands or queries (such as information retrieval queries transmitted to remote devices) to other devices (such as media players, transportation systems, smart home devices, etc.). In some implementations, Figure 1 The voice assistant application 156 is located away from Figure 8 The position of the audio analyzer 140, in which case the output data 808 may include Figure 1 150 of audio data.

[0130] Integrated circuit 802 can be used as a component in a system including a microphone to implement speaker-specific speech filtering for multiple users, such as a system for Figure 9 The depicted mobile phone or tablet, such as Figure 10 The wearable electronic devices depicted, such as Figure 11 The camera depicted, such as Figure 12 The depicted extended reality (e.g., virtual reality, mixed reality, or augmented reality) head-mounted device or Figure 2 or Figure 13 The means of transport depicted.

[0131] As an illustrative, non-limiting example, Figure 9 An implementation 900 is depicted in which the device 102 includes a mobile device 902, such as a phone or tablet. In a particular implementation, the integrated circuit 802 is integrated within the mobile device 902. Figure 9 , mobile device 902 includes microphone 104, audio transducer 164, and display screen 904. Components of processor 190, including audio analyzer 140, are integrated into mobile device 902 and are illustrated using dashed lines to indicate internal components that are generally not visible to a user of mobile device 902.

[0132] In a specific example, Figure 9 The audio analyzer 140 is as shown in the reference Figures 1 to 8 904 to selectively enable speaker-specific voice filtering for multiple users in a manner that improves the accuracy of speech recognition by the voice assistant application 156 and limits the ability of others to interrupt the voice assistant session. During the voice assistant session, responses from the voice assistant application may be provided as output to the user via the audio transducer 164, via the display screen 904, or both.

[0133] Figure 10 An implementation 1000 is depicted in which the device 102 includes a wearable electronic device 1002 (illustrated as a "smart watch"). In a particular implementation, the integrated circuit 802 is integrated within the wearable electronic device 1002. Figure 10 In FIG, the wearable electronic device 1002 includes a microphone 104 , an audio transducer 164 , and a display screen 1004 .

[0134] Components of the processor 190, including the audio analyzer 140, are integrated into the wearable electronic device 1002. In certain examples, Figure 10 The audio analyzer 140 is as shown in the reference Figures 1 to 8 1004 , to selectively enable speaker-specific voice filtering for multiple users in a manner that improves the accuracy of speech recognition by the voice assistant application 156 and limits the ability of others to interrupt the voice assistant session. During the voice assistant session, responses from the voice assistant application may be provided as output to the user via the audio transducer 164, via tactile feedback to the user, via the display screen 1004, or any combination thereof.

[0135] As one example of the operation of the wearable electronic device 1002, during a voice assistant session, the person initiating the voice assistant session may provide a voice requesting that a message (e.g., a text message, an email, etc.) transmitted to the person be displayed via the display screen 1004 of the wearable electronic device 1002. In this example, another person near the wearable electronic device 1002 may speak a wake word associated with the audio analyzer 140, and a new voice assistant session may be initiated (if permitted) without interrupting the voice assistant session because the audio data is filtered during the voice assistant session to attenuate a portion of the audio data that does not correspond to the voice of the person initiating the voice assistant session.

[0136] Figure 11 An implementation 1100 is depicted in which the device 102 includes a portable electronic device corresponding to a camera device 1102. In a particular implementation, the integrated circuit 802 is integrated within the camera device 1102. Figure 11 , the camera device 1102 includes a microphone 104 and an audio transducer 164. The camera device 1102 may also include Figure 11 The display screen on the side not shown in the figure.

[0137] Components of the processor 190, including the audio analyzer 140, are integrated into the camera device 1102. In certain examples, Figure 11 The audio analyzer 140 is as shown in the reference Figures 1 to 8 16. The system may operate as described in any of the figures in order to selectively enable speaker-specific voice filtering for multiple users in a manner that improves the accuracy of speech recognition by the voice assistant application 156 and limits the ability of others to interrupt the voice assistant session. During the voice assistant session, responses from the voice assistant application may be provided as output to the user via the audio transducer 164, via the display screen, or both.

[0138] As one example of the operation of the camera device 1102, during a voice assistant session, the person initiating the voice assistant session may provide a voice requesting that the camera device 1102 capture an image. In this example, another person near the camera device 1102 may speak a wake word associated with the audio analyzer 140, and a new voice assistant session may be initiated (if permitted) without interrupting the voice assistant session because the audio data is filtered during the voice assistant session to attenuate a portion of the audio data that does not correspond to the voice of the person initiating the voice assistant session.

[0139] Figure 12 An implementation 1200 is depicted in which the device 102 includes a portable electronic device corresponding to an extended reality (e.g., virtual reality, mixed reality, or augmented reality) head-mounted device 1202. In a particular implementation, the integrated circuit 802 is integrated within the head-mounted device 1202. Figure 12 , the head mounted device 1202 includes a microphone 104 and an audio transducer 164. Additionally, a visual interface device is positioned in front of the user's eyes to enable display of augmented reality, mixed reality, or virtual reality images or scenes to the user while the head mounted device 1202 is worn.

[0140] Components of the processor 190, including the audio analyzer 140, are integrated into the head mounted device 1202. In certain examples, Figure 12 The audio analyzer 140 is as shown in the reference Figures 1 to 8 16. The system may operate as described in any of the figures in order to selectively enable speaker-specific voice filtering for multiple users in a manner that improves the accuracy of speech recognition by the voice assistant application 156 and limits the ability of others to interrupt the voice assistant session. During the voice assistant session, responses from the voice assistant application may be provided as output to the user via the audio transducer 164, via a visual interface device, or both.

[0141] As one example of the operation of the head mounted device 1202, during a voice assistant session, the person initiating the voice assistant session may provide a voice requesting that specific media be displayed on the visual interface device of the head mounted device 1202. In this example, another person near the head mounted device 1202 may speak a wake word associated with the audio analyzer 140, and a new voice assistant session may be initiated (if permitted) without interrupting the voice assistant session because the audio data is filtered during the voice assistant session to attenuate a portion of the audio data that does not correspond to the voice of the person initiating the voice assistant session.

[0142] Figure 13 An implementation 1300 is depicted in which the device 102 corresponds to or is integrated within a vehicle 1302, illustrated as a manned or unmanned aerial vehicle (e.g., a package delivery drone). In a particular implementation, the integrated circuit 802 is integrated within the vehicle 1302. Figure 13 , vehicle 1302 also includes microphone 104 and audio transducer 164 .

[0143] Components of the processor 190, including the audio analyzer 140, are integrated into the vehicle 1302. In certain examples, Figure 13 The audio analyzer 140 is as shown in the reference Figures 1 to 8 164 to selectively enable speaker-specific voice filtering for multiple users by improving the accuracy of speech recognition by the voice assistant application 156 and limiting the ability of others to interrupt the voice assistant session. During the voice assistant session, responses from the voice assistant application may be provided as output to the user via the audio transducer 164.

[0144] As an example of the operation of vehicle 1302, during a voice assistant session, the person initiating the voice assistant session can provide a voice requesting vehicle 1302 to deliver a package to a specified location. In this example, other people near vehicle 1302 can speak a wake-up word associated with audio analyzer 140 and initiate a new voice assistant session (if allowed) without interrupting the voice assistant session because the audio data is filtered during the voice assistant session to attenuate the portion of the audio data that does not correspond to the voice of the person initiating the voice assistant session. Therefore, other people cannot redirect vehicle 1302 to a different delivery location.

[0145] Figure 14 is a block diagram of illustrative aspects of a system 1400 operable to perform speaker-specific speech filtering for multiple users according to some examples of the present disclosure. Figure 14 In FIG, the processor 190 includes an always-on power domain 1403 and a second power domain 1405, such as an on-demand power domain. The operation of the system 1400 is divided such that some operations are performed in the always-on power domain 1403 and other operations are performed in the second power domain 1405. For example, in Figure 14 In FIG, the audio preprocessor 118, the first stage speech processor 124 and the buffer 1460 are included in the normally-on power domain 1403 and are configured to operate in a normally-on mode. Figure 14 In the embodiment of the present invention, the second stage speech processor 154 is included in the second power domain 1405 and is configured to operate in the on-demand mode. The second power domain 1405 also includes an activation circuit 1430.

[0146] The audio data 116 received from the microphone 104 is stored in a buffer 1460. In certain implementations, the buffer 1460 is a circular buffer that stores the audio data 116 so that the most recent audio data 116 can be accessed for processing by other components, such as the audio pre-processor 118, the first stage speech processor 124, the second stage speech processor 154, or a combination thereof.

[0147] One or more components of the always-on power domain 1403 are configured to generate at least one of a wake-up signal 1422 or an interrupt 1424 to initiate one or more operations at the second power domain 1405. In one example, the wake-up signal 1422 is configured to transition the second power domain 1405 from a low-power mode 1432 to an active mode 1434 to activate one or more components of the second power domain 1405. As an example, the wake-up signal 1422 or the interrupt 1424 may be generated by the wake-up word detector 126 when a wake-up word is detected in the audio data 116.

[0148] In various implementations, the activation circuit 1430 includes or is coupled to a power management circuit, a clock circuit, a head switch or foot switch circuit, a buffer control circuit, or any combination thereof. The activation circuit 1430 can be configured to initiate powering up the second power domain 1405, such as by selectively applying or raising the voltage of a power source of the second power domain 1405. As another example, the activation circuit 1430 can be configured to selectively enable or disable a clock signal to the second power domain 1405, such as to prevent or enable circuit operation without removing the power source.

[0149] The output 1452 generated by the second stage speech processor 154 may be provided to an application 1454. The application 1454 may be configured to perform operations as directed by one or more instances of the voice assistant application 156. For purposes of illustration, the application 1454 may correspond to a vehicle navigation and entertainment application or a home automation system as illustrative, non-limiting examples.

[0150] In a particular implementation, when a voice assistant session is active, the second power domain 1405 can be activated. As an example of the operation of the system 1400, the audio pre-processor 118 operates in the always-on power domain 1403 to filter the audio data 116 accessed from the buffer 1460 and provide the filtered audio data to the first stage speech processor 124. In this example, when no voice assistant session is active, the audio pre-processor 118 operates in a non-speaker specific manner, such as by performing echo cancellation, noise suppression, etc.

[0151] When the wake-up word detector 126 detects a wake-up word in the filtered audio data from the audio preprocessor 118, the first stage speech processor 124 causes the speaker detector 128 to identify the person who said the wake-up word, transmits a wake-up signal 1422 or interrupt 1424 to the second power domain 1405, and causes the audio preprocessor 118 to obtain configuration data associated with the person who said the wake-up word.

[0152] Based on the configuration data, the audio preprocessor 118 begins operating in a speaker-specific mode to process the speech of the person who spoke the wake word, as shown in FIG. Figures 1 to 6In speaker-specific mode, the audio preprocessor 118 provides a speech output signal 152 corresponding to the voice of the person who spoke the wake word to the second-stage speech processor 154. The speech output signal 152 is filtered by a speaker-specific speech input filter to attenuate, attenuate, or remove portions of the audio data 116 that do not correspond to the voice of the specific person whose speech signature data is provided to the audio preprocessor 118 along with the configuration data. In some implementations, the audio preprocessor 118 also provides the speech output signal 152 to the first-stage speech processor 124 until the voice assistant session is terminated.

[0153] By selectively activating the second stage speech processor 154 based on the results of processing audio data at the first stage speech processor 124, the overall power consumption associated with speech processing may be reduced.

[0154] refer to Figure 15 , shows a specific implementation of the method 1500 for speaker-specific speech filtering for multiple users. In certain aspects, by Figure 1 At least one of the audio analyzer 140 , processor 190 , device 102 , system 100 , or a combination thereof, performs one or more operations of method 1500 .

[0155] Method 1500 includes detecting speech of a first user and a second user at one or more processors at block 1502. For example, audio analyzer 140 may detect speech of person 180A at speaker detector 128 by processing a portion of audio data 116 corresponding to utterance 108A from person 180A to determine a speech signature and comparing the speech signature to speech signature data 134. Audio analyzer 140 may also detect speech of person 180B at speaker detector 128 by processing a portion of audio data 116 corresponding to utterance 108B from person 180B to determine a speech signature and comparing the speech signature to speech signature data 134.

[0156] The method 1500 includes obtaining, at block 1504, at one or more processors, first voice signature data associated with a first user and second voice signature data associated with a second user. For example, the audio preprocessor 118 may obtain Figure 1 The audio preprocessor 118 may also obtain configuration data 132 that includes at least voice signature data 134A associated with person 180A. Figure 1 Configuration data 132 includes at least voice signature data 134B associated with person 180B.

[0157] Method 1500 includes, at block 1506, selectively enabling, at one or more processors, a first speaker-specific voice input filter based on the first voice signature data to generate a first voice output signal corresponding to the voice of the first user. For example, the voice signature data 134A may include: Figure 1 The configuration data 132 enables the speech input filter 120A of the audio preprocessor 118 to operate in a speaker-specific mode to enhance the speech of the first person 180A, to attenuate sounds other than the speech of the first person 180A (such as attenuating the speech of the second person 180B), or both. The first speech signature data may correspond to a first speaker embedding, such as speaker embedding 414, and enabling the first speaker-specific speech input filter may include providing the first speaker embedding as input to a speech enhancement model, such as Figures 4 to 6 Speech enhancement model 440.

[0158] Method 1500 includes, at block 1508, selectively enabling, at one or more processors, a second speaker-specific voice input filter based on the second voice signature data to generate a second voice output signal corresponding to the voice of the second user. For example, the voice signature data 134B may include: Figure 1 The configuration data 132 enables the speech input filter 120B of the audio preprocessor 118 to operate in a speaker-specific mode to enhance the speech of the second person 180B, attenuate sounds other than the speech of the second person 180B (such as attenuating the speech of the first person 180A), or both.

[0159] Method 1500 optionally includes activating a first voice assistant instance based on detecting a first wake word in a first voice output signal at block 1510, and activating a second voice assistant instance different from the first voice assistant instance based on detecting a second wake word in a second voice output signal at block 1512. For example, audio analyzer 140 may activate first voice assistant instance 158A based on detecting wake word 110A in voice output signal 152A, and may activate second voice assistant instance 158B based on detecting wake word 110B in voice output signal 152B.

[0160] Method 1500 optionally includes providing the first speech output signal as input to a first voice assistant instance at block 1514, and providing the second speech output signal as input to a second voice assistant instance different from the first voice assistant instance at block 1516. For example, audio analyzer 140 may provide speech output signal 152A to first voice assistant instance 158A and speech output signal 152B to second voice assistant instance 158B.

[0161] According to one aspect, generating a first voice output signal using a first speaker-specific voice input filter substantially prevents a second user's voice from interfering with the first user's voice assistant session. In some implementations, the first user's voice and the second user's voice temporally overlap, wherein the first speaker-specific voice input filter suppresses the second user's voice during generation of the first voice output signal, and the second speaker-specific voice input filter suppresses the first user's voice during generation of the second voice output signal.

[0162] One benefit of selectively enabling speaker-specific filtering of audio data for multiple users is that such filtering can improve the accuracy of speech recognition by the voice assistant application for each of the multiple users. Another benefit of selectively enabling speaker-specific filtering of audio data for multiple users is that such filtering can limit the ability of users to interrupt voice assistant sessions that they have not yet initiated, thereby enabling multiple voice assistant sessions to proceed simultaneously with each user's speech having minimal or no impact on the voice assistant sessions of other users.

[0163] Figure 15 The method 1500 may be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 15 The method 1500 may be performed by a processor executing instructions, such as reference Figure 18 described.

[0164] refer to Figure 16 , shows a specific implementation of the method 1600 for speaker-specific speech filtering for multiple users. In certain aspects, by Figure 1 The audio analyzer 140, the processor 190, the device 102, the system 100, Figure 2 At least one of the vehicles 250 or a combination thereof performs one or more operations of method 1600.

[0165] Method 1600 optionally includes, at box 1602, processing audio data received from one or more microphones in a vehicle. Processing the audio data optionally includes, at box 1604, generating a first zone audio signal that includes sounds originating from a first zone of a plurality of logical zones of the vehicle and at least partially attenuates sounds originating from outside the first zone, wherein the first zone includes a first seating position. Processing the audio data optionally also includes, at box 1606, generating a second zone audio signal that includes sounds originating from a second zone of the plurality of logical zones and at least partially attenuates sounds originating from outside the second zone, wherein the second zone includes a second seating position. For example, Figure 2 The audio preprocessor 118 processes the audio data from the microphone 104 of the vehicle 250 to generate a different regional audio signal 260 for each region 254 where sound is detected. For illustration, the AIC 208 processes the received audio data to attenuate or remove sounds originating outside of each regional audio signal 260.

[0166] Method 1600 includes detecting speech of the first user and the second user at block 1608. For example, audio analyzer 140 may detect speech of person 180A using speaker detector 128 by processing a portion of audio data 116 corresponding to utterance 108A from person 180A to determine a speech signature and comparing the speech signature to speech signature data 134. Audio analyzer 140 may also detect speech of person 180B by processing a portion of audio data 116 corresponding to utterance 108B from person 180B to determine a speech signature and comparing the speech signature to speech signature data 134.

[0167] Method 1600 optionally includes, at block 1610, detecting that a first user is in a first seated position and a second user is in a second seated position based on sensor data from one or more sensors of the vehicle. For example, the sensor data may correspond to audio data 116 from microphone 104, image data from one or more cameras, data from one or more weight sensors of seat 252, or one or more other types of sensor data for determining which user is in which seated position. For example, the sensor data may indicate that first person 180A is in first seat 252A corresponding to first zone 254A, and second person 180B is in second seat 252B corresponding to second zone 254B.

[0168] The method 1600 includes obtaining, at block 1612, at one or more processors, first voice signature data associated with the first user and second voice signature data associated with the second user. For example, the audio preprocessor 118 may obtain Figure 1 The audio preprocessor 118 may also obtain configuration data 132 that includes at least voice signature data 134A associated with person 180A. Figure 1 Configuration data 132 includes at least voice signature data 134B associated with person 180B.

[0169] Method 1600 includes, at block 1614, selectively enabling, at one or more processors, a first speaker-specific voice input filter based on the first voice signature data to generate a first voice output signal corresponding to the voice of the first user. For example, Figure 1 The configuration data 132 enables the speech input filter 120A of the audio preprocessor 118 to operate in a speaker-specific mode when processing the first region audio signal 260A to enhance the speech of the first person 180A, attenuate sounds other than the speech of the first person 180A (such as attenuating the speech of the second person 180B), or both during generation of the first speech output signal 152A.

[0170] Method 1600 includes, at block 1616, selectively enabling, at one or more processors, a second speaker-specific voice input filter based on the second voice signature data to generate a second voice output signal corresponding to the voice of the second user. For example, Figure 1 The configuration data 132 enables the speech input filter 120B of the audio preprocessor 118 to operate in a speaker-specific mode when processing the second region audio signal 260B to enhance the speech of the second person 180B, attenuate sounds other than the speech of the second person 180B (such as attenuating the speech of the first person 180A), or both.

[0171] One benefit of selectively enabling speaker-specific filtering of audio data is that such filtering can improve the accuracy of speech recognition by a voice assistant application. Another benefit of selectively enabling speaker-specific filtering of audio data is that such filtering can limit the ability of others to interrupt a voice assistant session, such as to enable multiple occupants of a vehicle to simultaneously participate in a voice assistant session without substantially disrupting the voice assistant sessions of other occupants.

[0172] Figure 16The method 1600 may be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 16 The method 1600 may be performed by a processor executing instructions, such as reference Figure 18 described.

[0173] refer to Figure 17 , shows a specific implementation of the method 1700 for speaker-specific speech filtering for multiple users. In certain aspects, by Figure 1 At least one of the audio analyzer 140, processor 190, device 102, system 100, or a combination thereof performs one or more operations of method 1700.

[0174] Method 1700 includes, at block 1702, performing a registration operation to register a first user. The registration operation includes, at block 1704, generating first voice signature data based on one or more utterances of the first user. For example, the first user may be instructed to recite a plurality of words or phrases, which are captured by microphone 104 and processed to determine first voice signature data, such as speaker embedding 414 of the first user. The registration operation also includes, at block 1706, storing the first voice signature data in a voice signature storage device. For example, processor 190 may store voice signature data 134A in memory 142 as part of stored registration data 136.

[0175] The method 1700 includes, after a registration operation, detecting speech of the first user and the second user at block 1708, and retrieving first speech signature data from a speech signature storage device based on identifying the presence of the first user at block 1710. For example, the speech of the first user and the second user may be detected via operation of the speaker detector 128 operating on the filtered audio data 122, and the speech signature data 134A may be included in the configuration data 132 provided to the audio preprocessor 118 in response to detecting the speech of the first user.

[0176] Method 1700 includes, at block 1712, enabling a speaker-specific speech input filter based on the first speech signature data to generate a first speech output signal corresponding to the speech of the first user. For example, audio analyzer 140 activates speech input filter 120A to operate as a speaker-specific speech input filter, thereby generating speech output signal 152A including the speech of the first user.

[0177] The method 1700 also includes generating a second speech output signal corresponding to the speech of the second user using the non-speaker-specific speech input filter at block 1720. For example, when the audio analyzer 140 determines that none of the speech signature data 134 in the enrollment data 136 matches a signature generated based on the speech of the second user, the speech input filter 120B may provide speech enhancement that is not specific to the second user.

[0178] Method 1700 includes, at block 1722, processing the second user's speech to generate second speech signature data corresponding to the second user. For example, processor 190 may store samples of the second user's speech and use the stored samples to train a machine learning model to generate speaker embedding 414 as speech signature data 134B corresponding to the second user. Processor 190 may periodically or occasionally update the second user's speech signature data 134B to enable speech input filter 120B to more accurately perform speaker-specific filtering on the second user's speech as processor 190 obtains more samples of the second user's speech.

[0179] The method 1700 includes storing the second voice signature data in the voice signature storage at block 1724. For example, the processor 190 may store the voice signature data 134B as part of the registration data 136 in the memory 142 so that it is available for retrieval the next time the second user uses the device 102 (e.g., riding in the vehicle 250).

[0180] Figure 17 The method 1700 may be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 17 The method 1700 may be performed by a processor executing instructions, such as reference Figure 18 described.

[0181] refer to Figure 18 , a block diagram of a particular exemplary implementation of a device is depicted and generally designated 1800. In various implementations, the device 1800 may have Figure 18 More or fewer components may be illustrated. In an exemplary implementation, device 1800 may correspond to device 102. In an exemplary implementation, device 1800 may execute the reference Figures 1 to 17 One or more operations described.

[0182] In certain implementations, the device 1800 includes a processor 1806 (e.g., a central processing unit (CPU)). The device 1800 may include one or more additional processors 1810 (e.g., one or more DSPs). In certain aspects, Figure 1 The processor 190 corresponds to the processor 1806, the processor 1810, or a combination thereof. The processor 1810 may include a speech and music coder-decoder (CODEC) 1808, which includes a speech coder ("vocoder") encoder 1836 and a vocoder decoder 1838. Figure 18 In the illustrated example, the processor 1810 also includes the audio pre-processor 118 , the first stage speech processor 124 , and optionally the second stage speech processor 154 .

[0183] Device 1800 may include memory 142 and codec 1834. In certain implementations, Figure 2 and Figure 7 The codec 204 corresponds to Figure 18 The memory 142 may include instructions 1856 that are executable by one or more additional processors 1810 (or processor 1806) to implement the functionality described with reference to the audio pre-processor 118, the first stage speech processor 124, the second stage speech processor 154, or a combination thereof. Figure 18 In the illustrated example, memory 142 also includes registration data 136 .

[0184] Device 1800 may include a display 1828 coupled to a display controller 1826. Audio transducer 164, microphone 104, or both may be coupled to codec 1834. Codec 1834 may include a digital-to-analog converter (DAC) 1802, an analog-to-digital converter (ADC) 1804, or both. In certain implementations, codec 1834 may receive an analog signal from microphone 104 and convert the analog signal into a digital signal (e.g., a digital-to-digital converter) using analog-to-digital converter 1804. Figure 1 The speech and music codec 1808 may process the digital signal, and the digital signal may be further processed by the audio preprocessor 118, the first stage speech processor 124, the second stage speech processor 154, or a combination thereof. In certain implementations, the speech and music codec 1808 may provide the digital signal to the codec 1834. The codec 1834 may convert the digital signal to an analog signal using the digital-to-analog converter 1802 and may provide the analog signal to the audio transducer 164.

[0185] In a particular implementation, device 1800 can be included in a system-in-package or system-on-chip device 1822. In a particular implementation, memory 142, processor 1806, processor 1810, display controller 1826, codec 1834, and modem 1854 are included in the system-in-package or system-on-chip device 1822. In a particular implementation, input device 1830 and power source 1844 are coupled to the system-in-package or system-on-chip device 1822. Additionally, in a particular implementation, as shown in FIG. Figure 18 As illustrated, the display 1828, input device 1830, audio transducer 164, microphone 104, antenna 1852, and power source 1844 are external to the system-in-package or system-on-chip device 1822. In a particular implementation, each of the display 1828, input device 1830, audio transducer 164, microphone 104, antenna 1852, and power source 1844 can be coupled to a component of the system-in-package or system-on-chip device 1822, such as an interface or controller.

[0186] In some implementations, device 1800 includes a modem 1854 coupled to antenna 1852 via transceiver 1850. In some such implementations, modem 1854 can be configured to transmit data associated with utterances from the first person (e.g., Figure 1 The second-stage speech processor 154 may be omitted from the device 1800; however, speaker-specific speech input filtering may be performed at the device 1800.

[0187] Device 1800 may include a smart speaker, a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a head-mounted device, an augmented reality head-mounted device, a mixed reality head-mounted device, a virtual reality head-mounted device, an aircraft, a home automation system, a voice-activated device, a wireless speaker and voice-activated device, a portable electronic device, an automobile, a computing device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0188] In conjunction with the described implementations, an apparatus includes means for detecting speech of a first user and a second user. For example, the means for detecting speech of the first user and the second user may correspond to device 102, microphone 104, processor 190, audio analyzer 140, audio preprocessor 118, speech input filter 120, first-stage speech processor 124, wake word detector 126, speaker detector 128, integrated circuit 802, processor 1806, processor 1810, one or more other circuits or components configured to detect speech of the first user and the second user, or any combination thereof.

[0189] The apparatus includes means for obtaining first voice signature data associated with a first user and second voice signature data associated with a second user. For example, the means for obtaining the first voice signature data and the second voice signature data may correspond to device 102, processor 190, audio analyzer 140, audio preprocessor 118, voice input filter 120, first-stage voice processor 124, speaker detector 128, integrated circuit 802, processor 1806, processor 1810, one or more other circuits or components configured to obtain voice signature data, or any combination thereof.

[0190] The apparatus also includes means for selectively enabling a first speaker-specific voice input filter based on the first voice signature data to generate a first voice output signal corresponding to the voice of the first user. For example, the means for selectively enabling the first speaker-specific voice input filter may correspond to device 102, processor 190, audio analyzer 140, audio preprocessor 118, voice input filter 120, first-stage voice processor 124, speaker detector 128, integrated circuit 802, processor 1806, processor 1810, one or more other circuits or components configured to selectively enable the first speaker-specific voice input filter, or any combination thereof.

[0191] The apparatus further includes means for selectively enabling a second speaker-specific voice input filter based on the second voice signature data to generate a second voice output signal corresponding to the voice of the second user. For example, the means for selectively enabling the second speaker-specific voice input filter may correspond to device 102, processor 190, audio analyzer 140, audio preprocessor 118, voice input filter 120, first-stage voice processor 124, speaker detector 128, integrated circuit 802, processor 1806, processor 1810, one or more other circuits or components configured to selectively enable the second speaker-specific voice input filter, or any combination thereof.

[0192] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 142) includes instructions (e.g., instructions 1856) that, when executed by one or more processors (e.g., processor(s) 190, processor(s) 1810, or processor 1806), cause the one or more processors to detect speech of a first user and a second user, obtain first speech signature data associated with the first user and second speech signature data associated with the second user, selectively enable a first speaker-specific speech input filter based on the first speech signature data to generate a first speech output signal corresponding to the speech of the first user, and selectively enable a second speaker-specific speech input filter based on the second speech signature data to generate a second speech output signal corresponding to the speech of the second user.

[0193] Specific aspects of the present disclosure are described below in various sets of related embodiments:

[0194] According to embodiment 1, a device includes: one or more processors, wherein the one or more processors are configured to: detect the voice of a first user and a second user; obtain first voice signature data associated with the first user and second voice signature data associated with the second user; selectively enable a first speaker-specific voice input filter based on the first voice signature data to generate a first voice output signal corresponding to the voice of the first user; and selectively enable a second speaker-specific voice input filter based on the second voice signature data to generate a second voice output signal corresponding to the voice of the second user.

[0195] Embodiment 2 includes an apparatus according to embodiment 1, wherein the one or more processors are implemented in a vehicle and are configured to: selectively enable the first speaker-specific voice input filter based on a first seating position of the first user within the vehicle; and selectively enable the second speaker-specific voice input filter based on a second seating position of the second user within the vehicle.

[0196] Embodiment 3 includes the apparatus of Embodiment 2, wherein the one or more processors are further configured to detect that the first user is in the first seating position and the second user is in the second seating position based on sensor data from one or more sensors of the vehicle.

[0197] Example 4 includes an apparatus according to Example 2 or Example 3, wherein the one or more processors are further configured to process audio data received from one or more microphones in the vehicle to: generate a first area audio signal, the first area audio signal including sounds originating from a first area of multiple logical areas of the vehicle and at least partially attenuating sounds originating from outside the first area, wherein the first area includes the first seating position; and generate a second area audio signal, the second area audio signal including sounds originating from a second area of the multiple logical areas and at least partially attenuating sounds originating from outside the second area, wherein the second area includes the second seating position.

[0198] Embodiment 5 includes a device according to embodiment 4, wherein the one or more processors are further configured to: as part of a first filtering operation of the first area audio signal, enable the first speaker-specific voice input filter to enhance the voice of the first user, attenuate sounds other than the voice of the first user, or both, thereby generating the first voice output signal; and as part of a second filtering operation of the second area audio signal, enable the second speaker-specific voice input filter to enhance the voice of the second user, attenuate sounds other than the voice of the second user, or both, thereby generating the second voice output signal.

[0199] Embodiment 6 includes a device according to any one of embodiments 1 to 5, wherein the one or more processors are further configured to: provide the first voice output signal as input to a first voice assistant instance; and provide the second voice output signal as input to a second voice assistant instance different from the first voice assistant instance.

[0200] Embodiment 7 includes the apparatus of embodiment 6, wherein generating the first speech output signal using the first speaker-specific speech input filter substantially prevents the speech of the second user from interfering with the voice assistant session of the first user.

[0201] Embodiment 8 includes a device according to embodiment 6 or embodiment 7, wherein the first voice assistant instance corresponds to a first instance of a first voice assistant application, and wherein the second voice assistant instance corresponds to a second instance of the first voice assistant application.

[0202] Embodiment 9 includes a device according to embodiment 6 or embodiment 7, wherein the first voice assistant instance corresponds to a first voice assistant application, and wherein the second voice assistant instance corresponds to a second voice assistant application that is different from the first voice assistant application.

[0203] Embodiment 10 includes a device according to any one of embodiments 6 to 9, wherein the one or more processors are further configured to: activate the first voice assistant instance based on detecting a first wake-up word in the first voice output signal; and activate the second voice assistant instance based on detecting a second wake-up word in the second voice output signal.

[0204] Embodiment 11 includes an apparatus according to any one of embodiments 1 to 10, wherein the speech of the first user and the speech of the second user overlap in time, wherein the first speaker-specific speech input filter suppresses the speech of the second user during generation of the first speech output signal, and wherein the second speaker-specific speech input filter suppresses the speech of the first user during generation of the second speech output signal.

[0205] Embodiment 12 includes an apparatus according to any one of embodiments 1 to 11, wherein the first speech signature data corresponds to a first speaker embedding, and wherein the one or more processors are configured to enable the first speaker-specific speech input filter by providing the first speaker embedding as input to a speech enhancement model.

[0206] Embodiment 13 includes an apparatus according to any one of embodiments 1 to 12, wherein the one or more processors are further configured to: during a registration operation: generate the first voice signature data based on one or more utterances of the first user; and store the first voice signature data in a voice signature storage device; and after the registration operation, retrieve the first voice signature data from the voice signature storage device based on identifying the presence of the first user.

[0207] Embodiment 14 includes the apparatus of any one of Embodiments 1 to 13, wherein the one or more processors are further configured to process the speech of the second user to generate the second speech signature data.

[0208] Embodiment 15 includes the device of any one of Embodiments 1 to 14, further comprising a microphone configured to capture the voice of the first user, the voice of the second user, or both.

[0209] Embodiment 16 includes the apparatus of any one of embodiments 1 to 15, further comprising a modem configured to transmit data associated with the first voice output signal to a remote voice assistant server.

[0210] Embodiment 17 includes the device of any one of Embodiments 1 to 16, further comprising a speaker configured to output a sound corresponding to a voice assistant response to the speech of the first user.

[0211] Embodiment 18 includes the device of any one of Embodiments 1 to 17, further comprising a display device configured to display data corresponding to the voice assistant response to the speech of the first user.

[0212] According to embodiment 19, a method includes: detecting the speech of a first user and a second user at one or more processors; obtaining first speech signature data associated with the first user and second speech signature data associated with the second user at the one or more processors; selectively enabling a first speaker-specific speech input filter based on the first speech signature data at the one or more processors to generate a first speech output signal corresponding to the speech of the first user; and selectively enabling a second speaker-specific speech input filter based on the second speech signature data at the one or more processors to generate a second speech output signal corresponding to the speech of the second user.

[0213] Embodiment 20 includes a method according to embodiment 19, wherein the first speaker-specific voice input filter is selectively enabled based on a first seating position of the first user within a vehicle, and wherein the second speaker-specific voice input filter is selectively enabled based on a second seating position of the second user within the vehicle.

[0214] Embodiment 21 includes the method of Embodiment 20, further comprising detecting, based on sensor data from one or more sensors of the vehicle, that the first user is in the first seating position and the second user is in the second seating position.

[0215] Embodiment 22 includes a method according to embodiment 20 or embodiment 21, the method further comprising processing audio data received from one or more microphones in the vehicle, comprising: generating a first area audio signal, the first area audio signal including sounds originating from a first area among multiple logical areas of the vehicle and at least partially attenuating sounds originating from outside the first area, wherein the first area includes the first seating position; and generating a second area audio signal, the second area audio signal including sounds originating from a second area among the multiple logical areas and at least partially attenuating sounds originating from outside the second area, wherein the second area includes the second seating position.

[0216] Embodiment 23 includes the method according to embodiment 22, further comprising: as part of a first filtering operation of the first area audio signal, enabling the first speaker-specific voice input filter to enhance the voice of the first user, attenuate sounds other than the voice of the first user, or both, thereby generating the first voice output signal; and as part of a second filtering operation of the second area audio signal, enabling the second speaker-specific voice input filter to enhance the voice of the second user, attenuate sounds other than the voice of the second user, or both, thereby generating the second voice output signal.

[0217] Embodiment 24 includes a method according to any one of embodiments 19 to 24, further comprising: providing the first voice output signal as input to a first voice assistant instance; and providing the second voice output signal as input to a second voice assistant instance different from the first voice assistant instance.

[0218] Embodiment 25 includes the method of embodiment 24, wherein generating the first speech output signal using the first speaker-specific speech input filter substantially prevents the speech of the second user from interfering with the voice assistant session of the first user.

[0219] Embodiment 26 includes a method according to embodiment 24 or embodiment 25, wherein the first voice assistant instance corresponds to a first instance of a first voice assistant application, and wherein the second voice assistant instance corresponds to a second instance of the first voice assistant application.

[0220] Embodiment 27 includes a method according to embodiment 24 or embodiment 26, wherein the first voice assistant instance corresponds to a first voice assistant application, and wherein the second voice assistant instance corresponds to a second voice assistant application that is different from the first voice assistant application.

[0221] Embodiment 28 includes a method according to any one of embodiments 24 to 27, further comprising: activating the first voice assistant instance based on detecting a first wake-up word in the first voice output signal; and activating the second voice assistant instance based on detecting a second wake-up word in the second voice output signal.

[0222] Embodiment 29 includes a method according to any one of embodiments 19 to 28, wherein the speech of the first user and the speech of the second user overlap in time, wherein the first speaker-specific speech input filter suppresses the speech of the second user during generation of the first speech output signal, and wherein the second speaker-specific speech input filter suppresses the speech of the first user during generation of the second speech output signal.

[0223] Embodiment 30 includes a method according to any one of embodiments 19 to 29, wherein the first speech signature data corresponds to a first speaker embedding, and wherein enabling the first speaker-specific speech input filter includes providing the first speaker embedding as input to a speech enhancement model.

[0224] Embodiment 31 includes the method according to any one of embodiments 19 to 30, further comprising: during a registration operation: generating the first voice signature data based on one or more utterances of the first user; and storing the first voice signature data in a voice signature storage device; and after the registration operation, retrieving the first voice signature data from the voice signature storage device based on identifying the presence of the first user.

[0225] Embodiment 32 includes the method of any one of Embodiments 19 to 31, further comprising processing the speech of the second user to generate the second speech signature data.

[0226] Embodiment 33 includes the method of any one of Embodiments 19 to 32, further comprising capturing the voice of the first user, the voice of the second user, or both via a microphone.

[0227] Embodiment 34 includes the method of any one of embodiments 19 to 33, further comprising transmitting data associated with the first voice output signal to a remote voice assistant server.

[0228] Embodiment 35 includes the method of any one of embodiments 19 to 34, further comprising outputting a sound corresponding to a voice assistant response to the speech of the first user.

[0229] Embodiment 36 includes the method of any one of embodiments 19 to 35, further comprising displaying data corresponding to the voice assistant response to the speech of the first user.

[0230] Embodiment 37 includes an apparatus comprising components for performing the method according to any one of embodiments 19 to 36.

[0231] Embodiment 38 includes a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform the method according to any one of embodiments 19 to 36.

[0232] Embodiment 39 includes a device comprising: a memory storing instructions; and a processor configured to execute the instructions to perform a method according to any one of embodiments 19 to 36.

[0233] According to embodiment 40, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to: detect the speech of a first user and a second user; obtain first speech signature data associated with the first user and second speech signature data associated with the second user; selectively enable a first speaker-specific speech input filter based on the first speech signature data to generate a first speech output signal corresponding to the speech of the first user; and selectively enable a second speaker-specific speech input filter based on the second speech signature data to generate a second speech output signal corresponding to the speech of the second user.

[0234] Embodiment 41 includes a non-transitory computer-readable medium according to embodiment 40, wherein the instructions are executable to further cause the one or more processors to: generate a first zone audio signal, the first zone audio signal including sounds originating from a first zone among a plurality of logical zones of a vehicle and at least partially attenuating sounds originating from outside the first zone, wherein the first zone includes a first seating position of the first user; and generate a second zone audio signal, the second zone audio signal including sounds originating from a second zone among the plurality of logical zones and at least partially attenuating sounds originating from outside the second zone, wherein the second zone includes a second seating position of the second user.

[0235] According to embodiment 42, an apparatus includes: a component for detecting the voice of a first user and a second user; a component for obtaining first voice signature data associated with the first user and second voice signature data associated with the second user; a component for selectively enabling a first speaker-specific voice input filter based on the first voice signature data to generate a first voice output signal corresponding to the voice of the first user; and a component for selectively enabling a second speaker-specific voice input filter based on the second voice signature data to generate a second voice output signal corresponding to the voice of the second user.

[0236] It will also be apparent to those skilled in the art that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in conjunction with the specific implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of the two. Various illustrative components, blocks, configurations, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. A skilled artisan may implement the described functionality in different ways for each specific application, and such specific implementation decisions should not be interpreted as causing a departure from the scope of this disclosure.

[0237] The steps of the method or algorithm described in conjunction with the specific implementation disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative, the storage medium may be integral with the processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In an alternative, the processor and the storage medium may reside in a computing device or a user terminal as discrete components.

[0238] The preceding description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but should be accorded the broadest scope possible consistent with the principles and novel features as defined by the following claims.

Claims

1. A device, comprising: One or more processors configured to: detecting the voices of the first user and the second user; obtaining first voice signature data associated with the first user and second voice signature data associated with the second user; selectively enabling a first speaker-specific speech input filter based on the first speech signature data to generate a first speech output signal corresponding to the speech of the first user; as well as A second speaker-specific speech input filter based on the second speech signature data is selectively enabled to generate a second speech output signal corresponding to the speech of the second user.

2. The apparatus of claim 1 , wherein the one or more processors are implemented in a vehicle and are configured to: selectively enabling the first speaker-specific voice input filter based on a first seating location of the first user within the vehicle; and The second speaker-specific voice input filter is selectively enabled based on a second seating location of the second user within the vehicle.

3. The apparatus of claim 2, wherein the one or more processors are further configured to detect that the first user is in the first seating position and the second user is in the second seating position based on sensor data from one or more sensors of the vehicle.

4. The apparatus of claim 2, wherein the one or more processors are further configured to process audio data received from one or more microphones in the vehicle to: generating a first zone audio signal that includes sounds originating from a first zone of a plurality of logical zones of the vehicle and at least partially attenuates sounds originating outside of the first zone, wherein the first zone includes the first seating position; and A second zone audio signal is generated that includes sound originating from a second zone of the plurality of logical zones and at least partially attenuates sound originating outside of the second zone, wherein the second zone includes the second seating position.

5. The apparatus of claim 4, wherein the one or more processors are further configured to: as part of a first filtering operation of the first region audio signal, enabling the first speaker-specific speech input filter to enhance the speech of the first user, attenuate sounds other than the speech of the first user, or both, thereby generating the first speech output signal; and As part of a second filtering operation of the second area audio signal, the second speaker-specific voice input filter is enabled to enhance the voice of the second user, attenuate sounds other than the voice of the second user, or both, thereby generating a second voice output signal.

6. The apparatus of claim 1 , wherein the one or more processors are further configured to: providing the first speech output signal as input to a first voice assistant instance; and The second speech output signal is provided as input to a second voice assistant instance that is different from the first voice assistant instance.

7. The apparatus of claim 6, wherein generating the first speech output signal using the first speaker-specific speech input filter substantially prevents the speech of the second user from interfering with the voice assistant session of the first user.

8. The apparatus of claim 6, wherein the first voice assistant instance corresponds to a first instance of a first voice assistant application, and wherein the second voice assistant instance corresponds to a second instance of the first voice assistant application.

9. The apparatus of claim 6, wherein the first voice assistant instance corresponds to a first voice assistant application, and wherein the second voice assistant instance corresponds to a second voice assistant application that is different from the first voice assistant application.

10. The apparatus of claim 6, wherein the one or more processors are further configured to: activating the first voice assistant instance based on detecting a first wake word in the first speech output signal; and The second voice assistant instance is activated based on detecting a second wake-up word in the second voice output signal.

11. The apparatus of claim 1 , wherein the speech of the first user and the speech of the second user overlap in time, wherein the first speaker-specific speech input filter suppresses the speech of the second user during generation of the first speech output signal, and wherein the second speaker-specific speech input filter suppresses the speech of the first user during generation of the second speech output signal.

12. The apparatus of claim 1 , wherein the first speech signature data corresponds to a first speaker embedding, and wherein the one or more processors are configured to enable the first speaker-specific speech input filter by providing the first speaker embedding as input to a speech enhancement model.

13. The apparatus of claim 1 , wherein the one or more processors are further configured to: During the registration operation: generating the first voice signature data based on one or more utterances of the first user; as well as storing the first voice signature data in a voice signature storage device; as well as After the registration operation, the first voice signature data is retrieved from the voice signature storage device based on identifying the presence of the first user.

14. The device of claim 1, wherein the one or more processors are further configured to process the speech of the second user to generate the second speech signature data.

15. The device of claim 1, further comprising a microphone configured to capture the voice of the first user, the voice of the second user, or both.

16. The device of claim 1, further comprising a modem configured to transmit data associated with the first voice output signal to a remote voice assistant server.

17. The device of claim 1, further comprising a speaker configured to output a sound corresponding to a voice assistant response to the speech of the first user.

18. The device of claim 1, further comprising a display device configured to display data corresponding to the voice assistant response to the speech of the first user.

19. A method comprising: detecting, at one or more processors, speech of a first user and a second user; obtaining, at the one or more processors, first voice signature data associated with the first user and second voice signature data associated with the second user; selectively enabling, at the one or more processors, a first speaker-specific speech input filter based on the first speech signature data to generate a first speech output signal corresponding to the speech of the first user; as well as A second speaker-specific speech input filter based on the second speech signature data is selectively enabled at the one or more processors to generate a second speech output signal corresponding to the speech of the second user.

20. The method of claim 19, wherein the first speaker-specific voice input filter is selectively enabled based on a first seating position of the first user within a vehicle, and wherein the second speaker-specific voice input filter is selectively enabled based on a second seating position of the second user within the vehicle.

21. The method of claim 20, further comprising detecting that the first user is in the first seating position and the second user is in the second seating position based on sensor data from one or more sensors of the vehicle.

22. The method of claim 20, further comprising processing audio data received from one or more microphones in the vehicle, comprising: generating a first zone audio signal that includes sounds originating from a first zone of a plurality of logical zones of the vehicle and at least partially attenuates sounds originating outside of the first zone, wherein the first zone includes the first seating position; as well as A second zone audio signal is generated that includes sound originating from a second zone of the plurality of logical zones and at least partially attenuates sound originating outside of the second zone, wherein the second zone includes the second seating position.

23. The method according to claim 19, further comprising: providing the first speech output signal as input to a first voice assistant instance; as well as The second speech output signal is provided as input to a second voice assistant instance that is different from the first voice assistant instance.

24. The method of claim 23, wherein generating the first speech output signal using the first speaker-specific speech input filter substantially prevents the speech of the second user from interfering with the voice assistant session of the first user.

25. The method according to claim 23, further comprising: activating the first voice assistant instance based on detecting a first wake-up word in the first voice output signal; as well as The second voice assistant instance is activated based on detecting a second wake-up word in the second voice output signal.

26. A method according to claim 19, wherein the speech of the first user and the speech of the second user overlap in time, wherein the first speaker-specific speech input filter suppresses the speech of the second user during generation of the first speech output signal, and wherein the second speaker-specific speech input filter suppresses the speech of the first user during generation of the second speech output signal.

27. The method of claim 19, wherein the first speech signature data corresponds to a first speaker embedding, and wherein enabling the first speaker-specific speech input filter comprises providing the first speaker embedding as input to a speech enhancement model.

28. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: detecting the voices of the first user and the second user; obtaining first voice signature data associated with the first user and second voice signature data associated with the second user; selectively enabling a first speaker-specific speech input filter based on the first speech signature data to generate a first speech output signal corresponding to the speech of the first user; as well as A second speaker-specific speech input filter based on the second speech signature data is selectively enabled to generate a second speech output signal corresponding to the speech of the second user.

29. The non-transitory computer readable medium of claim 28, wherein the instructions are executable to further cause the one or more processors to: generating a first zone audio signal that includes sound originating from a first zone of a plurality of logical zones of a vehicle and at least partially attenuates sound originating from outside the first zone, wherein the first zone includes a first seating location of the first user; and A second zone audio signal is generated that includes sound originating from a second zone of the plurality of logical zones and at least partially attenuates sound originating outside of the second zone, wherein the second zone includes a second seating location of the second user.

30. An apparatus comprising: means for detecting speech of a first user and a second user; means for obtaining first voice signature data associated with the first user and second voice signature data associated with the second user; means for selectively enabling a first speaker-specific speech input filter based on the first speech signature data to generate a first speech output signal corresponding to the speech of the first user; as well as means for selectively enabling a second speaker-specific speech input filter based on the second speech signature data to generate a second speech output signal corresponding to the speech of the second user.