Speech filter for speech processing

By detecting wake words in voice assistant devices, the speaker-specific voice input filter is enabled, and the problem of voice assistant's difficulty in voice separation in noisy environments and multi-person interactions is solved, achieving more efficient resource utilization and user experience improvement.

CN120390954APending Publication Date: 2025-07-29QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085129.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-11-27
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, voice assistant devices are difficult to effectively separate voice and background noise in noisy environments or interactions between multiple people, resulting in poor user experience and waste of resources.

Method used

Enable speaker-specific voice input filters by detecting wake words, enhance specific person's voice, weaken or remove other people's voice and background noise, and process only audio data related to specific people.

Benefits of technology

It improves user experience, reduces resource waste, enhances the accuracy of voice recognition, prevents interruptions, and improves the operation efficiency of voice assistants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390954A_ABST
    Figure CN120390954A_ABST
Patent Text Reader

Abstract

An apparatus includes one or more processors configured to obtain first speech signature data associated with a first person based on detecting a wake-up word in an utterance from the first person. The one or more processors are also configured to selectively enable a speaker-specific voice input filter based on the first voice signature data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] I. Cross - Reference to Related Applications

[0002] This application claims priority to co - owned U.S. Non - Provisional Patent Application No. 18 / 069,618, filed on December 21, 2022, the entire content of which is hereby incorporated by reference in its entirety. II. Technical Field

[0003] The present disclosure generally relates to selectively filtering audio data for voice processing.

[0004] III. Related Art

[0005] Technological advancements have led to smaller and more powerful computing devices. Many of these devices can communicate voice and data packets over wired or wireless networks. In addition, many such devices incorporate additional functions, such as digital still cameras, digital video cameras, digital recorders, and audio file players. Further, such devices can process executable instructions, including software applications, such as web browser applications, which can be used to access the Internet.

[0006] Many of these devices incorporate functionality for interacting with users via voice commands. For example, a computing device can include a voice assistant application and one or more microphones to generate audio data based on detected sounds. In this example, the voice assistant application is configured to perform various operations in response to a user's voice, such as transmitting commands to other devices, retrieving information, and so on.

[0007] While voice assistant applications can enable hands - free interaction with computing devices, using voice to control a computing device is not without complications. For example, when a computing device is in a noisy environment, it may be difficult to separate the voice from background noise. As another example, when multiple people are present, voices from multiple people may be detected, resulting in confused input to the computing device and an unsatisfactory user experience. IV. Summary of the Invention

[0008] According to one particular implementation of the present disclosure, a device includes one or more processors configured to obtain first voice signature data associated with a first person based on detecting a wake - up word in the speech of the first person. The one or more processors are further configured to selectively enable a speaker - specific voice input filter based on the first voice signature data.

[0009] According to another particular implementation of the present disclosure, a method includes obtaining first voice signature data associated with a first person based on detecting a wake - up word in the speech of the first person. The method further includes selectively enabling a speaker - specific voice input filter based on the first voice signature data.

[0010] In another specific implementation according to the present disclosure, a non-transitory computer-readable medium stores instructions that can be executed by one or more processors to cause the one or more processors to obtain first voice signature data associated with a first person based on detecting a wake word in the speech of the first person. These instructions can also be executed by the one or more processors to selectively enable a speaker-specific voice input filter based on the first voice signature data.

[0011] In another specific implementation according to the present disclosure, a device includes components for obtaining first voice signature data associated with a first person based on detecting a wake word in the speech of the first person. The device also includes components for selectively enabling a speaker-specific voice input filter based on the first voice signature data.

[0012] Other aspects, advantages, and features of the present disclosure will become apparent upon reviewing the entire application including the following sections: the description of the drawings, the detailed description, and the claims. V. Description of the Drawings

[0013] Figure 1 is a block diagram of specific illustrative aspects of a system operable to selectively filter audio data for voice processing according to some examples of the present disclosure.

[0014] Figure 2A is a diagram of illustrative aspects of operations associated with selectively filtering audio data for voice processing according to some examples of the present disclosure.

[0015] Figure 2B is a diagram of illustrative aspects of operations associated with selectively filtering audio data for voice processing according to some examples of the present disclosure.

[0016] Figure 2C is a diagram of illustrative aspects of operations associated with selectively filtering audio data for voice processing according to some examples of the present disclosure.

[0017] Figure 3 is a diagram of illustrative aspects of operations associated with selectively filtering audio data for voice processing according to some examples of the present disclosure.

[0018] Figure 4 is a diagram of illustrative aspects of operations associated with selectively filtering audio data for voice processing according to some examples of the present disclosure.

[0019] Figure 5 is a diagram of illustrative aspects of operations associated with selectively filtering audio data for voice processing according to some examples of the present disclosure.

[0020] Figure 6 Is a diagram of a first example of a vehicle operable to selectively filter audio data for speech processing according to some examples of the present disclosure.

[0021] Figure 7 Is a diagram of a voice control speaker system operable to selectively filter audio data for speech processing according to some examples of the present disclosure.

[0022] Figure 8 Illustrates an example of an integrated circuit operable to selectively filter audio data for speech processing according to some examples of the present disclosure.

[0023] Figure 9 Is a diagram of a mobile device operable to selectively filter audio data for speech processing according to some examples of the present disclosure.

[0024] Figure 10 Is a diagram of a wearable electronic device operable to selectively filter audio data for speech processing according to some examples of the present disclosure.

[0025] Figure 11 Is a diagram of a camera operable to selectively filter audio data for speech processing according to some examples of the present disclosure.

[0026] Figure 12 Is a diagram of a head-mounted device (such as a virtual reality, mixed reality, or augmented reality head-mounted device) operable to selectively filter audio data for speech processing according to some examples of the present disclosure.

[0027] Figure 13 Is a diagram of a second example of a vehicle operable to selectively filter audio data for speech processing according to some examples of the present disclosure.

[0028] Figure 14 Is according to some examples of the present disclosure Figure 1 A diagram of illustrative aspects of the operation of components of a system.

[0029] Figure 15 Is according to some examples of the present disclosure and can be performed by Figure 1 A diagram of a specific implementation of a method for selectively filtering audio data for speech processing that can be performed by a device.

[0030] Figure 16 Is according to some examples of the present disclosure and can be performed by Figure 1 A diagram of a specific implementation of a method for selectively filtering audio data for speech processing that can be performed by a device.

[0031] Figure 17 is a diagram of a specific implementation of a method for selectively filtering audio data for speech processing that can be performed by a device of Figure 1 the present disclosure.

[0032] Figure 18 is a block diagram of a specific illustrative example of a device operable to selectively filter audio data for speech processing according to some examples of the present disclosure. VI. DETAILED DESCRIPTION

[0033] According to specific aspects disclosed herein, a speaker-specific speech input filter is selectively used to generate a speech input for a voice assistant. For example, in some implementations, the speaker-specific speech input filter is enabled in response to detecting a wake word in the utterance from a specific person. In such implementations, the speaker-specific speech input filter, when enabled, is configured to process the received audio data to enhance the speech of the specific person. Enhancing the speech of the specific person may include, for example, reducing background noise in the audio data, removing the speech of one or more other persons from the audio data, and the like.

[0034] The voice assistant enables hands-free interaction with the computing device; however, when multiple people are present, the operation of the voice assistant may be interrupted or confused by the speech from multiple people. As an example, a first person may initiate an interaction with the voice assistant by sequentially speaking a wake word and a command. In this example, if a second person speaks while the first person is speaking to the voice assistant, the speech of the first person and the speech of the second person may overlap, such that the voice assistant cannot correctly interpret the command from the first person. This confusion results in an unsatisfactory user experience and waste (because the voice assistant processes the audio data without generating the requested result). By way of illustration, this confusion may lead to inaccurate speech recognition, which may cause the voice assistant to make inappropriate responses.

[0035] Another example may be referred to as an interruption. In the case of an interruption, a first person may initiate an interaction with the voice assistant by sequentially speaking a wake word and a first command. In this example, a second person may interrupt the interaction between the first person and the voice assistant by speaking a wake word (possibly followed by a second command) before the voice assistant finishes performing the operation associated with the first command. When the second person interrupts, the voice assistant may stop performing the operation associated with the first command to process the input from the second person (e.g., the second command). The interruption results in an unsatisfactory user experience and waste in a manner similar to confusion, because the voice assistant processes the audio data associated with the first command without generating the requested result.

[0036] According to certain aspects, selectively enabling a speaker-specific voice input filter implements an improved user experience and more efficient use of resources (e.g., power, processing time, bandwidth, etc.). For example, the speaker-specific voice input filter may be enabled in response to detecting a wake word in the utterance from a first person. In this example, the speaker-specific voice input filter is configured to provide filtered audio data corresponding to the voice from the first person to the voice assistant based on voice signature data associated with the first person. The speaker-specific voice input filter is configured to remove the voice from other people (e.g., a second person) from the filtered audio data provided to the voice assistant. Thus, the first person can conduct a voice assistant session without interruption, thereby improving resource utilization and the user experience.

[0037] Certain aspects of the present disclosure are described below with reference to the accompanying drawings. In this description, common features are designated by common reference numerals. As used herein, various terms are for the purpose of describing particular specific implementations only and are not intended to limit the specific implementations. For example, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Additionally, some features described herein are singular in some specific implementations and plural in other specific implementations. For purposes of illustration, Figure 1 a device 102 including one or more processors ( Figure 1 referred to as "processor" 190) is depicted, which indicates that in some specific implementations, the device 102 includes a single processor 190, while in other specific implementations, the device 102 includes multiple processors 190. For ease of reference herein, such features are typically introduced as "one or more" features and subsequently referred to in the singular form or an optional plural form (as typically indicated by "(s)") unless aspects related to the multiplicity of the features are being described.

[0038] In some of the drawings, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference numeral is used for each feature, and the different instances are distinguished by adding letters to the reference numeral. When a group or a type of features is referred to herein (e.g., when no particular feature among these features is being referred to), the reference numeral is used without the distinguishing letter. However, when a particular feature among multiple features of the same type is being referred to herein, the reference numeral is used together with the distinguishing letter. For example, reference Figure 6, which illustrates a plurality of microphones and these microphones are associated with reference numerals 104A through 104F. When referring to a specific one of these microphones (such as microphone 104A), the distinguishing letter "A" is used. However, when referring to any arbitrary one of these microphones or referring to these microphones as a group, the reference numeral 104 is used without a distinguishing letter.

[0039] As used herein, the term "comprise" may be used interchangeably with "include". Additionally, the term "wherein" may be used interchangeably with "where". As used herein, "exemplary" indicates examples, specific implementations, and / or aspects, and should not be construed as restrictive or indicating a preference or preferred specific implementation. As used herein, ordinal terms (e.g., "first", "second", "third", etc.) used to modify elements (such as structures, components, operations, etc.) do not themselves indicate any priority or order of the element relative to another element, but merely distinguish the element from another element with the same name (but using an ordinal term). As used herein, the term "set" refers to one or more particular elements among a particular group of elements, and the term "plurality" refers to more than one (e.g., two or more) particular elements.

[0040] As used herein, "coupled" may include "communicatively coupled", "electrically coupled", or "physically coupled", and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), etc. As an illustrative, non-limiting example, two devices (or components) that are electrically coupled may be included in the same device or in different devices, and may be connected via electronics, one or more connectors, or inductive coupling. In some specific implementations, two devices (or components) that are communicatively coupled (such as electrically connected) may directly or indirectly transmit and receive signals (e.g., digital signals or analog signals) via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without an intermediate component.

[0041] In the present disclosure, terms such as "determine", "calculate", "estimate", "shift", "adjust", etc. may be used to describe how to perform one or more operations. It should be noted that such terms should not be construed as restrictive, and other techniques may be utilized to perform similar operations. Additionally, as mentioned herein, "generate", "calculate", "estimate", "use", "select", "access", and "determine" may be used interchangeably. For example, "generate", "calculate", "estimate", or "determine" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or may refer to using, selecting, or accessing (such as by another component or device) a parameter (or signal) that has already been generated.

[0042] Figure 1 Illustrates a particular specific implementation of system 100, which is operable to selectively filter audio data provided to one or more voice assistant applications. System 100 includes device 102, which includes one or more processors 190 and memory 142. Device 102 is coupled to or includes: one or more microphones 104 coupled via input interface 114 of processor 190, and one or more audio transducers 162 (e.g., speakers) coupled via output interface 158 of processor 190.

[0043] In Figure 1 it, microphone 104 is disposed in an acoustic environment to receive sound 106. Sound 106 may include, for example, speech 108 from one or more persons 180, ambient sound 112, or both. Microphone 104 is configured to provide a signal to input interface 114 to generate audio data 116 representative of sound 106. The audio data 116 is provided to processor 190 for processing, as further described below.

[0044] In Figure 1In the illustrated example, the processor 190 includes an audio analyzer 140. The audio analyzer 140 includes an audio pre-processor 118 and a multi-stage speech processor that includes a first-stage speech processor 124 and a second-stage speech processor 154. In a particular implementation, the first-stage speech processor 124 is configured to perform wake-word detection, and the second-stage speech processor 154 is configured to perform more resource-intensive speech processing, such as speech-to-text conversion, natural language processing, and related operations. To conserve resources (e.g., power, processor time, etc.) associated with the resource-intensive speech processing performed at the second-stage speech processor 154, the first-stage speech processor 124 is configured to provide audio data 150 to the second-stage speech processor 154 after the first-stage speech processor 124 detects a wake word 110 in the utterance 108 from the person 180. In some implementations, the second-stage speech processor 154 remains in a low-power or standby state until the first-stage speech processor 124 signals the second-stage speech processor 154 to wake up or enter a high-power state to process the audio data 150. In some such implementations, the first-stage speech processor 124 operates in an always-on mode such that the first-stage speech processor 124 is always listening for the wake word 110. However, in other such implementations, the first-stage speech processor 124 is configured to be activated by some additional action, such as a button press. The technical advantage of such a multi-stage speech processor is that the most resource-intensive operations associated with speech processing can be offloaded to the second-stage speech processor 154, which is only active after the wake word 110 is detected, thereby conserving power, processor time, and other computational resources associated with the operation of the second-stage speech processor 154.

[0045] Although the second-stage speech processor 154 is illustrated as being included in the device 102 in Figure 1 some implementations, in some implementations, the second-stage speech processor 154 is remote from the device 102. For example, the second-stage speech processor 154 may be provided at a remote voice assistant server. In such implementations, after the first-stage speech processor 124 detects the wake word 110, the device 102 sends the audio data 150 to the second-stage speech processor 154 via one or more networks. The technical advantage of such an arrangement is that, since the audio data 150 sent to the second-stage speech processor 154 represents only a subset of the audio data 116 generated by the microphone 104, communication resources associated with sending the audio data to the second-stage speech processor 154 are conserved. Additionally, by not sending all of the audio data 116 to the remote voice assistant server, power, processor time, and other computational resources associated with the operation of the second-stage speech processor 154 at the remote voice assistant server are conserved.

[0046] InFigure 1 In this case, the audio pre-processor 118 includes one or more voice input filters 120. At least one of the voice input filters 120 can be configured to operate as a speaker-specific voice input filter. In this context, a "speaker-specific voice input filter" refers to a filter configured to enhance the voice of one or more designated persons. For example, the speaker-specific voice input filter associated with person 180A can be operable to enhance the voice of utterance 108A from person 180A. By way of illustration, enhancing the voice of person 180A can include attenuating (or removing) portions (or components) of the audio data 116 that do not correspond to the voice from person 180A, such as the portion of the audio data 116 representing ambient sound 112, the portion of the audio data 116 representing the utterance 108B of person 180B, or both. In Figure 1 In the illustrated specific implementation, the voice input filter 120 is configured to receive the audio data 116 and output filtered audio data 122, where portions or components of the audio data 116 that do not correspond to the voice from person 180A are attenuated or removed.

[0047] In a particular specific implementation, the voice input filter 120 is configured to operate as a speaker-specific voice input filter based on detecting a wake word 110. For example, in response to detecting the wake word 110 in the utterance 108A from person 180A, the voice input filter 120 retrieves the voice signature data 134A associated with person 180A. In this example, the voice input filter 120 uses the voice signature data 134A to generate the filtered audio data 122 based on the audio data 116. As a simplified example, the voice input filter 120 compares the input audio data (e.g., the audio data 116) with the voice signature data 134A to generate output audio data (e.g., the filtered audio data 122), which attenuates (e.g., removes) portions or components of the input audio data that do not correspond to the voice from person 180A. In some specific implementations, the voice input filter 120 includes one or more trained models, as further described in reference Figures 3 to 5 and the voice signature data 134 includes one or more speaker embeddings, which are provided as inputs to the voice input filter 120 together with the audio data 116 to customize the voice input filter 120 to operate as a speaker-specific voice input filter.

[0048] In a particular specific implementation, the audio analyzer 140 includes a speaker detector 128, which is operable to determine the speaker identifier 130 of the person 180 whose voice is detected or who is detected to have spoken the wake word 110. For example, in Figure 1In this example, the audio preprocessor 118 is configured to provide the filtered audio data 122 to the first-stage speech processor 124. In this example, before the wake word 110 is detected (e.g., when no voice assistant session is in progress), the audio preprocessor 118 may perform non-speaker-specific filtering operations, such as noise suppression, echo cancellation, etc. In this example, the first-stage speech processor 124 includes a wake word detector 126 and a speaker detector 128. The wake word detector 126 is configured to detect one or more wake words, such as the wake word 110, in the utterance 108 from one or more persons 180.

[0049] In response to detecting the wake word 110, the wake word detector 126 causes the speaker detector 128 to determine an identifier of the person 180 associated with the utterance 108 in which the wake word 110 is detected (e.g., the speaker identifier 130). In a particular implementation, the speaker detector 128 may be operative to generate voice signature data based on the utterance 108 and compare the voice signature data with the voice signature data 134 in the memory 142. The voice signature data 134 in the memory 142 may be included within the registration data 136 associated with a set of registered users associated with the device 102. In this example, the speaker detector 128 provides the speaker identifier 130 to the audio preprocessor 118, and the audio preprocessor 118 retrieves the configuration data 132 based on the speaker identifier 130. The configuration data 132 may include, for example, the voice signature data 134 of the person 180 associated with the utterance 108 in which the wake word 110 is detected.

[0050] In some implementations, in addition to the voice signature data 134 of the person 180 associated with the utterance 108 in which the wake word 110 is detected, the configuration data 132 further includes other information. For example, the configuration data 132 may include the voice signature data 134 associated with multiple persons 180. In such implementations, the configuration data 132 enables the voice input filter 120 to generate the filtered audio data 122 based on the voices of two or more specific persons.

[0051] Therefore, in Figure 1 the illustrated example, after the wake word 110 is detected in the utterance 108 from a specific person 180, the voice input filter 120 is configured to operate a speaker-specific voice input filter associated at least with the specific person 180 who uttered the wake word 110. The portion of the audio data 116 after the wake word 110 is processed by the speaker-specific voice input filter such that the audio data 150 provided to the second-stage speech processor 154 includes the voice 152 of the specific person 180 and other portions of the audio data 116 are omitted or attenuated.

[0052] The second - level voice processor 154 includes one or more voice assistant applications 156 that are configured to perform voice assistant operations in response to commands detected within the voice 152. For example, the voice assistant operations may include accessing information from the memory 142 or from another memory (such as the memory of a remote server device). By way of illustration, the voice 152 may include an inquiry about local weather conditions, and in response to the inquiry, the voice assistant application 156 may determine the location of the device 102 and transmit a query to a weather database based on the location of the device 102. As another example, the voice assistant operations may include instructions for controlling other devices (e.g., smart home devices), outputting media content, or other similar instructions. At an appropriate time, the voice assistant application 156 may generate a voice assistant response 170, and the processor 190 may transmit an output audio signal 160 to the audio transducer 162 to output the voice assistant response 170. Although Figure 1 the example of Figure 1 illustrates the voice assistant response 170 provided via the audio transducer 162, in other embodiments, the voice assistant response 170 may be provided via a display device or another output device coupled to the output interface 158.

[0053] The technical beneficial effect of filtering the audio data 116 to remove or attenuate portions of the audio data 116 other than the voice 152 of the specific person 180 who speaks the wake - up word 110 is that such audio filtering operations prevent (or reduce the likelihood) of others interrupting the voice assistant session. For example, when person 180A speaks the wake - up word 110, the device 102 initiates a voice assistant session associated with person 180A and configures the voice input filter 120 to attenuate portions of the audio data 116 other than the voice of person 180A. In this example, another person 180B cannot interrupt the voice assistant session because the portion of the audio data 116 associated with the utterance 108B of person 180B is not provided to the first - level voice processor 124, not provided to the second - level voice processor 154, or both. Reducing interruptions improves the user experience associated with the voice assistant application 156. Additionally, when the utterance 108B of person 180B is not relevant to the voice assistant session associated with person 180A, reducing interruptions can save resources of the second - level voice processor 154. For example, if the audio data 150 provided to the second - level voice processor 154 includes the irrelevant speech of person 180B, the voice assistant application 156 uses computing resources to process the irrelevant speech. Moreover, the irrelevant speech may cause the voice assistant application 156 to misinterpret the voice of person 180A associated with the voice assistant session, resulting in person 180A having to repeat the voice and the voice assistant application 156 having to repeat operations to analyze the voice. Additionally, the irrelevant speech may reduce the accuracy of the speech recognition operations performed by the voice assistant application 156.

[0054] In some specific implementations, when the interrupting voice is relevant to the ongoing voice assistant session, the voice may be allowed. For example, as further described in Figure 5 When the audio data 116 includes an "interrupting voice" (e.g., a voice not associated with the person 180 who uttered the wake word 110 to initiate the voice assistant session), the interrupting voice is processed to determine a relevance score, and only the interrupting voice associated with a relevance score that meets the relevance criteria is provided to the voice assistant application 156.

[0055] As an example of the operation of the system 100, the microphone 104 detects sound 106 and provides the audio data 116 to the processor 190. Before the wake word 110 is detected, the audio pre-processor 118 performs non-speaker-specific audio pre-processing operations such as echo cancellation, noise reduction, etc. Additionally, in some specific implementations, the second-stage speech processor 154 remains in a low-power state before the wake word 110 is detected. In some such specific implementations, the first-stage speech processor 124 operates in an always-on mode, and the second-stage speech processor 154 operates in a standby mode or a low-power mode until it is activated by the first-stage speech processor 124. The audio pre-processor 118 provides the filtered audio data 122 to the first-stage speech processor 124, which executes the wake word detector 126 to process the filtered audio data 122 to detect the wake word 110.

[0056] When the wake word detector 126 detects the wake word 110 in the utterance 108A from the person 180A, the speaker detector 128 determines the speaker identifier 130 associated with the person 180A. In some specific implementations, the speaker detector 128 provides the speaker identifier 130 to the audio pre-processor 118, and the audio pre-processor 118 obtains the voice signature data 134A associated with the person 180A. In other specific implementations, the speaker detector 128 provides the voice signature data 134A as the speaker identifier 130 to the audio pre-processor 118. The voice signature data 134A and optional other configuration data 132 are provided to the voice input filter 120 to enable the voice input filter 120 to operate as a speaker-specific voice input filter 120 associated with the first person 180A.

[0057] Additionally, based on detecting wake word 110, wake word detector 126 activates second-stage speech processor 154 and causes audio data 150 to be provided to second-stage speech processor 154. Audio data 150 includes a portion of audio data 116 after being processed by speaker-specific speech input filter 120. For example, audio data 150 may include the entire utterance 108 containing wake word 110 based on the processing of audio data 116 by speaker-specific speech input filter 120. By way of illustration, audio analyzer 140 may store audio data 116 in a buffer and, in response to detecting wake word 110, cause the audio data 116 stored in the buffer to be processed by speaker-specific speech input filter 120. In this illustrative example, portions of audio data 116 received before the speech input filter 120 is configured to be speaker-specific may still be filtered by speaker-specific speech input filter 120 before being provided to second-stage speech processor 154.

[0058] In a particular embodiment, when speech input filter 120 is configured to operate as a speaker-specific speech input filter 120 associated with person 180A, speech from person 180B is not provided to wake word detector 126 and is not provided to voice assistant application 156. In such embodiments, person 180B cannot interact with device 102 in a manner that interrupts the voice assistant session between person 180A and voice assistant application 156. In such embodiments, when wake word detector 126 detects wake word 110 in utterance 108A from person 180A, a voice assistant session is initiated between person 180A and voice assistant application 156, and the voice assistant session continues until a termination condition is met. For example, the termination condition may be met when a specific duration of the voice assistant session has elapsed, when a voice assistant operation that does not require a response or further interaction with person 180A is performed, or when person 180A commands the termination of the voice assistant session.

[0059] In some embodiments, during a voice assistant session associated with person 180A, speech from person 180B may be analyzed to determine whether the speech is relevant to speech 152 provided by person 180A to voice assistant application 156. In such embodiments, relevant speech from person 180B may be provided to voice assistant application 156 during the voice assistant session.

[0060] In some specific implementations, the configuration data 132 provided to the audio pre-processor 118 to configure the voice input filter is based on the voice signature data 134 associated with multiple persons. In such specific implementations, the configuration data 132 enables the voice input filter 120 to operate as a speaker-specific voice input filter 120 associated with multiple persons. For illustration, when the configuration data 132 is based on the voice signature data 134A associated with person 180A and the voice signature data 134B associated with person 180B, the voice input filter 120 can be configured to operate as a speaker-specific voice input filter 120 associated with person 180A and person 180B. Examples of specific implementations that can use voice signature data 134 based on the voices of multiple persons include the case where person 180A is a child and person 180B is a parent. In this case, based on the configuration data 132, the parent can have the permission to interrupt any voice assistant session initiated by the child.

[0061] In a particular specific implementation, the voice signature data 134 associated with a particular person 180 includes a speaker embedding. For example, during a registration operation, the microphone 104 can capture the voice of person 180, and the speaker detector 128 (or another component of device 102) can generate a speaker embedding. The speaker embedding can be stored at the memory 142 as registration data 136 together with other data, such as the speaker identifier of the particular person 180. In Figure 1 the illustrated example, the registration data 136 includes three sets of voice signature data 134, including voice signature data 134A, voice signature data 134B, and voice signature data 134N. However, in other specific implementations, the registration data 136 includes more than three sets or less than three sets of voice signature data 134. The registration data 136 optionally further includes information specifying the sets of voice signature data 134 to be used together, such as in the above example where the voice signature data 134 of the parent is provided to the audio pre-processor 118 together with the voice signature data 134 of the child.

[0062] Figures 2A to 2C Aspects of operations associated with selectively filtering audio data for voice processing in accordance with some examples of the present disclosure are illustrated. Referring to Figure 2A , a first example 200 is illustrated. In the first example 200, the configuration data 132 for configuring the voice input filter 120 to operate as a speaker-specific voice input filter 210 includes first voice signature data 206. The first voice signature data 206 includes, for example, a speaker embedding associated with a first person (such as Figure 1 person 180A).

[0063] In the first example 200, the audio data 116 provided as input to the speaker-specific voice input filter 210 includes ambient sound 112 and speech 204. The speaker-specific voice input filter 210 is operable to generate audio data 150 as output based on the audio data 116. In the first example 200, the audio data 150 includes speech 204 and does not include or attenuates the ambient sound 112. For example, the speaker-specific voice input filter 210 is configured to compare the audio data 116 with the first voice signature data 206 to generate the audio data 150. The audio data 150 attenuates the portion of the audio data 116 that does not correspond to the speech 204 from the person associated with the first voice signature data 206.

[0064] In Figure 2A the illustrated first example 200, as part of a voice assistant session, the audio data 150 representing the speech 204 is provided to the voice assistant application 156. Additionally, a portion of the audio data 116 representing the ambient sound 112 is attenuated or omitted in the audio data 150 provided to the voice assistant application 156. The technical beneficial effect of filtering the audio data 116 to attenuate or omit the ambient sound 112 from the audio data 150 is that such filtering enables the voice assistant application 156 to more accurately recognize the speech in the audio data 150, which reduces the error rate of the voice assistant application 156 and improves the user experience.

[0065] Reference Figure 2B , a second example 220 is illustrated. In the second example 220, the configuration data 132 for configuring the voice input filter 120 includes Figure 2A the first voice signature data 206. For example, the first voice signature data 206 includes a speaker embedding associated with a first person (such as Figure 1 person 180A).

[0066] In the second example 220, the audio data 116 provided as input to the speaker-specific voice input filter 210 includes multi-person speech 222, such as Figure 1 the speech of person 180A and the speech of person 180B. The speaker-specific voice input filter 210 is operable to generate audio data 150 as output based on the audio data 116. In the second example 220, the audio data 150 includes single-person speech 224, such as the speech of person 180A. In this example, the speech of one or more other people, such as the speech of person 180B, is omitted or attenuated in the audio data 150. For example, the audio data 150 attenuates the portion of the audio data 116 that does not correspond to the speech from the person associated with the first voice signature data 206.

[0067] In Figure 2BIn the second example 220 illustrated, as part of a voice assistant conversation, audio data 150 representing single-person speech 224 (e.g., the speech of the person who initiated the voice assistant conversation) is provided to a voice assistant application 156. Additionally, a portion of the audio data 116 representing the speech of other people (e.g., the speech of a person who did not initiate the voice assistant conversation) is attenuated or omitted from the audio data 150 provided to the voice assistant application 156. A technical advantage of filtering the audio data 116 to attenuate or omit the speech of people who did not initiate a particular voice assistant conversation is that such filtering limits the ability of such other people to chime in on the voice assistant conversation.

[0068] Although Figure 2B the ambient sound 112 in the audio data 116 provided to the speaker-specific voice input filter 210 is not specifically illustrated, in some specific implementations, the audio data 116 in the second example 220 also includes the ambient sound 112. In such specific implementations, the speaker-specific voice input filter 210 performs both speaker separation (e.g., to distinguish single-person speech 224 from multi-person speech 222) and noise reduction (e.g., to remove or attenuate the ambient sound 112).

[0069] Referring Figure 2C , a third example 240 is illustrated. In the third example 240, the configuration data 132 for configuring the voice input filter 120 includes first voice signature data 206 and second voice signature data 242. For example, the first voice signature data 206 includes a speaker embedding associated with a first person (such as Figure 1 person 180A), and the second voice signature data 242 includes a speaker embedding associated with a second person (such as Figure 1 person 180B).

[0070] In the third example 240, the audio data 116 provided as input to the speaker-specific voice input filter 210 includes the ambient sound 112 and speech 244. The speech 244 may include the speech of the first person, the speech of the second person, the speech of one or more other people, or any combination thereof. The speaker-specific voice input filter 210 is operable to generate audio data 150 as output based on the audio data 116. In the third example 240, the audio data 150 includes speech 246. The speech 246 includes the speech of the first person (if present in the audio data 116), the speech of the second person (if present in the audio data 116), or both. Additionally, in the audio data 150, the ambient sound 112 and the speech of other people are attenuated (e.g., attenuated or removed). That is, in the audio data 150, the portion of the audio data 116 that does not correspond to the speech of the first person associated with the first voice signature data 206 or the speech of the second person associated with the second voice signature data 242 is attenuated.

[0071] In Figure 2C In the third example 240 illustrated, as part of a voice assistant session, audio data 150 representing speech 246 is provided to a voice assistant application 156. Additionally, in the audio data 150 provided to the voice assistant application 156, the representation of ambient sound 112 or a portion of the speech of other people in the audio data 116 is attenuated or omitted. The technical beneficial effect of filtering the audio data 116 to attenuate or omit the speech of some people (e.g., people not associated with the first voice signature data 206 or the second voice signature data 242) while still allowing multi-person speech (e.g., speech from people associated with the first voice signature data 206 or the second voice signature data 242) to pass to the voice assistant application 156 is that such filtering limits the ability of specific users to interrupt. For example, multiple members of a household may be allowed to interrupt each other's voice assistant sessions while preventing others from interrupting a voice assistant session initiated by a member of the household.

[0072] Figure 3 A specific example of the voice input filter 120 is illustrated. In Figure 3 the example illustrated, the voice input filter 120 includes or corresponds to one or more voice enhancement models 340. The voice enhancement models 340 include one or more machine learning models that are configured and trained to perform voice enhancement operations such as denoising, speaker separation, and the like. In Figure 3 the example illustrated, the voice enhancement model 340 includes a dimensionality reduction network 310, a combiner 316, and a dimensionality expansion network 318. The dimensionality reduction network 310 includes multiple layers (e.g., neural network layers) that are arranged to perform convolution, pooling, concatenation, etc. to generate a latent space representation 312 based on the audio data 116. In one example, the audio data 116 is input to the dimensionality reduction network 310 as a series of input feature vectors, where each input feature vector in the series represents one or more audio data samples (e.g., frames or another portion) of the audio data 116, and the dimensionality reduction network 310 generates a latent space representation 312 associated with each input feature vector. The input feature vectors may include, for example, values representing spectral features (e.g., complex spectrum, magnitude spectrum, Mel spectrum, Bark spectrum, etc.) of a time-windowed portion of the audio data 116, cepstral features (e.g., Mel frequency cepstral coefficients, Bark frequency cepstral coefficients, etc.) of the time-windowed portion of the audio data 116, or other data representing the time-windowed portion of the audio data 116.

[0073] The combiner 316 is configured to combine the speaker embedding 314 and the latent space representation 312 to generate a combined vector 317 as the input to the dimension expansion network 318. In one example, the combiner 316 includes a concatenator that is configured to concatenate the speaker embedding 314 to the latent space representation 312 of each input feature vector to generate the combined vector 317.

[0074] The dimension expansion network 318 includes one or more recurrent layers (e.g., one or more gated recurrent unit (GRU) layers) and a plurality of additional layers (e.g., neural network layers) that are arranged to perform convolution, pooling, concatenation, etc. to generate the audio data 150 based on the combined vector 317.

[0075] Optionally, the speech enhancement model 340 may further include one or more skip connections 319. Each skip connection 319 connects the output of one layer in the dimension reduction network 310 to the input of a corresponding layer in the dimension expansion network 318.

[0076] During operation, the audio data 116 (or the feature vector representing the audio data 116) is provided as an input to the speech enhancement model 340. The audio data 116 may include speech 302, ambient sound 112, or both. The speech 302 may include the speech of a single person or multiple people.

[0077] The dimension reduction network 310 processes each feature vector of the audio data 116 through a series of convolution operations, pooling operations, activation layers, recurrent layers, other data manipulation operations, or any combination thereof, based on the architecture and training of the dimension reduction network 310, to generate the latent space representation 312 of the feature vectors of the audio data 116. In Figure 3 the illustrated example, the generation of the latent space representation 312 of the feature vectors is performed independently of the voice signature data 134. Thus, the same operations are performed regardless of who initiates the voice assistant session.

[0078] The speaker embedding 314 is speaker-specific and is selected based on the specific person (or persons) whose voice is to be enhanced. Each latent space representation 312 is combined with the speaker embedding 314 to generate a corresponding combined vector 317, and the combined vector 317 is provided as an input to the dimension expansion network 318. As described above, the dimension expansion network 318 includes at least one recurrent layer, such as a GRU layer, such that each output vector of the audio data 150 depends on a series (e.g., more than one) of combined vectors 317. In some embodiments, the dimension expansion network 318 is configured (and trained) to generate a person-specific enhanced voice 320 as the audio data 150. In such embodiments, the specific person whose voice is enhanced is the person whose voice is represented by the speaker embedding 314. In some embodiments, the dimension expansion network 318 is configured (and trained) to generate enhanced voices 320 for more than one specific person as the audio data 150. In such embodiments, the specific persons whose voices are enhanced are the persons associated with the speaker embedding 314.

[0079] The dimension expansion network 318 can be considered a generative network that is configured and trained to recreate that portion of the input audio data stream (e.g., audio data 116) that is similar to the voice of a specific person (e.g., the person associated with the speaker embedding 314). Thus, the voice enhancement model 340 can use a set of machine learning operations to perform both noise reduction and speaker separation to generate the enhanced voice 320.

[0080] Figure 4 Another specific example of the voice input filter 120 is illustrated. In Figure 4 the illustrated example, the voice input filter 120 includes or corresponds to one or more voice enhancement models 340. As Figure 3 shown, the voice enhancement model 340 includes one or more machine learning models that are configured (and trained) to perform voice enhancement operations, such as denoising, speaker separation, etc. In Figure 4 the illustrated example, the voice enhancement model 340 includes a dimension reduction network 310 coupled to a switch 402. The switch 402 can include, for example, a logic switch that is configured to select which of a plurality of subsequent processing paths to execute. The dimension reduction network 310 operates as described with reference to Figure 3 to generate a latent space representation 312 associated with each input feature vector of the audio data 116.

[0081] In Figure 4In the illustrated example, switch 402 is coupled to a first processing path that includes combiner 404 and dimensionality expansion network 408, and switch 402 is also coupled to a second processing path that includes combiner 412 and multi-person dimensionality expansion network 418. In this example, the first processing path is configured (and trained) to perform operations associated with enhancing the speech of a single person, and the second processing path is configured (and trained) to perform operations associated with enhancing the speech of multiple people. Thus, switch 402 is configured to select the first processing path when Figure 1 the configuration data 132 of Figure 1 includes a single speaker embedding 406 or otherwise indicates that the speech of a single identified speaker is to be enhanced to generate enhanced speech 410 of a single person. In contrast, switch 402 is configured to select the second processing path when Figure 1 the configuration data 132 of Figure 1 includes multiple speaker embeddings (such as first speaker embedding 414 and second speaker embedding 416) or otherwise indicates that the speech of multiple identified speakers is to be enhanced to generate enhanced speech 420 of multiple people.

[0082] Combiner 404 is configured to combine speaker embedding 406 and latent space representation 312 to generate a combined vector as the input to dimensionality expansion network 408. Dimensionality expansion network 408 is configured to process the combined vector, as described with reference to Figure 3 to generate enhanced speech 410 of a single person.

[0083] Combiner 412 is configured to combine two or more speaker embeddings (e.g., first speaker embedding 414 and second speaker embedding 416) and latent space representation 312 to generate a combined vector as the input to multi-person dimensionality expansion network 418. Multi-person dimensionality expansion network 418 is configured to process the combined vector, as described with reference to Figure 3 to generate enhanced speech 420 of multiple people. Although the first processing path and the second processing path perform similar operations, different processing paths are used in the Figure 4 illustrated example because the combined vectors generated by combiners 404, 412 have different dimensions. Thus, dimensionality expansion network 408 and multi-person dimensionality expansion network 418 have different architectures to accommodate combined vectors of different dimensions.

[0084] Alternatively, in some embodiments, different processing paths are used in Figure 4 to account for the different operations performed by combiners 404, 412. For example, combiner 412 may be configured to combine speaker embeddings 414, 416 in an element-by-element manner such that the combined vectors generated by combiners 404, 412 have the same dimension. By way of illustration, combiner 412 may add or average the value of each element of first speaker embedding 414 with the corresponding element value of second speaker embedding 416.

[0085] Figure 5 Illustrates another specific example of the voice input filter 120. In Figure 5 the illustrated example, the voice input filter 120 includes or corresponds to one or more voice enhancement models 340. As Figure 3 and Figure 4 shown, the voice enhancement model 340 includes one or more machine learning models that are configured (and trained) to perform voice enhancement operations such as denoising, speaker separation, etc. In Figure 5 the illustrated example, the voice enhancement model 340 includes a dimensionality reduction network 310 that operates as described with reference to Figure 3 to generate a latent space representation 312 associated with each input feature vector of the audio data 116.

[0086] In Figure 5 the illustrated example, the dimensionality reduction network 310 is coupled to a first processing path including a combiner 502 and a dimensionality expansion network 506, and is coupled to a second processing path including a combiner 510 and a dimensionality expansion network 514. In this example, the first processing path is configured (and trained) to perform operations associated with enhancing the voice of a first person (e.g., the person who initiated a particular voice assistant session), and the second processing path is configured (and trained) to perform operations associated with enhancing the voice of one or more second persons (e.g., persons who are approved to interrupt the voice assistant session in certain cases based on Figure 1 the configuration data 132).

[0087] The combiner 502 is configured to combine the speaker embedding 504 (e.g., the speaker embedding associated with the person who spoke the wake word 110 to initiate the voice assistant session) and the latent space representation 312 to generate a combined vector as the input to the dimensionality expansion network 506. The dimensionality expansion network 506 is configured to process the combined vector as described with reference to Figure 3 to generate the enhanced voice 508 of the first person. Since the first person is the one who initiated the voice assistant session, the enhanced voice 508 of the first person is provided to the voice assistant application 156 for processing.

[0088] The combiner 510 is configured to combine the speaker embedding 512 (e.g., the speaker embedding associated with a second person who did not speak the wake word 110 to initiate the voice assistant session) and the latent space representation 312 to generate a combined vector as the input to the dimensionality expansion network 514. The dimensionality expansion network 514 is configured to process the combined vector as described with reference to Figure 3 (or in the case where the speaker embedding 512 corresponds to multiple persons (collectively referred to as the "second person") Figure 4) As described, to generate the enhanced speech 516 of the second person. Note that at any given time, the latent space representation 312 may include the speech of the first person, the speech of the second person, neither, or both. Thus, in some embodiments, each latent space representation 312 may be processed via both the first processing path and the second processing path.

[0089] The second person has conditional access rights to the voice assistant conversation. Therefore, the enhanced speech 516 of the second person is further analyzed to determine whether the conditions for providing the speech 516 of the second person to the voice assistant application 156 are met. In Figure 5 the illustrated example, the enhanced speech 516 of the second person is provided to the natural language processing (NLP) engine 520. Additionally, the context data 522 associated with the enhanced speech 508 of the first person is provided to the NLP engine 520. The context data 522 may include, for example, the enhanced speech 508 of the first person, data summarizing the enhanced speech 508 of the first person (e.g., keywords from the enhanced speech 508 of the first person), results generated by the voice assistant application 156 in response to the enhanced speech 508 of the first person, other data indicating the content of the enhanced speech 508 of the first person, or any combination thereof.

[0090] The NLP engine 520 is configured to determine whether the speech of the second person (as represented in the enhanced speech 516 of the second person) is contextually relevant to a voice assistant request, command, query, or other content of the speech of the first person as indicated by the context data 522. As an example, the NLP engine 520 may perform a context-aware semantic embedding of the context data 522, the enhanced speech 516 of the second person, or both, to determine the value of a relevance metric associated with the enhanced speech 516 of the second person. In this example, the context-aware semantic embedding may be used to map the enhanced speech 516 of the second person to a feature space, in which semantic similarity may be estimated based on the distance between two points (e.g., cosine distance, Euclidean distance, etc.), and the relevance metric may correspond to the value of the distance metric. If the relevance metric meets a threshold, the content of the enhanced speech 516 of the second person may be considered relevant to the virtual assistant conversation.

[0091] If the content of the enhanced speech 516 of the second person is considered relevant to the virtual assistant conversation, the NLP engine 520 provides the relevant speech 524 of the second person to the voice assistant application 156. Otherwise, if the content of the enhanced speech 516 of the second person is considered not relevant to the virtual assistant conversation, the enhanced speech 516 of the second person is discarded or ignored.

[0092] Figure 6FIG. is a diagram of a first example of a vehicle 650 operable to selectively filter audio data for voice processing according to some examples of the present disclosure. In Figure 6 system 100 or portions thereof are integrated within vehicle 650, and in Figure 6 example, the vehicle is illustrated as an automobile including a plurality of seats 652A - 652E. Although vehicle 650 is illustrated as an automobile in Figure 6 in other specific implementations, vehicle 650 is a bus, train, airplane, ship, or another type of vehicle configured to transport one or more passengers (which may optionally include a vehicle operator).

[0093] Vehicle 650 includes an audio analyzer 140 and one or more audio sources 602. Audio analyzer 140 and audio source 602 are coupled to microphone 104, audio transducer 162, or both via codec 604. Figure 6 Vehicle 650 of

[0094] also includes one or more vehicle systems 660, some or all of which may be coupled to audio analyzer 140 to enable voice assistant application 156 to control various operations of vehicle systems 660. Figure 6 In Figure 6 vehicle 650 includes a plurality of microphones 104A - 104F. For example, in Figure 6 each microphone 104 is positioned near a respective one of seats 652A - 652E. In Figure 6 example, the positioning of microphones 104 relative to seats 652 enables audio analyzer 140 to distinguish audio zones 654 of vehicle 650. In

[0095] There is a one - to - one relationship between audio zones 654 and seats 652. In some other specific implementations, one or more of audio zones 654 include more than one seat 652. For illustration, seats 652C - 652E may be associated with a single "rear seat" audio zone. Figure 6 Although vehicle 650 of

[0096] is illustrated as including a plurality of microphones 104A - 104F arranged to detect sounds within vehicle 650 and optionally enable audio analyzer 140 to distinguish which audio zone 654 includes the source of the sound, in other specific implementations, vehicle 650 includes only a single microphone 104. In other specific implementations, vehicle 650 includes a plurality of microphones 104 and audio analyzer 140 does not distinguish audio zones 654. Figure 6In [the figure], the audio analyzer 140 includes an audio pre-processor 118, a first-stage speech processor 124, and a second-stage speech processor 154, each of which operates as described in any of the figures as Figures 1 to 5 shown. In the Figure 6 specific example illustrated, the audio pre-processor 118 includes a voice input filter 120 that is configurable to operate as a speaker-specific voice input filter to selectively filter audio data for speech processing.

[0097] Figure 6 The audio pre-processor 118 in [the figure] also includes an echo cancellation and noise suppression (ECNS) unit 606 and an adaptive interference canceller (AIC) 608. The ECNS unit 606 and the AIC 608 are operable to filter audio data from the microphone 104 independently of the voice input filter 120. For example, the ECNS unit 606, the AIC 608, or both may perform non-speaker-specific audio filtering operations. By way of illustration, the ECNS unit 606 is operable to perform an echo cancellation operation, a noise suppression operation (e.g., adaptive noise filtering), or both. The AIC 608 is configured to distinguish audio regions 654 and, optionally, limit the audio data provided to the first-stage speech processor 124, the second-stage speech processor 154, or both to audio from a particular one or more of the audio regions within the audio regions 654. By way of illustration, based on the configuration of the audio analyzer 140, the AIC 608 may only allow audio from a person in one of the front seats 652A, 652B to be provided to the wake word detector 126, the voice assistant application 156, or both.

[0098] During operation, one or more of the microphones 104 may detect sounds within the vehicle 650 and provide audio data representative of the sounds to the audio analyzer 140. When no voice assistant session is in progress, the ECNS unit 606, the AIC 608, or both process the audio data to generate filtered audio data (e.g., filtered audio data 122) and provide the filtered audio data to the wake word detector 126. If the wake word detector 126 detects a wake word in the filtered audio data (e.g., Figure 1If the wake word 110 is detected, the wake word detector 126 signals the speaker detector 128 to identify the person who spoke the wake word. Additionally, the wake word detector 126 activates the second-stage speech processor 154 to initiate a voice assistant session. The speaker detector 128 provides an identifier of the person who spoke the wake word (e.g., speaker identifier 130) to the audio pre-processor 118, and the audio pre-processor 118 obtains configuration data (e.g., configuration data 132) to activate the voice input filter 120 as a speaker-specific voice input filter. In some embodiments, the wake word detector 126 may also provide information to the AIC 608 to indicate which audio region 654 the wake word originated from, and the AIC 608 may filter the audio data provided to the voice input filter 120 based on the audio region 654 from which the wake word originated.

[0099] The speaker-specific voice input filter is used to filter the audio data and provide the filtered audio data to the voice assistant application 156, as described in any of the figures referenced Figures 1 to 5 above. Based on the speech content represented in the filtered audio data, the voice assistant application 156 may control the operation of the audio source 602, control the operation of the vehicle system 660, or perform other operations, such as retrieving information from a remote data source.

[0100] A response from the voice assistant application 156 (e.g., voice assistant response 170) may be played to the occupants of the vehicle 650 via the audio transducer 162. In Figure 6 the example illustrated above, the audio transducer 162 is disposed near or within a specific audio region within the audio region 654, which enables the voice assistant application 156 to provide a response to a specific occupant of the vehicle 650 (e.g., the occupant who initiated the voice assistant session) or multiple occupants.

[0101] The selective operation of the voice input filter 120 as a speaker-specific voice input filter enables the voice assistant application 156 to perform more accurate speech recognition because noise and irrelevant speech are removed from the audio data provided to the voice assistant application 156. Additionally, the selective operation of the voice input filter 120 as a speaker-specific voice input filter limits the ability of other occupants in the vehicle 650 to interrupt the voice assistant session. For example, if the driver of the vehicle 650 initiates a voice assistant session to request driving directions, the voice assistant session may be associated only with the driver (or, as described above, with one or more other people), such that other occupants of the vehicle 650 cannot interrupt the voice assistant session.

[0102] Figure 7is a specific implementation in which system 100 is integrated within a wireless speaker and voice-activated device 700. The wireless speaker and voice-activated device 700 may have wireless network connectivity and is configured to perform voice assistant operations. In Figure 7 the audio analyzer 140, audio source 602, and codec 604 are included within the wireless speaker and voice-activated device 700. The wireless speaker and voice-activated device 700 also includes an audio transducer 162 and a microphone 104.

[0103] During operation, one or more microphones in the microphone 104 may detect sounds near the wireless speaker and voice-activated device 700 (such as in the room in which the wireless speaker and voice-activated device 700 is disposed). The microphone 104 provides audio data representative of the sounds to the audio analyzer 140. When no voice assistant session is in progress, the ECNS unit 606, AIC 608, or both process the audio data to generate filtered audio data (e.g., filtered audio data 122), and provide the filtered audio data to the wake word detector 126. If the wake word detector 126 detects a wake word (e.g., Figure 1 wake word 110) in the filtered audio data, the wake word detector 126 signals the speaker detector 128 to identify the person who spoke the wake word. Additionally, the wake word detector 126 activates the second-stage speech processor 154 to initiate a voice assistant session. The speaker detector 128 provides an identifier of the person who spoke the wake word (e.g., speaker identifier 130) to the audio pre-processor 118, and the audio pre-processor 118 obtains configuration data (e.g., configuration data 132) to activate the voice input filter 120 as a speaker-specific voice input filter. In some specific implementations, the wake word detector 126 may also provide information to the AIC 608 to indicate the direction from which the wake word originated, and the AIC 608 may perform beamforming or other directional audio processing to filter the audio data provided to the voice input filter 120 based on the direction from which the wake word originated.

[0104] The speaker-specific voice input filter is used to filter the audio data and provide the filtered audio data to the voice assistant application 156, as described in any of the figures referenced in Figures 1 to 5 Based on the speech content represented in the filtered audio data, the voice assistant application 156 performs one or more voice assistant operations, such as transmitting commands to smart home devices, playing media, or performing other operations, such as retrieving information from remote data sources. Responses from the voice assistant application 156 (e.g., voice assistant response 170) may be played via the audio transducer 162.

[0105] The selective operation of the voice input filter 120 as a speaker-specific voice input filter enables the voice assistant application 156 to perform more accurate speech recognition because noise and irrelevant speech are removed from the audio data provided to the voice assistant application 156. Additionally, the selective operation of the voice input filter 120 as a speaker-specific voice input filter limits the ability of others in a room equipped with the wireless speaker and voice activation device 700 to chime into the voice assistant conversation.

[0106] Figure 8 Illustrated is a particular implementation 800 of the device 102 as an integrated circuit 802 including one or more processors 190, the one or more processors including one or more components of the audio analyzer 140. The integrated circuit 802 also includes an input circuit 804 (such as one or more bus interfaces) to enable receipt of audio data 116 for processing. The integrated circuit 802 also includes an output circuit 806 (such as a bus interface) to enable transmission of output data 808 from the integrated circuit 802. For example, the output data 808 may include Figure 1 a voice assistant response 170. As another example, the output data 808 may include commands or queries (such as an information retrieval query transmitted to a remote device) to other devices (such as a media player, a transportation system, a smart home device, etc.). In some particular implementations, Figure 1 the voice assistant application 156 is located remotely from Figure 8 the audio analyzer 140, in which case the output data 808 may include Figure 1 audio data 150.

[0107] The integrated circuit 802 is capable of selectively filtering audio data for speech processing as a component in a system including a microphone, such as a mobile phone or tablet as depicted in Figure 9 , a wearable electronic device as depicted in Figure 10 , a camera as depicted in Figure 11 , an extended reality (e.g., virtual reality, mixed reality, or augmented reality) headset as depicted in Figure 12 or a transportation vehicle as depicted in Figure 6 or Figure 13 .

[0108] As an illustrative, non-limiting example, Figure 9 illustrated is a particular implementation 900 in which the device 102 includes a mobile device 902 (such as a phone or tablet). In a particular implementation, the integrated circuit 802 is integrated within the mobile device 902. In Figure 9In [the figure], mobile device 902 includes microphone 104, audio transducer 162, and display screen 904. Components of processor 190, including audio analyzer 140, are integrated in mobile device 902 and are illustrated using dashed lines to indicate internal components that are generally not visible to the user of mobile device 902.

[0109] In a specific example, Figure 9 audio analyzer 140 of [the figure] operates as described in any of the figures referenced in Figures 1 to 8 to selectively enable speaker-specific voice input filtering in a manner that improves the accuracy of speech recognition of voice assistant application 156 and limits the ability of others to interrupt a voice assistant session. During a voice assistant session, responses from the voice assistant application can be provided to the user as output via audio transducer 162, via display screen 904, or both.

[0110] Figure 10 Embodiment 1000 is depicted in which device 102 includes wearable electronic device 1002 (illustrated as a "smartwatch"). In a specific embodiment, integrated circuit 802 is integrated within wearable electronic device 1002. In Figure 10 [the figure], wearable electronic device 1002 includes microphone 104, audio transducer 162, and display screen 1004.

[0111] Components of processor 190, including audio analyzer 140, are integrated in wearable electronic device 1002. In a specific example, Figure 10 audio analyzer 140 of [the figure] operates as described in any of the figures referenced in Figures 1 to 8 to selectively enable speaker-specific voice input filtering in a manner that improves the accuracy of speech recognition of voice assistant application 156 and limits the ability of others to interrupt a voice assistant session. During a voice assistant session, responses from the voice assistant application can be provided to the user as output via audio transducer 162, via tactile feedback to the user, via display screen 1004, or any combination thereof.

[0112] As an example of the operation of wearable electronic device 1002, during a voice assistant session, a person initiating the voice assistant session can provide a voice request to display a message (e.g., a text message, an email, etc.) transmitted to that person via display screen 1004 of wearable electronic device 1002. In this example, others near wearable electronic device 1002 can speak a wake word associated with audio analyzer 140 without interrupting the voice assistant session because audio data is filtered during the voice assistant session to attenuate a portion of the audio data that does not correspond to the voice of the person initiating the voice assistant session.

[0113] Figure 11Depicts a specific implementation 1100 in which device 102 includes a portable electronic device corresponding to camera device 1102. In a particular implementation, integrated circuit 802 is integrated within camera device 1102. In Figure 11 it, camera device 1102 includes microphone 104 and audio transducer 162. Camera device 1102 may also include Figure 11 a display screen on a side not illustrated in

[0114] Components of processor 190, including audio analyzer 140, are integrated within camera device 1102. In a particular example, Figure 11 the audio analyzer 140 of Figures 1 to 8 operates as described in any of the figures referenced in to selectively enable speaker-specific voice input filtering in a manner that improves the accuracy of speech recognition of voice assistant application 156 and limits the ability of others to interrupt the voice assistant session. During a voice assistant session, responses from the voice assistant application may be provided to the user as output via audio transducer 162, via the display screen, or both.

[0115] As an example of the operation of camera device 1102, during a voice assistant session, the person initiating the voice assistant session may provide a voice request for camera device 1102 to capture an image. In this example, others near camera device 1102 may speak a wake word associated with audio analyzer 140 without interrupting the voice assistant session because audio data is filtered during the voice assistant session to attenuate a portion of the audio data that does not correspond to the voice of the person initiating the voice assistant session.

[0116] Figure 12 Depicts a specific implementation 1200 in which device 102 includes a portable electronic device corresponding to an extended reality (e.g., virtual reality, mixed reality, or augmented reality) head-mounted device 1202. In a particular implementation, integrated circuit 802 is integrated within head-mounted device 1202. In Figure 12 it, head-mounted device 1202 includes microphone 104 and audio transducer 162. Additionally, a visual interface device is positioned in front of the user's eyes to enable the display of augmented reality, mixed reality, or virtual reality images or scenes to the user when wearing head-mounted device 1202.

[0117] Components of processor 190, including audio analyzer 140, are integrated within head-mounted device 1202. In a particular example, Figure 12 the audio analyzer 140 of Figures 1 to 8operate as depicted in any of the figures to selectively enable speaker - specific voice input filtering in a manner that improves the accuracy of speech recognition of the voice assistant application 156 and limits the ability of others to interrupt the voice assistant session. During a voice assistant session, responses from the voice assistant application can be provided as output to the user via the audio transducer 162, via the visual interface device, or both.

[0118] As an example of the operation of the head - mounted device 1202, during a voice assistant session, the person initiating the voice assistant session can provide a voice request to display a specific media on the visual interface device of the head - mounted device 1202. In this example, others near the head - mounted device 1202 can say the wake - up word associated with the audio analyzer 140 without interrupting the voice assistant session because the audio data is filtered during the voice assistant session to attenuate a portion of the audio data that does not correspond to the voice of the person initiating the voice assistant session.

[0119] Figure 13 Illustrates a particular implementation 1300 in which the device 102 corresponds to or is integrated within a vehicle 1302 (exemplified as a manned or unmanned aerial device (e.g., a package - delivery drone)). In a particular implementation, the integrated circuit 802 is integrated within the vehicle 1302. In Figure 13 this, the vehicle 1302 also includes a microphone 104 and an audio transducer 162.

[0120] Components of the processor 190, including the audio analyzer 140, are integrated within the vehicle 1302. In a particular example, Figure 13 the audio analyzer 140 of Figures 1 to 8 operates as depicted in any of the figures to selectively enable speaker - specific voice input filtering in a manner that improves the accuracy of speech recognition of the voice assistant application 156 and limits the ability of others to interrupt the voice assistant session. During a voice assistant session, responses from the voice assistant application can be provided as output to the user via the audio transducer 162.

[0121] As an example of the operation of the vehicle 1302, during a voice assistant session, the person initiating the voice assistant session can provide a voice request for the vehicle 1302 to deliver a package to a specified location. In this example, others near the vehicle 1302 can say the wake - up word associated with the audio analyzer 140 without interrupting the voice assistant session because the audio data is filtered during the voice assistant session to attenuate a portion of the audio data that does not correspond to the voice of the person initiating the voice assistant session. Thus, others cannot redirect the vehicle 1302 to a different delivery location.

[0122] Figure 14 is a block diagram illustrating exemplary aspects of a system 1400 operable to selectively filter audio data for speech processing in accordance with some examples of the present disclosure. In Figure 14 , the processor 190 includes a always-on power domain 1403 and a second power domain 1405, such as a power-on-demand domain. The operation of the system 1400 is partitioned such that some operations are performed in the always-on power domain 1403 and other operations are performed in the second power domain 1405. For example, in Figure 14 , the audio pre-processor 118, the first-stage speech processor 124, and the buffer 1460 are included in the always-on power domain 1403 and are configured to operate in an always-on mode. Additionally, in Figure 14 , the second-stage speech processor 154 is included in the second power domain 1405 and is configured to operate in a power-on-demand mode. The second power domain 1405 also includes an activation circuit 1430.

[0123] The audio data 116 received from the microphone 104 is stored in the buffer 1460. In a particular implementation, the buffer 1460 is a circular buffer that stores the audio data 116 such that the most recent audio data 116 is accessible for processing by other components, such as the audio pre-processor 118, the first-stage speech processor 124, the second-stage speech processor 154, or a combination thereof.

[0124] One or more components of the always-on power domain 1403 are configured to generate at least one of a wake signal 1422 or an interrupt 1424 to initiate one or more operations at the second power domain 1405. In one example, the wake signal 1422 is configured to transition the second power domain 1405 from a low-power mode 1432 to an active mode 1434 to activate one or more components of the second power domain 1405. As an example, when a wake word is detected in the audio data 116, the wake word detector 126 may generate the wake signal 1422 or the interrupt 1424.

[0125] In various implementations, the activation circuit 1430 includes or is coupled to a power management circuit, a clock circuit, a head switch or a foot switch circuit, a buffer control circuit, or any combination thereof. The activation circuit 1430 may be configured to initiate the power-on of the second power domain 1405, such as by selectively applying or raising the voltage of the power source of the second power domain 1405. As another example, the activation circuit 1430 may be configured to selectively gate or ungate the clock signal to the second power domain 1405, such as to prevent or enable circuit operation without removing the power source.

[0126] The output 1452 generated by the second - stage speech processor 154 can be provided to the application 1454. The application 1454 can be configured to perform operations as directed by the voice assistant application 156. By way of illustration, as an illustrative, non - limiting example, the application 1454 can correspond to a vehicle navigation and entertainment application or a home automation system.

[0127] In a particular embodiment, when a voice assistant session is active, the second power domain 1405 can be activated. As an example of the operation of the system 1400, the audio pre - processor 118 operates in the always - on power domain 1403 to filter the audio data 116 accessed from the buffer 1460 and provide the filtered audio data to the first - stage speech processor 124. In this example, when there is no voice assistant session active, the audio pre - processor 118 operates in a non - speaker - specific manner, such as by performing echo cancellation, noise suppression, etc.

[0128] When the wake - word detector 126 detects a wake - word in the filtered audio data from the audio pre - processor 118, the first - stage speech processor 124 causes the speaker detector 128 to identify the person who uttered the wake - word, transmits a wake - up signal 1422 or an interrupt 1424 to the second power domain 1405, and causes the audio pre - processor 118 to obtain configuration data associated with the person who uttered the wake - word.

[0129] Based on the configuration data, the audio pre - processor 118 begins to operate in a speaker - specific mode, as described in any of the figures referenced Figures 1 to 5 In the speaker - specific mode, the audio pre - processor 118 provides the audio data 150 to the second - stage speech processor 154. The audio data 150 is filtered by a speaker - specific voice input filter to attenuate, dampen, or remove portions of the audio data 116 that do not correspond to the speech of a particular person, and the voice signature data of that particular person is provided to the audio pre - processor 118 along with the configuration data. In some embodiments, the audio pre - processor 118 also provides the audio data 150 to the first - stage speech processor 124 until the voice assistant session terminates.

[0130] By selectively activating the second - stage speech processor 154 based on the results of processing audio data at the first - stage speech processor 124, the total power consumption associated with speech processing can be reduced.

[0131] Referencing Figure 15 illustrates a particular embodiment of a method 1500 for selectively filtering audio data for speech processing. In a particular aspect, one or more operations of the method 1500 are performed by at least one of the audio analyzer 140, processor 190, device 102, system 100, or a combination thereof. Figure 1

[0132] Method 1500 includes: at block 1502, obtaining first voice signature data associated with a first person based on detecting a wake word in the speech from the first person. For example, the audio pre-processor 118 may obtain Figure 1 configuration data 132 based on the wake word detector 126 detecting the wake word 110 in the speech 108A from person 180A. In this example, the configuration data 132 at least includes the voice signature data 134A associated with person 180A.

[0133] Method 1500 includes: at block 1504, selectively enabling a speaker-specific voice input filter based on the first voice signature data. For example, Figure 1 the configuration data 132 enables the voice input filter 120 of the audio pre-processor 118 to operate in a speaker-specific mode. One beneficial effect of selectively enabling speaker-specific filtering of audio data is that such filtering can improve the accuracy of speech recognition for a voice assistant application. Another beneficial effect of selectively enabling speaker-specific filtering of audio data is that such filtering can limit the ability of others to interrupt a voice assistant session.

[0134] Figure 15 The method 1500 can be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 15 the method 1500 can be executed by a processor that executes instructions, such as described with reference to Figure 18 as such.

[0135] With reference to Figure 16 , a specific implementation of method 1600 for selectively filtering audio data for speech processing is shown. In a particular aspect, one or more operations of method 1600 are performed by at least one of Figure 1 the audio analyzer 140, the processor 190, the device 102, the system 100, or a combination thereof.

[0136] Method 1600 includes: at block 1602, obtaining first voice signature data associated with a first person based on detecting a wake word in the speech from the first person. For example, the audio pre-processor 118 may obtain Figure 1 configuration data 132 based on the wake word detector 126 detecting the wake word 110 in the speech 108A from person 180A. In this example, the configuration data 132 at least includes the voice signature data 134A associated with person 180A. In Figure 16In the illustrated example, obtaining first voice signature data associated with a first person includes, at block 1604, selecting the first voice signature data from a set of voice signature data associated with multiple persons based on a comparison of the characteristics of the utterance with the enrollment data. For example, the speaker detector 128 can compare the characteristics of the utterance 108A with the characteristics of the enrollment data 136 (e.g., with the voice signature data 134) to determine the speaker identifier 130 for selecting the voice signature data 134A associated with the person 180A who spoke the wake word 110.

[0137] Method 1600 includes, at block 1606, selectively enabling a speaker-specific voice input filter based on the first voice signature data. For example, Figure 1 the configuration data 132 enables the voice input filter 120 of the audio pre-processor 118 to operate in a speaker-specific mode.

[0138] Method 1600 further includes, at block 1608, comparing the input audio data with the first voice signature data to generate output audio data that attenuates portions of the input audio data that do not correspond to the voice of the first person. In Figure 16 the illustrated example, attenuating portions of the input audio data that do not correspond to the voice of the first person can include, at block 1610, separating the voice of the first person from the voices of one or more other persons; at block 1612, removing or attenuating sounds from the audio data that are not associated with the voice of the first person; or both. For example, as described with reference to Figures 2A to 2C the configuration data 132 can include voice signature data for one or more persons and can be provided as input to the voice input filter 120 to enable the voice input filter 120 to operate as a speaker-specific voice input filter 210. The speaker-specific voice input filter 210 can attenuate portions of the audio data 116 that do not correspond to the voice of one or more persons, such as by removing or attenuating the ambient sound 112, etc.

[0139] Method 1600 further includes, at block 1614, initiating a voice assistant session after detecting the wake word. For example, the first-stage voice processor 124 can initiate a voice assistant session by providing the configuration data 132 to the audio pre-processor 118 and causing the audio data 150 to be provided to the second-stage voice processor 154. In some embodiments, the first-stage voice processor 124 can cause the second-stage voice processor 154 to be activated, such as described with reference to Figure 14 the description.

[0140] Method 1600 further includes, at block 1616, providing the voice of the first person to one or more voice assistant applications. For example, Figure 1The audio data 150 is provided to the voice assistant application 156. In this example, the audio data 150 includes the portion of the audio data 116 corresponding to the voice of the person 180A who spoke the wake word to initiate the voice assistant session.

[0141] Method 1600 further includes, at block 1618, disabling the speaker-specific voice input filter based on determining that a voice assistant session associated with a first person has ended. For example, one or more components of the audio analyzer 140, such as the audio pre-processor 118, the first-stage speech processor 124, or the second-stage speech processor 154, can determine when a termination condition associated with the voice assistant session is met. The termination condition can be met based on the elapsed time associated with the voice assistant session, the time elapsed since speech was provided via the audio data 150, a termination instruction in the audio data 150, and the like.

[0142] One beneficial effect of selectively enabling speaker-specific filtering of audio data is that such filtering can improve the accuracy of speech recognition by the voice assistant application. Another beneficial effect of selectively enabling speaker-specific filtering of audio data is that such filtering can limit the ability of others to interrupt the voice assistant session.

[0143] Figure 16 Method 1600 can be implemented by a field-programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 16 Method 1600 can be executed by a processor that executes instructions, such as described with reference to Figure 18 as described.

[0144] Referring to Figure 17 illustrates a particular specific implementation of method 1700 for selectively filtering audio data for speech processing. In a particular aspect, one or more operations of method 1700 are performed by at least one of the audio analyzer 140, the processor 190, the device 102, the system 100, or a combination thereof. Figure 1 The audio analyzer 140, the processor 190, the device 102, the system 100, or a combination thereof.

[0145] Method 1700 includes, at block 1702, obtaining first voice signature data associated with a first person based on detecting a wake word in the utterance from the first person. For example, the audio pre-processor 118 can obtain Figure 1 configuration data 132 based on the wake word detector 126 detecting the wake word 110 in the utterance 108A from the person 180A. In this example, the configuration data 132 includes at least the voice signature data 134A associated with the person 180A.

[0146] Method 1700 includes, at block 1704, selectively enabling a speaker-specific voice input filter based on first voice signature data. For example, Figure 1 configuration data 132 of Figure 1 enables the voice input filter 120 of the audio pre-processor 118 to operate in a speaker-specific mode.

[0147] Method 1700 further includes, at block 1706, initiating a voice assistant session based on detection of a wake word. For example, the first-stage voice processor 124 may initiate a voice assistant session by providing the configuration data 132 to the audio pre-processor 118 and causing the audio data 150 to be provided to the second-stage voice processor 154. In some embodiments, the first-stage voice processor 124 may cause the second-stage voice processor 154 to be activated, as described with reference to Figure 14 as described.

[0148] Method 1700 further includes, at block 1708, providing the voice of a first person to one or more voice assistant applications. For example, Figure 5 the enhanced voice 508 of the first person of Figure 5 is provided to the voice assistant application 156. In this example, the enhanced voice 508 of the first person includes the portion of the audio data 116 corresponding to the voice of the person who spoke the wake word to initiate the voice assistant session.

[0149] Method 1700 further includes, at block 1710, receiving audio data including a second utterance from a second person. For example, in Figure 5 Figure 5 , the audio data 116 includes multi-person speech 302, which may include the voice of the person who spoke the wake word to initiate the voice assistant session and the voices of one or more other people.

[0150] Method 1700 further includes, at block 1712, determining whether to provide the content of the second utterance to the voice assistant application based on whether the content of the second utterance is contextually relevant to a voice assistant request received from the first person. For example, in Figure 5 Figure 5 , the latent space representation 312 associated with the second utterance is processed via a second processing path to generate the enhanced voice 516 of the second person. In this example, the enhanced voice 516 of the second person is provided to the NLP engine 520 together with the context data 522 associated with the voice assistant session. The NLP engine 520 determines whether the enhanced voice 516 of the second person is relevant to the voice assistant session, and if appropriate, provides the relevant voice 524 of the second person to the voice assistant application 156.

[0151] Figure 17The method 1700 can be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 17 the method 1700 can be executed by a processor that executes instructions, such as those described with reference to Figure 18 as described.

[0152] With reference to Figure 18 , a block diagram depicting a particular illustrative embodiment of a device is shown and generally designated as 1800. In various embodiments, the device 1800 may have more or fewer components than those Figure 18 illustrated. In an illustrative embodiment, the device 1800 may correspond to the device 102. In an illustrative embodiment, the device 1800 may perform one or more operations described with reference to Figures 1 to 17 as described.

[0153] In a particular embodiment, the device 1800 includes a processor 1806 (e.g., a central processing unit (CPU)). The device 1800 may include one or more additional processors 1810 (e.g., one or more DSPs). In a particular aspect, Figure 1 the processor 190 corresponds to the processor 1806, the processor 1810, or a combination thereof. The processor 1810 may include a voice and music codec, which includes a voice codec (“vocoder”) encoder 1836 and a vocoder decoder 1838. In the Figure 18 illustrated example, the processor 1810 also includes an audio pre-processor 118, a first stage voice processor 124, and optionally a second stage voice processor 154.

[0154] The device 1800 may include a memory 142 and a codec 1834. In a particular embodiment, Figure 6 and Figure 7 the codec 604 corresponds to Figure 18 the codec 1834. The memory 142 may include instructions 1856 that can be executed by one or more additional processors 1810 (or the processor 1806) to implement the functionality described with reference to the audio pre-processor 118, the first stage voice processor 124, the second stage voice processor 154, or a combination thereof. In the Figure 18 illustrated example, the memory 142 also includes registration data 136.

[0155] Device 1800 may include a display 1828 coupled to a display controller 1826. Audio transducer 162, microphone 104, or both may be coupled to codec 1834. Codec 1834 may include a digital-to-analog converter (DAC) 1802, an analog-to-digital converter (ADC) 1804, or both. In a particular implementation, codec 1834 may receive an analog signal from microphone 104, convert the analog signal to a digital signal (e.g., Figure 1 audio data 116), and provide the digital signal to a voice and music codec 1808. Voice and music codec 1808 may process the digital signal, and the digital signal may be further processed by an audio pre-processor 118, a first-stage voice processor 124, a second-stage voice processor 154, or a combination thereof. In a particular implementation, voice and music codec 1808 may provide the digital signal to codec 1834. Codec 1834 may convert the digital signal to an analog signal using DAC 1802 and may provide the analog signal to audio transducer 162.

[0156] In a particular implementation, device 1800 may be included in a system-in-package or system-on-chip device 1822. In a particular implementation, memory 142, processor 1806, processor 1810, display controller 1826, codec 1834, and modem 1854 are included in system-in-package or system-on-chip device 1822. In a particular implementation, input device 1830 and power source 1844 are coupled to system-in-package or system-on-chip device 1822. Additionally, in a particular implementation, as Figure 18 illustrated, display 1828, input device 1830, audio transducer 162, microphone 104, antenna 1852, and power source 1844 are external to system-in-package or system-on-chip device 1822. In a particular implementation, each of display 1828, input device 1830, audio transducer 162, microphone 104, antenna 1852, and power source 1844 may be coupled to a component (such as an interface or a controller) of system-in-package or system-on-chip device 1822.

[0157] In some implementations, device 1800 includes a modem 1854 coupled to an antenna 1852 via a transceiver 1850. In some such implementations, modem 1854 may be configured to transmit data associated with the speech of a first person (e.g., Figure 1At least a portion of the audio data 116) is transmitted to the remote voice assistant server 1840. In such specific implementations, the voice assistant application 156 is executed at the voice assistant server 1840. In such specific implementations, the second-stage voice processor 154 may be omitted from the device 1800; however, speaker-specific voice input filtering may be performed at the device 1800 based on wake word detection at the device 1800.

[0158] The device 1800 may include a smart speaker, a speaker bar, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet computer, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a head-mounted device, an augmented reality head-mounted device, a mixed reality head-mounted device, a virtual reality head-mounted device, an aircraft, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, an automobile, a computing device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0159] In connection with the described specific implementations, a device includes components for obtaining first voice signature data associated with a first person based on detecting a wake word in the utterance from the first person. For example, the components for obtaining the first voice signature data may correspond to the device 102, the processor 190, the audio analyzer 140, the audio preprocessor 118, the voice input filter 120, the first-stage voice processor 124, the speaker detector 128, the integrated circuit 802, the processor 1806, the processor 1810, one or more other circuits or components configured to obtain voice signature data, or any combination thereof.

[0160] The device further includes components for selectively enabling a speaker-specific voice input filter based on the first voice signature data. For example, the components for selectively enabling the speaker-specific voice input filter may correspond to the device 102, the processor 190, the audio analyzer 140, the audio preprocessor 118, the voice input filter 120, the first-stage voice processor 124, the speaker detector 128, the integrated circuit 802, the processor 1806, the processor 1810, one or more other circuits or components configured to selectively enable the speaker-specific voice input filter, or any combination thereof.

[0161] In some specific implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 142) includes instructions (e.g., instruction 1856) that, when executed by one or more processors (e.g., one or more processors 190, one or more processors 1810, or processor 1806), cause the one or more processors to obtain first voice signature data associated with a first person based on detecting a wake word in the speech of the first person, and selectively enable a speaker-specific voice input filter based on the first voice signature data.

[0162] Specific aspects of the present disclosure are described below in a collection of related embodiments:

[0163] According to Embodiment 1, a device includes one or more processors configured to: obtain first voice signature data associated with a first person based on detecting a wake word in the speech of the first person; and selectively enable a speaker-specific voice input filter based on the first voice signature data.

[0164] Embodiment 2 includes the device according to Embodiment 1, wherein the one or more processors are further configured to process audio data including speech from multiple persons to detect the wake word.

[0165] Embodiment 3 includes the device according to Embodiment 1 or Embodiment 2, wherein obtaining the first voice signature data includes: selecting the first voice signature data from a set of voice signature data associated with multiple persons based on a comparison of features of the speech with registered data.

[0166] Embodiment 4 includes the device according to any one of Embodiments 1 to 3, wherein the speaker-specific voice input filter is configured to separate the speech of the first person from the speech of one or more other persons and provide the speech of the first person to one or more voice assistant applications.

[0167] Embodiment 5 includes the device according to any one of Embodiments 1 to 4, wherein the speaker-specific voice input filter is configured to remove or attenuate sounds in the audio data that are not associated with the speech of the first person.

[0168] Embodiment 6 includes the device according to any one of Embodiments 1 to 5, wherein the speaker-specific voice input filter is configured to compare input audio data with the first voice signature data to generate output audio data, and the output audio data weakens portions of the input audio data that do not correspond to the speech of the first person.

[0169] Example 7 includes the apparatus according to any one of Examples 1 to 6, wherein the one or more processors are further configured to, based on detection of the wake word: obtain second voice signature data associated with at least a second person based on configuration data; and configure the speaker-specific voice input filter based on the first voice signature data and the second voice signature data.

[0170] Example 8 includes the apparatus according to any one of Examples 1 to 6, wherein the one or more processors are further configured to, after enabling the speaker-specific voice input filter based on the first voice signature data: receive audio data including a second utterance from a second person; and determine whether to provide the content of the second utterance to a voice assistant application based on whether the content of the second utterance is contextually relevant to a voice assistant request received from the first person.

[0171] Example 9 includes the apparatus according to any one of Examples 1 to 8, wherein the one or more processors are further configured to: when the speaker-specific voice input filter is enabled, provide first audio data to a first voice enhancement model based on the first voice signature data; and when the speaker-specific voice input filter is not enabled, provide second audio data to a second voice enhancement model based on second voice signature data.

[0172] Example 10 includes the apparatus according to Example 9, wherein the second voice signature data represents the voices of multiple people.

[0173] Example 11 includes the apparatus according to any one of Examples 1 to 10, wherein the one or more processors are further configured to disable the speaker-specific voice input filter based on determining that a voice assistant session associated with the first person has ended after enabling the speaker-specific voice input filter.

[0174] Example 12 includes the apparatus according to Example 11, wherein the one or more processors are further configured to, during the voice assistant session: receive first audio data representing the voices of multiple people; generate second audio data representing the voice of a single person based on the speaker-specific voice input filter; and provide the second audio data to a voice assistant application.

[0175] Example 13 includes the apparatus according to any one of Examples 1 to 12, wherein the first voice signature data corresponds to a first speaker embedding, and wherein the one or more processors are configured to enable the speaker-specific voice input filter by providing the first speaker embedding as an input to a voice enhancement model.

[0176] Example 14 includes the apparatus according to any one of Examples 1 to 13, wherein the one or more processors are integrated into a vehicle.

[0177] Example 15 includes the apparatus according to any one of Examples 1 to 13, wherein the one or more processors are integrated into at least one of the following: a smart speaker, a speaker bar, a smart phone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a tuner, a camera, a navigation device, a head-mounted device, an augmented reality head-mounted device, a mixed reality head-mounted device, a virtual reality head-mounted device, a home automation system, a voice-activated device, a wireless speaker and voice-activated device, a portable electronic device, a communication device, an Internet of Things (IoT) device, an extended reality (XR) device, a base station, or a mobile device.

[0178] Example 16 includes the apparatus according to any one of Examples 1 to 15, the apparatus further including a microphone configured to capture sound including the utterance from the first person.

[0179] Example 17 includes the apparatus according to any one of Examples 1 to 16, the apparatus further including a modem configured to transmit data associated with the utterance from the first person to a remote voice assistant server.

[0180] Example 18 includes the apparatus according to any one of Examples 1 to 17, the apparatus further including an audio transducer configured to output sound corresponding to a voice assistant response to the first person.

[0181] According to Example 19, a method includes: obtaining first voice signature data associated with a first person based on detecting a wake word in an utterance from the first person; and selectively enabling a speaker-specific voice input filter based on the first voice signature data.

[0182] Example 20 includes the method according to Example 19, the method further including processing audio data including speech from multiple people to detect the wake word.

[0183] Example 21 includes the method according to Example 19 or Example 20, wherein obtaining the first voice signature data includes: selecting the first voice signature data from a set of voice signature data associated with multiple people based on a comparison of features of the utterance with registration data.

[0184] Example 22 includes the method according to any one of Examples 19 to 21, the method further comprising: separating, by the speaker-specific voice input filter, the voice of the first person from the voices of one or more other persons; and providing the voice of the first person to one or more voice assistant applications.

[0185] Example 23 includes the method according to any one of Examples 19 to 22, the method further comprising removing or attenuating, by the speaker-specific voice input filter, sounds in the audio data that are not associated with the voice of the first person.

[0186] Example 24 includes the method according to any one of Examples 19 to 23, the method further comprising comparing, by the speaker-specific voice input filter, the input audio data with the first voice signature data to generate output audio data that attenuates portions of the input audio data that do not correspond to the voice of the first person.

[0187] Example 25 includes the method according to any one of Examples 19 to 24, the method further comprising, based on detection of the wake word: obtaining, based on configuration data, second voice signature data associated with at least one second person; and configuring the speaker-specific voice input filter based on the first voice signature data and the second voice signature data.

[0188] Example 26 includes the method according to any one of Examples 19 to 24, the method further comprising, after enabling the speaker-specific voice input filter based on the first voice signature data: receiving audio data including a second utterance from a second person; and determining whether to provide the content of the second utterance to a voice assistant application based on whether the content of the second utterance is contextually relevant to a voice assistant request received from the first person.

[0189] Example 27 includes the method according to any one of Examples 19 to 26, the method further comprising: when the speaker-specific voice input filter is enabled, providing first audio data to a first voice enhancement model based on the first voice signature data; and when the speaker-specific voice input filter is not enabled, providing second audio data to a second voice enhancement model based on second voice signature data.

[0190] Example 28 includes the method according to Example 27, wherein the second voice signature data represents the voices of multiple persons.

[0191] Example 29 includes the method according to any one of Examples 19 to 28, the method further including, after enabling the speaker-specific voice input filter, disabling the speaker-specific voice input filter based on determining that a voice assistant session associated with the first person has ended.

[0192] Example 30 includes the method according to Example 29, the method further including, during the voice assistant session: receiving first audio data representing multi-person speech; generating second audio data representing single-person speech based on the speaker-specific voice input filter; and providing the second audio data to a voice assistant application.

[0193] Example 31 includes the method according to any one of Examples 19 to 30, wherein the first voice signature data corresponds to a first speaker embedding, and wherein enabling the speaker-specific voice input filter includes providing the first speaker embedding as an input to a voice enhancement model.

[0194] According to Example 32, a non-transitory computer-readable medium stores instructions that can be executed by one or more processors to cause the one or more processors to: obtain first voice signature data associated with a first person based on detecting a wake word in a speech from the first person; and selectively enable a speaker-specific voice input filter based on the first voice signature data.

[0195] Example 33 includes the non-transitory computer-readable medium according to Example 32, wherein the instructions can further be executed to cause the one or more processors to process audio data including speech from multiple people to detect the wake word.

[0196] Example 34 includes the non-transitory computer-readable medium according to Example 32 or Example 33, wherein obtaining the first voice signature data includes: selecting the first voice signature data from a set of voice signature data associated with multiple people based on a comparison of features of the speech with registered data.

[0197] Example 35 includes the non-transitory computer-readable medium according to any one of Examples 32 to 34, wherein the speaker-specific voice input filter is configured to separate the speech of the first person from the speech of one or more other people and provide the speech of the first person to one or more voice assistant applications.

[0198] Example 36 includes the non-transitory computer-readable medium according to any one of Examples 32 to 35, wherein the speaker-specific voice input filter is configured to remove or attenuate sounds in the audio data that are not associated with the speech of the first person.

[0199] Example 37 includes the non-transitory computer-readable medium according to any one of Examples 32 to 36, wherein the speaker-specific voice input filter is configured to compare the input audio data with the first voice signature data to generate output audio data, and the output audio data weakens a portion of the input audio data that does not correspond to the voice of the first person.

[0200] Example 38 includes the non-transitory computer-readable medium according to any one of Examples 32 to 37, wherein the instructions are further executable to cause the one or more processors, based on the detection of the wake word: obtain second voice signature data associated with at least a second person based on configuration data; and configure the speaker-specific voice input filter based on the first voice signature data and the second voice signature data.

[0201] Example 39 includes the non-transitory computer-readable medium according to any one of Examples 32 to 37, wherein the instructions are further executable to cause the one or more processors, after enabling the speaker-specific voice input filter based on the first voice signature data: receive audio data including a second utterance from a second person; and determine whether to provide the content of the second utterance to the voice assistant application based on whether the content of the second utterance is contextually relevant to the voice assistant request received from the first person.

[0202] Example 40 includes the non-transitory computer-readable medium according to any one of Examples 32 to 39, wherein the instructions are further executable to cause the one or more processors: when the speaker-specific voice input filter is enabled, provide first audio data to a first voice enhancement model based on the first voice signature data; and when the speaker-specific voice input filter is not enabled, provide second audio data to a second voice enhancement model based on second voice signature data.

[0203] Example 41 includes the non-transitory computer-readable medium according to Example 40, wherein the second voice signature data represents the voices of multiple people.

[0204] Example 42 includes the non-transitory computer-readable medium according to any one of Examples 32 to 41, wherein the instructions are further executable to cause the one or more processors to disable the speaker-specific voice input filter based on determining that a voice assistant session associated with the first person has ended after enabling the speaker-specific voice input filter.

[0205] Example 43 includes the non-transitory computer-readable medium according to Example 42, wherein the instructions are further executable to cause the one or more processors, during the voice assistant session: receive first audio data representing multiple-person speech; generate second audio data representing single-person speech based on the speaker-specific voice input filter; and provide the second audio data to a voice assistant application.

[0206] Example 44 includes the non-transitory computer-readable medium according to any one of Examples 32 to 43, wherein the first voice signature data corresponds to a first speaker embedding, and wherein the instructions are further executable to cause the one or more processors to enable the speaker-specific voice input filter by providing the first speaker embedding as an input to a voice enhancement model.

[0207] According to Example 45, a device includes: means for obtaining first voice signature data associated with a first person based on detecting a wake word in a speech from the first person; and means for selectively enabling a speaker-specific voice input filter based on the first voice signature data.

[0208] Example 46 includes the device according to Example 45, the device further including means for processing audio data including speech from multiple persons to detect the wake word.

[0209] Example 47 includes the device according to Example 45 or Example 46, wherein the means for obtaining the first voice signature data includes: means for selecting the first voice signature data from a set of voice signature data associated with multiple persons based on a comparison of features of the speech with registration data.

[0210] Example 48 includes the device according to any one of Examples 45 to 47, wherein the speaker-specific voice input filter includes: means for separating the speech of the first person from the speech of one or more other persons; and means for providing the speech of the first person to one or more voice assistant applications.

[0211] Example 49 includes the device according to any one of Examples 45 to 48, wherein the speaker-specific voice input filter includes means for removing or attenuating sounds in the audio data that are not associated with the speech of the first person.

[0212] Example 50 includes the apparatus according to any one of Examples 45 to 49, wherein the speaker-specific voice input filter includes components for comparing input audio data with the first voice signature data to generate output audio data that attenuates portions of the input audio data that do not correspond to the voice of the first person.

[0213] Example 51 includes the apparatus according to any one of Examples 45 to 50, the apparatus further including: components for obtaining second voice signature data associated with at least one second person based on configuration data; and components for configuring the speaker-specific voice input filter based on the first voice signature data and the second voice signature data.

[0214] Example 52 includes the apparatus according to any one of Examples 45 to 50, the apparatus further including: components for receiving audio data including a second utterance from a second person when the speaker-specific voice input filter is enabled based on the first voice signature data; and components for determining whether to provide the content of the second utterance to a voice assistant application based on whether the content of the second utterance is contextually relevant to a voice assistant request received from the first person.

[0215] Example 53 includes the apparatus according to any one of Examples 45 to 52, the apparatus further including: components for providing first audio data to a first voice enhancement model based on the first voice signature data when the speaker-specific voice input filter is enabled; and components for providing second audio data to a second voice enhancement model based on second voice signature data when the speaker-specific voice input filter is not enabled.

[0216] Example 54 includes the apparatus according to Example 53, wherein the second voice signature data represents the voices of multiple people.

[0217] Example 55 includes the apparatus according to any one of Examples 45 to 54, the apparatus further including components for disabling the speaker-specific voice input filter based on determining that a voice assistant session associated with the first person has ended.

[0218] Example 56 includes the apparatus according to Example 55, the apparatus further including: components for receiving first audio data representing the voices of multiple people during the voice assistant session; components for generating second audio data representing the voice of a single person based on the speaker-specific voice input filter; and components for providing the second audio data to a voice assistant application.

[0219] Example 57 includes the apparatus according to any one of Examples 45 to 56, wherein the first voice signature data corresponds to a first speaker embedding, and wherein the component for selectively enabling the speaker-specific voice input filter includes a component for providing the first speaker embedding as an input to a voice enhancement model.

[0220] Example 58 includes the apparatus according to any one of Examples 45 to 57, wherein the component for obtaining the first voice signature data associated with the first person and the component for selectively enabling the speaker-specific voice input filter are integrated into a vehicle.

[0221] Example 59 includes the apparatus according to any one of Examples 45 to 57, wherein the component for obtaining the first voice signature data associated with the first person and the component for selectively enabling the speaker-specific voice input filter are integrated into at least one of the following: a smart speaker, a speaker bar, a smart phone, a cellular phone, a laptop computer, a computer, a tablet computer, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a tuner, a camera, a navigation device, a head-mounted device, an augmented reality head-mounted device, a mixed reality head-mounted device, a virtual reality head-mounted device, a home automation system, a voice-activated device, a wireless speaker and voice-activated device, a portable electronic device, a communication device, an Internet of Things (IoT) device, an extended reality (XR) device, a base station, or a mobile device.

[0222] Example 60 includes the apparatus according to any one of Examples 45 to 59, the apparatus further including a component for capturing sound including the utterance from the first person.

[0223] Example 61 includes the apparatus according to any one of Examples 45 to 60, the apparatus further including a component for transmitting data associated with the utterance from the first person to a remote voice assistant server.

[0224] Example 62 includes the apparatus according to any one of Examples 45 to 61, the apparatus further including a component for outputting sound corresponding to a voice assistant response to the first person.

[0225] Those skilled in the art will also appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the specific implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various illustrative components, blocks, configurations, modules, circuits, and steps have been described generally above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends upon the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, and such implementation decisions will not be interpreted as causing a departure from the scope of the present disclosure.

[0226] The steps of a method or algorithm described in connection with the specific implementations disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module can reside in random access memory (RAM), flash memory, read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), registers, a hard disk, a removable disk, a compact disc read only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or a user terminal.

[0227] The foregoing description of the disclosed aspects is provided to enable any person skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features as defined by the following claims.

Claims

1. An apparatus, the apparatus comprising: one or more processors configured to: obtain first voice signature data associated with a first person based on detecting a wake word in speech from the first person; and selectively enable a speaker-specific voice input filter based on the first voice signature data.

2. The apparatus according to claim 1, wherein the one or more processors are further configured to process audio data including speech from multiple persons to detect the wake word.

3. The device according to claim 1, wherein obtaining the first voice signature data comprises: Select the first voice signature data from a set of voice signature data associated with multiple persons based on a comparison of features of the speech with registration data.

4. The apparatus according to claim 1, wherein the speaker-specific voice input filter is configured to separate the speech of the first person from the speech of one or more other persons and provide the speech of the first person to one or more voice assistant applications.

5. The apparatus according to claim 1, wherein the speaker-specific voice input filter is configured to remove or attenuate sounds in the audio data that are not associated with the speech of the first person.

6. The apparatus according to claim 1, wherein the speaker-specific voice input filter is configured to compare the input audio data with the first voice signature data to generate output audio data that attenuates portions of the input audio data that do not correspond to the speech of the first person.

7. The apparatus according to claim 1, wherein the one or more processors are further configured based on the detection of the wake word: obtain second voice signature data associated with at least one second person based on configuration data; and configure the speaker-specific voice input filter based on the first voice signature data and the second voice signature data.

8. The apparatus according to claim 1, wherein the one or more processors are further configured after enabling the speaker-specific voice input filter based on the first voice signature data: receive audio data including a second utterance from a second person; and determine whether to provide the content of the second utterance to a voice assistant application based on whether the content of the second utterance is contextually relevant to a voice assistant request received from the first person.

9. The apparatus according to claim 1, wherein the one or more processors are further configured to: when the speaker-specific voice input filter is enabled, provide first audio data to a first voice enhancement model based on the first voice signature data; and when the speaker-specific voice input filter is not enabled, provide second audio data to a second voice enhancement model based on second voice signature data.

10. The apparatus according to claim 9, wherein the second voice signature data represents the speech of multiple persons.

11. The device according to claim 1, wherein the one or more processors are further configured to disable the speaker-specific voice input filter based on determining that a voice assistant session associated with the first person has ended after enabling the speaker-specific voice input filter.

12. The device according to claim 11, wherein the one or more processors are further configured during the voice assistant session: Receive first audio data representing multi-person speech; Generate second audio data representing single-person speech based on the speaker-specific voice input filter; and Provide the second audio data to a voice assistant application.

13. The device according to claim 1, wherein the first voice signature data corresponds to a first speaker embedding, and wherein the one or more processors are configured to enable the speaker-specific voice input filter by providing the first speaker embedding as an input to a voice enhancement model.

14. The device according to claim 1, wherein the one or more processors are integrated into a vehicle.

15. The device according to claim 1, wherein the one or more processors are integrated into at least one of the following: a smart speaker, a speaker bar, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a tuner, a camera, a navigation device, a head-mounted device, an augmented reality head-mounted device, a mixed reality head-mounted device, a virtual reality head-mounted device, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, a communication device, an Internet of Things (IoT) device, an extended reality (XR) device, a base station, or a mobile device.

16. The device according to claim 1, the device further comprising a microphone configured to capture sound including the utterance from the first person.

17. The device according to claim 1, the device further comprising a modem configured to transmit data associated with the utterance from the first person to a remote voice assistant server.

18. The device according to claim 1, the device further comprising one or more audio transducers configured to output sound corresponding to a voice assistant response to the first person.

19. A method, the method comprising: Obtaining first voice signature data associated with the first person based on detecting a wake word in an utterance from the first person; And Selectively enabling a speaker-specific voice input filter based on the first voice signature data.

20. The method according to claim 19, wherein obtaining the first voice signature data comprises: Selecting the first voice signature data from a set of voice signature data associated with multiple persons based on a comparison of the features of the utterance with registration data.

21. The method according to claim 19, the method further comprising: Separating the voice of the first person from the voices of one or more other persons by the speaker-specific voice input filter; And Provide the speech of the first person to one or more voice assistant applications.

22. The method according to claim 19, the method further comprising removing or attenuating from the audio data sounds not associated with the speech of the first person by the speaker-specific voice input filter.

23. The method according to claim 19, the method further comprising comparing the input audio data with the first voice signature data by the speaker-specific voice input filter to generate output audio data, the output audio data attenuating portions of the input audio data that do not correspond to the speech of the first person.

24. The method according to claim 19, the method further comprising after enabling the speaker-specific voice input filter based on the first voice signature data: Receiving audio data including a second utterance from a second person; and Determining whether to provide the content of the second utterance to a voice assistant application based on whether the content of the second utterance is contextually relevant to a voice assistant request received from the first person.

25. The method according to claim 19, the method further comprising: When the speaker-specific voice input filter is enabled, providing first audio data to a first voice enhancement model based on the first voice signature data; And When the speaker-specific voice input filter is not enabled, providing second audio data to a second voice enhancement model based on second voice signature data.

26. The method according to claim 19, wherein the first voice signature data corresponds to a first speaker embedding, and wherein enabling the speaker-specific voice input filter includes providing the first speaker embedding as an input to a voice enhancement model.

27. A non-transitory computer-readable medium storing instructions that can be executed by one or more processors to cause the one or more processors to: Obtain first voice signature data associated with the first person based on detecting a wake word in an utterance from the first person; and Selectively enable a speaker-specific voice input filter based on the first voice signature data.

28. An apparatus, the apparatus comprising: [[ID=!5]]Components for obtaining first voice signature data associated with the first person based on detecting a wake word in an utterance from the first person; And Components for selectively enabling a speaker-specific voice input filter based on the first voice signature data.