Voice filtering other speakers from calls and audio messages
By using voice filtering technology, audio data is processed using data processing hardware and memory hardware to generate enhanced audio data, user voice is isolated and background noise is eliminated, thus solving the problem of background noise interference in audio communication in noisy environments and achieving clear audio communication.
Patent Information
- Application Number
- CN202180074499.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-30
- Filing Date
- 2021-10-26
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2041-10-26
AI Technical Summary
In noisy environments, it is difficult for the receiver of audio communication to hear or understand the content, and existing technologies are not effective in removing background noise to isolate the user's voice.
Using voice filtering technology, data processing hardware and memory hardware are used to receive and process raw audio data, generate enhanced audio data, isolate user voice and eliminate background noise, and transmit it to the receiving device.
This invention achieves clear transmission of user voice in noisy environments, with the receiving device audibly outputting clear and consistent audio-based communication results. It solves the background noise interference problem that is difficult to solve in existing technologies. By implementing the aforementioned technical means, the effect of clear transmission of user voice in noisy environments is achieved.
Smart Images

Figure CN116420188B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to voice filtering other speakers from calls and audio messages. BACKGROUND
[0002] Voice-enabled environments allow users to simply speak out queries or commands, and an automated assistant will address and answer the queries and / or cause the commands to be executed. Voice-enabled environments (e.g., homes, workplaces, schools, etc.) can be implemented using a network of connected microphone devices distributed throughout various rooms and / or areas of the environment. Thus, the connected microphone devices can implement an automated assistant, and a user can interact with the automated assistant by providing spoken utterances, which the automated assistant can respond to by performing actions, controlling another device, and / or providing responsive content (e.g., visual and / or audible natural language output).
[0003] An automated assistant can convert audio data corresponding to a user’s spoken utterance into corresponding text (or other semantic representation). For example, an automated assistant can include a speech recognition engine that attempts to recognize various characteristics of a spoken utterance, such as the sounds produced (e.g., phonemes), the order of pronunciation, the rhythm of speech, the intonation, etc., and then identify the text words or phrases represented by these characteristics. An automated assistant can employ voice filtering techniques as a preprocessing step performed on a user’s spoken utterance to help the speech recognition engine focus on the voice of the user speaking the utterance. SUMMARY
[0004] One aspect of the disclosure provides a method for activating voice filtering in an audio-based communication. The method includes receiving, at data processing hardware, a first instance of raw audio data, the first instance of raw audio data corresponding to a voice-based command of an assistant-enabled device to facilitate an audio-based communication between a user of the assistant-enabled device and a recipient. The voice-based command is spoken by the user and captured by the assistant-enabled device. The method further includes receiving, at the data processing hardware, a second instance of raw audio data, the second instance of raw audio data corresponding to an utterance of audible content for the audio-based communication spoken by the user and captured by the assistant-enabled device. The second instance of raw audio data captures one or more additional sounds that are not spoken by the user. The method further includes executing, by the data processing hardware, a voice filtering recognition routine to determine, based on the first instance of raw audio data, whether to activate voice filtering for speech of at least the user in the audio-based communication. When the voice filtering recognition routine determines to activate voice filtering for at least the speech of the user, the method includes obtaining, by the data processing hardware, a respective speaker embedding of the user representing voice characteristics of the user; and processing, by the data processing hardware, the second instance of raw audio data using the respective speaker embedding of the user to generate enhanced audio data for the audio-based communication, the enhanced audio data isolating the utterance of audible content spoken by the user and excluding at least a portion of the one or more additional sounds that are not spoken by the user. The method further includes transmitting, by the data processing hardware, the enhanced audio data to a recipient device associated with the recipient. The enhanced audio data, when received by the recipient device, causes the recipient device to audibly output the utterance of audible content spoken by the user.
[0005] Another aspect of the present disclosure provides a system for activating voice filtering in an audio-based communication. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations comprising: receiving a first instance of raw audio data, the first instance of raw audio data corresponding to a voice-based command of an assistant-enabled device to facilitate an audio-based communication between a user of the assistant-enabled device and a recipient. The voice-based command is spoken by the user and captured by the assistant-enabled device. The operations further include receiving a second instance of raw audio data, the second instance of raw audio data corresponding to an utterance of audible content for the audio-based communication spoken by the user and captured by the assistant-enabled device. The second instance of raw audio data captures one or more additional sounds that are not spoken by the user. The operations further include performing a voice filtering recognition routine to determine whether to activate voice filtering for at least the user’s voice in the audio-based communication based on the first instance of raw audio data. When the voice filtering recognition routine determines to activate voice filtering for at least the user’s voice, the operations further include: obtaining a respective speaker embedding of the user that represents voice characteristics of the user; and, processing the second instance of raw audio data using the respective speaker embedding of the user to generate enhanced audio data for the audio-based communication, the enhanced audio data isolating the utterance of audible content spoken by the user and excluding at least a portion of the one or more additional sounds that are not spoken by the user. The operations further include transmitting the enhanced audio data to a recipient device associated with the recipient. The enhanced audio data, when received by the recipient device, causes the recipient device to audibly output the utterance of audible content spoken by the user.
[0006] The details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF DRAWINGS
[0007] Figure 1A and Figure 1B is an example system for activating voice filtering to focus on one or more voices in an audio-based communication.
[0008] Figure 2 is an example voice filtering recognition routine.
[0009] Figure 3 is an example voice filtering engine that includes a voice filtering model for generating enhanced audio data.
[0010] Figure 4 is a flow diagram of an example arrangement of operations of a method of activating voice filtering in an audio-based communication.
[0011] Figure 5is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.
[0012] Like reference numbers in the various drawings indicate like elements. DETAILED DESCRIPTION
[0013] A user can use an automated assistant to transmit an audio communication, such as sending / receiving an audio message and placing a telephone call (e.g., audio and / or video) with a remote recipient. When the user is in a noisy environment (e.g., in a busy place, in a car, or in a noisy home), a recipient of the audio communication can have difficulty hearing or understanding the content of the audio communication due to a high level of background noise.
[0014] Implementations herein relate to applying speech filtering to focus on one or more voices in an audio-based communication transmitted to (or received from) another user by removing unwanted background noise from the audio communication. When audio data captured by an assistant-enabled device includes utterances spoken by a user conveying audible content of an audio-based communication as well as unwanted noise, applying speech filtering can generate an enhanced version of the audio data by removing the unwanted background noise so that a clear and consistent audio-based communication is received by an end recipient. As used herein, an audio-based communication can refer to an audio message, a telephone call, a video call (e.g., an audio-video call), or a broadcast audio. For example, an assistant-enabled device can record content of an audio message spoken by a user and then send the audio message to a recipient through a messaging or email platform. Speech filtering can be applied to remove unwanted background noise from audio data conveying the audio message at the assistant-enabled device, at a cloud-based intermediary node while the audio message is in route to the recipient, or at a recipient client device once the audio message is received. Thus, when the recipient wishes to play back the audio message, the recipient client device audibly outputs an enhanced version of the audio message that does not include the unwanted background noise originally captured when the user spoke utterances conveying the content of the audio message. Likewise, an assistant-enabled device can facilitate a telephone call and apply speech filtering to remove unwanted background noise in real-time. As with an audio message, speech filtering can be applied to remove unwanted noise from audio data of the telephone call at the assistant-enabled device locally or at any point along the communication path to the recipient device.
[0015] Figure 1A and Figure 1BAn example system 100 is shown that performs voice filtering by removing unwanted background noise from audio-based communications 150 to at least focus on the voice of a user 102 in audio-based communications 150 transmitted to (or received from) another user 103. The system 100 includes an assistant-enabled device (AED) 104 that executes a digital assistant 109 with which the user 102 can interact by voice. In the illustrated example, the AED 104 corresponds to a smart speaker. However, the AED 104 can include other computing devices such as, but not limited to, a smartphone, a tablet, a smart display, a desktop / laptop computer, a smart watch, a smart appliance, a headset, or a vehicle infotainment device. The AED 104 includes data processing hardware 10 and memory hardware 12 storing instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The AED 104 includes an array of one or more microphones 16 configured to capture acoustic sound, such as voice directed at the AED 104. The AED 104 can also include or be in communication with an audio output device (e.g., a speaker) 16 that can output audio, such as audible content from audio-based communications 150 received from other users 103 and / or synthetic voice from the digital assistant 109.
[0016] Figure 1AThe illustration shows user 102 speaking the first utterance 106, “Ok Computer, send the following audio message to Bob,” near AED 104. Microphone 16 of AED 104 receives the utterance 106 and processes the raw audio data corresponding to the first utterance 106. Initial processing of the audio data may involve filtering the audio data and converting it from an analog signal to a digital signal. While AED 104 processes the audio data, it can store the audio data in a buffer of memory hardware 12 for additional processing. Using the audio data in the buffer, AED 104 can use hot word detector 108 to detect whether the raw audio data includes hot words 110. Hot word detector 108 is configured to identify hot words included in the audio data without performing speech recognition on the audio data. In the example shown, if the hot word detector 108 detects an acoustic feature in the audio data that is a characteristic of hot word 110, then the hot word detector 108 can determine that the utterance 106 “Ok Computer, send the following audio message to Bob” includes hot word 110 “ok computer”. The acoustic feature may be Mel frequency cepstral coefficients (MFCCs) representing the short-term power spectrum of utterance 106, or it may be the Mel scale filter group energy of utterance 106.
[0017] When the hot word detector 108 determines that the audio data corresponding to utterance 106 includes hot word 110, AED 104 can trigger a wake-up process to initiate speech recognition of the audio data corresponding to utterance 106. For example, an Automatic Speech Recognizer (ASR) 116 running on AED 104 can perform speech recognition and semantic interpretation on the audio data corresponding to utterance 106. ASR 116 can process at least the portion of the original audio data following hot word 110 to generate a speech recognition result of the received original audio data, and perform semantic interpretation on the speech recognition result to determine that utterance 106 includes a voice-based command 118 to facilitate audio-based communication 150 between user 102 and receiver 103. In this example, ASR 116 can process a first instance of the original audio data “send the following audio message to Bob” and recognize the voice-based command 118.
[0018] In some implementations, the ASR 116 is located on the server 120 in addition to or instead of the AED 104. Upon the hotword detector 108 triggering the AED 104 to wake up in response to detecting the hotword 110 in the utterance 106, the AED 104 can transmit a first instance of the raw audio data corresponding to the utterance 106 to the server 120 over the network 132. The AED 104 can transmit to the server 120 the portion of the audio data that includes the hotword 110 to confirm the presence of the hotword 110. Alternatively, the AED 104 can transmit to the server 120 only the portion of the audio data corresponding to the portion of the utterance 106 that follows the hotword 110. The server 120 executes the ASR 116 to perform speech recognition and return the speech recognition results (e.g., transcription) of the audio data to the AED 104. The AED 104 in turn recognizes the words in the utterance 106, and the AED 104 performs semantic interpretation to identify the voice command 118. The AED 104 (and / or the server 120) can identify the voice-based command 118 for the digital assistant 109 to facilitate the audio-based communication 150 of the audible message from the AED 104 to the recipient device 105 associated with the recipient 103 over the network 132. Thereafter, the AED 104 keeps the microphone 16 open and receives a second instance of the raw audio data that corresponds to the utterance 124 of the audible content 126 of the audio message 150 spoken by the user 102 and captured by the AED 104. In the illustrated example, the utterance 124 of the audible content 126 includes “Hi Bob, how are you?” The second instance of the raw audio data also captures one or more additional sounds 128 that are not spoken by the user 102, such as background noise.
[0019] Before or after receiving the second instance of raw audio data corresponding to the utterance of the audible content, the AED 104 executes a voice filtering recognition routine (“routine”) 200 to determine whether to activate voice filtering at least for the user’s 102 voice in the audio-based communication (e.g., audio message) 150 based on the first instance of raw audio data corresponding to the voice-based command 118. When the routine 200 determines not to activate voice filtering, the AED 104 will simply transmit the second instance of raw audio data corresponding to the utterance 124 of the audible content 126 of the audible message 155 to the recipient device 105. Here, the recipient device 105 will simply play back the audible content 126 of the utterance 124 of the audible content 126 including “Hi Bob, how are you?” to Bob recipient 103 along with any background noise captured by the second instance of raw audio data. When the routine 200 determines to activate voice filtering, the AED 104 uses a voice filter engine 300 to generate enhanced audio data 152 for the audio-based communication 150 that isolates the utterance 124 of the audible content 126 spoken by the user 102 and excludes at least a portion of one or more additional sounds that are not spoken by the user 102. That is, when the routine 200 determines to activate voice filtering for other individuals in addition to the user 102, and when the at least a portion of the one or more additional sounds includes an additional utterance of audible content spoken by another individual, the voice filter engine 300 will generate enhanced audio data 152 that does not exclude the additional utterance of audible content. Otherwise, if the routine 200 determines to activate voice filtering only for the user 102, then the voice filter engine 300 will generate enhanced audio data 152 that isolates only the user’s 102 voice and excludes any other sounds captured by the second instance of raw audio data that are not spoken by the user 102.
[0020] While the following refers to Figure 3 In more detail, when the routine 200 determines to activate voice filtering for the user’s 102 voice, the AED 104 (or server 120) instructs the voice filter engine to obtain a respective speaker embedding 318 for the user 102 that represents voice characteristics of the user Figure 3 ), and use the respective speaker embedding to process the second instance of raw audio data corresponding to the utterance 124 of the audible content 126 to generate enhanced audio data 152 for the audio message 150 that isolates the utterance 124 of the audible content 126 spoken by the user 102 and excludes one or more additional sounds such as background noise 128 that are not spoken by the user 102. Figure 1AThe AED 104 (or the server 120) is shown transmitting the augmented audio data 152 of the audio message 150 to the recipient device 105 of the recipient 103, whereby the recipient device 105 audibly outputs the augmented audio data 152 to allow the recipient 103 to hear the utterance 124 of the audible content 126 "Hi Bob, how are you?" spoken by the user 102 without hearing the background noise 128 originally captured in the environment of the AED 104.
[0021] In some examples, the audio message 150 is not transmitted to the recipient device 105, but is stored on the AED 104 for later retrieval by the intended recipient. In these examples, the recipient 103 can invoke the AED 104 to audibly playback the recorded audio message 150 with the augmented audio data 152 generated by the voice filter engine 300 to isolate the voice of the user 102 conveying the audible content of the audio message 150. In other examples, the functionality of the routine 200 and the voice filter engine 300 can be performed on the recipient device 105, such that the recipient device 105 only receives the raw audio data in the audio-based communication 150. In these examples, the recipient device 105 can determine to activate voice filtering for at least the voice of the sender 102 and process the raw audio data to isolate the voice of the sender 102 conveying the audible content of the audio-based communication 150. In some further examples, the AED 104 sends both the raw audio data 301 Figure 3 ) and the augmented audio data 152 to the recipient device 105 to allow the recipient 103 to select playback of the raw audio data with voice filtering activated for at least the voice of the user 102 to hear the audible content of the audio-based communication without voice filtering, or to select playback of the augmented audio data 152 to hear the audible content of the audio-based communication 150 with voice filtering activated for at least the voice of the user 102. Further, the AED 104 can send multiple versions of the augmented audio data 152, each version associated with voice filtering applied to a different combination of voices. In this way, the recipient 103 can switch between different versions of the augmented audio data 152 to hear different combinations of isolated voices.
[0022] The recipient device 105 and / or the AED 104 can display a graphical indicator in a graphical user interface (GUI) indicating whether speech filtering is currently activated at least for the user's 102 speech. The GUI can further render one or more controls for activating / deactivating speech filtering at least for the user's speech. Here, the user can select the controls to select between playing back the original audio data without speech filtering to listen to the audible content of the audio-based communication and playing back the enhanced audio data 152 with speech filtering activated at least for the user's 102 speech to listen to the audible content of the audio-based communication 150. The user input indication of the control selection can be provided as user feedback 315 for training the classification model 210 of the speech filtering recognition routine 200 discussed below. The AED 104 can also include a physical button that can be selected to activate or deactivate speech filtering. However, the recipient device will not be provided these types of controls for activating or deactivating speech filtering.
[0023] Figure 2 An example of a speech filtering recognition routine 200 executed on the AED 104 (or the server 120) is shown for determining whether to activate speech filtering at least for the user's 102 speech in the audio-based communication 150. Executing the speech filtering recognition routine 200 can include executing a classification model 210 configured to receive context inputs 202 associated with the audio-based communication 150 and generate a classification result 212 as output indicating one of: speech filtering is activated for one or more of the speech in the audio-based communication; or no speech filtering is activated for any of the speech. When the classification result 212 based on the context inputs 202 is to activate speech filtering for one or more of the speech, the result 212 can specify each of the one or more of the speech.
[0024] In some examples, one or more of the context inputs 202 are derived from performing semantic interpretation from a speech recognition result for a first instance of the original audio data corresponding to the speech-based command 118. Here, the ASR 116 Figure 1A and Figure 1B) can generate speech recognition results corresponding to the first instance of the raw audio data of the voice-based command 118 and perform semantic interpretation on the speech recognition results to identify / determine one or more of the contextual inputs 202, such as the recipient 103 of the audio-based communication 150. These contextual inputs 202 can include the identity of the recipient 103 of the audio-based communication 150 and / or an explicit instruction for at least voice-activated speech filtering of the user’s 102 voice in the audio-based communication. The classification model 210 can determine whether the identified recipient 103 includes a particular recipient type that indicates that voice-activated speech filtering is appropriate for the audio-based communication. For example, when the identified recipient 103 includes a business, the classification model 210 can determine to activate speech filtering. On the other hand, when the identified recipient 103 includes a friend or family member of the user 102, the classification model 210 can determine to not activate speech filtering. In additional examples, when the voice-based command 118 includes an explicit instruction to activate speech filtering, the classification model 210 determines to activate voice-activated speech filtering for at least the user’s voice. For example, a voice-based command 118 that says “call the plumber and cancel out the background noise” includes an explicit command to activate speech filtering and identifies a recipient (e.g., the plumber) that includes a particular recipient type where speech filtering can be appropriate. In another example, a voice-based command 118 that says “call mom so that the kids can speak with her” identifies that the “kids” are also participants in the audio-based communication and identifies a recipient (e.g., mom) that includes a particular recipient type where speech filtering can not be appropriate. In this example, the classification result 212 can be to activate voice-activated speech filtering for each of the user’s 102 kids during the subsequent audio-based communication (e.g., phone call) between the kids and mom.
[0025] In additional examples, the AED 104 (or the server 120) otherwise processes the first instance of the raw audio data (e.g., Figure 1A the utterance 106 in the audio-based communication 150, or Figure 1BThe utterance 156 in the speech data is used to derive contextual input 202 that may be meaningful to routine 200 to determine whether activating speech filtering is appropriate. For example, when the first instance of the raw audio data includes preamble audio and / or hot words 110, 160 preceding voice commands 118, 168, routine 200 can extract audio features from the preamble audio and / or hot words to determine the background noise level of the environment of AED 104 at the time the voice command is initiated. Here, the background noise level can be used as contextual input 202 fed to classification model 210, indicating the likelihood that a subsequent second instance of the raw audio data corresponding to the audible content of the audio-based communication 150 will capture background noise. For example, a higher background noise level may indicate that activating speech filtering is more appropriate than when the background noise level is low.
[0026] Similarly, contextual input 202 can include the location of AED 104. In this case, an AED 104 located in user 102's home or office environment is likely less likely to activate voice filtering than if AED 104 were located in a public place such as a train station. Classification model 210 can also consider the type of AED 104 as contextual input when determining whether voice filtering is activated. Here, certain types of AEDs may be more suitable for activating voice filtering than others. For example, a shared AED 104 such as a smart speaker in a multi-user environment may be more suitable for activating voice filtering than a personal AED 104 such as a telephone, because a shared AED 104 is more likely to capture background sounds than a telephone held close to user 102's mouth.
[0027] refer to Figure 1B and Figure 2 In some implementations, one of the scene inputs 202 includes image data 20 captured by an image capture device 18 implemented at or otherwise communicating with the AED 104. Figure 1B ).For example, Figure 1B A first example is shown of AED 104 capturing raw audio data of spoken words 156, corresponding to a voice command 168 of AED 104 for facilitating a video call 150 as audio-based communication between user 102 and receiver Bob 103 (i.e., via a digital assistant 109 executed on AED 104). AED 104 may include a tablet or smart display configured for voice calls, so that image capture device 18 can capture image data 20 indicating that at least user 102 is in an image frame and thus engaged in the video call. Figure 1BImage data 20 is shown capturing the user 102 and another individual 107 (e.g., the daughter of the user 102). The AED 104 receives a first instance of raw audio data capturing the utterance 156 of the user 102 speaking "Ok Computer, video call Bob," where the hotword 160 "Ok Computer" precedes the voice command 168 "video call Bob." So far, the contextual input 202 fed to the classification model 210 of the voice filtering recognition routine 200 can include the recipient "Bob" identified as the user's brother, the type of AED 104 (such as a shared smart display configured for video calling), the environment of the AED 104, the background noise level derived from the audio features extracted from the preamble and / or hotword 160, and the image data 20 indicating that the user 102 and another individual 107 can be participants in the video call 150 with the recipient 103 to occur. The contextual input 202 can further indicate that the semantic interpretation of the recognition results of the utterance 156 did not identify any explicit instruction for activating voice filtering.
[0028] Based on the received instruction voice command 168 for AED 104 to facilitate the video call 150 with the recipient Bob 103, the AED 104 can initiate the video call 150 by first establishing a connection with the recipient device 105 associated with the recipient 103 via the network 132. Thereafter, the AED 104 keeps the microphone 16 open and receives a second instance of raw audio data corresponding to the utterance 176 of the user speaking the audible content 178 of the video call 150, which is captured by the AED 104. In the illustrated example, the utterance 176 of the audible content 178 includes "Hi Uncle Bob." The second instance of raw audio data also captures additional sounds that are not spoken by the user 102, such as background noise 179 and an additional utterance 180 spoken by the other individual 107, which includes the audible content "We miss you" after the audible content "Hi Uncle Bob" of the video call 150. While recognized as additional sounds that are not spoken by the user 102, the additional utterance 180 is spoken by the other individual 107 indicated by the image data 20 as a potential participant in the voice call, and thus contains audible content intended to be heard by the recipient 103. Accordingly, when the execution of the routine 200 results in the classification model 210 generating a classification result 212 indicating voice activation of voice filtering for the user 102 and the other individual 107, the voice filtering engine 300 will apply voice filtering to generate the enhanced audio data 152 that excludes the background noise 179 and isolates the voices of the user 102 and the other individual 107 in the video call 150.
[0029] While the following is for purposes of explanation, numerous modifications and changes can be made to the example of FIGURE 1 by those skilled in the art without departing from the scope of this disclosure. Figure 3 In more detail, when routine 200 determines that the speech of user 102 and another individual 107 is activated for speech filtering, AED 104 (or server 120) instructs speech filtering engine 300 to obtain respective speaker embeddings 318 for each of user 102 and another individual 107. Figure 3 The respective speaker embedding 318 for user 102 can be obtained by processing audio features of a first instance of raw audio data (e.g., hotword 160) to generate a verification embedding and matching it to a stored speaker embedding 318. If no stored speaker embedding 318 is available (e.g., user 102 is not registered with AED 104), the respective speaker embedding 318 can be used directly as a verification embedding for applying speech filtering to the speech of user 102 in subsequent utterances. When individual 107 is a registered user of the AED, the respective speaker embeddings 318 for individual 107 and optionally user 102 can be obtained by recognizing individual 107 based on image data 20 via facial recognition. Optionally, a facial image of individual 107 can be extracted from image data 20, and a speaker embedding 318 can be resolved by extracting audio features from audio that is synchronized with the lips of the individual moving in the extracted facial image. Speech filtering engine 300 uses the respective speaker embeddings 318 to process a second instance of raw audio data to generate enhanced audio data 152 for video call 150 that isolates utterance 176 (spoken by user 102) and additional utterance 180 (spoken by another individual 107) and excludes background noise 179. Thus, in conjunction with image data 20, AED 104 (or server 120) can transmit enhanced audio data 152 to a recipient device 105 of recipient 103 during video call 150. Recipient device 105 can audibly output enhanced audio data 152 to allow recipient 103 to hear utterance 178 “Hi Uncle Bob” spoken by user 102 and additional utterance 180 “We miss you” spoken by another individual (e.g., user’s daughter) 107 without hearing background noise 179 that was originally captured in the environment of AED 104.
[0030] With continued reference to Figure 2While doing so, the routine 200 can dynamically adjust which voices are active for speech filtering during an audio-based communication session that is ongoing between the AED 104 and the recipient device 105. For example, the classification model 210 can initially generate a classification result 212 that indicates that speech filtering is active only for the voice of the user 102, such that the speech filter engine 300 generates enhanced audio data 152 that isolates only the user's voice and excludes all other sounds that are not spoken by the user. However, upon receiving a second instance of the raw audio data that conveys audible content of an audio message, the ASR 116, through speech recognition and semantic interpretation, can indicate that a speech recognition result of the audible content identifies at least one other individual who is a participant in the audio-based communication 150. In one example, the audible content can include the user 102 saying "Hi Bob, it's me and Alex," whereby recognition of the utterance and subsequent semantic interpretation can identify that Alex is also a participant in the audio-based communication 150 in addition to the user. Accordingly, the classification model 210 can receive the contextual input 202 that the user 102 and Alex are participants, and generate an updated classification result 212 that activates speech filtering for the voices of the user 102 and Alex. Without this update based on the contextual input 202, any utterances spoken by Alex would be excluded from the audio-based communication, even though these utterances can contain audible content that is intended for the recipient 103 to hear. In some examples, during a current voice-based communication session, the speech filtering recognition routine 200 simply determines to re-activate speech filtering that was activated for a previous audio-based communication for the same voices in the current outgoing audio-based communication.
[0031] Executing the voice filter recognition routine 200 can include executing a classification model 210 as either a heuristic-based model or a trained machine learning model. In some implementations, when the classification model 210 is a trained machine learning model, the trained machine learning model is retrained / adjusted based on user feedback 215 received after the voice filter engine 300 applies voice filtering to the audio-based communication based on the classification results 212 generated by the model 210 for the same particular context input 202 to adaptively learn how to activate voice filtering for the particular context input 202. Here, the user feedback 215 can indicate acceptance of the voice for which the voice filter was active or can indicate subsequent user input indications indicating adjustments to the voice for which the voice filter was active. For example, if voice filtering is applied to isolate only the user’s voice, the user can provide user input indications to indicate that the user does not want to isolate particular voices that are not spoken by the user and / or other sounds from the communication-based audio. Accordingly, the AED 104 can perform a training process that continuously preserves the machine learning classification model 210 based on the context input 202, the associated classification results 212, and the obtained user feedback 215 so that the classification model 210 adaptively learns to output voice filtering classification results 212 that are personalized to the user 102 based on past user behavior / reactions in similar contexts.
[0032] Referring now to Figure 3 When the voice filter recognition routine 200 determines that voice filtering is activated for at least the voice of the user 102, the voice filter engine 300 can use a frequency transformer 303 (which can be implemented at the ASR 116) to generate a frequency representation 302 of the received raw audio data 301 captured by the AED 104. Here, the raw audio data 301 can include one or more utterances of audible content for an audio-based communication. The frequency representation 302 can be, for example, streaming audio data that is processed in an online fashion (e.g., in real-time or near real-time, such as in a telephone or video call) or non-streaming audio data that has been previously recorded (such as in an audio message) and provided to the voice filter engine. The voice filter engine also receives a speaker embedding 318 from a speaker embedding engine 317.
[0033] The speaker embedding 318 is an embedding of a given human speaker and can be obtained based on processing one or more instances of audio data from the given speaker using a speaker embedding model. As described herein, in some implementations, the speaker embedding 318 is previously generated by a speaker embedding engine based on prior instances of audio data from the given speaker. In some of those implementations, the speaker embedding 318 is associated with an account of the given speaker and / or a client device of the given speaker, and can be provided for utilization with the frequency representation 302 based on the frequency representation 302 from an AED 104 where the account has been authorized for utilization. The speaker embedding engine 317 can determine respective speaker embeddings 318 representing voice characteristics of each of one or more human speakers identified by the routine 200 for activating voice filtering. In some implementations, the speaker embedding engine 317 processes portions of the captured raw audio data 301 using a speaker embedding model (not depicted) to generate speaker embeddings. Additionally or alternatively, the speaker embedding engine 317 can select a pre-generated speaker embedding (e.g., a speaker embedding previously generated using a registration process) using voiceprint recognition, image recognition, passwords, and / or other verification techniques to determine a currently active human speaker, and thus, a speaker embedding of the currently active human speaker. In many implementations, the normalization engine 312 normalizes each of the one or more selected speaker embeddings 318.
[0034] The voice filter engine 300 can optionally process the frequency representation 302 using a power compression process to generate a power compression 304. In many implementations, the power compression process equalizes (or partially equalizes) the importance of quieter sounds relative to louder sounds in the audio data. Additionally or alternatively, the voice filter engine 300 can optionally process the frequency representation 302 using a normalization process to generate a normalization 306, and can optionally process the speaker embedding 318 using a normalization process to generate a normalization 312.
[0035] The voice filter engine 300 can include a voice filter model 112 trained to process the frequency representation 302 of the raw audio data 301 and the speaker embedding 318 corresponding to a human speaker to generate a predicted mask 322, where the frequency representation can be processed with the predicted mask 322 to generate a revised frequency representation 310 that isolates speech of the human speaker. Instead of using a predicted mask 322, other types of voice filter models 112 are possible without departing from the scope of the present disclosure. For example, an end-to-end voice filter model or a generative adversarial network (GAN) based model can directly generate a filtered spectrogram.
[0036] More specifically, the frequency representation 302 can be applied as input to a convolutional neural network (CNN) portion 314 of the voice filter model 112. In some implementations, the CNN portion 314 is a one-dimensional convolutional neural network. In many implementations, the convolutional output generated by the CNN portion 314, as well as the speaker embedding 318, are applied as input to a recurrent neural network (RNN) portion 316 of the voice filter model 112. Here, the RNN portion 316 can include unidirectional memory units (e.g., long short-term memory units (LSTMs), gated recurrent units (GRUs), and / or additional memory units). Additionally or alternatively, the RNN output generated by the RNN portion 316 can be applied as input to a fully connected feedforward neural network portion 320 of the voice filter model 112 to generate the predicted mask 322. In some examples, the CNN portion 314 is omitted and both the frequency representation 302 and the speaker embedding 318 are applied as input to the RNN 316.
[0037] The engine 300 can process the frequency representation 302 with the predicted mask 322 to generate a revised frequency representation 310. For example, the frequency representation 302 can be convolved 308 with the predicted mask 322 to generate the revised frequency representation 310. A waveform synthesizer 324 can apply an inverse frequency transform to the revised frequency representation 310 to generate the enhanced audio data 152, which is a human speaker’s utterance for playback. The enhanced audio data 152 can be: the same as the original audio data 301 when the original audio data 301 only captures the utterance from the speaker corresponding to the speaker embedding 318; empty / zero when the original audio data 301 lacks the utterance from the speaker corresponding to the speaker embedding 318; or, when the original audio data 301 includes the utterance from the speaker and additional sounds (e.g., overlapping utterances of other human speakers and / or additional background noise), exclude the additional sounds while preserving the utterance from the speaker corresponding to the speaker embedding 318.
[0038] Figure 4 A flowchart of an example method 400 for activating voice filtering in an audio-based communication 150 to focus on at least a voice of a user 102 is provided. At operation 402, the method 400 includes receiving a first instance of original audio data corresponding to a voice-based command 118 to facilitate an audio-based communication 150 between a user 102 and a recipient 103 of an assistant-enabled device 104 to enable the assistant-enabled device 104. The voice-based command 118 is spoken by the user 102 and captured by the assistant-enabled device 104.
[0039] At operation 404, the method 400 includes receiving a second instance of the raw audio data, the second instance corresponding to the utterance 124 of the audible content 126 for the audio-based communication 150 spoken by the user 102 and captured by the assistant-enabled device 104. The second instance of the raw audio data captures one or more additional sounds that are not spoken by the user 102.
[0040] At operation 406, the method 400 includes executing the voice filter recognition routine 200 to determine whether voice filtering is activated for at least the user’s voice in the audio-based communication 150 based on the first instance of the raw audio data. At operation 408, when the voice filter recognition routine determines that voice filtering is activated for at least the user’s voice, the method 400 further includes obtaining a respective speaker embedding 318 for the user 102 that represents voice characteristics of the user. At operation 410, the method 400 includes processing the second instance of the raw audio data using the speaker embedding 318 to generate enhanced audio data 152 for the audio-based communication 150 that is of the utterance 124 of the audible content spoken by the user 102 and that excludes at least a portion of the one or more additional sounds that are not spoken by the user.
[0041] At operation 412, the method 400 includes transmitting the enhanced audio data 152 to a recipient device 105 associated with the recipient 103. The enhanced audio data 152 when received by the recipient device 105 causes the recipient device 105 to audibly output the utterance 124 of the audible content 126 spoken by the user 102.
[0042] A software application (i.e., a software resource) can refer to computer software that causes a computing device to perform a task. In some examples, a software application can be referred to as an “application,” an “app,” or a “program.” Example applications include, without limitation, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0043] A non-transitory memory can be a physical device that is used on a temporary or permanent basis to store a program (e.g., a sequence of instructions) or data (e.g., program state information) for use by a computing device. A non-transitory memory can be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, without limitation, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware such as a boot program). Examples of volatile memory include, without limitation, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks or tapes.
[0044] Figure 5 is a schematic diagram of an example computing device 500 that can be used to implement the systems and methods described in this document. The computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions, are meant to be examples for the purposes of this document only, and are not meant to limit implementations of the applications described and / or claimed in this document.
[0045] The computing device 500 includes a processor 510, memory 520, a storage device 530, a high-speed interface / controller 540 connecting the memory 520 and the high-speed expansion ports 550 to the processor 510, and a low speed interface / controller 560 connecting the low speed bus 570 and the storage device 530 to the processor 510. Each of the components 510, 520, 530, 540, 550, and 560 are interconnected using various busses, and can be mounted on a common motherboard or in other ways as appropriate. The processor 510 can process instructions for execution within the computing device 500, including instructions stored in the memory 520 or on the storage device 530 to display graphical information for a graphical user interface (GUI) on an external input / output device, such as display 580 coupled to high speed interface 540. In other implementations, multiple processors and / or multiple buses can be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices 500 can be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).
[0046] The storage device 530 stores information within the computing device 500. The storage device 530 can be a computer-readable medium, a volatile memory unit or a non-volatile memory unit, among others. The non-transitory memory 520 can be a physical device that is used to store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 500 on a temporary or permanent basis. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware, such as a Basic Input-Output System (BIOS)). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and magnetic or optical disks.
[0047] The storage device 530 is capable of providing mass storage for the computing device 500. In some implementations, the storage device 530 is a computer- readable medium. In various different implementations, the storage device 530 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 520, the storage device 530, or memory on processor 510.
[0048] The high-speed controller 540 manages bandwidth-intensive operations for the computing device 500, while the low-speed controller 560 manages lower bandwidth-intensive operations. Such allocation of functions is exemplary only. In some implementations, the high-speed controller 550 is coupled to the memory 520, the display 580 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 550, which can accept various expansion cards (not shown). In some implementations, the low-speed controller 560 is coupled to the storage device 530 and the low-speed expansion port 590. The low-speed expansion port 590, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device (e.g., a switch or router), through the network adapter.
[0049] The computing device 500 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a standard server 500a or multiple times in a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0050] Various implementations of the systems and techniques described here can be realized in digital electronic and / or optical circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0051] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer readable medium, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0052] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0053] To provide for interaction with a user, one or more aspects of the disclosure can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen for displaying information to the user and optionally a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device of the user (e.g., a Web browser) using a data signal that can be modulated onto a carrier wave, e.g., using a modem to send the documents to and receive the documents from the Web browser.
[0054] Many implementations have been described. However, various modifications can be made without departing from their spirit and scope. Therefore, other implementations are within the scope of the following claims.
Claims
1. One method includes: A first instance of receiving raw audio data at data processing hardware, the first instance corresponding to a voice command for an assistant-enabled device, used to facilitate audio-based communication between a user and a receiver of the assistant-enabled device, the voice command being spoken by the user and captured by the assistant-enabled device; A second instance of the raw audio data is received at the data processing hardware, the second instance corresponding to the audible content of the audio-based communication spoken by the user and captured by the assistant-enabled device, the second instance of the raw audio data capturing one or more additional sounds not spoken by the user; The data processing hardware executes a speech filtering identification routine to determine, based on a first instance of the raw audio data, whether to activate speech filtering at least for the user's speech in the audio-based communication, wherein the speech filtering identification routine is executed based on contextual input associated with the audio-based communication, the contextual input including at least one of the following: information for identifying the receiver, the type of the assistant-enabled device, the environment in which the assistant-enabled device is located, the background noise level of the environment of the assistant-enabled device, and image data obtained from the image capture device of the assistant-enabled device; When the speech filtering recognition routine determines that speech filtering is activated at least for the user's speech: The data processing hardware obtains the user's corresponding speaker embedding, representing the user's voice characteristics; as well as The data processing hardware uses the second instance of the original audio data to process the corresponding speaker embedding of the user to generate enhanced audio data for the audio-based communication, which isolates the utterance of the audible content spoken by the user and excludes at least a portion of the one or more additional sounds not spoken by the user. as well as The enhanced audio data is transmitted by the data processing hardware to a receiver device associated with the receiver. When the enhanced audio data is received by the receiver device, the receiver device outputs the speech, which is audible, as spoken by the user.
2. The method according to claim 1, further comprising: The first instance of the raw audio data is processed by the data processing hardware using a speech recognition device to generate a speech recognition result; as well as The data processing hardware performs a semantic interpretation on the speech recognition result of the first instance of the raw audio data to determine that the first instance of the raw audio data includes the voice command used to facilitate the audio-based communication between the user and the receiver.
3. The method of claim 2, wherein, Executing the speech filtering identification routine to determine whether to activate speech filtering, at least for the user's speech in the audio-based communication, includes: The receiver of the audio-based communication is identified by the semantic interpretation performed based on the speech recognition result of the first instance of the original audio data; Determine whether the identified receivers of the audio-based communication include a specific receiver type, which indicates that activating the voice filtering is appropriate at least for the user's voice in the audio-based communication; and When the identified receiver of the audio-based communication includes the specific receiver type, it is determined that voice filtering is activated at least for the user's voice.
4. The method of claim 3, wherein, The recipient type includes enterprises.
5. The method of any one of claims 2-4, wherein, Executing the speech filtering identification routine to determine whether to activate speech filtering, at least for the user's speech in the audio-based communication, includes: Based on the semantic interpretation performed on the speech recognition result of the first instance of the original audio data, it is determined whether the voice command includes an explicit instruction to activate speech filtering at least for the user's speech; as well as When the voice command includes the explicit instruction to activate voice filtering at least for the user's voice, it is determined that voice filtering is activated at least for the user's voice.
6. The method of claim 5, further comprising, upon executing the speech filtering identification routine and determining that the speech command comprises an explicit instruction to activate speech filtering for the speech of the user and another individual: For the other entity, the data processing hardware obtains a corresponding speaker embedding representing the speech characteristics of the other entity. in: The one or more additional sounds captured by the second instance of the original audio data that are not spoken by the user include additional utterances of the audible content of the audio-based communication spoken by the other entity and background noise that is not spoken by either the user or the other entity. as well as The second instance of processing the original audio data to generate the enhanced audio data includes: using the corresponding speaker embeddings of the user and the other individual to process the original audio data to generate the enhanced audio data of the audio-based communication, wherein the enhanced audio data isolates the additional utterances and the utterances of the audible content and excludes the background noise.
7. The method according to any one of claims 1-6, wherein: The first instance of the raw audio data includes preamble audio and hot words for the assistant-enabled device preceding the voice command used to facilitate the audio-based communication; as well as Executing the speech filtering identification routine to determine whether to activate speech filtering, at least for the user's speech in the audio-based communication, includes: Audio features are extracted from the preamble audio and / or hot words to determine the background noise level of the environment of the assistant-enabled device; as well as Based on the background noise level of the environment of the device with the assistant enabled, determine to activate voice filtering at least for the user's voice in the audio-based communication.
8. The method according to any one of claims 1-7, further comprising: The data processing hardware determines the type of the device that enables the assistant. The execution of the voice filtering recognition routine to determine whether voice filtering is activated at least for the user's voice is further based on the type of the assistant-enabled device.
9. The method according to any one of claims 1-8, further comprising: The data processing hardware determines the environment in which the assistant-enabled device is located. The execution of the voice filtering recognition routine to determine whether voice filtering is activated, at least for the user's voice, is further based on the environment in which the assistant-enabled device is located.
10. The method according to any one of claims 1-9, further comprising, when the audio-based communication facilitated by the assistant-enabled device includes a video call: At the data processing hardware, image data indicating that at least the user is participating in the video call is received from the image capture device of the assistant-enabled device. wherein Executing the voice filtering identification routine to determine whether voice filtering is activated at least for the user's voice is further based on the image data indicating that at least the user is participating in the video call.
11. The method of claim 10, further comprising, when executing the speech filtering identification routine to activate speech filtering based on the image data indicating that the user and at least one other individual are participating in the video call: For the at least one other individual, the data processing hardware obtains a corresponding speaker embedding representing the speech characteristics of the at least one other individual. in: The one or more additional sounds captured by the second instance of the original audio data that are not spoken by the user include additional utterances of the audible content of the video call spoken by the at least one other individual and background noise that is not spoken by the user or any of the at least one other individual; as well as The second instance of processing the original audio data to generate the enhanced audio data includes: using the corresponding speaker embeddings of the user and the at least one other individual to process the original audio data to generate the enhanced audio data of the video call, the enhanced audio data isolating the additional utterances and the utterances of the audible content and excluding the background noise.
12. The method according to any one of claims 1-11, further comprising: The second instance of the raw audio data is processed by the data processing hardware using a voice recognition device to generate a voice recognition result of the audible content of the audio-based communication; as well as The data processing hardware performs semantic interpretation on the voice recognition results of the audible content of the audio-based communication. The execution of the speech filtering recognition routine to determine whether speech filtering is activated at least for the user's speech is further based on the semantic interpretation performed on the speech recognition result of the audible content of the audio-based communication.
13. The method of claim 12, further comprising, when the semantic interpretation performed by the speech filtering identification routine based on the speech identification result of the audible content instructs the audible content to identify at least one other individual involved in the audio-based communication between the user and the receiver to determine speech filtering for the user and at least one other individual: For the at least one other individual, the data processing hardware obtains a corresponding speaker embedding representing the speech characteristics of the at least one other individual. in: The one or more additional sounds captured by the second instance of the original audio data that are not spoken by the user include additional utterances of the audible content of the audio-based communication spoken by the at least one other individual and background noise that is not spoken by the user or any of the at least one other individual; as well as The second instance of processing the original audio data to generate the enhanced audio data includes: using the corresponding speaker embeddings of the user and the at least one other individual to process the original audio data to generate the enhanced audio data of the audio-based communication, the enhanced audio data isolating the additional utterances and the utterances of the audible content and excluding the background noise.
14. The method of any one of claims 1-13, wherein, The audio-based communication includes one of the following: audio call, telephone call, video call, audio message, or broadcast audio.
15. The method according to any one of claims 1-14, further comprising: displaying by the data processing hardware in a graphical user interface (GUI) on a screen in communication with the data processing hardware: A graphical indicator for indicating whether voice filtering is currently activated, at least for the user's voice; and A control for activating / deactivating voice filtering for at least the user's voice.
16. A system comprising: Data processing hardware; as well as Memory hardware that communicates with the data processing hardware, the memory hardware storing instructions, which, when executed on the data processing hardware, cause the data processing hardware to perform operations, the operations including: A first instance of receiving raw audio data, the first instance corresponding to a voice command for an assistant-enabled device, used to facilitate audio-based communication between a user of the assistant-enabled device and a receiver, the voice command being spoken by the user and captured by the assistant-enabled device; A second instance of the original audio data is received, the second instance corresponding to the audible content of the audio-based communication spoken by the user and captured by the assistant-enabled device, the second instance of the original audio data capturing one or more additional sounds not spoken by the user; A speech filtering identification routine is executed to determine, based on a first instance of the original audio data, whether speech filtering is activated at least for the user's speech in the audio-based communication, wherein the speech filtering identification routine is executed based on contextual input associated with the audio-based communication, the contextual input including at least one of the following: information for identifying the receiver, the type of the assistant-enabled device, the environment in which the assistant-enabled device is located, the background noise level of the environment of the assistant-enabled device, and image data obtained from the image capture device of the assistant-enabled device; When the speech filtering recognition routine determines that speech filtering is activated at least for the user's speech: Obtain the corresponding speaker embedding of the user representing the user's speech characteristics; and The second instance of processing the original audio data using the corresponding speaker embedding of the user is used to generate enhanced audio data for the audio-based communication, the enhanced audio data isolating the utterance of audible content spoken by the user and excluding at least a portion of the one or more additional sounds not spoken by the user; and The enhanced audio data is transmitted to a receiver device associated with the receiver, and when the enhanced audio data is received by the receiver device, the receiver device outputs the speech of the user in an audible manner.
17. The system of claim 16, wherein, The operation further includes: The first instance of processing the raw audio data using a speech recognition device to generate a speech recognition result; and Perform semantic interpretation on the speech recognition result of the first instance of the original audio data to determine that the first instance of the original audio data includes the voice command used to facilitate the audio-based communication between the user and the receiver.
18. The system of claim 17, wherein, Executing the speech filtering identification routine to determine whether to activate speech filtering, at least for the user's speech in the audio-based communication, includes: The receiver of the audio-based communication is identified by the semantic interpretation performed based on the speech recognition result of the first instance of the original audio data; Determine whether the identified receivers of the audio-based communication include a specific receiver type, which indicates that activating the voice filtering is appropriate at least for the user's voice in the audio-based communication; and When the identified receiver of the audio-based communication includes the specific receiver type, it is determined that voice filtering is activated at least for the user's voice.
19. The system according to claim 18, wherein, The recipient type includes enterprises.
20. The system according to any one of claims 17-19, wherein, Executing the speech filtering identification routine to determine whether to activate speech filtering, at least for the user's speech in the audio-based communication, includes: Based on the semantic interpretation performed on the speech recognition result of the first instance of the original audio data, it is determined whether the voice command includes an explicit instruction to activate speech filtering at least for the user's speech; as well as When the voice command includes the explicit instruction to activate voice filtering at least for the user's voice, it is determined that voice filtering is activated at least for the user's voice.
21. The system according to claim 20, wherein, The operation further includes, when executing the speech filtering recognition routine and determining that the speech command includes an explicit instruction to activate speech filtering for the speech of the user and another individual: For the other individual, obtain the corresponding speaker embedding representing the speech characteristics of the other individual. in: The one or more additional sounds captured by the second instance of the original audio data that are not spoken by the user include additional utterances of the audible content of the audio-based communication spoken by the other entity and background noise not spoken by either the user or the other entity; and The second instance of processing the original audio data to generate the enhanced audio data includes: using the corresponding speaker embeddings of the user and the other individual to process the original audio data to generate the enhanced audio data of the audio-based communication, wherein the enhanced audio data isolates the additional utterances and the utterances of the audible content and excludes the background noise.
22. The system according to any one of claims 16-21, wherein: The first instance of the raw audio data includes preamble audio and hot words for the assistant-enabled device preceding the voice command used to facilitate the audio-based communication; as well as Executing the speech filtering identification routine to determine whether to activate speech filtering, at least for the user's speech in the audio-based communication, includes: Audio features are extracted from the preamble audio and / or hot words to determine the background noise level of the environment of the assistant-enabled device; as well as Based on the background noise level of the environment of the device with the assistant enabled, determine to activate voice filtering at least for the user's voice in the audio-based communication.
23. The system according to any one of claims 16-22, wherein, The operation further includes: Determine the type of the device that enables the assistant. The execution of the voice filtering recognition routine to determine whether voice filtering is activated at least for the user's voice is further based on the type of the assistant-enabled device.
24. The system according to any one of claims 16-23, wherein, The operation further includes: Determine the environment in which the assistant-enabled device is located. The execution of the voice filtering recognition routine to determine whether voice filtering is activated, at least for the user's voice, is further based on the environment in which the assistant-enabled device is located.
25. The system according to any one of claims 16-24, wherein, The operation further includes, when the audio-based communication facilitated by the assistant-enabled device includes a video call: Image data indicating that at least the user is participating in the video call is received from the image capture device of the assistant-enabled device. The execution of the voice filtering identification routine to determine whether voice filtering is activated at least for the user's voice is further based on the image data indicating that at least the user is participating in the video call.
26. The system according to claim 25, wherein, The operation further includes activating voice filtering when the voice filtering identification routine is executed to determine, based on the image data indicating that the user and at least one other individual are participating in the video call, voice filtering for the user and at least one other individual: For the at least one other individual, obtain the corresponding speaker embedding representing the speech characteristics of the at least one other individual. in: The one or more additional sounds captured by the second instance of the original audio data that are not spoken by the user include additional utterances of the audible content of the video call spoken by the at least one other individual and background noise not spoken by the user or any of the at least one other individual; and The second instance of processing the original audio data to generate the enhanced audio data includes: using the corresponding speaker embeddings of the user and the at least one other individual to process the original audio data to generate the enhanced audio data of the video call, the enhanced audio data isolating the additional utterances and the utterances of the audible content and excluding the background noise.
27. The system according to claim 26, wherein, The operation further includes: The second instance of processing the raw audio data using a voice recognition device to generate a voice recognition result of the audible content of the audio-based communication; and Perform semantic interpretation on the speech recognition results of the audible content of the audio-based communication. The execution of the speech filtering recognition routine to determine whether speech filtering is activated at least for the user's speech is further based on the semantic interpretation performed on the speech recognition result of the audible content of the audio-based communication.
28. The system according to claim 27, wherein, The operation further includes, when the semantic interpretation performed by the speech filtering identification routine based on the speech recognition result of the audible content instructs the audible content to identify at least one other individual involved in the audio-based communication between the user and the receiver to determine and activate speech filtering for the user and at least one other individual: For the at least one other individual, obtain the corresponding speaker embedding representing the speech characteristics of the at least one other individual. in: The one or more additional sounds captured by the second instance of the original audio data that are not spoken by the user include additional utterances of the audible content of the audio-based communication spoken by the at least one other individual and background noise not spoken by the user or any of the at least one other individual; and The second instance of processing the original audio data to generate the enhanced audio data includes: using the corresponding speaker embeddings of the user and the at least one other individual to process the original audio data to generate the enhanced audio data of the audio-based communication, the enhanced audio data isolating the additional utterances and the utterances of the audible content and excluding the background noise.
29. The system according to any one of claims 16-28, wherein, The audio-based communication includes one of the following: audio call, telephone call, video call, audio message, or broadcast audio.
30. The system according to any one of claims 16-29, wherein, The operation further includes displaying in a graphical user interface (GUI) on a screen communicating with the data processing hardware: A graphical indicator used to indicate whether voice filtering is currently activated, at least for the user's voice; as well as A control for activating / deactivating voice filtering, at least for the user's voice.
Citation Information
Patent Citations
Voice output device, voice output method and program
JP2019148780A
Speaker diarization
US20190115029A1