Audio processing using sound sources
By using a sound source representation-based processing system, combined sound source representations are generated using neural networks and audio analyzers. This solves the problem of processing interesting sounds in multi-sound-source environments, achieving efficient resource utilization and improved user experience.
Patent Information
- Application Number
- CN202380026687.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-18
- Filing Date
- 2023-03-16
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2043-03-16
AI Technical Summary
In multi-sound-source environments, existing technologies struggle to effectively distinguish and process sound sources that are of interest to others, leading to wasted resources and a poor user experience.
By using a sound source representation-based processing system, neural networks and audio analyzers are employed to generate combined sound source representations, selectively retaining or removing sounds from multiple sound sources, thereby enhancing the perceptibility of sounds of interest.
It improves the perceptibility of interesting sounds, reduces the waste of computing resources and bandwidth, and enhances the user experience.
Smart Images

Figure CN118891674B_ABST
Abstract
Description
[0001] I. Cross-referencing of related applications
[0002] This application claims the benefit of priority from jointly owned U.S. nonprovisional patent application No. 17 / 655,511, filed March 18, 2022, the entire contents of which are expressly incorporated herein by reference.
[0003] II. Technical Field
[0004] This disclosure relates in general to audio representation processing based on sound sources.
[0005] III. Related Fields
[0006] Technological advancements have led to smaller and more powerful computing devices. For example, a wide variety of portable personal computing devices exist, including small, lightweight, and easily portable cordless phones (such as mobile and smartphones, tablets, and laptops). These devices can transmit voice and data packets over wireless networks. Furthermore, many of these devices incorporate additional functionality, such as digital still cameras, digital camcorders, digital recorders, and audio file players. Moreover, such devices can process executable instructions, including software applications such as web browser applications that can be used to access the internet. Accordingly, these devices can include significant computing power.
[0007] Such computing devices often incorporate the ability to receive audio signals from one or more microphones. For example, the audio signal may include sound from multiple sound sources. In some cases, only the sound from some of these sound sources is of interest. In such cases, audio from uninteresting sound sources may be processed or transmitted along with audio from interested sound sources, which can result in a less than satisfactory user experience, inefficient resource utilization (e.g., processor time or transmission bandwidth), or both.
[0008] IV. Summary of the Invention
[0009] According to one embodiment of this disclosure, an apparatus includes one or more processors configured to receive an input audio signal. The one or more processors are further configured to process the input audio signal based on a combined representation of multiple sound sources to generate an output audio signal. The combined representation is used to selectively retain or remove sounds from the multiple sound sources from the input audio signal. The one or more processors are further configured to provide the output audio signal to a second device.
[0010] According to another embodiment of this disclosure, a method includes receiving an input audio signal at a first device. The method further includes processing the input audio signal based on a combined representation of multiple sound sources to generate an output audio signal. The combined representation is used to selectively retain or remove sounds from the multiple sound sources from the input audio signal. The method also includes providing the output audio signal to a second device.
[0011] According to another embodiment of this disclosure, a non-transitory computer-readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to receive an input audio signal at a first device. When executed by the one or more processors, the instructions further cause the one or more processors to process the input audio signal based on a combined representation of multiple sound sources to generate an output audio signal. The combined representation is used to selectively retain or remove sounds from the multiple sound sources from the input audio signal. When executed by the one or more processors, the instructions further cause the one or more processors to provide the output audio signal to a second device.
[0012] According to another embodiment of this disclosure, an apparatus includes means for receiving an input audio signal at a first device. The apparatus also includes means for processing the input audio signal based on a combined representation of multiple sound sources to generate an output audio signal. This combined representation is used to selectively retain or remove sounds from the input audio signal from the multiple sound sources. The apparatus further includes means for providing the output audio signal to a second device.
[0013] Other aspects, advantages, and features of this disclosure will become apparent upon examination of the entire application, which includes the description of the drawings, detailed description, and claims.
[0014] V. Illustrations
[0015] Figure 1 This is a block diagram illustrating specific exemplary aspects of a system capable of operating to process audio based on a sound source representation, according to some examples of this disclosure.
[0016] Figure 2A These are illustrations of exemplary aspects of a sound source encoder capable of generating a representation of a sound source, based on some examples of this disclosure.
[0017] Figure 2B This is an illustration of another exemplary aspect of a sound source encoder capable of generating a representation of a sound source, based on some examples of this disclosure.
[0018] Figure 2C This is an illustration of another exemplary aspect of a sound source encoder capable of generating a representation of a sound source, based on some examples of this disclosure.
[0019] Figure 3 Based on some examples of this disclosure Figure 1 An illustrative diagram of the neural network aspect of the system.
[0020] Figure 4 Based on some examples of this disclosure and Figure 1 An illustrative diagram illustrating the associated operations of the joint training of the neural network and the sound source encoder of the system.
[0021] Figure 5 Based on some examples of this disclosure and Figure 1 Another example of an illustrative aspect of the joint training of the neural network and the sound source encoder of the system is shown in the diagram.
[0022] Figure 6 Examples of devices that are operable to process audio received from another device using a sound source representation are illustrated according to some examples of this disclosure.
[0023] Figure 7 Examples of devices, according to some examples of this disclosure, capable of operating to send audio based on sound source representation processing to another device are illustrated.
[0024] Figure 8 Based on some examples of this disclosure Figure 1 The diagram illustrates the illustrative aspects of the operation of the system's components.
[0025] Figure 9 Examples of integrated circuits capable of operating to process audio based on sound source representation are illustrated according to some examples of this disclosure.
[0026] Figure 10 This is an illustration of a mobile device capable of operating to process audio based on a sound source representation, according to some examples of this disclosure.
[0027] Figure 11 This is an illustration of a headset capable of operating to process audio based on a sound source representation, according to some examples of this disclosure.
[0028] Figure 12 This is an illustration of a wearable electronic device capable of operating to process audio based on a sound source representation, according to some examples of this disclosure.
[0029] Figure 13 This is a diagram illustrating a voice-controlled speaker system capable of operating to process audio based on a sound source representation, according to some examples of this disclosure.
[0030] Figure 14 This is an illustration of a camera capable of operating to process audio based on a sound source representation, according to some examples of this disclosure.
[0031] Figure 15 This is an illustration of a headset (such as a virtual reality, mixed reality, or augmented reality headset) capable of operating to process audio based on a sound source representation, according to some examples of this disclosure.
[0032] Figure 16 This is a diagram illustrating a first example of a vehicle capable of operating to process audio based on a sound source representation, according to some examples of this disclosure.
[0033] Figure 17 This is a diagram illustrating a second example of a vehicle capable of operating to process audio based on a sound source representation, according to some examples of this disclosure.
[0034] Figure 18 Based on some examples of this disclosure, it is possible to... Figure 1 A diagram illustrating a specific implementation of a method for processing audio based on a sound source representation, performed by the device.
[0035] Figure 19 This is a block diagram of a particular exemplary example of a device capable of operating to process audio based on a sound source representation, according to some examples of this disclosure.
[0036] VI. Detailed Implementation
[0037] The input audio signal can include sounds from multiple sound sources. Only some of these sources may be of interest. For example, during a conference call, the voices of the participants may be of interest, while the voices of people speaking in the background, non-voice noise, etc., may be distracting. In another example, the audio signal includes the voices of one or more known users and the voice of an unknown person, and the voice of the unknown person is of interest. Retaining or removing the voices from selected sound sources can enhance the perceptibility of the sounds of interest.
[0038] Systems and methods for processing audio based on sound source representations are disclosed. For example, an input audio signal is processed based on sound representations of one or more sound sources to generate an output audio signal. Sound source representations are used to retain or remove the sounds of one or more sound sources from the input audio signal. For example, the sound source representation of a participant's voice in a teleconference can be used to retain the participant's voice in the input audio signal while retaining other sounds. Similarly, the sound source representation of a known user can be used to remove the voice of a known user from the input audio signal while removing other sounds, including the voice of an unknown user. Therefore, the perceptibility of sounds of interest (such as teleconference participants or unknown individuals) is enhanced in the output audio signal.
[0039] Specific aspects of this disclosure are described below with reference to the accompanying drawings. In this specification, common features are designated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular embodiments and are not intended to limit the scope of the embodiments. For example, the singular forms “a,” “an,” and “the” are intended to also include the plural forms unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some embodiments and plural in others. For example, Figure 1 It describes a system that includes one or more processors ( Figure 1 The device 102 (with the “processor 190”) indicates that in some embodiments, device 102 includes a single processor 190, while in other embodiments, device 102 includes multiple processors 190. For ease of reference herein, such a feature is generally introduced as “one or more” features and is referred to thereafter in the singular unless an aspect relating to multiple features of the feature is described.
[0040] As used herein, the term “comprising” may be used interchangeably with “including”. Additionally, the term “wherein” may be used interchangeably with “where”. As used herein, “exemplary” indicates an example, specific implementation, and / or aspect, and should not be construed as restrictive or indicating a preference or preferred implementation. As used herein, ordinal terms used to modify elements (such as structures, components, operations, etc.) (e.g., “first,” “second,” “third,” etc.) do not themselves indicate any priority or order of that element relative to another element, but merely distinguish that element from another element with the same name (but using ordinal terms). As used herein, the term “set” refers to one or more specific elements among specific elements, while the term “multiple” refers to multiple (e.g., two or more) specific elements.
[0041] As used herein, “coupling” can include “communicationally coupled,” “electrically coupled,” or “physically coupled,” and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof). As an illustrative, non-limiting example, two electrically coupled devices (or components) may be included in the same device or in different devices, and may be connected via electronics, one or more connectors, or inductive coupling. In some specific implementations, two communicationally coupled (e.g., electrically connected) devices (or components) may directly or indirectly transmit and receive signals (e.g., digital or analog signals) via one or more wires, buses, networks, etc. As used herein, “direct coupling” can include two devices coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) without intermediate components.
[0042] In this disclosure, terms such as “determine,” “calculate,” “estimate,” “shift,” and “adjust” can be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as restrictive, and similar operations can be performed using other techniques. Furthermore, as mentioned herein, “generate,” “calculate,” “estimate,” “use,” “select,” “access,” and “determine” are used interchangeably. For example, “generating,” “calculating,” “estimate,” or “determining” a parameter (or signal) can refer to actively generating, estimating, calculating, or determining the parameter (or signal), or it can refer to using, selecting, or accessing a parameter (or signal) that has already been generated (e.g., by another component or device).
[0043] refer to Figure 1 This document discloses specific exemplary aspects of a system configured to process audio based on a sound source representation, and designates it generally as 100. System 100 includes a device 102 coupled to one or more microphones 120, one or more speakers 160, or combinations thereof. Device 102 includes one or more processors 190 coupled to memory 132.
[0044] Memory 132 stores one or more sound source representations (SSRs) 154, such as sound source representation (SSR) 154A, sound source representation 154B, sound source representation 154C, one or more additional sound source representations, or combinations thereof. For example, sound source representation 154A represents the sound of sound source 184A (e.g., speech, non-speech sound, or both). In certain aspects, sound source representation 154A is based on the sound of sound source 184A or on the sound of a specific sound source of the same sound source type as sound source 184A, as referenced. Figure 2AFurther described.
[0045] In some aspects, memory 132 stores one or more combined sound source representations 147, such as combined sound source representation 147A, combined sound source representation 147B, combined sound source representation 147C, one or more additional combined sound source representations, or combinations thereof. For example, combined sound source representation 147A represents the sounds (e.g., speech, non-speech sounds, or both) of multiple sound sources (such as sound source 184A, sound source 184B, and sound source 184C). In a particular aspect, combined sound source representation 147A is based on sound source representations 154A, 154B, and 154C, respectively representing the sounds of sound source 184A, sound source 184B, and sound source 184C. In another aspect, combined sound source representation 147A is based on sounds from sound source 184A, sound source 184B, and sound source 184C, or sounds from sound sources of the same type as sound source 184A, sound source 184B, and sound source 184C, as referenced. Figures 2B to 2C Further described.
[0046] In some respects, a sound source representation (e.g., sound source representation 154A or combined sound source representation 147B) can represent the sound of a sound source in a specific environment, as referenced. Figure 2C Further described. For example, a combined sound source representation 147B may represent a sound (e.g., wind, road noise, etc.) captured in a specific type of vehicle (e.g., brand, model, sedan, SUV, van, electric vehicle, gasoline-powered vehicle, hybrid vehicle, etc.) while traveling on a highway. In some aspects, one or more sound source representations in sound source representation 154, one or more combined sound source representations 147, or both, are associated with metadata indicating the corresponding sound source type information (such as demographic information, object type information, vehicle type information, person identifier, sound source identifier, environmental information, etc.).
[0047] One or more processors 190 include an audio analyzer 140 configured to process audio based on a sound source representation. The audio analyzer 140 includes a configurator 144 coupled to an audio adjuster 148. The configurator 144 is configured to determine adjuster configuration settings 143 based on context 149.
[0048] Adjuster configuration setting 143 indicates the value of the retention flag 145 and indicates the selected sound source 162. For example, a first value of the retention flag 145 (e.g., 1) indicates that the sound of the selected sound source 162 will be retained. A second value of the retention flag 145 (e.g., 0) indicates that the sound of the selected sound source 162 will be removed.
[0049] Context 149 indicates that when the detected condition 159 matches the activation condition 139, the audio adjuster 148 will be activated using the adjuster configuration setting 143 determined based on the sound source selection criterion 157 and the reservation flag criterion 137. In some aspects, the activation condition 139, the sound source selection criterion 157, the reservation flag criterion 137, or a combination thereof, are based on a default configuration, user input 103, configuration input from an application, a configuration request from another device, or a combination thereof. The configurator 144 is configured to determine the selected sound source 162 based on the sound source selection criterion 157 and to determine the value of the reservation flag 145 based on the reservation flag criterion 137.
[0050] Additionally, in some aspects, the configurator 144 is configured to provide the audio adjuster 148 with a combined SSR 147A representing the sound of the selected sound source 162. In some examples, the configurator 144 is configured to provide the SSR generator 146 with one or more selected sound source representations 156 of the selected sound source 162. The SSR generator 146 is configured to generate the combined SSR 147A based on one or more selected sound source representations 156 and store the combined sound source representation 147A in memory 132. In these examples, the SSR generator 146 is configured to provide the combined SSR 147A to the audio adjuster 148.
[0051] Audio adjuster 148 is configured to use neural network 150 to process input audio signal 126 based on combined sound source representation 147A to generate output audio signal 135. For example, audio adjuster 148 is configured to use neural network 150 to generate mask 151, as referenced. Figure 3 Further described. The audio adjuster 148 generates an output audio signal 135 by applying a mask 151 to the input audio signal 126 based on the value of a retention flag 145 to retain or remove sounds from selected sound sources 162. By selectively retaining or removing sounds corresponding to the selected sound sources 162, the perceptibility of the sounds of interest is enhanced in the output audio signal 135. Compared to processing the input audio signal 126 using a combined sound source representation 147A that can represent multiple sound sources, generating the mask 151 can be faster and use less computation. In some aspects, the neural network 150 includes a convolutional neural network (CNN), an autoregressive (AR) generative network, an audio generation network (AGN), an attention network (AN), a long short-term memory (LSTM) network, or a combination thereof.
[0052] In some specific implementations, device 102 corresponds to or is included in one of various types of devices. In an exemplary example, one or more processors 190 are integrated into a headset device that includes one or more microphones 120, one or more speakers 160, or combinations thereof, such as reference emulators. Figure 11 Further described. In other examples, one or more processors 190 are integrated in at least one of the following: as described in the reference Figure 10 The mobile phone or tablet computer device mentioned above, as referenced Figure 12 The wearable electronic device, as referenced Figure 13 The aforementioned voice-controlled speaker system, as referenced Figure 14 The camera device mentioned above, or as referenced Figure 15 The aforementioned virtual reality, mixed reality, or augmented reality headset. In another illustrative example, one or more processors 190 are integrated into a vehicle, which also includes one or more microphones 120, one or more speakers 160, or combinations thereof, such as reference... Figure 16 and Figure 17 Further described.
[0053] During operation, configurator 144 determines a context 149 that indicates the audio adjuster 148 will be activated when the detected condition 159 matches activation condition 139 to retain or remove the sound of the selected sound source 162 from the input audio signal 126. Context 149 indicates that the sound source satisfying sound source selection criterion 157 will be the selected sound source 162, and the sound will be retained or removed based on retention flag criterion 137. Context 149 (e.g., activation condition 139, sound source selection criterion 157, and retention flag criterion 137) is based on default configuration, user input 103, configuration input from the application, configuration request from another device, or a combination thereof.
[0054] In some examples, user 101 provides user input 103 to audio analyzer 140. Audio analyzer 140 determines context 149 based at least in part on user input 103. In a first exemplary example, user input 103 corresponds to the creation or acceptance of a meeting invitation for a scheduled meeting. In one aspect, user input 103 indicates that the voices of one or more participants in the scheduled meeting will be preserved. In another aspect, a default configuration indicates that the voices of any participant in the scheduled meeting will be preserved. Configurator 144 generates context 149 in response to receiving user input 103 to indicate activation condition 139, which indicates that audio adjuster 148 will be activated to initiate processing of input audio signal 126 in response to the start of the scheduled meeting. Configurator 144 also generates context 149 to indicate sound source selection criteria 157 and reservation flag criteria 137, the sound source selection criteria indicating that the selected sound source 162 will correspond to a participant in the scheduled meeting, and the reservation flag criteria indicating that reservation flag 145 will have a first value (e.g., 1).
[0055] In the second exemplary example, user input 103 corresponds to a user selection in an audio processing application to activate audio adjuster 148 to remove the voices of known users (e.g., sound sources 184A, 184B, and 184C) from input audio signal 126. Configurator 144 generates context 149 in response to receiving user input 103 to indicate that the detected condition 159 (e.g., receiving user input 103) matches activation condition 139 of activating audio adjuster 148. Configurator 144 also updates context 149 to indicate that the selected sound sources 162 will include the sound source selection criterion 157 of known users (e.g., sound sources 184A, 184B, and 184C) and the retention flag criterion 137 will have a second value (e.g., 0). For ease of description, the operation of audio adjuster 148 is described herein with reference to the first and second exemplary examples. It should be understood that various other examples of the operation of audio adjuster 148 in retaining or removing sound are possible. For example, in one instance, user input 103 corresponds to the activation of a recording application. Configurator 144 generates context 149 in response to receiving user input 103 to indicate that the detected condition 159 (e.g., receiving user input 103) matches the activation condition 139 of the activated audio adjuster 148. Configurator 144 also updates context 149 to indicate that the selected sound source 162 will correspond to the sound source selection criterion 157 of the estimated recording target and the reservation flag 145 will have a first value (e.g., 1) of the reservation flag criterion 137.
[0056] When activated, audio adjuster 148 initiates processing of input audio signal 126. For example, audio adjuster 148 uses the sound source representation of selected sound source 162 to retain or remove sound from input audio signal 126, as further described below. In some specific implementations, prior to activation of audio adjuster 148, configurator 144 retrieves the sound source representation of at least one sound source that may be included in selected sound source 162 (e.g., a sound source expected to satisfy sound source selection criterion 157). For example, in response to determining that selected sound source 162 may include sound source 184A, configurator 144 retrieves the sound source representation 154A (e.g., representing speech) of sound source 184A from a server and stores the sound source representation 154A locally at device 102 (e.g., in device memory 132). In the first exemplary example, configurator 144 determines, in response to determining that a meeting invitation has been sent to audio source 184A for the scheduled meeting, that audio source 184A may satisfy audio source selection criterion 157 (e.g., may be included in the selected audio source 162). In the second exemplary example, configurator 144 determines, in response to determining that a known user includes audio source 184A, that audio source 184A may satisfy audio source selection criterion 157 (e.g., may be included in the selected audio source 162).
[0057] In some implementations, prior to the activation of the audio adjuster 144, the configurator 144 generates (or updates) the SSR of at least one sound source that may be included in the selected sound sources 162. For example, in response to determining that the selected sound sources 162 may include sound source 184B, the configurator 144 generates a sound source representation 154B based on the input audio signal representing the sound of sound source 184B, as referenced. Figure 2A Further described. For example, receiving an input audio signal during a teleconference between user 101 and audio source 184B.
[0058] In some implementations, the configurator 144 generates (or updates) the SSR of one or more sound sources 184 (e.g., at least one of the selected sound sources 162) when the audio adjuster 148 is activated. For example, the configurator 144 generates (or updates) a sound source representation 154B based on a portion of the input audio signal 126 that corresponds to a single speaker and that the single speaker is sound source 184B, as referenced. Figure 2A Further described.
[0059] The input audio signal 126 corresponds to the sound 186 of a sound source 184 (such as sound source 184A, sound source 184B, sound source 184C, sound source 184D, sound source 184E, one or more additional sound sources, or combinations thereof). The sound source 184 may include one or more of the following: a vehicle, an emergency vehicle, traffic, wind, echo, channel distortion, birds, animals, an alarm, another non-voice sound source, a person, an authorized user, another voice source, or an audio player.
[0060] Audio analyzer 140 receives input audio signal 126 from another device (e.g., a server or storage device), one or more microphones 120, memory 132, or a combination thereof. In a first exemplary example, audio analyzer 140 receives input audio signal 126 from a server during a scheduled meeting. For example, the input audio signal 126 is based on audio data received from another device (e.g., a server), as referenced... Figure 6 Further described. In a second exemplary example, the audio analyzer 140 receives an input audio signal 126 from one or more microphones 120. For example, the one or more microphones 120 capture sound 186, and the input audio signal 126 is based on the microphone output of the one or more microphones 120.
[0061] In some examples, the input audio signal 126 is based on audio data retrieved from memory 132. For example, the input audio signal 126 is based on audio data generated by an application of one or more processors 190, such as a music application, game application, graphics application, augmented reality application, communication application, entertainment application, or a combination thereof. In some examples, the audio adjuster 148 processes the input audio signal 126 in real time when the audio analyzer 140 receives it. In other examples, the input audio signal 126 corresponds to a previously generated audio signal.
[0062] Configurator 144 determines that context 149 indicates that when detected condition 159 matches activation condition 139, audio adjuster 148 will be activated with adjuster configuration settings 143 based on sound source selection criterion 157 and retention flag criterion 137. In a first exemplary example, configurator 144 activates audio adjuster 148 in response to determining that detected condition 159 (e.g., the start of a scheduled meeting) matches activation condition 139 (e.g., the start of a scheduled meeting). Configurator 144 determines adjuster configuration settings 143 based on context 149. For example, configurator 144, in response to determining that sound source selection criterion 157 indicates that the selected sound source 162 will correspond to a participant in the scheduled meeting and retention flag criterion 137 indicates that retention flag 145 will have a first value (e.g., 1), designates a participant (e.g., expected participant, detected participant, or a combination thereof) as the selected sound source 162 and sets retention flag 145 to have a first value (e.g., 1) indicating that the corresponding sound will be retained.
[0063] The selected sound source 162 can be determined statically (e.g., by anticipated participants), dynamically (e.g., by detected participants), or in both ways using sound source selection criterion 157. In the first exemplary example, the static determination of the selected sound source 162 may include anticipated participants of the scheduled meeting, and detected additional participants (e.g., participants who join the meeting, even if they were not initially invited) may be dynamically added to the selected sound source 162.
[0064] In the second exemplary example, configurator 144 activates audio adjuster 148 in response to determining that context 149 indicates that detected condition 159 (e.g., receiving user input 103 indicating user selection in an audio processing application) matches activation condition 139. Configurator 144 determines adjuster configuration settings 143 based on context 149. For example, configurator 144 activates audio adjuster 148 in response to determining that sound source selection criterion 157 indicates that the selected sound source 162 will correspond to a known user (e.g., sound source 184A, sound source 184B, and sound source 184C) and reservation flag criterion 137 indicates that reservation flag 145 will have a second value (e.g., 0), designates the known user as the selected sound source 162, and sets reservation flag 145 to have a second value (e.g., 0) indicating that the corresponding sound will be removed.
[0065] The selected sound source 162 can be determined statically (e.g., all known users among known users), dynamically (e.g., detected known users among known users), or in both ways using sound source selection criterion 157. In the second exemplary example, the static determination of the selected sound source 162 may include all known users among known users, and one or more known users (e.g., whose voices were not detected for a threshold duration) may be dynamically removed from the selected sound source 162.
[0066] The value of the retention flag 145 can be determined statically or dynamically using retention flag criterion 137. In a first exemplary example, retention flag criterion 137 corresponds to the static determination of the value of retention flag 145 (e.g., retaining a first value of the sound). In a second exemplary example, retention flag criterion 137 corresponds to the static determination of the value of retention flag 145 (e.g., removing a second value of the sound). In various aspects, the value of retention flag 145 can be determined dynamically using retention flag criterion 137. In one example, user input 103 corresponds to the activation of a recording application. Configurator 144 generates context 149 in response to receiving user input 103 to indicate that the detected condition 159 (e.g., receiving user input 103) matches the activation condition 139 of activating audio adjuster 148. Configurator 144 also updates context 149 to indicate sound source selection criterion 157, i.e., if the sound source representation of the estimated target is available, the selected sound source 162 will correspond to one or more estimated targets of the recording. Sound source selection criterion 157 indicates that, if the sound source representation of the estimated target is unavailable, and if the sound source representation of the interfering sound source is available, the selected sound source 162 will correspond to one or more interfering sound sources. Retention mark criterion 137 indicates that, if the selected sound source 162 corresponds to the estimated target, retention mark 145 will have a first value for retaining the sound (e.g., 1). Retention mark criterion 137 also indicates that, if the selected sound source 162 corresponds to an interfering sound source, retention mark 145 will have a second value for removing the sound (e.g., 0). In this example, sound source selection criterion 157 is used to dynamically determine the selected sound source 162, and retention mark criterion 137 is used to dynamically determine the value of retention mark 145.
[0067] When the audio adjuster 148 is activated, the configurator 144 concurrently provides the audio adjuster 148 with the value of the reserved flag 145 and a combined sound source representation 147A of the selected sound source 162. In some examples, the configurator 144 has access to the combined sound source representation 147A and provides it to the audio adjuster 148 (e.g., bypassing the SSR generator 146). In other examples, the configurator 144 provides the SSR generator 146 with one or more selected sound source representations 156 of the selected sound source 162, and the SSR generator 146 provides the combined sound source representation 147A to the audio adjuster 148.
[0068] In some implementations, the adjuster configuration setting 143 includes a combo setting 164. A first value (e.g., 0) of the combo setting 164 indicates that the configurator 144 will provide the SSR generator 146 with individual sound source representations of the selected sound sources 162, regardless of whether a combined sound source representation of two or more of the selected sound sources 162 is available. A second value (e.g., 1) of the combo setting 164 indicates that the configurator 144 will bypass the SSR generator 146 and provide the audio adjuster 148 with a combined sound source representation of all the selected sound sources 162 when a combined sound source representation is available.
[0069] In some implementations, the combination setting 164 is based on context 149. For example, context 149 includes combination criteria 141 for determining the combination setting 164. Combination criteria 141 is based on a default configuration, user input 103, configuration input from an application, a configuration request from another device, or a combination thereof. Combination criteria 141 may indicate that the combination setting 164 will have a specific value (e.g., 0 or 1). In certain aspects, combination criteria 141 may be used to determine the combination setting 164 statically, dynamically, or both. For example, the static determination of the combination setting 164 may have a first value (e.g., 0) that the configurator 144 can dynamically update to a second value (e.g., 1) in response to a detected combination condition (e.g., remaining battery life is less than a threshold).
[0070] As an illustrative example, the selected sound sources 162 include sound sources 184A, 184B, and 184C. The configurator 144, in response to determining that the combination setting 164 has a second value (e.g., 1), determines whether a combined sound source representation representing all selected sound sources in the selected sound sources 162 is available (e.g., in memory 132 or another device). The configurator 144, in response to determining that the sound source type information of the combined sound source representation 147A matches the sound source type information of each selected sound source in the selected sound sources 162, determines that the combined sound source representation 147A represents all selected sound sources in the selected sound sources 162, and bypasses the SSR generator 146 to provide the combined sound source representation 147A to the audio adjuster 148.
[0071] Alternatively, configurator 144, in response to determining that a combined sound source representation of all selected sound sources in the selected sound sources 162 is unavailable, determines whether an SSR for a plurality of selected sound sources in the selected sound sources 162 is available (e.g., in memory 132 or another device). Configurator 144, in response to determining that combined setting 164 has a second value (e.g., 1) and that a combined sound source representation 147B of the sounds of sound source representations 154B and 154C is available, adds combined sound source representation 147B to one or more selected sound source representations 156.
[0072] Configurator 144 determines whether a separate SSR for sound source 184A is available in response to determining that combination setting 164 has a first value (e.g., 0) or that sound source 184A of selected sound source 162 is not represented by an SSR included in any combination of one or more selected sound source representations 156. For example, configurator 144 selects sound source representation 154A to represent sound source 184A in response to determining that first sound source type information of sound source representation 154A matches second sound source type information of sound source 184A, and adds sound source representation 154A to one or more selected sound source representations 156. Configurator 144 provides one or more selected sound source representations 156 to SSR generator 146.
[0073] SSR generator 146 generates a combined sound source representation 147A based on one or more selected sound source representations 156. In some embodiments, the SSR includes feature values representing audio features of the sound source (e.g., short-term spectral features, speech source features, spectral temporal features, prosodic features, high-level features, or combinations thereof). For example, sound source representation 154A includes a first feature value of a first audio feature, sound source representation 154B includes a second feature value of the first audio feature, and sound source representation 154C includes a third feature value of the first audio feature. In some embodiments, the SSR indicates feature values of 512 audio features. In an exemplary example, the SSR is represented by a multi-dimensional (e.g., 512-dimensional) vector.
[0074] The combined sound source representation 147B (representing the sounds of sound sources 184B and 184C) includes a first specific feature value of a first audio feature, which is based on a second feature value and a third feature value. In a particular embodiment, the combined sound source representation 147B corresponds to a sound source representation 154B cascaded with a sound source representation 154C, the average value of sound source representation 154B and sound source representation 154C, or both. In this example, the first specific feature value corresponds to a cascade of the second and third feature values (or a list including the second and third feature values), the average value of the second and third feature values, or both. The combined sound source representation 147A includes a second specific feature value of the first audio feature, which is based on the first feature value (e.g., the first feature value of sound source representation 154A) and the first specific feature value (e.g., the first specific feature value of combined sound source representation 147B) (e.g., a list, an average value, or both).
[0075] One or more selected sound source representations 156, including sound source representation 154A and combined sound source representation 147B, are provided as exemplary examples. In other examples, one or more selected sound source representations 156 may include an independent SSR for each of the selected sound sources 162 or multiple combined SSRs for various combinations of the selected sound sources 162. For example, one or more selected sound source representations 156 may include sound source representation 154A representing the sound of sound source 184A, sound source representation 154B representing the sound of sound source 184B, and sound source representation 154C representing the sound of sound source 184C. As another example, one or more selected sound source representations 156 may include a combined sound source representation 147B representing the sound of sound sources 184B and 184C, and a combined sound source representation 147C representing the sound of sound sources 184A and 184C.
[0076] In some implementations, configurator 144 dynamically updates adjuster configuration settings 143 and combined sound source representation 147A while audio adjuster 148 processes input audio signal 126. For example, configurator 144 dynamically updates selected sound source 162 (and the combined sound source representation 147A provided to audio adjuster 148) in response to determining sound source selection criterion 157, sound source 184 satisfying sound source selection criterion 157, or both having changed, context 149 indicating that selected sound source 162 will correspond to a participant in the scheduled meeting, and detecting an update of the participant (e.g., due to someone leaving or joining a conference call). In a first exemplary example, configurator 144 dynamically updates selected sound source 162 to include the detected participant (and updates the combined sound source representation 147A provided to audio adjuster 148) in response to determining that sound source selection criterion 157 indicates that selected sound source 162 will correspond to a detected participant in the scheduled meeting, and detecting an update of the participant (e.g., due to someone leaving or joining a conference call). In the second exemplary example, the configurator 144 dynamically removes a known user from the selected sound source 162 (and updates the combined sound source representation 147A provided to the audio adjuster 148) in response to determining that the context 149 indicates that the selected sound source 162 corresponds to a known user (e.g., sound source 184A, sound source 184B, and sound source 184C) and detecting the voice of a known user (e.g., sound source 184A) that was not detected in the input audio signal 126 within a threshold time. If the voice of a known user is later detected in the input audio signal 126, the configurator 144 may subsequently add the removed known user to the selected sound source 162 (and update the combined sound source representation 147A provided to the audio adjuster 148).
[0077] When activated, the audio adjuster 148 processes the input audio signal 126 based on the values of the combined sound source representation 147A and the retention flag 145 to retain or remove sound to generate the output audio signal 135. For example, the audio adjuster 148 uses a neural network 150 to process the input audio signal 126 based on one or more combined sound source representations 147 to generate a mask 151, as referenced. Figure 3 Further described.
[0078] Audio adjuster 148, in response to a retention flag 145 having a first value (e.g., 1), applies a filter corresponding to mask 151 to the input audio signal 126 to retain the sound of the selected sound source 162 and remove the remaining sound. In a first exemplary example, audio adjuster 148 retains the sounds of the participants in the scheduled meeting (e.g., sound sources 184A, 184B, and 184C) in the input audio signal 126 to generate an output audio signal 135. Remaining sounds (such as sound source 184D (e.g., speech noise, such as non-participant speech in the background) and sound source 184E (e.g., non-speech noise, such as passing traffic)) are not included (or are reduced) in the output audio signal 135. Alternatively, audio adjuster 148, in response to a retention flag 145 having a second value (e.g., 0), applies a filter corresponding to the inverse of mask 151 to the input audio signal 126 to remove the sound of the selected sound source 162 and retain the remaining sound. In a second exemplary example, audio adjuster 148 removes the voices of known users (e.g., sound sources 184A, 184B, and 184C) from the input audio signal 126 to generate an output audio signal 135. The remaining sounds (such as sound source 184D (e.g., the voice of an unknown person) and sound source 184E (e.g., non-voice sounds, such as emergency vehicles) are included in (or relatively enhanced) the output audio signal 135.
[0079] Compared to filtering the sound based on individual sound source representations for each of the selected sound sources 162, filtering the sound of the input audio signal 126 using a mask 151 based on a combined sound source representation 147A improves efficiency (e.g., reduces computation). In some respects, this increased efficiency enables real-time processing of the input audio signal 126. Another advantage of using a mask 151 based on a combined sound source representation 147A to filter the sound of the input audio signal 126 is improved accuracy in filtering portions of the input audio signal 126 that include overlapping sounds from multiple selected sound sources 162, compared to filtering sequentially based on individual sound source representations.
[0080] Audio analyzer 140 provides an output audio signal 135 to one or more speakers 160. The one or more speakers 160 output sound 196 based on the output audio signal 135. The sound of interest is more perceptible in sound 196 compared to sound 186 from sound source 184. In some examples, audio analyzer 140 provides audio data based on the output audio signal 135 to another device, as shown in the reference. Figure 7Further described. In some examples, the audio analyzer 140 stores audio data based on the output audio signal 135 in memory 132, another storage device, or both.
[0081] System 100 thus enhances the perception of the sound of interest in the output audio signal 135 by retaining the sound of interest or removing the remaining sound. Using a combined sound source representation 147A representing the sounds of the selected sound source 162 can improve the efficiency and accuracy of processing the input audio signal 126 to generate the output audio signal 135.
[0082] Although one or more microphones 120, one or more speakers 160, or combinations thereof are shown coupled to device 102, in other embodiments, one or more microphones 120, one or more speakers 160, or combinations thereof may be integrated into device 102. Although device 102 is shown to include configurator 144, SSR generator 146, and audio adjuster 148, in other embodiments, configurator 144, SSR generator 146, and audio adjuster 148 may be included in two or more separate devices. As an illustrative example, configurator 144 and SSR generator 146 may be included in a user device, and audio adjuster 148 may be included in a headset.
[0083] refer to Figure 2A Figure 200 illustrates an exemplary aspect of a sound source encoder 202 capable of operating to generate a sound source representation 154. In some specific embodiments, Figure 1 Device 102 includes a sound source encoder 202. In other embodiments, the sound source encoder 202 is included in a second device that provides a sound source representation 154 to device 102.
[0084] Input audio signal 226 represents sound 286 from sound source 284. In some aspects, input audio signal 226 is based on the microphone output of one or more microphones capturing sound 286. In other aspects, input audio signal 226 is based on audio data received from another device, audio data generated by an application, or a combination thereof.
[0085] The sound source encoder 202 uses various techniques to process the input audio signal 226 to generate a sound source representation 154 of sound 286. In a particular aspect, the sound source encoder 202 performs a Fast Fourier Transform (FFT) on a portion (e.g., a 20-millisecond time window) of the input audio signal 226 to determine the time-frequency information of the input audio signal 226. In this aspect, the sound source representation 154 (e.g., a spectrogram) is based on a first FFT feature associated with a first time window, a second FFT feature associated with a second time window, and so on. In some specific implementations, the sound source representation 154 includes the time-domain envelope of the input audio signal 226, the frequency composition of the input audio signal 226, a time-frequency representation of the input audio signal 226 (e.g., one or more spectrograms, one or more cochlear maps, one or more correlation maps, etc.) or a combination thereof.
[0086] The sound source encoder 202 also generates metadata for the sound source representation 154, which includes sound source type information for the sound source 284 (e.g., demographic information, object type information, vehicle type information, person identifier, sound source identifier, environmental information, etc.). Figure 1 The configurator 144 determines that sound source representation 154 represents the sound of sound source 184 in response to determining that sound source 184 and sound source 284 are of the same sound source type. The configurator 144 determines that sound source 184 and sound source 284 are of the same sound source type in response to determining that the sound source type information associated with sound source representation 154 matches the sound source type information of sound source 184.
[0087] In some implementations, sound source type information is sorted by matching priority. In response to selecting multiple sound source representations with corresponding sound source type information that matches the sound source type information of sound source 184, configurator 144 selects the sound source representation with the matching sound source type information and the highest matching priority to represent sound source 184. For example, a sound source identifier has a higher matching priority than other types of sound source type information (such as demographic information). For instance, sound source representation 154 with the same source identifier as sound source 184 has a higher priority than sound source representations with other types of matching type information (e.g., it's a closer match) because the matching sound source identifier indicates that sound source 184 is the same as sound source 284.
[0088] In some respects, although sound source 284 is not the same as sound source 184, sound source 284 is of the same type as sound source 184. In a particular example, sound source 184 is a first person and sound source 284 is a second person having one or more second demographic characteristics (e.g., age, location, race, ethnicity, sex, or combinations thereof) that match one or more first demographic characteristics of the first person. In some examples, sound source 284 is an object of the same type as sound source 184 (e.g., a vehicle, a bird, or broken glass). As another example, sound source 284 is a vehicle of the same type as sound source 184 (e.g., an ambulance, fire truck, police car, or airplane). In one example, sound source 284 is an alarm of the same type (e.g., a fire alarm, a manufacturing alarm, a smoke detection alarm, etc.).
[0089] In some examples, configurator 144 updates sound source representation 154 based on the sound of sound source 184 in response to determining that sound source 184 is the same as sound source 284. In some examples, configurator 144 updates sound source representation 154 based on the sound of sound source 184 in response to determining that although sound source 184 is different from sound source 284, sound source 184 and sound source 284 are of the same sound source type.
[0090] refer to Figure 2B Figure 230 shows an exemplary aspect of a sound source encoder 202 that is capable of operating to generate a combined sound source representation 147.
[0091] The input audio signal 236 represents sound 296 from multiple sound sources, such as sound source 284A and sound source 284B. In some aspects, the input audio signal 236 is based on the microphone output of one or more microphones capturing sound 296. In some aspects, the input audio signal 236 is based on audio data received from another device, audio data generated by an application, or a combination thereof.
[0092] The sound source encoder 202 uses various techniques to process the input audio signal 236 to generate a combined sound source representation 147 of sound 296. For example, the combined sound source representation 147 includes the time-domain envelope of the input audio signal 236, the frequency composition of the input audio signal 236, the time-frequency representation of the input audio signal 236 (e.g., spectrogram, cochlear graph, correlogram, etc.), or a combination thereof.
[0093] The sound source encoder 202 also generates metadata for the combined sound source representation 147, which includes a first sound source type information set (e.g., first sound source type information of sound source 284A and second sound source type information of sound source 284B). Figure 1The configurator 144 determines a combined sound source representation 147 representing the sounds of the multiple sound sources (e.g., sound source 184A and sound source 184B) in response to determining that there is a one-to-one match between each first sound source type information set in a first sound source type information set and each second sound source type information set in a second sound source type information set. For example, Figure 1 In response to determining that sound source 284A and sound source 184A are of the same sound source type and that sound source 284B and sound source 184B are of the same sound source type, the configurator 144 determines a combined sound source representation 147 representing the sounds of sound source 184A and sound source 184B.
[0094] In some examples, configurator 144 updates the combined sound source representation 147 based on the sound of any of a plurality of sound sources (e.g., sound source 184A and sound source 184B). For example, configurator 144 updates the combined sound source representation 147 based on the sound of sound sources 184A and 184B in response to determining that sound sources 284A and 184B are the same as sound sources 184A and 184B, respectively. In another example, configurator 144 updates the combined sound source representation 147 based on the sound of sound sources 184A and 184B in response to determining that although sound sources 184A and 184B are different from sound sources 284A and 284B, respectively, sound sources 184A and 184B are of the same sound source type as sound sources 284A and 284B, respectively.
[0095] refer to Figure 2C Figure 240 shows an exemplary aspect of a sound source encoder 202 that is operable to generate a combined sound source representation 147 representing the sound of environment 204.
[0096] Input audio signal 246 represents sound 298 of environment 204. Environment 204 corresponds to multiple sound sources 284 that may include various elements affecting sound 298. For example, environment 204 may include the interior of a particular vehicle during vehicle operation. In this example, environment 204 may correspond to sound source 284, which includes elements such as wind, traffic, tires, roads, operating conditions (e.g., speed, half-open windows, fully open windows, etc.), interior shape, exterior shape, other acoustic characteristics, or combinations thereof. As another example, environment 204 may correspond to the interior of a particular manufacturing facility. In this example, environment 204 may correspond to sound source 284, which includes elements such as machine noise, specific alarms, windows, doors, operating conditions (e.g., machine speed, open windows, closed windows, open doors, closed doors), facility interior shape, facility exterior shape, other acoustic characteristics, or combinations thereof.
[0097] Vehicles and manufacturing facilities are provided as exemplary, non-limiting examples of environment 204. Environment 204 may include a variety of other types of environments, such as an aircraft environment, a stock exchange environment, other types of indoor environments, a beach, a concert hall, a shopping mall, other types of outdoor environments, virtual environments, augmented environments, etc. In some examples, sound source 284 corresponds to background noise in environment 204 (e.g., sound 298).
[0098] In some respects, the input audio signal 246 is based on the microphone output of one or more microphones 206 that capture sound 298. In other respects, the input audio signal 246 is based on audio data received from another device, audio data generated by an application, or a combination thereof.
[0099] The sound source encoder 202 uses various techniques to process the input audio signal 246 to generate a combined sound source representation 147 of sound 298. For example, the combined sound source representation 147 includes the time-domain envelope of the input audio signal 246, the frequency composition of the input audio signal 246, the time-frequency representation of the input audio signal 246 (e.g., spectrogram, cochlear graph, correlogram, etc.), or a combination thereof.
[0100] The sound source encoder 202 also generates metadata for the combined sound source representation 147, which includes sound source type information of the environment 204 (e.g., environment identifier, environment type, vehicle type, tire type, operating conditions, internal shape type, external shape type, building type, location type, operating status, event type, or a combination thereof). Figure 1 The configurator 144 determines that the combined sound source representation 147 represents sound source 184 (e.g., an environment with one or more matching elements of environment 204). For example, Figure 1 The configurator 144 can determine that the combined sound source representation 147 represents a first environment (e.g., a vehicle) with second sound source type information matching environment 204 (e.g., a second vehicle).
[0101] In some examples, configurator 144 updates the combined sound source representation 147 based on the sound of sound source 184 (e.g., a first environment) in response to determining that sound source 184 is the same as sound source 284. In some examples, configurator 144 updates the combined sound source representation 147 based on the sound of sound source 184 (e.g., a first environment) in response to determining that although sound source 184 is different from sound source 284 (e.g., environment 204), sound source 184 and sound source 284 are of the same sound source type (e.g., the same type of vehicle).
[0102] refer to Figure 3This illustrates an exemplary example of a neural network 150. Figure 3 In the example shown, neural network 150 includes a feature extractor 350 coupled to CNN 352. CNN 352 is coupled to an initial LSTM network layer, such as LSTM354A. The output of LSTM 354A is coupled to the input of a fully connected (FC) layer 356.
[0103] In some implementations, the output of LSTM 354A is coupled to the input of fully connected layer 356 independently (e.g., without requiring) any additional intermediary LSTM network layer. In some implementations, the output of LSTM 354A is coupled to the input of fully connected layer 356 via one or more LSTM combiner layers. The LSTM combiner layer includes an LSTM network layer coupled to the combiner. For example, a first LSTM combiner layer includes an LSTM 354B coupled to combiner 384A, and a second LSTM combiner layer includes an LSTM 354C coupled to combiner 384B. The output of LSTM 354A is coupled to the input of the LSTM network layer of the initial LSTM combiner layer. For example, the output of LSTM 354A is coupled to the input of LSTM 354B.
[0104] The combiner in the LSTM combiner layer combines the input of the LSTM network layer with the output of the LSTM network layer to generate the combiner's output. For example, combiner 384A combines the input of LSTM 354B (e.g., the output of LSTM 354A) and the output of LSTM 354B to generate the output. Each subsequent LSTM combiner layer receives the output of the previous LSTM combiner layer. For example, LSTM 354C receives the output of combiner 384A. Combiner 384B combines the input of LSTM 354C (e.g., the output of combiner 384A) and the output of LSTM 354C to generate the output of combiner 384B.
[0105] Fully connected layer 356 processes the output of the last LSTM combiner layer to generate the output of neural network 150. For example, fully connected layer 356 processes the output of combiner 384B to generate the output of neural network 150.
[0106] During operation, feature extractor 350 receives a first portion of the input audio signal 126 (e.g., one or more audio frames of the input audio signal). Feature extractor 350 extracts features 351 from the first portion of the input audio signal 126. In an exemplary example, feature extractor 350 generates a spectrogram of the first portion of the input audio signal 126. In a particular aspect, feature extractor 350 performs an FFT on a sub-part of the first portion (e.g., corresponding to a 20-millisecond time window) to determine the time-frequency information of the first portion. In this aspect, feature 351 is based on a first FFT feature associated with a first time window, a second FFT feature associated with a second time window, and so on. In some aspects, feature 351 includes representations of... Figure 1 The short-term spectral features, speech source features, spectral temporal features, prosodic features, high-level features, or combinations thereof of the sound 186.
[0107] CNN 352 processes features 351 to generate convolutional features 353. CNN 352 considers temporal dependencies on the temporally ordered portion of the input audio signal 126. For example, CNN 352 includes one or more convolutional layers that apply weights to the feature set received from feature extractor 350 to generate convolutional features 353. In some specific implementations, higher weights are applied to the most recently received feature set (e.g., feature 351) to generate convolutional features 353.
[0108] In certain aspects, CNN 352 includes a one-dimensional CNN (e.g., 1D CNN) or a two-dimensional CNN (e.g., 2D CNN). For example, a 1D CNN reduces the latency between receiving a portion of the input audio signal 126, generating a mask 151, and outputting a portion of the output audio signal 135 using the mask 151. In some real-time low-latency examples (e.g., voice calls), CNN 352 includes a 1D CNN. In some aspects, a 2D CNN corresponds to improved accuracy of the output audio signal 135 in preserving or removing sounds corresponding to the selected sound source 162 from the input audio signal 126. In some high-latency examples (e.g., voice user interfaces), CNN 352 includes a 2D CNN.
[0109] The LSTM 354A processes convolutional features 353 and combined sound source representation 147A. In some examples, the input to the LSTM 354A corresponds to convolutional features 353 concatenated with the combined sound source representation 147A. In some specific implementations, the output of the LSTM 354A indicates one or more features in the convolutional features 353 that have feature values that match the corresponding feature values of one or more features in the combined sound source representation 147A. The output of the LSTM 354A is fed to a fully connected layer 356 via any LSTM combiner layer to generate a mask 151.
[0110] refer to Figure 4 Figure 400 illustrates an exemplary aspect of the operation associated with the joint training of the neural network 150 and the sound source encoder 202.
[0111] Audio combiner 450 is coupled to sound source encoder 202 via audio adjuster 148. Sound source encoder 202 is coupled to speaker detector 472. Network trainer 462 is coupled to audio adjuster 148, and classification network trainer 482 is coupled to sound source encoder 202 and speaker detector 472.
[0112] During operation, audio combiner 450 receives a speech audio signal 424A from sound source 484A, an interfering speech audio signal 424B from sound source 484B, and a noise audio signal 424C from sound source 484C. Audio combiner 450 generates an input audio signal 426 based on a combination of the speech audio signal 424A, the interfering speech audio signal 424B, the noise audio signal 424C, one or more additional audio signals, or combinations thereof. In some aspects, audio combiner 450 performs channel distortion enhancement 452, reverberation enhancement 454, or both to update the input audio signal 426. In some aspects, the input audio signal 426 approximates various noise conditions that may be present when the speech audio signal of interest is received (e.g., interfering speech, noise, channel distortion, reverberation, or combinations thereof).
[0113] Audio combiner 450 provides input audio signal 426 to audio adjuster 148. Audio adjuster 148 receives sound source representation 154 representing sound source 484A. Neural network 150 performs operations as described in reference... Figure 3 The described similar operation processes the first input portion of the input audio signal 426 and the sound source representation 154 to generate a mask 451, wherein the input audio signal 426, the sound source representation 154, and the mask 451 respectively correspond to Figure 3 The input audio signal 126, the combined sound source representation 147A, and the mask 151 are used. The audio adjuster 148 applies the mask 451 to the first input section to generate the first output section of the output audio signal 435. The audio adjuster 148 applies the mask 451 to retain the sound of the sound source 484A from the first input section and remove (or reduce) the remaining sound of the sound source 484B, the sound source 484C, one or more additional audio signals, channel distortion enhancement 452, reverberation enhancement 454, or combinations thereof to generate the first output section.
[0114] Audio adjuster 148 provides a first output portion of the output audio signal 435 to sound source encoder 202 and network trainer 462. Sound source encoder 202 processes the first output portion (e.g., corresponding to the sound of sound source 484A retained from the first input portion) to generate an updated version of sound source representation 154 of sound source 484A. Sound source encoder 202 provides the updated version of sound source representation 154 to audio adjuster 148 to process a second input portion of the input audio signal 426. Sound source encoder 202 also provides the updated version of sound source representation 154 to speaker detector 472.
[0115] The speaker detector 472 uses the classification network 474 to process the updated version of the sound source representation 154 to generate an estimated speaker identifier 475. The speaker detector 472 provides the estimated speaker identifier 475 to the classification network trainer 482.
[0116] Network trainer 462 generates a noise reduction loss metric 464 based on a comparison between a first output portion of the output audio signal 435 and a corresponding first speech portion of the speech audio signal 424A. The first output portion corresponds to the sound of sound source 484A retained from the input audio signal 426, and the first speech portion corresponds to the original sound of sound source 484A. Network trainer 462 generates update data 463 based on the noise reduction loss metric 464. For example, update data 463 indicates updates to the weights, biases, or combinations thereof of the neural network 150 to reduce the noise reduction loss metric 464 in subsequent iterations. Network trainer 462 thus trains neural network 150 over time to improve noise reduction and reduce the difference between the output audio signal 435 and the speech audio signal 424A.
[0117] The classification network trainer 482 determines a classification loss metric 486 based on the estimated speaker identifier 475 and the speaker identifier 481 of the sound source 484A. For example, the classification network trainer 482 retrieves a first sound source representation of the sound source (e.g., a person) associated with the estimated speaker identifier 475, retrieves a second sound source representation of the sound source 484A, and generates the classification loss metric 486 based on a comparison of the first and second sound source representations. In a particular aspect, the second sound source representation was previously generated based on the sound of the sound source 484A.
[0118] The classification network trainer 482 generates updated data 483 and updated data 485 based on the classification loss metric 486. For example, updated data 483 indicates updates to the weights, bias values, or combinations thereof of the sound source encoder 202 to reduce the classification loss metric 486. The classification network trainer 482 thus trains the sound source encoder 202 over time to generate a sound source representation 154 that more closely matches the second sound source representation. Similarly, updated data 485 indicates updates to the weights, bias values, or combinations thereof of the classification network 474 to reduce the classification loss metric 486. The classification network trainer 482 thus trains the classification network 474 over time to generate an estimated speaker identifier 475 that corresponds to a speaker with a sound source representation that more closely matches the second sound source representation. The network trainer 462 and the classification network trainer 482 thus jointly train the neural network 150, the sound source encoder 202, and the classification network 474.
[0119] In some implementations, neural network 150, sound source encoder 202, classification network 474, or combinations thereof, are trained during the training phase and designated for availability upon completion of training. For example, network trainer 462 determines that training of neural network 150 is complete in response to determining that noise reduction loss metric 464 satisfies a training criterion (e.g., less than a loss threshold). Similarly, classification network trainer 482 determines that training of sound source encoder 202 and classification network 474 is complete in response to determining that classification loss metric 486 satisfies a training criterion (e.g., less than a loss threshold).
[0120] In some respects, the audio adjuster 148 may be available when the network trainer 462 determines that the training of the neural network 150 is complete. In some respects, the sound source encoder 202 and the classification network 474 may be available when the classification network trainer 482 determines that the training of the sound source encoder 202 and the classification network 474 is complete.
[0121] In some specific implementations, the audio combiner 450, the network trainer 462, the speaker detector 472, the classification network trainer 482, or a combination thereof are integrated into Figure 1 In device 102. In other embodiments, audio combiner 450, network trainer 462, classification network trainer 482, or combinations thereof are integrated into a second device. In these embodiments, the second device provides neural network 150, audio combiner 148, sound source encoder 202, speaker detector 472, classification network 474, or combinations thereof to device 102 in response to determining that audio adjuster 148, sound source encoder 202, classification network 474, or combinations thereof are available.
[0122] In some implementations, the neural network 150, the sound source encoder 202, the classification network 474, or a combination thereof, are trained (e.g., dynamically updated) during use. In these implementations, the input audio signal 426 corresponds to... Figure 1 The input audio signal 426 is used as the input audio signal. In an exemplary example, network trainer 462 receives a speech audio signal 424A from a first microphone closer to the sound source 484A (e.g., a known user with speaker identifier 481) and an input audio signal 426 from one or more second microphones. Network trainer 462 trains neural network 150 such that output audio signal 435 matches speech audio signal 424A (e.g., noise reduction loss metric 464 is less than a loss threshold), and uses output audio signal 435 to generate sound source representation 154 of sound source 484A. Classification network trainer 482 uses sound source representation 154 to train sound source encoder 202 and classification network 474. Network trainer 462 can use the sound source representation 154 and the trained version of neural network 150 to retain or remove sound from sound source 484A to generate output audio signal 135.
[0123] refer to Figure 5 Figure 500 illustrates an exemplary aspect of the operation associated with the joint training of the neural network 150 and the sound source encoder 202.
[0124] During operation, audio combiner 450 receives voice audio signal 524A from sound sources 484A and 484D, interfering voice audio signal 424B from sound source 484B, and noise audio signal 424C from sound source 484C. Audio combiner 450 generates input audio signal 526 based on a combination of voice audio signal 524A, interfering voice audio signal 424B, noise audio signal 424C, one or more additional audio signals, or combinations thereof. In some aspects, audio combiner 450 performs channel distortion enhancement 452, reverberation enhancement 454, or both to update input audio signal 526. In some aspects, input audio signal 526 approximates various noise conditions that may be present when the voice audio signal of interest is received (e.g., interfering voice, noise, channel distortion, reverberation, or combinations thereof).
[0125] Audio combiner 450 provides input audio signal 526 to audio adjuster 148. Audio adjuster 148 receives a combined sound source representation 147 representing sound source 484A and sound source 484D. Neural network 150 performs operations as described in the reference. Figure 3 The described similar operation processes the first input portion of the input audio signal 526 and the combined sound source representation 147 to generate a mask 551, wherein the input audio signal 526, the combined sound source representation 147, and the mask 551 respectively correspond to Figure 3 The input audio signal 126, the combined sound source representation 147A, and the mask 151 are used. The audio adjuster 148 applies the mask 551 to the first input section to generate the first output section of the output audio signal 535. The audio adjuster 148 applies the mask 551 to retain the sound of sound sources 484A and 484D from the first input section and remove (or reduce) the remaining sound of sound sources 484B, 484C, one or more additional audio signals, channel distortion enhancement 452, reverberation enhancement 454, or combinations thereof to generate the first output section.
[0126] Audio adjuster 148 provides a first output portion of the output audio signal 535 to sound source encoder 202 and to network trainer 462. Sound source encoder 202 processes the first output portion (e.g., corresponding to the sounds of sound sources 484A and 484D retained from the first input portion) to generate an updated version of the sound source representation 147 of sound sources 484A and 484D. Sound source encoder 202 provides the audio adjuster 148 with a second input portion combining the updated version of the sound source representation 147 to process the input audio signal 526. Sound source encoder 202 also provides the updated version of the combined sound source representation 147 to speaker detector 472.
[0127] The speaker detector 472 uses a classification network 574 to process an updated version of the combined sound source representation 147 to generate an estimated speaker identifier 575. The speaker detector 472 provides the estimated speaker identifier 575 to the classification network trainer 482.
[0128] Network trainer 462 generates a noise reduction loss metric 564 based on a comparison between a first output portion of the output audio signal 535 and a corresponding first speech portion of the speech audio signal 524A. The first output portion corresponds to the sounds of sound sources 484A and 484D retained from the input audio signal 526, and the first speech portion corresponds to the original sounds of sound sources 484A and 484D. Network trainer 462 generates update data 563 based on the noise reduction loss metric 564. For example, update data 563 indicates updates to the weights, biases, or combinations thereof of the neural network 150 to reduce the noise reduction loss metric 564 in subsequent iterations. Network trainer 462 thus trains neural network 150 over time to improve noise reduction and reduce the difference between the output audio signal 535 and the speech audio signal 524A.
[0129] The classification network trainer 482 determines a classification loss metric 586 based on the estimated speaker identifier 575 and speaker identifiers 581 of sound sources 484A and 484D. For example, the classification network trainer 482 retrieves a first combined sound source representation of a first pair of sound sources (e.g., two people) associated with the estimated speaker identifier 575, and a second combined sound source representation of sound sources 484A and 484D. The classification network trainer 482 generates the classification loss metric 486 based on a comparison of the first and second combined sound source representations. In a particular aspect, the second combined sound source representation was previously generated based on the sounds of sound sources 484A and 484D.
[0130] The classification network trainer 482 generates updated data 583 and updated data 585 based on the classification loss metric 586. For example, updated data 583 indicates updates to the weights, bias values, or combinations thereof of the sound source encoder 202 to reduce the classification loss metric 586. The classification network trainer 482 thus trains the sound source encoder 202 over time to generate a combined sound source representation 147 that more closely matches the second combined sound source representation. Similarly, updated data 585 indicates updates to the weights, bias values, or combinations thereof of the classification network 574 to reduce the classification loss metric 586. The classification network trainer 482 thus trains the classification network 574 over time to generate an estimated speaker identifier 575 that corresponds to a pair of speakers having a combined sound source representation that more closely matches the second combined sound source representation. The network trainer 462 and the classification network trainer 482 therefore jointly train the neural network 150, the sound source encoder 202, and the classification network 574.
[0131] In some implementations, neural network 150, sound source encoder 202, classification network 474, or combinations thereof, are trained during the training phase and designated for use upon completion of training. For example, network trainer 462 determines that training of neural network 150 is complete in response to determining that noise reduction loss metric 564 satisfies a training criterion (e.g., less than a loss threshold). Similarly, classification network trainer 482 determines that training of sound source encoder 202 and classification network 574 is complete in response to determining that classification loss metric 586 satisfies a training criterion (e.g., less than a loss threshold).
[0132] In some respects, the audio adjuster 148 may be available when the network trainer 462 determines that the training of the neural network 150 is complete. In some respects, the sound source encoder 202 and the classification network 574 may be available when the classification network trainer 482 determines that the training of the sound source encoder 202 and the classification network 574 is complete.
[0133] In some specific implementations, the audio combiner 450, the network trainer 462, the speaker detector 472, the classification network trainer 482, or a combination thereof are integrated into Figure 1 In device 102. In other embodiments, audio combiner 450, network trainer 462, classification network trainer 482, or combinations thereof are integrated into a second device. In these embodiments, the second device provides neural network 150, audio combiner 148, sound source encoder 202, speaker detector 472, classification network 574, or combinations thereof to device 102 in response to determining that audio adjuster 148, sound source encoder 202, classification network 574, or combinations thereof are available.
[0134] In some implementations, the neural network 150, the sound source encoder 202, the classification network 474, or a combination thereof, are trained (e.g., dynamically updated) during use. In these implementations, the input audio signal 526 corresponds to... Figure 1 The input audio signal is 126. In an exemplary example, network trainer 462 receives a speech audio signal 524A from a first microphone closer to sound sources 484A and 484D (e.g., a pair of known sound sources with speaker identifier 581), and receives the input audio signal 526 from one or more second microphones. Network trainer 462 trains neural network 150 such that output audio signal 535 matches speech audio signal 524A (e.g., noise reduction loss metric 564 is less than a loss threshold), and uses output audio signal 535 to generate a combined sound source representation 147 of sound sources 484A and 484D. Classification network trainer 482 uses combined sound source representation 147 to train sound source encoder 202 and classification network 574. Network trainer 462 can use sound source representation 154 and the trained version of neural network 150 to retain or remove the sounds of sound sources 484A and 484D to generate output audio signal 135.
[0135] refer to Figure 6 Figure 600 illustrates a device 102 capable of operating to receive audio data 626 from device 650 and processing the audio data 626 using a sound source representation. Device 102 includes an audio analyzer 140 coupled to receiver 640.
[0136] During operation, receiver 640 receives audio data 626 from device 650. Audio data 626 represents input audio signal 126. For example, audio data 626 represents sound 186 from sound source 184. In some specific implementations, audio data 626 corresponds to encoded audio data, and a decoder of device 102 decodes audio data 626 to generate input audio signal 126.
[0137] For reference Figure 1 As described, the audio analyzer 140 processes the input audio signal 126 based on the combined sound source representation 147A to generate an output audio signal 135. The audio analyzer 140 provides the output audio signal 135 to one or more speakers 160 to output sound 196.
[0138] refer to Figure 7 Figure 700 illustrates a device 102 capable of operating to process an input audio signal 126 using a sound source representation to generate an output audio signal 135, to generate audio data 726 based on the output audio signal 135, and to transmit the audio data 726 to a device 702. Device 102 includes an audio analyzer 140 coupled to a transmitter 740.
[0139] During operation, the audio analyzer 140 receives input audio signals 126 corresponding to the microphone outputs of one or more microphones 120. The input audio signals 126 represent sound 186 captured by the one or more microphones 120 from a sound source 184. (See reference...) Figure 1 As described, the audio analyzer 140 uses a combined sound source representation 147A to process the input audio signal 126 to generate an output audio signal 135.
[0140] Transmitter 740 sends audio data 726 to device 702. Audio data 726 is based on output audio signal 135. In some examples, audio data 726 corresponds to encoded audio data, and the encoder of device 102 encodes the output audio signal 135 to generate audio data 726.
[0141] The decoder of device 702 decodes audio data 726 to generate a decoded audio signal. Device 702 provides the decoded audio signal to one or more speakers 160 to output sound 196.
[0142] In some implementations, device 702 is the same as device 650. For example, device 102 receives audio data 626 from the second device and outputs sound 196 via one or more speakers 160, while concurrently capturing sound 186 via one or more microphones 120 and transmitting the audio data 726 to the second device.
[0143] Figure 8 Based on some examples of this disclosure Figure 1The diagram illustrates an illustrative aspect of the operation of the system components. The neural network 150 is configured to receive a sequence 810 of audio data samples (e.g., a sequence of consecutively captured frames of the input audio signal 126), shown as a first frame (F1) 812, a second frame (F2) 814, and one or more other frames including an Nth frame (FN) 816 (where N is an integer greater than two). The neural network 150 is also configured to receive a sequence 840 of sound source representations, such as sequences of combined sound source representations (including combined sound source representation 147A), shown as a first sound source representation (S1) 842, a second sound source representation (S2) 844, and one or more additional sound source representations (including an Rth sound source representation (SR) 846) (where R is an integer greater than two and less than or equal to N). The neural network 150 is configured as a sequence 820 of output masks (including a first mask (M1) 822, a second mask (M2) 824 and one or more additional masks (including an Nth mask (MN) 826)).
[0144] The neural network 150 is configured to receive a sequence 810 of audio data samples and adaptively generate a mask (e.g., a second mask 824) of sequence 810 corresponding to a frame (e.g., a second frame (F2) 814) of sequence 820, based at least in part on a sequence 810 of audio data samples (e.g., a first frame (F1) 812). As an illustrative and non-limiting example, the neural network 150 may include a CNN.
[0145] The audio adjuster 148 is configured to apply a mask of sequence 810 (e.g., a first mask (M1) 822) to the corresponding frame of sequence 820 (e.g., a first frame (F1) 812) to generate frames of sequence 830 of audio data samples (e.g., a first frame (O1) 832), such as a sequence of consecutive frames of output audio signal 135, which are shown as a first frame (O1) 832, a second frame (O2) 834 and one or more additional frames (including a Nth frame (ON) 836) (where N is an integer greater than two).
[0146] During operation, neural network 150 processes first frame (F1) 812 using a first sound source representation (S1) 842 to generate a first mask (M1) 822, and audio adjuster 148 applies the first mask (M1) 822 to first frame (F1) 812 to generate the first frame (O1) 832 of the audio data sample sequence 830. Neural network 150 processes second frame (F2) 814 using a second sound source representation (S2) 844 to generate a second mask (M2) 824, and audio adjuster 148 applies the second mask (M2) 824 to second frame (F2) 814 to generate the second frame (O2) 834 of the audio data sample sequence 830. In some specific implementations, the second mask (M2) 824 is based on second frame (F2) 814 and at least partially based on the first frame (F1) 812 of the audio data samples. This process continues, including the neural network 150 using the Nth sound source representation (SR) 846 to process the Nth frame (FN) 816 to generate the Nth mask (MN) 826, and the audio adjuster 148 applying the Nth mask (MN) 826 to the Nth frame (FN) 816 to generate the Nth frame (ON) 836 of the sequence 830 of audio data samples.
[0147] In some implementations, configurator 144 provides SSR to neural network 150 to process each frame of sequence 810 of audio data samples (e.g., integer R equals integer N). In some examples, if one or more selected sound sources 162 remain unchanged from the first frame (F1) 812 to the second frame (F2) 814, then the second sound source representation (S2) 844 is the same as the first sound source representation (S1) 842 or an updated (e.g., dynamically trained) version of the first sound source representation (S1) 842.
[0148] In some implementations, configurator 144 provides SSRs to neural network 150 at a rate different from the rate at which neural network 150 processes frames of the sequence 810 of audio data samples. For example, configurator 144 provides an SSR to neural network 150 for every four frames of the sequence 810 of audio data samples (e.g., an integer N equal to four times the integer R). For instance, neural network 150 uses a first sound source representation (S1) 842 to process the first four frames of the sequence 810 of audio data samples, uses a second sound source representation (S2) 844 to process the last four frames, and so on. If one or more selected sound sources 162 are the same for the first and fifth frames, then the second sound source representation (S2) 844 is the same as the first sound source representation (S1) 842 or an updated (e.g., dynamically trained) version of the first sound source representation (S1) 842.
[0149] In some implementations, the configurator 144 provides an SSR to the neural network 150 in response to changes in one or more selected sound sources 162. For example, if one or more selected sound sources 162 remain unchanged when processing a sequence 810 of audio data samples, the configurator 144 provides the neural network 150 with only a first sound source representation (S1) 842 (e.g., an integer R equal to 1), and the neural network 150 uses the first sound source representation (S1) 842 to process all frames of the sequence 810 of audio data samples.
[0150] In some implementations, the Nth mask (MN) 826 is based on the Nth frame (FN) 816 and at least partially on one or more previous frames of the audio data samples of sequence 810. By dynamically generating the mask based on one or more previous frames of the audio data samples, the accuracy of audio adjustment by the audio adjuster 148 for audio (e.g., speech, music, etc.) that may span multiple frames of audio data can be improved.
[0151] Figure 9 The device 102 is depicted as a specific implementation 900 of an integrated circuit 902 including one or more processors 190. The integrated circuit 902 also includes an audio input section 904 (such as one or more bus interfaces) to enable receiving input audio signals 126 for processing. The integrated circuit 902 also includes a signal output section 906 (such as a bus interface) to enable transmitting output signals, such as output audio signals 135. The integrated circuit 902 enables audio processing based on sound source representation to be implemented as a component in a system including a microphone, such as in… Figure 10 The mobile phone or tablet shown, such as in Figure 11 The headphones shown, as in Figure 12 The wearable electronic devices shown, such as in Figure 13 The voice-controlled speaker system shown, as in Figure 14 The camera shown, as in Figure 15 The virtual reality, mixed reality, or augmented reality headsets shown are as follows: Figure 16 or Figure 17 The vehicles shown.
[0152] As an illustrative and non-restrictive example, Figure 10 A specific implementation 1000 of which device 102 includes mobile device 1002 (such as a telephone or tablet) is depicted. Mobile device 1002 includes one or more microphones 120, one or more speakers 160, and a display screen 1004. Components of one or more processors 190 (including an audio analyzer 140) are integrated into mobile device 1002 and are shown using dashed lines to indicate internal components that are not normally visible to the user of mobile device 1002.
[0153] In a particular example, audio analyzer 140 operates to process input audio signal 126 using a sound source representation to generate output audio signal 135. In a specific aspect, input audio signal 126 is based on the microphone output of one or more microphones 120, and audio analyzer 140 outputs audio signal 135 to another device (not shown) such as... Figure 7 The device 702 provides audio data.
[0154] In a particular aspect, the input audio signal 126 can be based on a signal from another device (not shown) such as... Figure 6 The device 650 receives audio data. In some embodiments, the audio analyzer 140 provides an output audio signal 135 to one or more speakers 160 for playback. In some embodiments, the output audio signal 135 is processed to perform one or more operations at the mobile device 1002, such as launching a graphical user interface or otherwise displaying other information associated with the user's voice on the display screen 1004 (e.g., via an integrated "smart assistant" application).
[0155] Figure 11 A specific implementation 1100 of the device 102, including a headset device 1102, is depicted. The headset device 1102 includes one or more microphones 120, one or more speakers 160, or combinations thereof. Components of one or more processors 190, including an audio analyzer 140, are integrated into the headset device 1102. In a particular example, the audio analyzer 140 operates to process an input audio signal 126 using a sound source representation to generate an output audio signal 135. In a particular aspect, the input audio signal 126 may be based on the microphone output of one or more microphones 120, and the audio analyzer 140 outputs the audio signal 135 to another device (not shown), such as... Figure 7 The device 702 provides audio data for further processing.
[0156] In a particular aspect, the input audio signal 126 can be based on a signal from another device (not shown) such as... Figure 6 The device 650 receives audio data. In some embodiments, the audio analyzer 140 provides an output audio signal 135 to one or more speakers 160 for playback. In some embodiments, the output audio signal 135 is processed to perform one or more operations at the headphone device 1102.
[0157] Figure 12A specific implementation 1200 of a device 102, including a wearable electronic device 1202 (shown as a "smartwatch"), is depicted. An audio analyzer 140, one or more microphones 120, and one or more speakers 160 are integrated into the wearable electronic device 1202. In a particular example, the audio analyzer 140 operates to process an input audio signal 126 using a sound source representation to generate an output audio signal 135. In a particular aspect, the input audio signal 126 is based on the microphone output of one or more microphones 120, and the audio analyzer 140 outputs the audio signal 135 to another device (not shown), such as... Figure 7 The device 702 provides audio data.
[0158] In a particular aspect, the input audio signal 126 can be based on a signal from another device (not shown) such as... Figure 6 The device 650 receives audio data. In some embodiments, the audio analyzer 140 provides an output audio signal 135 to one or more speakers 160 for playback. In some embodiments, the output audio signal 135 is processed to perform one or more operations at the wearable electronic device 1202, such as launching a graphical user interface or otherwise displaying additional information associated with the user's voice on a display screen 1204 of the wearable electronic device 1202. For example, the wearable electronic device 1202 may include a display screen configured to display notifications based on user voice detected by the wearable electronic device 1202 in the output audio signal 135. In a particular example, the wearable electronic device 1202 includes a haptic device that provides haptic notifications (e.g., vibration) in response to the detection of user speech activity. For example, a haptic notification may cause the user to look at the wearable electronic device 1202 to see a displayed notification indicating that a keyword spoken by the user has been detected. Thus, the wearable electronic device 1202 may alert a user with hearing impairment or a user wearing headphones to the detection of the user's speech activity.
[0159] Figure 13 This is a specific implementation 1300 of which device 102 includes a wireless speaker and a voice-activated device 1302. The wireless speaker and voice-activated device 1302 may have wireless network connectivity and is configured to perform auxiliary operations. One or more processors 190, including an audio analyzer 140, one or more microphones 120, one or more speakers 160, or combinations thereof, are included in the wireless speaker and voice-activated device 1302.
[0160] In a particular example, audio analyzer 140 operates to process input audio signal 126 using a sound source representation to generate output audio signal 135. In a specific aspect, input audio signal 126 may be based on microphone outputs from one or more microphones 120, and audio analyzer 140 outputs audio signal 135 to another device (not shown), such as... Figure 7 The device 702 provides audio data for further processing.
[0161] In a particular aspect, the input audio signal 126 can be based on a signal from another device (not shown) such as... Figure 6 The device 650 receives audio data. In some embodiments, the audio analyzer 140 provides an output audio signal 135 to one or more speakers 160 for playback. In some embodiments, the output audio signal 135 is processed to perform one or more operations at the wireless speaker and voice-activated device 1302. For example, in response to the detection of a spoken command in the output audio signal 135, the wireless speaker and voice-activated device 1302 can perform an auxiliary operation, such as via the execution of a voice-activated system (e.g., an integrated assistance application). The auxiliary operation may include adjusting the temperature, playing music, turning on a light, etc. For example, the auxiliary operation is performed in response to receiving a command after a keyword or key phrase (e.g., “Hello, assistant”).
[0162] Figure 14 A specific embodiment 1400 is depicted in which device 102 includes a portable electronic device corresponding to camera device 1402. An audio analyzer 140, one or more microphones 120, one or more speakers 160, or combinations thereof are included in camera device 1402.
[0163] In a particular example, audio analyzer 140 operates to process input audio signal 126 using a sound source representation to generate output audio signal 135. In a specific aspect, input audio signal 126 may be based on microphone outputs from one or more microphones 120, and audio analyzer 140 outputs audio signal 135 to another device (not shown), such as... Figure 7 The device 702 provides audio data for further processing.
[0164] In a particular aspect, the input audio signal 126 can be based on a signal from another device (not shown) such as... Figure 6The device 650 receives audio data. In some embodiments, the audio analyzer 140 provides an output audio signal 135 to one or more speakers 160 for playback. In some embodiments, the output audio signal 135 is processed to perform one or more operations at the camera device 1402. For example, as an illustrative example, in response to the detection of a verbal command in the output audio signal 135, the camera device 1402 may perform operations such as adjusting image or video capture settings, image or video playback settings, or image or video capture instructions in response to a spoken user command.
[0165] Figure 15 A specific implementation 1500 is depicted in which device 102 includes a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality headset 1502. An audio analyzer 140, one or more microphones 120, one or more speakers 160, or combinations thereof, are integrated into the headset 1502. In a particular aspect, the headset 1502 includes a first microphone 120 and a second microphone 120, the first microphone being positioned to primarily capture the user's voice and the second microphone being positioned to primarily capture ambient sound.
[0166] In a particular example, audio analyzer 140 operates to process input audio signal 126 using a sound source representation to generate output audio signal 135. In a specific aspect, input audio signal 126 may be based on microphone outputs from one or more microphones 120, and audio analyzer 140 outputs audio signal 135 to another device (not shown), such as... Figure 7 The device 702 provides audio data for further processing.
[0167] In a particular aspect, the input audio signal 126 can be based on a signal from another device (not shown) such as... Figure 6 The device 650 receives audio data. In some embodiments, the audio analyzer 140 provides an output audio signal 135 to one or more speakers 160 for playback. In some embodiments, the output audio signal 135 is processed to perform one or more operations at the headset 1502. The visual interface device is positioned in front of the user's eyes to enable the display of augmented reality, mixed reality, or virtual reality images or scenes to the user while wearing the headset 1502. In a particular example, the visual interface device is configured to display a notification indicating user voice detected in the output audio signal 135.
[0168] Figure 16The illustration depicts a device 102 corresponding to a vehicle 1602 (shown as a manned or unmanned aerial device (e.g., a package delivery drone)) or a specific implementation 1600 integrated within the vehicle. An audio analyzer 140, one or more microphones 120, one or more speakers 160, or combinations thereof are integrated into the vehicle 1602.
[0169] In a particular example, audio analyzer 140 operates to process input audio signal 126 using a sound source representation to generate output audio signal 135. In a specific aspect, input audio signal 126 may be based on microphone outputs from one or more microphones 120, and audio analyzer 140 outputs audio signal 135 to another device (not shown), such as... Figure 7 The device 702 provides audio data for further processing.
[0170] In a particular aspect, the input audio signal 126 can be based on a signal from another device (not shown) such as... Figure 6 The device 650 receives audio data. In some embodiments, the audio analyzer 140 provides an output audio signal 135 to one or more speakers 160 for playback. In some embodiments, the output audio signal 135 is processed to perform one or more operations at the vehicle 1602. For example, user voice activity detection may be performed based on the output audio signal 135, such as delivery instructions to an authorized user from the vehicle 1602.
[0171] Figure 17 The illustration depicts a device 102 corresponding to a vehicle 1702 (shown as an automobile) or another specific embodiment 1700 integrated within such a vehicle. The vehicle 1702 includes one or more processors 190, including an audio analyzer 140. The vehicle 1702 also includes one or more microphones 120, one or more speakers 160, or combinations thereof.
[0172] In a particular example, audio analyzer 140 operates to process input audio signal 126 using a sound source representation to generate output audio signal 135. In a specific aspect, input audio signal 126 may be based on microphone outputs from one or more microphones 120, and audio analyzer 140 outputs audio signal 135 to another device (not shown), such as... Figure 7 The device 702 provides audio data for further processing. In an exemplary example, the combined sound source representation 147 represents sound 298 from the sound source 284 of the first environment, as shown in the reference. Figure 2CAs described. Audio analyzer 140, in response to determining a second environment of vehicle 1702 and a first environment associated with combined sound source representation 147, uses combined sound source representation 147 to process input audio signal 126 to remove background noise and generate output audio signal 135.
[0173] In a particular aspect, the audio analyzer 140 determines that the first environment matches the second environment in response to determining that the first environment is associated with a first vehicle of the matching vehicle 1702, a first operating state of the first vehicle of the matching vehicle 1702, a first external condition of one or more second external conditions of the matching vehicle 1702, or a combination thereof.
[0174] In one aspect, the audio analyzer 140 determines that the first vehicle matches vehicle 1702 in response to determining that the first vehicle is the same as vehicle 1702. In another aspect, the audio analyzer 140 determines that the first vehicle is the same vehicle model (e.g., The first vehicle is identified by the registered trademarks of Polestar Holding GmbH (Sweden), the same year of manufacture (e.g., 2022), the same type of vehicle (e.g., electric SUV), or a combination thereof, as vehicle 1702. In a particular aspect, the operating state of the vehicle may include speed, reversing, turning, braking, etc. In a particular aspect, the external conditions of the vehicle may include road type (e.g., highway, muddy road, suburban road, urban road), weather conditions (e.g., wind, rain, storm), traffic conditions (e.g., heavy traffic, medium traffic, no traffic), etc.
[0175] In a particular aspect, the input audio signal 126 can be based on a signal from another device (not shown) such as... Figure 6The device 650 receives audio data. In some embodiments, the audio analyzer 140 provides an output audio signal 135 to one or more speakers 160 for playback. In some embodiments, the output audio signal 135 is processed to perform one or more operations at the vehicle 1702. For example, user voice activity detection may be performed based on the output audio signal 135. In some embodiments, user voice activity detection may be performed based on input audio signals 126 received from internal microphones (e.g., one or more microphones 120), such as voice commands from authorized passengers. For example, input audio signals 126 include sounds from sound source 184A, such as voice commands from an authorized user (e.g., a parent) of the vehicle 1702 to set the volume to 5 or to set the destination of the autonomous vehicle. Input audio signals 126 also include sounds from sound source 184B, such as voice commands from an unauthorized user (e.g., a child) of the vehicle 1702 to set the volume to 9. The input audio signal 126 may also include sound from sound source 184C (such as other passengers discussing another location).
[0176] The audio analyzer 140 can process the input audio signal 126 to retain the sound from the sound source 184A and remove the sound from the sound sources 184B and 184C to generate an output audio signal 135, and can perform user voice activity detection on the output audio signal 135 to detect voice commands from authorized users (e.g., voice commands from parents to set the volume to 5 or to set the destination of an autonomous vehicle).
[0177] In some implementations, user voice activity detection can be performed based on input audio signals 126 received from external microphones (e.g., one or more microphones 120) (such as an authorized user of the vehicle). In a particular implementation, the voice activation system initiates one or more operations of the vehicle 1702 based on one or more keywords detected in the output audio signal 135 (e.g., “unlock,” “start engine,” “play music,” “display weather forecast,” or another voice command), such as by providing feedback or information via display 1720 or one or more speakers 160.
[0178] refer to Figure 18 This illustrates a specific implementation of a method 1800 for processing audio based on sound source representation. In a particular aspect, one or more operations of method 1800 are performed by a configurator 144, an SSR generator 146, an audio adjuster 148, a neural network 150, an audio analyzer 140, one or more processors 190, and a device 102. Figure 1System 100, Feature Extractor 350, CNN 352, LSTM 354A, LSTM 354B, LSTM 354C, Figure 3 Fully connected layer 356 Figure 4 Audio combiner 450 Figure 6 Receiver 640 Figure 7 The transmitter 740 or at least one of their combinations shall be used to perform the operation.
[0179] At 1802, method 1800 includes receiving an input audio signal at a first device. For example, audio analyzer 140 receives input audio signal 126 at device 102. In some embodiments, input audio signal 126 is based on microphone outputs from one or more microphones 120. In some embodiments, input audio signal 126 is based on audio data 626 received from device 650, as referenced. Figure 6 As described.
[0180] At 1804, method 1800 further includes processing the input audio signal based on a combined representation of multiple sound sources to generate an output audio signal, wherein the combined representation is used to selectively retain or remove sounds from the input audio signal from the multiple sound sources. For example, audio analyzer 140 processes input audio signal 126 based on a combined sound source representation 147A of sound sources 184A and 184B to generate output audio signal 135, as referenced. Figure 1 As described. The audio analyzer 140 selectively uses a combined sound source representation 147A based on the retention flag 145 of the adjuster configuration setting 143 to retain or remove the sounds of sound source 184A and sound source 184B from the input audio signal 126.
[0181] At 1806, method 1800 also includes providing an output audio signal to a second device. For example, audio adjuster 148 provides an output audio signal 135 to one or more speakers 160, as referenced. Figure 1 As described. For example, audio adjuster 148 provides output audio signal 135 to transmitter 740 to send audio data 726 to device 702. Audio data 726 is based on output audio signal 135 (e.g., an encoded version of the output audio signal).
[0182] Method 1800 enhances the perception of the sound of interest in the output audio signal 135 by retaining the sound of interest or removing the remaining sound. Using a combined sound source representation 147A representing the sounds of the selected sound source 162 can improve the efficiency and accuracy of processing the input audio signal 126 to generate the output audio signal 135.
[0183] Figure 18Method 1800 can be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a digital signal processor (DSP), a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 18 Method 1800 can be executed by a processor that executes instructions, such as reference... Figure 19 As described.
[0184] refer to Figure 19 A block diagram depicting a specific, exemplary embodiment of the device is provided, and is generally designated as 1900. In various embodiments, device 1900 may have... Figure 19 The number of components shown may be more or less. In an exemplary embodiment, device 1900 may correspond to device 102. In an exemplary embodiment, device 1900 may perform reference... Figures 1 to 18 One or more operations as described.
[0185] In a particular implementation, device 1900 includes a processor 1906 (e.g., a CPU). Device 1900 may include one or more additional processors 1910 (e.g., one or more DSPs). In a particular aspect, Figure 1 One or more processors 190 correspond to processor 1906, processor 1910, or combinations thereof. Processor 1910 may include a voice and music encoder-decoder (codec) 1908, which includes a speech decoder (“phonecoder”) encoder 1936, a phonecoder decoder 1938, an audio analyzer 140, or combinations thereof.
[0186] Device 1900 may include memory 1986 and codec 1934. Memory 1986 may include instructions 1956, which can be executed by one or more additional processors 1910 (or processor 1906) to implement the functionality described in reference audio analyzer 140. Device 1900 may include a modem 1970 coupled to antenna 1952 via transceiver 1950. In a particular aspect, memory 1986 includes... Figure 1 The memory 132. In a particular aspect, the transceiver 1950 includes... Figure 6 Receiver 640 Figure 7 The transmitter 740 or both.
[0187] Device 1900 may include a display 1928 coupled to display controller 1926. One or more microphones 120, one or more speakers 160, or combinations thereof may be coupled to codec 1934. Codec 1934 may include digital-to-analog converter (DAC) 1902, analog-to-digital converter (ADC) 1904, or both. In a particular embodiment, codec 1934 may receive analog signals from one or more microphones 120, use ADC 1904 to convert the analog signals to digital signals, and provide digital signals to voice and music codec 1908. Voice and music codec 1908 may process digital signals, and the digital signals may be further processed by audio analyzer 140. In a particular embodiment, audio analyzer 140 of voice and music codec 1908 may provide digital signals to codec 1934. Codec 1934 may use ADC 1902 to convert digital signals to analog signals and may provide analog signals to one or more speakers 160.
[0188] In a particular embodiment, device 1900 may be included in a system-in-package or system-on-a-chip device 1922. In a particular embodiment, memory 1986, processor 1906, processor 1910, display controller 1926, codec 1934, and modem 1970 are included in a system-in-package or system-on-a-chip device 1922. In a particular embodiment, input device 1930 and power supply 1944 are coupled to system-on-a-chip device 1922. Furthermore, in a particular embodiment, such as Figure 19 As shown, the display 1928, input device 1930, one or more microphones 120, one or more speakers 160, antenna 1952, and power supply 1944 are located external to the system-on-a-chip device 1922. In a particular embodiment, each of the display 1928, input device 1930, one or more microphones 120, one or more speakers 160, antenna 1952, and power supply 1944 may be coupled to components of the system-on-a-chip device 1922, such as an interface or controller.
[0189] Device 1900 may include smart speakers, speaker bars, mobile communication devices, smartphones, cellular phones, laptops, computers, tablets, personal digital assistants, display devices, televisions, game consoles, music players, radios, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, vehicles, head-mounted devices, augmented reality head-mounted devices, mixed reality head-mounted devices, virtual reality head-mounted devices, aircraft, home automation systems, voice-activated devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, computing devices, communication devices, Internet of Things (IoT) devices, virtual reality (VR) devices, base stations, mobile devices, or any combination thereof.
[0190] In conjunction with the described specific embodiments, an apparatus includes a component for receiving an input audio signal at a first device. For example, the component for receiving may correspond to an audio adjuster 148, a neural network 150, an audio analyzer 140, one or more processors 190, or a device 102. Figure 1 System 100 Figure 6 Receiver 640, codec 1934, digital-to-analog converter 1904, voice and music codec 1908, antenna 1952, transceiver 1950, modem 1970, processor 1906, one or more processors 1910, Figure 19 The device 1900 is configured to receive one or more other circuits or components or any combination thereof for receiving input audio signals.
[0191] The apparatus also includes components for processing the input audio signal based on a combined representation of multiple sound sources to generate an output audio signal, wherein the combined representation is used to selectively retain or remove sounds from the input audio signal from multiple sound sources. For example, the components for processing may correspond to an audio adjuster 148, a neural network 150, an audio analyzer 140, one or more processors 190, or a device 102. Figure 1 System 100, voice and music codecs 1908, processor 1906, one or more processors 1910, Figure 19 The device 1900 is configured to process one or more other circuits or components, or any combination thereof, based on a combination of multiple sound sources to represent the input audio signal.
[0192] The device also includes components for providing an output audio signal to a second device. For example, the components for providing the signal may correspond to an audio adjuster 148, a neural network 150, an audio analyzer 140, one or more processors 190, or a device 102. Figure 1 System 100 Figure 7Transmitter 740, codec 1934, digital-to-analog converter 1902, voice and music codec 1908, antenna 1952, transceiver 1950, modem 1970, processor 1906, one or more processors 1910, Figure 19 The device 1900, and one or more other circuits or components or any combination thereof configured to provide an output audio signal to a second device.
[0193] In some embodiments, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 132 or memory 1986) stores instructions (e.g., instruction 1956) that, when executed by one or more processors (e.g., one or more processors 190, one or more processors 1910, or processor 1906), cause one or more processors to receive an input audio signal (e.g., input audio signal 126) at a first device (e.g., device 102). When executed by one or more processors, the instructions also cause one or more processors to process the input audio signal based on a combined representation (e.g., combined sound source representation 147A) of multiple sound sources (e.g., sound source 184A and sound source 184B) to generate an output audio signal (e.g., output audio signal 135). The combined representation is used to selectively retain or remove sounds from the input audio signal from the multiple sound sources. When executed by one or more processors, the instructions further cause one or more processors to provide the output audio signal to a second device (e.g., one or more speakers 160 or device 702).
[0194] Specific aspects of this disclosure are described below in various sets of related embodiments:
[0195] According to Embodiment 1, an apparatus includes: one or more processors configured to: receive an input audio signal; process the input audio signal based on a combined representation of a plurality of sound sources to generate an output audio signal, wherein the combined representation is used to selectively retain or remove sounds from the plurality of sound sources from the input audio signal; and provide the output audio signal to a second apparatus.
[0196] Example 2 includes the device according to Example 1, wherein one or more processors are configured to use the combined representation based on a retention flag having a first value to retain the sounds of the plurality of sound sources from the input audio signal and remove other sounds from one or more additional sound sources.
[0197] Example 3 includes the device according to Example 2, wherein one or more processors are configured to set the retention flag to a first value indicating that the sound of the plurality of sound sources will be retained in response to a detected condition indicating that processing of the input audio signal will be initiated, wherein the first value of the retention flag is based on user input, default configuration, configuration input from an application, configuration request from another device, or a combination thereof.
[0198] Example 4 includes the device according to any one of Examples 1 to 3, wherein the plurality of sound sources include one or more authorized users.
[0199] Example 5 includes the device according to any one of Examples 1 to 4, wherein the plurality of sound sources include emergency vehicles.
[0200] Example 6 includes a device according to any one of Examples 1 to 5, wherein the one or more processors are configured to use the combined representation based on a retention flag having a second value to remove the sounds of the plurality of sound sources from the input audio signal and retain other sounds of one or more additional sound sources.
[0201] Example 7 includes the device according to Example 6, wherein one or more processors are configured to set the reservation flag to a second value indicating that the sound from the plurality of sound sources will be removed in response to a detected condition indicating that processing of the input audio signal will be initiated, wherein the second value of the reservation flag is based on user input, default configuration, configuration input from an application, configuration request from another device, or a combination thereof.
[0202] Example 8 includes the device according to any one of Examples 1 to 7, wherein the plurality of sound sources include traffic, wind, echo, channel distortion, another non-voice sound source, a person, or a combination thereof.
[0203] Example 9 includes the device according to any one of Examples 1 to 8, wherein the plurality of sound sources are associated with background noise in a particular environment.
[0204] Example 10 includes the device according to Example 9, wherein the specific environment corresponds to the interior of a specific type of vehicle.
[0205] Example 11 includes a device according to any one of Examples 1 to 10, wherein the combination represents a specific sound based on a specific sound source, and wherein the specific sound source is of the same sound source type as one of the plurality of sound sources.
[0206] Example 12 includes a device according to any one of Examples 1 to 11, wherein the one or more processors are further configured to update the combined representation based on the sound of any of the plurality of sound sources.
[0207] Example 13 includes a device according to any one of Examples 1 to 12, wherein the one or more processors are further configured to generate the combined representation based on the individual representations of the plurality of sound sources based on a combined setting.
[0208] Example 14 includes the device according to Example 13, wherein one or more processors are further configured to update the combined settings based on user input, detected conditions, or both.
[0209] Example 15 includes a device according to any one of Examples 1 to 14, wherein the plurality of sound sources include at least a first sound source and a second sound source, wherein a first representation of the first sound source indicates a first value of a specific feature, wherein a second representation of the second sound source indicates a second value of the specific feature, and wherein the value of the specific feature indicated by the combined representation is based on the first value and the second value.
[0210] Example 16 includes the device according to Example 15, wherein the first representation includes one or more spectrograms, the one or more spectrograms being based on sounds from a specific sound source of the same type as the first sound source.
[0211] Example 17 includes the device according to Example 15 or Example 16, wherein the combination represents a cascade of a first representation corresponding to the first sound source and a second representation corresponding to the second sound source.
[0212] Example 18 includes a device according to any one of Examples 1 to 17, wherein the input audio signal is processed using a neural network to generate the output audio signal.
[0213] Example 19 includes the device according to Example 18, wherein the neural network includes a convolutional neural network (CNN), an autoregressive (AR) generative network, an audio generative network (AGN), an attention network (AN), a long short-term memory (LSTM) network, or a combination thereof.
[0214] Example 20 includes the device according to Example 18 or Example 19, the device further including a sound source encoder configured to process sounds from one or more sound sources to generate representations of the one or more sound sources, wherein the sound source encoder and the neural network are jointly trained.
[0215] Example 21 includes a device according to any one of Examples 1 to 20, the device further including a receiver configured to receive audio data representing the input audio signal.
[0216] Example 22 includes a device according to any one of Examples 1 to 21, the device further including a transmitter configured to send audio data to the second device, the audio data being based on the output audio signal.
[0217] According to embodiment 23, a method includes: receiving an input audio signal at a first device; processing the input audio signal based on a combined representation of a plurality of sound sources to generate an output audio signal, wherein the combined representation is used to selectively retain or remove sounds from the plurality of sound sources from the input audio signal; and providing the output audio signal to a second device.
[0218] Example 24 includes the method according to Example 23, wherein the combined representation is used based on a retention flag having a first value to retain the sounds of the plurality of sound sources from the input audio signal and remove other sounds from one or more additional sound sources.
[0219] Example 25 includes the method according to Example 24, the method further including setting the retention flag to a first value indicating that the sounds of the plurality of sound sources will be retained in response to a detected condition indicating that processing of the input audio signal will be initiated, wherein the first value of the retention flag is based on user input, default configuration, configuration input from an application, configuration request from another device, or a combination thereof.
[0220] Example 26 includes the method according to any one of Examples 23 to 25, wherein the plurality of sound sources include one or more authorized users.
[0221] Example 27 includes the method according to any one of Examples 23 to 26, wherein the plurality of sound sources include emergency vehicles.
[0222] Example 28 includes the method according to any one of Examples 23 to 27, wherein the combination representation is used based on a retention flag having a second value to remove the sounds of the plurality of sound sources from the input audio signal and retain other sounds of one or more additional sound sources.
[0223] Example 29 includes the method according to Example 28, the method further including setting the reservation flag to a second value indicating that the sound of the plurality of sound sources will be removed in response to a detected condition indicating that processing of the input audio signal will be initiated, wherein the second value of the reservation flag is based on user input, default configuration, configuration input from an application, configuration request from another device, or a combination thereof.
[0224] Example 30 includes the method according to any one of Examples 23 to 29, wherein the plurality of sound sources include traffic, wind, echo, channel distortion, another non-voice sound source, a person, or a combination thereof.
[0225] Example 31 includes the method according to any one of Examples 23 to 30, wherein the plurality of sound sources are associated with background noise in a particular environment.
[0226] Example 32 includes the method according to Example 31, wherein the specific environment corresponds to the interior of a specific type of vehicle.
[0227] Example 33 includes the method according to any one of Examples 23 to 32, wherein the combination represents a specific sound from a specific sound source, and wherein the specific sound source is of the same sound source type as one of the plurality of sound sources.
[0228] Example 34 includes the method according to any one of Examples 23 to 33, the method further comprising updating the combined representation based on the sound of any of the plurality of sound sources.
[0229] Example 35 includes the method according to any one of Examples 23 to 34, the method further comprising generating the combined representation based on the individual representations of the plurality of sound sources based on a combined setting.
[0230] Example 36 includes the method according to Example 35, the method further including updating the combined settings based on user input, detected conditions, or both.
[0231] Example 37 includes a method according to any one of Examples 23 to 36, wherein the plurality of sound sources include at least a first sound source and a second sound source, wherein a first representation of the first sound source indicates a first value of a specific feature, wherein a second representation of the second sound source indicates a second value of the specific feature, and wherein the value of the specific feature indicated by the combined representation is based on the first value and the second value.
[0232] Example 38 includes the method according to Example 37, wherein the first representation includes one or more spectrograms, the one or more spectrograms being based on sounds from a specific sound source of the same type as the first sound source.
[0233] Example 39 includes the method according to Example 37 or Example 38, wherein the combination represents a cascade of a first representation corresponding to the first sound source and a second representation corresponding to the second sound source.
[0234] Example 40 includes the method according to any one of Examples 23 to 39, wherein the input audio signal is processed using a neural network to generate the output audio signal.
[0235] Example 41 includes the method according to Example 40, wherein the neural network includes a convolutional neural network (CNN), an autoregressive (AR) generative network, an audio generative network (AGN), an attention network (AN), a long short-term memory (LSTM) network, or a combination thereof.
[0236] Example 42 includes the method according to Example 40 or Example 41, the method further including using a sound source encoder to process sound from one or more sound sources to generate a representation of the one or more sound sources, wherein the sound source encoder and the neural network are jointly trained.
[0237] According to embodiment 43, an apparatus includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform a method according to any one of embodiments 23 to 42.
[0238] According to Embodiment 44, a non-transitory computer-readable medium storage instruction, when executed by a processor, causes the processor to perform the method according to any one of Embodiments 23 to 42.
[0239] According to embodiment 45, an apparatus includes components for performing the method according to any one of embodiments 23 to 42.
[0240] According to Embodiment 46, a non-transitory computer-readable medium storage instruction, when executed by one or more processors, causes the one or more processors to: receive an input audio signal at a first device; process the input audio signal based on a combined representation of a plurality of sound sources to generate an output audio signal, wherein the combined representation is used to selectively retain or remove sounds from the plurality of sound sources from the input audio signal; and provide the output audio signal to a second device.
[0241] Example 47 includes a non-transitory computer-readable medium according to Example 46, wherein the input audio signal is processed using a neural network to generate the output audio signal.
[0242] According to embodiment 48, an apparatus includes: means for receiving an input audio signal at a first device; means for processing the input audio signal based on a combined representation of a plurality of sound sources to generate an output audio signal, wherein the combined representation is used to selectively retain or remove sounds from the plurality of sound sources from the input audio signal; and means for providing the output audio signal to a second device.
[0243] Example 49 includes the apparatus according to Example 48, wherein the component for receiving, the component for processing, and the component for providing are integrated into at least one of: a smart speaker, a speaker bar, a computer, a tablet computer, a display device, a television, a game console, a music player, a radio, a digital video player, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aircraft, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, or a mobile device.
[0244] Those skilled in the art will also understand that the various exemplary logic blocks, configurations, modules, circuits, and algorithm steps described in connection with the specific embodiments disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various exemplary components, blocks, configurations, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, and such implementation decisions shall not be construed as departing from the scope of this disclosure.
[0245] The steps of the methods or algorithms described in conjunction with the specific embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compressed optical disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or user terminal. Alternatively, the processor and storage medium may reside as discrete components in a computing device or user terminal.
[0246] The prior description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but should be granted the broadest scope that may be consistent with the principles and novel features as defined by the following claims.
Claims
1. A device for audio processing using sound source representations, the device comprising: one or more processors configured to: receive an input audio signal, the input audio signal including sound from a plurality of sound sources; receive a combined representation of a number of sound sources, the combined representation including a representation of sound from a selected one or more sound sources of the plurality of sound sources from the input audio signal; process the input audio signal based on the combined representation of the number of sound sources to generate an output audio signal, wherein: based on a reservation flag having a first value, the combined representation is used to reserve sound from the selected one or more sound sources of the plurality of sound sources from the input audio signal; and based on the reservation flag having a second value, the combined representation is used to remove sound from the selected one or more sound sources of the plurality of sound sources from the input audio signal; and provide the output audio signal to a second device.
2. The device of claim 1, wherein the one or more processors are configured to, based on the reservation flag having the first value, use the combined representation to reserve sound from the selected one or more sound sources of the plurality of sound sources from the input audio signal and remove one or more additional sound sources.
3. The device of claim 2, wherein the one or more processors are configured to set the reservation flag to have the first value indicating that the number of sound sources are to be reserved in response to a detected condition indicating that processing of the input audio signal is to be initiated, wherein the first value of the reservation flag is based on a user input, a default configuration, a configuration input from an application, a configuration request from another device, or a combination thereof.
4. The device of claim 1, wherein the number of sound sources includes one or more authorized users.
5. The device of claim 1, wherein the number of sound sources includes an emergency vehicle.
6. The device of claim 1, wherein the one or more processors are configured to, based on the reservation flag having the second value, use the combined representation to remove sound from the selected one or more sound sources of the plurality of sound sources from the input audio signal and reserve one or more additional sound sources.
7. The device of claim 6, wherein the one or more processors are configured to set the reservation flag to have the second value indicating that the number of sound sources are to be removed in response to a detected condition indicating that processing of the input audio signal is to be initiated, wherein the second value of the reservation flag is based on a user input, a default configuration, a configuration input from an application, a configuration request from another device, or a combination thereof.
8. The device of claim 1, wherein the plurality of sound sources includes traffic, wind, reverberation, channel distortion, another non-vocal sound source, a person, or a combination thereof.
9. The device of claim 1, wherein the plurality of sound sources are associated with background noise in a particular environment.
10. The device of claim 9, wherein the particular environment corresponds to an interior of a particular type of vehicle.
11. The device of claim 1, wherein the combined representation is based on a particular sound from a particular sound source, and wherein the particular sound source is of a same sound source type as one of the number of sound sources.
12. The device of claim 1, wherein the one or more processors are further configured to update the combined representation based on the sound of any of the number of sound sources.
13. The device of claim 1, wherein the one or more processors are further configured to generate the combined representation based on individual representations of the number of sound sources based on a combination setting.
14. The device of claim 13, wherein the one or more processors are further configured to update the combination setting based on user input, a detected condition, or both.
15. The device of claim 1, wherein the plurality of sound sources includes at least a first sound source and a second sound source, wherein a first representation of the first sound source indicates a first value of a particular characteristic, wherein a second representation of the second sound source indicates a second value of the particular characteristic, and wherein a value of the particular characteristic indicated by the combined representation is based on the first value and the second value.
16. The device of claim 15, wherein the first representation includes one or more spectrograms based on sound from a particular sound source of a same type as the first sound source.
17. The device of claim 15, wherein the combined representation corresponds to a concatenation of the first representation of the first sound source and the second representation of the second sound source.
18. The device of claim 1, wherein the one or more processors are configured to process the input audio signal using a neural network to generate the output audio signal.
19. The device of claim 18, wherein the neural network includes a convolutional neural network (CNN), an autoregressive (AR) generation network, an audio generation network (AGN), an attention network (AN), a long short-term memory (LSTM) network, or a combination thereof.
20. The device of claim 18, further comprising a sound source encoder configured to process sound from one or more sound sources to generate a representation of the one or more sound sources, wherein the sound source encoder and the neural network are jointly trained.
21. The device of claim 1, further comprising a receiver configured to receive audio data representative of the input audio signal.
22. The device of claim 1, further comprising a transmitter configured to transmit audio data to the second device, the audio data based on the output audio signal.
23. A method for audio processing using sound source representations, the method comprising: receiving, at a first device, an input audio signal, the input audio signal comprising sound from a plurality of sound sources; receiving a combined representation of a number of sound sources, the combined representation comprising a representation of sound from a selected one or more sound sources of the plurality of sound sources from the input audio signal; processing the input audio signal based on the combined representation of the number of sound sources to generate an output audio signal, wherein: based on a preservation flag having a first value, using the combined representation to preserve sound from the selected one or more sound sources of the plurality of sound sources from the input audio signal; and based on the preservation flag having a second value, using the combined representation to remove sound from the selected one or more sound sources of the plurality of sound sources from the input audio signal; and providing the output audio signal to a second device.
24. The method of claim 23, wherein based on the preservation flag having the first value, using the combined representation to preserve sound from the selected one or more sound sources of the plurality of sound sources from the input audio signal and remove one or more additional sound sources.
25. The method of claim 23, wherein the number of sound sources are associated with background noise in a particular environment.
26. The method of claim 25, wherein the particular environment corresponds to an interior of a particular type of vehicle.
27. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: receive, at a first device, an input audio signal, the input audio signal comprising sound from a plurality of sound sources; receive a combined representation of a number of sound sources, the combined representation comprising a representation of sound from a selected one or more sound sources of the plurality of sound sources from the input audio signal; process the input audio signal based on the combined representation of the number of sound sources to generate an output audio signal, wherein: based on a preservation flag having a first value, using the combined representation to preserve sound from the selected one or more sound sources of the plurality of sound sources from the input audio signal; and based on the preservation flag having a second value, using the combined representation to remove sound from the selected one or more sound sources of the plurality of sound sources from the input audio signal; and provide the output audio signal to a second device.
28. The non-transitory computer-readable medium of claim 27, wherein the input audio signal is processed using a neural network to generate the output audio signal.
29. An apparatus for audio processing using sound source representations, the apparatus comprising: means for receiving, at a first device, an input audio signal, the input audio signal comprising sound from a plurality of sound sources; means for receiving a combined representation of a number of sound sources, the combined representation comprising a representation of sound from a selected one or more sound sources of the plurality of sound sources of the input audio signal; means for processing the input audio signal based on the combined representation of the number of sound sources to generate an output audio signal, wherein: based on the reservation flag having a first value, using the combined representation to reserve sound from the selected one or more sound sources of the plurality of sound sources of the input audio signal; and based on the reservation flag having a second value, using the combined representation to remove sound from the selected one or more sound sources of the plurality of sound sources of the input audio signal; and means for providing the output audio signal to a second device.
30. The apparatus of claim 29, wherein the means for receiving, the means for processing, and the means for providing are integrated into at least one of the following: a smart speaker, a speaker bar, a computer, a tablet computer, a display device, a television, a game console, a music player, a radio, a digital video player, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice-activated device, a portable electronic device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, or a mobile device.
Citation Information
Patent Citations
Ambient sound activated headphone
US10361673B1