Context-based noise reduction during voice call
By dynamically configuring a noise reducer based on audio context detection during voice calls, the device reduces power consumption and extends battery life while maintaining audio quality.
Patent Information
- Application Number
- PCT/US2024/055524
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-12
- Publication Date
- 2025-05-22
AI Technical Summary
The use of audio enhancements during voice calls increases power consumption, as they prevent processing components from entering a low-power state, leading to shorter battery life and a negative user experience.
A device with an audio processor that performs audio context detection during voice calls to dynamically configure a noise reducer, allowing it to process audio inputs efficiently and transition to a low-power state after processing is complete.
This approach reduces power consumption by allowing the audio processing components to stay in a low-power state for longer periods without compromising audio quality, thereby extending battery life and improving user experience.
Smart Images

Figure US2024055524_22052025_PF_FP_ABST
Abstract
Description
CONTEXT-BASED NOISE REDUCTION DURING VOICE CALLI. Cross-Reference to Related Applications
[0001] The present application claims the benefit of priority from the commonly owned Indian Provisional Patent Application No. 202341078080, filed November 17, 2023, the contents of which are expressly incorporated herein by reference in their entirety.IL Field
[0002] The present disclosure is generally related to performing noise reduction based on a detected context during a voice call.III. Description of Related Art
[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there currently exist a variety of portable personal computing devices, including wireless telephones such as mobile and smart phones, tablets and laptop computers that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. Further, many such devices incorporate additional functionality such as a digital still camera, a digital video camera, a digital recorder, and an audio file player. Also, such devices can process executable instructions, including software applications, such as a web browser application, that can be used to access the Internet. As such, these devices can include significant computing capabilities.
[0004] Such computing devices often incorporate functionality to capture user speech from one or more microphones and encode the user speech for transmission to a remote device during a voice call. In some cases, power consumption associated with the voice call can be reduced by having components associated with the voice call, such as a modem and a processor that encodes the user’s speech for transmission, enter a low- power state during periods of the voice call where uplink and downlink communications are not scheduled to occur.
[0005] Audio enhancements can improve audio quality by reducing noise. However, such audio processing can be time consuming and can prevent the processing components from being able to enter the low-power state that would otherwise be available during a voice call. As a result, the use of audio enhancements can result in higher power consumption during a voice call, which can increase the discharge rate of a battery of a mobile communication device, decrease the usage time of the mobile communication device before having to recharge the battery, and negatively impact a user experience.IV. Summary
[0006] According to one embodiment of the present disclosure, a device includes a memory configured to buffer an audio input from an audio source, where the audio input corresponds to a voice call. The device also includes an audio processor configured to, responsive to a transition from a low-power state to an active state during the voice call, perform audio context detection (ACD) on the audio input to obtain an ACD indicator, configure a noise reducer based on a value of the ACD indicator, and process the audio input using the configured noise reducer to generate output audio. The audio processor is also configured to, after generating the output audio, transition from the active state to the low-power state.
[0007] According to another embodiment of the present disclosure, a method includes, responsive to transitioning from a low-power state to an active state during a voice call, obtaining, at an audio processor, an audio input from an audio source, where the audio input corresponds to the voice call. The method also includes performing, at the audio processor, audio context detection (ACD) on the audio input to obtain an ACD indicator. The method further includes configuring, at the audio processor, a noise reducer based on a value of the ACD indicator. The method also includes processing, at the audio processor, the audio input using the configured noise reducer to generate output audio. The method further includes after generating the output audio, transitioning from the active state to the low-power state.
[0008] According to another embodiment of the present disclosure, a non-transitory computer-readable medium stores instructions that, when executed by one or moreprocessors, cause the one or more processors to, responsive to transitioning from a low- power state to an active state during a voice call, obtain an audio input from an audio source, where the audio input corresponds to a voice call. The instructions, when executed by the one or more processors, also cause the one or more processors to perform audio context detection (ACD) on the audio input to obtain an ACD indicator. The instructions, when executed by the one or more processors, further cause the one or more processors to configure a noise reducer based on a value of the ACD indicator. The instructions, when executed by the one or more processors, also cause the one or more processors to process the audio input using the configured noise reducer to generate output audio. The instructions, when executed by the one or more processors, further cause the one or more processors to, after generating the output audio, transition from the active state to the low-power state.
[0009] According to another embodiment of the present disclosure, an apparatus includes means for obtaining, from an audio source, an audio input corresponding to a voice call, where the audio input is obtained responsive to transitioning from a low- power state to an active state during the voice call. The apparatus also includes means for performing audio context detection (ACD) on the audio input to obtain an ACD indicator. The apparatus further includes means for configuring a noise reducer based on a value of the ACD indicator. The apparatus also includes means for processing the audio input using the configured noise reducer to generate output audio. The apparatus further includes means for transitioning from the active state to the low-power state after generating the output audio.
[0010] Other aspects, advantages, and features of the present disclosure will become apparent after review of the entire application, including the following sections: Brief Description of the Drawings, Detailed Description, and the Claims.V. Brief Description of the Drawings
[0011] FIG. l is a diagram that includes a block diagram and a timing diagram of a particular illustrative aspect of a system operable to perform context-based noise reduction during a voice call, in accordance with some examples of the present disclosure.
[0012] FIG. 2 is a diagram of particular aspects of the system of FIG. 1, in accordance with some examples of the present disclosure.
[0013] FIG. 3 is a diagram of particular aspects of the system of FIG. 1, in accordance with some examples of the present disclosure.
[0014] FIG. 4 is a diagram of relative audio processing duration associated with various types of audio enhancement that can be performed by the system of FIG. 1, in accordance with some examples of the present disclosure.
[0015] FIG. 5 illustrates an example of an integrated circuit operable to perform context-based noise reduction during a voice call, in accordance with some examples of the present disclosure.
[0016] FIG. 6 is a diagram of a mobile device operable to perform context-based noise reduction during a voice call, in accordance with some examples of the present disclosure.
[0017] FIG. 7 is a diagram of a headset operable to perform context-based noise reduction during a voice call, in accordance with some examples of the present disclosure.
[0018] FIG. 8 is a diagram of a wearable electronic device operable to perform contextbased noise reduction during a voice call, in accordance with some examples of the present disclosure.
[0019] FIG. 9 is a diagram of a voice-controlled speaker system operable to perform context-based noise reduction during a voice call, in accordance with some examples of the present disclosure.
[0020] FIG. 10 is a diagram of a vehicle operable to perform context-based noise reduction during a voice call, in accordance with some examples of the present disclosure.
[0021] FIG. 11 is a diagram of a particular embodiment of a method of context-based noise reduction during a voice call that may be performed by the system of FIG. 1, in accordance with some examples of the present disclosure.
[0022] FIG. 12 is a block diagram of a particular illustrative example of a device that is operable to perform context-based noise reduction during a voice call, in accordance with some examples of the present disclosure.VI Detailed Description
[0023] Audio enhancements can improve audio quality by suppressing noise and cancelling echo. However, a problem with such audio processing includes preventing the processing components from being able to enter the low-power state that would otherwise be available during a voice call. As a result, the use of audio enhancements can result in higher power consumption during a voice call, which can increase the discharge rate of a battery of a mobile communication device, decrease the usage time of the mobile communication device before having to recharge the battery, and negatively impact a user experience.
[0024] Aspects described herein provide solutions to these, and other, problems by using context-based noise reduction in a transmit path, a receive path, or both, during a voice call. For example, according to a particular aspect, an audio context can indicate background noise levels or whether silence is detected in an audio environment during a voice call over a network (e.g., a long-term evolution new radio (LTE / NR) network). A noise reducer is dynamically configured based on the detected audio context and the configured noise reducer is used to process input audio data to generate output audio data.
[0025] The noise reducer can be configured to perform higher performance noise reduction when the audio context indicates greater than threshold noise. For example, the noise reducer is configured to use a neural network noise suppression engine in addition to a filter based noise suppression engine. Alternatively, the noise reducer can be configured to perform lower latency noise reduction when the audio context indicates lower than threshold noise. For example, the noise reducer is configured to use the filterbased noise suppression engine and to bypass the neural network noise suppression engine. Bypassing the neural network noise suppression engine avoids redundant processing and reduces computational resource usage (e.g., offloading to an embedded neural processor unit (ENPU)) for the neural network noise suppression engine. In some examples, when the audio context indicates no noise, the noise reducer is configured to bypass noise suppression (e.g., the neural network and the filter based noise suppression engines) and to use an echo cancellation engine. The echo cancellation engine can reduce (e.g., remove) any echoes in the audio, further improving the audio quality.
[0026] Restricting to lower latency noise reduction processing at the noise reducer solves the problems of redundant processing by the higher performance noise reduction in low noise contexts and delaying entry into the low-power state. Configuring the noise reducer based on the audio context enables the higher performance noise reduction to be available when the higher performance noise reduction is likely to provide a greater benefit, e.g., when greater than threshold noise is detected. When the higher performance noise reduction is likely to provide less benefit, e.g., when lower than threshold noise is detected, the lower latency noise reduction can be used for the audio processing components to enter the low-power state earlier. The technical advantage of such dynamic configuration of the noise reducer based on the audio context includes increasing the duration that the audio processing components can stay in a low-power state with limited reduction (e.g., no reduction) in audio quality. The longer duration that the audio processing components are in the low-power state reduces power consumption. Thus, the usage time of the communication device between battery charges, and the user experience, are improved.
[0027] In accordance with some aspects, the audio data is processed at an audio processor, such as a digital signal processor. Alignment of audio processing for a voice call with a modem sleep / wake cycle can be achieved, for example, by having the voice call subscribe to a static entity, referred to as a voice timer, that is responsible for scheduling threads of the voice call according to voice call timing criteria. As a result, the audio processing begins at a start timestamp for each sleep / wake cycle that is defined by the call timing criteria. A central sleep manager can track the active / idleduration of all threads running on the audio processor and trigger entry into a low power island mode once all of the threads transition to an idle state, enabling the audio processor to enter a power collapse mode.
[0028] Particular aspects of the present disclosure are described below with reference to the drawings. In the description, common features are designated by common reference numbers. As used herein, various terminology is used for the purpose of describing particular embodiments only and is not intended to be limiting of embodiments. For example, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Further, some features described herein are singular in some embodiments and plural in other embodiments. To illustrate, FIG. 1 depicts a device 102 including one or more audio processors (“audio processor(s)” 106 of FIG. 1), which indicates that in some embodiments the device 102 includes a single audio processor 106 and in other embodiments the device 102 includes multiple audio processors 106. For ease of reference herein, such features are generally introduced as “one or more” features and are subsequently referred to in the singular or optional plural (as indicated by “(s)” in the name of the feature) unless aspects related to multiple of the features are being described.
[0029] In some drawings, multiple instances of a particular type of feature are used. Although these features are physically and / or logically distinct, the same reference number is used for each, and the different instances are distinguished by addition of a letter to the reference number. When the features as a group or a type are referred to herein e.g., when no particular one of the features is being referenced, the reference number is used without a distinguishing letter. However, when one particular feature of multiple features of the same type is referred to herein, the reference number is used with the distinguishing letter. For example, referring to FIG. 1, multiple time periods in which a modem is in an active state are illustrated and associated with reference numbers 162 A and 162B. When referring to a particular one of these time periods, such as a time period 162A, the distinguishing letter “A” is used. However, when referring to any arbitrary one of these time periods or to these time periods as a group, the reference number 162 is used without a distinguishing letter.
[0030] As used herein, the terms “comprise,” “comprises,” and “comprising” may be used interchangeably with “include,” “includes,” or “including.” Additionally, the term “wherein” may be used interchangeably with “where.” As used herein, “exemplary” indicates an example, an embodiment, and / or an aspect, and should not be construed as limiting or as indicating a preference or a preferred embodiment. As used herein, an ordinal term (e.g., “first,” “second,” “third,” etc.) used to modify an element, such as a structure, a component, an operation, etc., does not by itself indicate any priority or order of the element with respect to another element, but rather merely distinguishes the element from another element having a same name (but for use of the ordinal term). As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to multiple (e.g., two or more) of a particular element.
[0031] As used herein, “coupled” may include “communicatively coupled,” “electrically coupled,” or “physically coupled,” and may also (or alternatively) include any combinations thereof. Two devices (or components) may be coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) directly or indirectly via one or more other devices, components, wires, buses, networks (e.g., a wired network, a wireless network, or a combination thereof), etc. Two devices (or components) that are electrically coupled may be included in the same device or in different devices and may be connected via electronics, one or more connectors, or inductive coupling, as illustrative, non-limiting examples. In some embodiments, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive signals (e.g., digital signals or analog signals) directly or indirectly, via one or more wires, buses, networks, etc. As used herein, “directly coupled” may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without intervening components.
[0032] In the present disclosure, terms such as “determining,” “calculating,” “estimating,” “shifting,” “adjusting,” etc. may be used to describe how one or more operations are performed. It should be noted that such terms are not to be construed as limiting and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, “obtaining,” “generating,” “calculating,” “estimating,” “using,” “selecting,” “accessing,” and “determining” may be usedinterchangeably. For example, “obtaining,” “generating,” “calculating,” “estimating,” or “determining” a parameter (or a signal) may refer to actively generating, estimating, calculating, or determining the parameter (or the signal) or may refer to using, selecting, or accessing the parameter (or signal) that is already generated, such as by another component or device.
[0033] Referring to FIG. 1, a particular illustrative aspect of a system 100 and a timing diagram 104 associated with performing context-based noise reduction during a voice call are shown. In the example illustrated in FIG. 1, the system 100 includes a device 102 configured to process audio input 124 for transmission to, and playout at, another device 154 during the voice call. In an illustrative example, the device 154 corresponds to a mobile phone, a headset device, etc., to enable telephonic communication between a user of the device 102 and the device 154 over one or more wired or wireless communication networks (e.g., long-term evolution (LTE), New Radio (NR), etc.) (LTE is a trademark of European Telecommunications Standards Institute). In a particular aspect, the voice call is over a LTE network. In a particular aspect, the voice call is over at least one of an NR network, a 5G network, or a beyond 5G wireless network.
[0034] The device 102 includes one or more audio processors 106 coupled to a modem 150. The audio processor 106 includes a digital signal processor (DSP), one or more other types of processor, or a combination thereof. The audio processor 106 is configured to transition between active and low-power states substantially concurrently with corresponding transitions of the modem 150 that are based on timing criteria associated with the voice call. As a result, power consumption associated with audio processing during the voice call can be reduced.
[0035] The audio processor 106 is configured, responsive to transitioning from a low- power state to an active state during the voice call, to obtain an audio input 124 (e.g., a user’s voice data) of the voice call. According to an aspect, the audio input 124 corresponds to one or more frames of voice data that are received for processing at the audio processor 106 during a voice call. In an illustrative embodiment, the audio input 124 is received from an audio source, such as via a microphone that is implemented inor coupled to the device 102. The audio input 124 can be processed for transmission to the device 154 as voice content of the voice call.
[0036] The audio processor 106 is also configured to generate output audio 142 based on the audio input 124. To illustrate, the audio processor 106 includes an audio enhancer 130 that includes an audio context detector 132 and a noise reducer 134. The audio context detector 132 is configured to detect an audio context associated with the audio input 124. For example, the audio context indicates a level of noise detected in the audio input 124 or whether silence is detected in the audio input 124. The noise reducer 134 includes one or more noise reduction engines, such as a high performance noise suppression engine, a low latency noise suppression engine, an echo cancellation engine, one or more other types of noise reduction engines, or a combination thereof.
[0037] The audio enhancer 130 configures the noise reducer 134 based on the detected audio context to bypass noise reduction or to use particular ones of the one or more noise reduction engines, as further described with reference to FIGS. 2-3. In an example, the audio enhancer 130, based on the detected audio context indicating a greater than threshold noise, configures the noise reducer 134 to use the high performance noise suppression engine, the low latency noise suppression engine, and the echo cancellation engine. Alternatively, the audio enhancer 130, based on the detected audio context indicating noise that is lower than a threshold, configures the noise reducer 134 to bypass the high performance noise suppression engine and to use the low latency noise suppression engine and the echo cancellation engine. In another example, the audio enhancer 130, based on the detected audio context indicating no noise, configures the noise reducer 134 to bypass noise suppression and to use the echo cancellation engine. In yet another example, the audio enhancer 130, based on the detected audio context indicating no sound, configures the noise reducer 134 to bypass noise reduction.
[0038] The configured noise reducer 134 processes the audio input 124 to generate output audio 126. For example, based on the configuration, the noise reducer 134 bypasses noise reduction or uses particular ones of the one or more of the noise reduction engines to generate the output audio 126. An audio processing durationchanges based on the configuration. For example, using fewer than all of the noise reduction engines reduces the audio processing duration.
[0039] According to an aspect, the audio processor 106 is further configured to encode the output audio 126 at a codec 140 to generate the output audio 142. After generating the output audio 142, the audio processor 106 is configured to transition from the active state back to the low-power state based on timing criteria associated with the voice call. The modem 150 is also configured to transition between a low-power state and an active state based on the timing criteria associated with the voice call, and is configured to initiate transmission of an output signal 152 based on the output audio 142 while in the active state.
[0040] According to some aspects, a voice timer is responsible for scheduling threads of a voice call according to the voice call timing criteria. In a particular aspect, the voice timer can correspond to a software thread of the audio processor 106 that assigns resources, such as clocks and memory bandwidth, to the various subscribed threads so that resources are allocated to audio processing threads during awake periods and deallocated from the audio processing threads during the low-power periods. A central sleep manager is configured to track processing threads at the audio processor 106 and control transitions of the audio processor 106 between an active state and a low power island state. In a particular example, the central sleep manager corresponds to a duty cycle manager and is configured to trigger entry into a low power island state in response to detecting that the audio processing threads are idle. By using the voice timer to schedule the audio processing threads, the voice timer can ensure that all of the audio processing threads are idle at the audio processor 106 as early as possible during the low-power periods 164, enabling the central sleep manager to trigger entry into the low power island state and resulting in power savings.
[0041] The timing diagram 104 illustrates an example of operation of the device 102 in which transitions between an active state and a low-power state of the modem 150 are aligned with the transitions between the active state and the low-power state of the audio processor 106 to enable synchronized processing and power savings using a low power island mechanism (e.g., a low power island memory). The timing diagram 104depicts modem operations 160 and audio processing operations 170 during multiple cycles 158 associated with the voice call, including a first cycle (“cycle 1”) 158 A and a second cycle (“cycle 2”) 158B. In each cycle 158, an awake period 162 indicates a time period in which the modem 150 is in an active state, and a low-power period 164 indicates a time period in which the modem 150 is not active and can enter a low-power state (e.g., a Deep / Light Sleep (“DLS”) mode) to conserve power. In a particular embodiment, the voice call is a connected mode discontinuous reception (CDRx) call, and timing criteria associated with the cycles 158 (e.g., the length of the awake period 162 and the length of the low-power period 164) are based on a CDRx cycle configuration. In an illustrative, non-limiting example, the duration of each cycle 158 is 40 milliseconds (ms), the duration of the awake period 162 is 20 ms, and the duration of the low-power period 164 is 20 ms. The low-power period 164 having the same duration as the awake period 162 is provided as an illustrative example, in other examples the low-power period 164 can be shorter or longer than the awake period 162 based on a cycle configuration.
[0042] The first cycle 158 A begins with an awake period 162 A, during which the modem 150 and the audio processor 106 transition from a low-power state to an active state. During the awake period 162, the modem 150 performs one or more uplink transmissions, one or more downlink transmissions, or a combination thereof, associated with the voice call. The audio processor 106 performs audio processing operations during an audio processing period 172A, e.g., data loading, context detection, noise reducer configuration, audio enhancement, and encoding operations of the audio input 124. In an illustrative example, the audio input 124 represents 40 ms of audio content. In an example, a first portion of the audio input 124 includes microphone data that was buffered while the audio processor 106 was in the low-power state and retrieved upon the audio processor 106 transitioning to the active state. In an example, a second portion of the audio input 124 includes microphone data that was at least partially buffered subsequent to the audio processor 106 transitioning to the active state. In another example, all of the audio input 124 can be buffered while the audio processor 106 was in the low-power state. In yet another example, all of the audio input 124 can be added to the buffer subsequent to the audio processor 106 transitioning tothe active state. To illustrate, the audio processor 106 can retrieve portions of the audio input 124 that are being written to the buffer in the active state, that have previously been written to the buffer in the low-power state, or a combination thereof.
[0043] The device 102 thus performs audio data retrieval, context detection, noise reducer configuration, audio enhancement, and encoding to generate the output audio 142 at the audio processor 106, and also performs transmission of the output signal 152 via the modem 150, during the awake period 162A. Upon completion of the awake period 162A, the modem 150 and the audio processor 106 halt operations and enter a low-power state during a low-power period 164A. To illustrate, the modem 150 ceases uplink and downlink activity and transitions to a sleep mode (or other low-power state) for the remainder of the first cycle 158 A, and the audio processor 106 ceases processing of the audio input 124 and transitions to a low-power state for the remainder of the first cycle 158 A.
[0044] Upon completion of the low-power period 164 A of the first cycle 158 A, the second cycle 158B commences with an awake period 162B, during which the modem 150 and the audio processor 106 each transition from a low-power state to an active state. During the awake period 162B, the modem 150 resumes uplink and / or downlink activity associated with the voice call, and the audio processor 106 resumes processing of the audio input 124 to generate a next set of output audio 142 for transmission to the device 154 via the modem 150.
[0045] To illustrate, the audio processor 106 performs audio processing operations 170 during an audio processing period 172B of the awake period 162B, e.g., a data loading, context detection, noise reducer configuration, audio enhancement, and encoding operations of one or more additional portions of the audio input 124. In an example, a third portion of the audio input 124 includes microphone data that was buffered during the low-power period 164 A and retrieved upon the audio processor 106 transitioning to the active state.
[0046] Upon completion of the awake period 162B, the modem 150 and the audio processor 106 halt operations and enter a low-power state during a low-power period 164B. To illustrate, the modem 150 ceases uplink and downlink activity and transitionsto a sleep mode (or other low-power state) for the remainder of the second cycle 158B, and the audio processor 106 ceases processing of the audio input 124 and transitions to a low-power state for the remainder of the second cycle 158B.
[0047] Synchronization of the modem operations 160 with the audio processing operations 170 can be performed using a voice call timer to schedule audio processing threads at the audio processor 106 according to timing criteria of the voice call. A central sleep manager can be configured to trigger entry into a low power island state in response to detecting that the audio processing threads are idle.
[0048] By reducing a duration of the audio processing operations 170 associated with a voice call based on a detected context, the audio processor 106 can enter the low-power state earlier during the low-power periods 164 associated with the sleep / wake cycle of the modem 150 and defined by the call timing criteria. The audio processing duration is reduced selectively when lower latency audio processing is less likely to impact audio quality. As a result, power consumption of the audio processor 106 is reduced with low or no adverse impact on audio quality as compared to conventional systems in which either entry into the low-power state is prevented by higher latency audio processing or audio quality is reduced by using lower performance audio processing.
[0049] The audio enhancer 130 included in a transmit path to generate the output signal 152 that can be transmitted to the device 154 is provided as an illustrative example. In some examples, an audio enhancer 130 can additionally, or in the alternative, be included in a receive path to process audio data received from the device 154. For example, the modem 150 provides audio data received from the device 154 to the codec 140 and the codec 140 decodes the audio data to generate decoded audio data. The audio context detector 132 determines an audio context of the decoded audio data. The audio enhancer 130 configures the noise reducer 134 based on the detected audio context, as further described with reference to FIGS. 2-3. The configured noise reducer 134 processes the decoded audio data to generate output audio that can be provided to a speaker coupled to the device 102 for playback, stored in a memory device coupled to the device 102, processed (e.g., for speech recognition) by another component of the device 102, or a combination thereof. After generating the output audio, the audioprocessor 106 is configured to transition from the active state back to the low-power state based on timing criteria associated with the voice call.
[0050] In some examples, the audio processor 106 can perform audio processing operations in the receive path concurrently with performing audio processing operation in the transmit path during the audio processing period 172 A. For example, the audio processor 106 can perform decoding operations, context detection, noise reducer configuration, and audio enhancement of audio data received from the device 154 concurrently with performing the data loading, context detection, noise reducer configuration, audio enhancement, and encoding operations of the audio input 124 obtained from a microphone coupled to the device 102. One or more context detection, noise reducer configuration, and audio enhancement operations can be performed during an audio processing period 172. One or more encoding operations for voice call data to be transmitted via the modem 150 can be performed during an audio processing period 172. Similarly, one or more decoding operations for voice call data received via the modem 150 can be performed during an audio processing period 172.
[0051] FIG. 2 is a diagram of particular aspects of the system of FIG. 1, in accordance with some examples of the present disclosure. In particular, FIG. 2 highlights an example of components of the audio enhancer 130. In the example illustrated in FIG. 2, the audio enhancer 130 includes a configurer 220 that is coupled to the audio context detector 132 and to the noise reducer 134. In a particular optional embodiment, the audio enhancer 130 also includes a pre-processor 240 coupled to the noise reducer 134.
[0052] The noise reducer 134 includes multiple engines configured to perform noise reduction. In FIG. 2, the multiple engines include a low latency noise suppression engine 202, a high performance noise suppression engine 204, and an echo cancellation engine 206. The low latency noise suppression engine 202 and the high performance noise suppression engine 204 are each configured to perform noise suppression. In a particular aspect, the low latency noise suppression engine 202 is faster (e.g., has fewer computation cycles) than the high performance noise suppression engine 204, whereas the high performance noise suppression engine 204 is configured to remove more noise than the low latency noise suppression engine 202. In a particular aspect, the lowlatency noise suppression engine 202 is a filter based noise suppression engine, and the high performance noise suppression engine 204 is a neural network noise suppression engine. The echo cancellation engine 206 is configured to perform echo cancellation.
[0053] The audio enhancer 130 operates substantially as described with reference to FIG. 1. In particular, the audio enhancer 130 is configured to, responsive to a transition from a low-power state to an active state during a voice call, process audio input 224 to generate output audio 226. In some aspects, the audio input 224 corresponds to the audio input 124 of FIG. 1 and is received from one or more microphones that are integrated in, or coupled to, the device 102. In these aspects, the output audio 226 corresponds to the output audio 126 of FIG. 1 that is encoded and made accessible to the modem 150 to transmit to the device 154. In other aspects, the audio input 224 (e.g., decoded audio data) is based on audio data received from the device 154 of FIG. 1. In these aspects, the output audio 226 can be provided for playback to a speaker coupled to the device 102, provided for storage to a memory device coupled to the device 102, provided for additional processing to one or more components of the device 102, or a combination thereof.
[0054] In a particular optional embodiment, the pre-processor 240 performs various preprocessing operations on the audio input 224 to generate audio input 242. The preprocessing operations can include conversion between time domain and frequency domain, sound source separation, other types of audio processing, or a combination thereof. In another embodiment, the audio enhancer 130 provides the audio input 224 as the audio input 242 to the noise reducer 134 for processing.
[0055] The audio context detector 132 performs audio context detection (ACD) on the audio input 224 to obtain an ACD indicator 210. In some examples, the ACD indicator 210 indicates a level of noise detected in the audio input 224, whether silence is detected in the audio input 224, or both. The audio context detector 132 provides the ACD indicator 210 to the configurer 220.
[0056] The configurer 220 configures the noise reducer 134 based on a comparison of a value of the ACD indicator 210 to one or more thresholds 222, as further described with reference to FIG. 3. For example, the configurer 220, based on a comparison of thevalue of the ACD indicator 210 to the one or more thresholds 222, generates a command 212 to set configuration parameters 214 of the noise reducer 134.
[0057] In a particular aspect, the configurer 220, based on a comparison of the value of the ACD indicator 210 to the one or more thresholds 222, generates the command 212 to set configuration parameters 214 to disable the noise reducer 134, as further described with reference to FIG. 3. In a particular aspect, the configurer 220, based on a comparison of the value of the ACD indicator 210 to the one or more thresholds 222, generates the command 212 to set configuration parameters 214 to selectively enable one or more of the multiple engines of the noise reducer 134.
[0058] In a particular embodiment, the configurer 220, based on a comparison of the value of the ACD indicator 210 to the one or more thresholds 222, configures the noise reducer 134 to use a neural network noise suppression engine (e.g., the high performance noise suppression engine 204). In a particular embodiment, the configurer 220, based on a comparison of the value of the ACD indicator 210 to the one or more thresholds 222, configures the noise reducer 134 to bypass a neural network noise suppression engine (e.g., the high performance noise suppression engine 204).
[0059] In a particular embodiment, the configurer 220, based on a comparison of the value of the ACD indicator 210 to the one or more thresholds 222, configures the noise reducer 134 to use a low latency noise suppression engine (e.g., the low latency noise suppression engine 202) and to bypass a high performance noise suppression engine (e.g., the high performance noise suppression engine 204). In a particular embodiment, the configurer 220, based on a comparison of the value of the ACD indicator 210 to the one or more thresholds 222, configures the noise reducer 134 to restrict noise reduction processing to using an echo cancellation engine (e.g., the echo cancellation engine 206). For example, the configurer 220 configures the noise reducer 134 to bypass noise suppression (e.g., disables the low latency noise suppression engine 202 and disables the high performance noise suppression engine 204).
[0060] If the noise reducer 134 is disabled, the noise reducer 134 outputs the audio input 242 as the output audio 226. Alternatively, the noise reducer 134 uses one ormore of the multiple engines that are enabled to process the audio input 242 to generate the output audio 226, as further described with reference to FIG. 3.
[0061] A technical advantage of selectively enabling one or more of multiple engines of the noise reducer 134 based on the ACD indicator 210 includes dynamically adjusting the duration of audio enhancement processing that is performed by the noise reducer 134. For example, higher performance audio enhancement with a higher performance engine, more engines, or a combination thereof, can be performed when the ACD indicator 210 indicates an audio context (e.g., noisy audio) that is likely to benefit more from the higher performance audio enhancement. Lower latency audio enhancement with a lower latency engine, fewer engines, or a combination thereof, can be performed when the ACD indicator 210 indicates an audio context (e.g., less noise audio, or silence) that is likely to benefit less from high performance audio enhancement. Use of the lower latency audio enhancement can reduce the duration of audio enhancement processing, which can cause the audio enhancer 130 to complete processing prior to the end of the current awake period 162. Thus, the audio processor 106 and the modem 150 may enter the low-power state substantially concurrently, resulting in reduced power consumption. In some examples, even when the audio enhancement processing is not completed prior to the end of the current awake period 162, reducing the duration of the audio enhancement processing by using the lower latency audio enhancement still provides a power saving advantage as compared to higher-latency audio enhancement, at least in part by causing the audio enhancer 130 to stop blocking the audio processor 106 of FIG. 1 from entering the low-power state earlier after the current awake period 162.
[0062] The noise reducer 134 including the low latency noise suppression engine 202, the high performance noise suppression engine 204, and the echo cancellation engine 206 is provided as an illustrative example. In other examples, the noise reducer 134 can include one or more of the low latency noise suppression engine 202, the high performance noise suppression engine 204, and the echo cancellation engine 206, one or more additional noise reduction engines, one or more different noise engines, or a combination thereof.
[0063] FIG. 3 is a diagram 300 of particular aspects of the system of FIG. 1, in accordance with some examples of the present disclosure. In particular, the diagram 300 highlights an example of components of the noise reducer 134.
[0064] In the example illustrated in FIG. 3, the noise reducer 134 includes a linear echo cancellation (EC) engine 302 and an echo cancellation post-processing (ECPP) engine 308 that correspond to the echo cancellation engine 206 of FIG. 2. In a particular aspect, the high performance noise suppression engine 204 includes neural network noise suppression engine. For example, the high performance noise suppression engine 204 includes an artificial intelligence (Al) based noise suppression (AINS) engine 304 couped to an embedded neural processing unit (ENPU) 310. In a particular aspect, the low latency noise suppression engine 202 includes a non-AI based noise suppression (NSPP) engine 306.
[0065] The configuration parameters 214 include an EC bit 312, an AINS bit 314, an NSPP bit 316, and an ECPP bit 318 indicating whether the linear EC engine 302, the AINS engine 304, the NSPP engine 306, and the ECPP engine 308, respectively, are enabled. In a particular aspect, the AINS engine 304 and the ENPU 310 are enabled or disabled together. In a particular embodiment, a first value (e.g., 0) of the EC bit 312, the AINS bit 314, the NSPP bit 316, or the ECPP bit 318 indicates that a corresponding engine is disabled, and a second value (e.g., 1) of the EC bit 312, the AINS bit 314, the NSPP bit 316, or the ECPP bit 318 indicates that a corresponding engine is enabled. The configuration parameters 214 corresponding to four noise reduction engines is provided as an illustrative example, in other examples the configuration parameters 214 can correspond to fewer than four or more than four noise reduction engines.
[0066] The configurer 220, based on a comparison of a value of the ACD indicator 210 and the one or more thresholds 222 sends a command 212 to set the configuration parameters 214. In an example 350, the configurer 220 determines whether sound is detected based on a comparison of the value of the ACD indicator 210 and a first threshold (Tl), at 352. The configurer 220, at 354, disables the noise reducer 134 in response to determining that the value of the ACD indicator 210 is less than or equal to the first threshold (Tl) indicating that no sound is detected. For example, the configurer220 sends the command 212 to set each of the configuration parameters 214 (e.g., the EC bit 312, the AINS bit 314, the NSPP bit 316, the ECPP bit 318, or a combination thereof) to the first value (e.g., 0) to disable the noise reducer 134.
[0067] The configurer 220, in response to determining that the value of the ACD indicator 210 is greater than the first threshold (Tl) indicating that sound is detected, at 352, determines whether the value of the ACD indicator 210 is greater than a third threshold (T3), at 358. The configurer 220, in response to determining that the value of the ACD indicator 210 is greater than the third threshold (T3) indicating a higher level of detected noise, sends the command 212 to set each of the configuration parameters 214 (e.g., the EC bit 312, the AINS bit 314, the NSPP bit 316, the ECPP bit 318, or a combination thereof) to the second value (e.g., 1) to enable each of the multiple noise reduction engines of the noise reducer 134.
[0068] The configurer 220, in response to determining that the value of the ACD indicator 210 is less than or equal to the third threshold (T3), at 358, determines whether the value of the ACD indicator 210 is greater than a second threshold (T2), at 360. The configurer 220, in response to determining that the value of the ACD indicator 210 is greater than the second threshold (T2), at 360, sends the command 212 to set one or more first bits of the configuration parameters 214 to the first value (e.g., 0) to disable a first subset of the multiple noise reduction engines of the noise reducer 134 and to set one or more second bits of the configuration parameters 214 to the second value (e.g., 1) to enable a second subset of the multiple noise reduction engines of the noise reducer 134.
[0069] In an example, the value of the ACD indicator 210 greater than the second threshold (T2) and less than or equal to the third threshold (T3) indicates an intermediate level of noise. In this example, the configurer 220 sends the command 212 to set the AINS bit 314 to the first value (e.g., 0) and to set each of the EC bit 312, the NSPP bit 316, and the ECPP bit 318 to the second value (e.g., 1). To illustrate, the noise reducer 134 is configured to bypass the neural network noise suppression engine and to enable a filter based noise suppression engine and an echo cancellation engine.
[0070] The configurer 220, in response to determining that the value of the ACD indicator 210 is less than or equal the second threshold (T2), at 360, sends the command 212 to set one or more third bits of the configuration parameters 214 to the first value (e.g., 0) to disable a third subset of the multiple noise reduction engines of the noise reducer 134 and one or more fourth bits of the configuration parameters 214 to the second value (e.g., 1) to enable a fourth subset of the multiple noise reduction engines of the noise reducer 134.
[0071] In an example, the value of the ACD indicator 210 greater than the first threshold (Tl) and less than or equal to the second threshold (T2) indicates a lower level of noise. In this example, the configurer 220 sends the command 212 to set each of the AINS bit 314 and to the NSPP bit 316 to the first value (e.g., 0) and to set each of the EC bit 312 and the ECPP bit 318 to the second value (e.g., 1). To illustrate, the noise reducer 134 to bypass noise suppression and to restrict noise reduction processing to using an echo cancellation engine.
[0072] The audio enhancer 130 uses the configured noise reducer 134 to process the audio input 242 to generate the output audio 226. In an example, when the linear EC engine 302 is enabled, the linear EC engine 302 performs first echo cancellation operations (e.g., linear echo cancellation) on the audio input 242 based on a reference input 342 to generate an output 303. In a particular aspect, when the linear EC engine 302 is disabled, the first echo cancellation operations are bypassed and the audio input 242 is used as the output 303.
[0073] In a particular aspect, the audio input 242 corresponds to audio data received via the modem 150 from the device 154 of FIG. 1 during a call and the reference input 342 corresponds to microphone output from a microphone coupled to the device 102 of FIG.1. In a particular aspect, the audio input 242 corresponds to first microphone output from a first microphone and the reference input 342 corresponds to second microphone output from a second microphone. In some aspects, the first microphone, the second microphone, or both, are coupled to the device 102 of FIG. 1.
[0074] When the high performance noise suppression engine 204 (e.g., the AINS engine 304) is enabled, the high performance noise suppression engine 204 performs first noisesuppression operations on the output 303 to generate an output 305. In a particular aspect, when the high performance noise suppression engine 204 (e.g., the AINS engine 304) is disabled, the first noise suppression operations are bypassed and the output 303 is used as the output 305. In a particular aspect, the high performance noise suppression engine 204 corresponds to a neural network noise suppression engine and at least a portion of audio processing of the neural network noise suppression engine is performed at the ENPU 310.
[0075] When the low latency noise suppression engine 202 (e.g., the NSPP engine 306) is enabled, the low latency noise suppression engine 202 performs second noise suppression operations on the output 305 to generate an output 307. In a particular aspect, when the low latency noise suppression engine 202 (e.g., the NSPP engine 306) is disabled, the second noise suppression operations are bypassed and the output 305 is used as the output 307. In a particular aspect, the low latency noise suppression engine 202 corresponds to a filter based noise suppression engine. In a particular aspect, the low latency noise suppression engine 202 is configured to use procedural processing to perform noise suppression on the output 305 to generate the output 307.
[0076] When the ECPP engine 308 is enabled, the ECPP engine 308 performs second echo cancellation operations (e.g., non-linear echo cancellation) on the output 307 to generate the output audio 226. In a particular aspect, when the ECPP engine 308 is disabled, the second echo cancellation operations are bypassed and the output 307 is used as the output audio 226.
[0077] The echo cancellation engine 206 including the linear EC engine 302 and the ECPP engine 308 is provided as an illustrative example, in other examples the echo cancellation engine 206 can include a single one of the linear EC engine 302 or the ECPP engine 308, one or more additional echo cancellation engines, or a combination thereof. The high performance noise suppression engine 204 including the AINS engine 304 and the ENPU 310 is provided as an illustrative example, in other examples the high performance noise suppression engine 204 can include a single one of the AINS engine 304 or the ENPU 310, one or more additional high performance noise suppression engines 204, or a combination thereof.
[0078] The low latency noise suppression engine 202 including the non-AI based noise suppression (NSPP) engine 306 is provided as an illustrative example. In other examples, the low latency noise suppression engine 202 can include one or more additional low latency noise suppression engines, one or more different low latency noise suppression engines, or a combination thereof.
[0079] FIG. 4 is a diagram 400 of relative audio processing duration associated with various types of audio enhancement that can be performed by the system 100 of FIG. 1, in accordance with some examples of the present disclosure.
[0080] In particular, the diagram 400 depicts a portion of a cycle 158 used by the audio enhancer 130 to perform the audio processing operations 170 with the audio enhancer 130 transitioning to a low-power state during a remainder portion of the cycle 158.
[0081] An example 402 depicts an audio processing (AP) portion of the cycle 158 during which the audio enhancer 130 uses the AINS engine 304, the NSPP engine 306, the linear EC engine 302, and the ECPP engine 308 to perform the audio processing operations 170 to generate the output audio 226, as described with reference to FIG. 3.
[0082] An example 404 depicts an audio processing portion of the cycle 158 during which the audio enhancer 130 uses the NSPP engine 306, the linear EC engine 302, and the ECPP engine 308 to perform the audio processing operations 170 to generate the output audio 226, as described with reference to FIG. 3. In the example 404, the neural network noise suppression engine is bypassed, reducing the duration of the audio processing portion relative to the example 402.
[0083] An example 406 depicts an audio processing portion of the cycle 158 during which the audio enhancer 130 uses the linear EC engine 302 and the ECPP engine 308 to perform the audio processing operations 170 to generate the output audio 226, as described with reference to FIG. 3. In the example 406, noise suppression is bypassed and the noise reduction is restricted to echo cancellation, reducing the duration of the audio processing portion relative to the example 404.
[0084] An example 408 depicts an audio processing portion of the cycle 158 during which the audio enhancer 130 bypasses the noise reducer 134 to generate the output audio 226, as described with reference to FIG. 3. In some examples, the audio processor 106 performs one or more audio processing operations 170 that are unrelated to operations of the noise reducer 134. In the example 408, the noise reducer 134 is bypassed, reducing the duration of the audio processing portion relative to the example 406.
[0085] Dynamically configuring the noise reducer 134 based on the ACD indicator 210 of FIG. 2 enables the audio processing portion of the cycle 158 to be adjusted. For example, all of the engines of the noise reducer 134 can be enabled during a high noise context when high performance noise reduction is likely to have a greater impact on user experience, and one or more of the engines of the noise reducer 134 can be disabled when high performance noise reduction is likely to have a lower impact on user experience to enable an earlier transition to the low-power state.
[0086] Upon completion of the audio processing, the modem 150 and the audio processor 106 halt operations and enter an idle state, and a central sleep manager initiates the transition from an active state to a low power island state during a low- power period 164 of FIG. 1. According to some aspects, the modem 150 ceases uplink and downlink activity and transitions to a sleep mode (or other low-power state) for the remainder of the cycle 158, and the audio processor 106 ceases processing of the audio input 124 and transitions to a low-power state for the remainder of the cycle 158. Upon completion of the cycle 158, another cycle 158 commences during which the central sleep manager transitions the audio processor 106 from the low power island state to the active state and the audio processor 106 obtains another portion of the audio input 124 for processing and the modem 150 also resumes uplink and / or downlink activity associated with the voice call.
[0087] FIG. 5 depicts an embodiment 500 of the device 102 as an integrated circuit 502 that includes one or more processors 510. The one or more processors 510 include the audio processor 106 and optionally include the modem 150, a central sleep manager,and a voice timer. The audio processor 106 includes the audio context detector 132 and the noise reducer 134 and optionally includes the codec 140.
[0088] The integrated circuit 502 also includes a data input 504, such as one or more microphone inputs and / or bus interfaces, to enable audio data 508 to be received for processing. To illustrate, the audio data 508 can correspond to the audio input 124, the audio input 224, the reference input 342, or a combination thereof, as illustrative, nonlimiting examples. The integrated circuit 502 also includes a signal output 506, such as a bus interface, to enable sending of an output signal 512, such as the output audio 142, the output signal 152, or the output audio 226, as illustrative, non-limiting examples. The integrated circuit 502 enables the audio processor 106 to be integrated (e.g., included as a component) in a system that includes microphones, such as a mobile phone or tablet computer device as depicted in FIG. 6, a headset device that includes a microphone configured to provide the audio data 508, as depicted in FIG. 7, a wearable electronic device as depicted in FIG. 8, a voice-controlled speaker system as depicted in FIG. 9, or a vehicle as depicted in FIG. 10.
[0089] FIG. 6 depicts an embodiment 600 in which the device 102 includes a mobile device 602, such as a phone or tablet computer device, as illustrative, non-limiting examples. The mobile device 602 includes a microphone 630 and a display screen 604. The one or more processors 510 including the audio processor 106 are integrated in the mobile device 602 and are illustrated using dashed lines to indicate internal components that are not generally visible to a user of the mobile device 602. In a particular example, the audio processor 106 is configured to, responsive to user instructions (e.g., received via a graphical user interface at the display screen 604), initiate context-based noise reduction during a voice call and to align active periods of audio processing with call timing criteria to enable low-power operation (e.g., to support a low power island state) during the voice call.
[0090] FIG. 7 depicts an embodiment 700 in which the device 102 includes a headset device 702. The headset device 702 includes a microphone 710, and the one or more processors 510 including the audio processor 106 are integrated in the headset device 702. In a particular example, the audio processor 106 is configured to, responsive touser instructions (e.g., received via one or more user controls of the headset device 702, or via a speech interface, as non-limiting examples), initiate context-based noise reduction during a voice call and to align the active periods of audio processing with call timing criteria to enable low-power operation (e.g., to support a low power island state) during the voice call. Although illustrated as an audio headset, in other embodiments the headset device 702 can correspond to an extended reality headset, such as a virtual reality, mixed reality, or augmented reality headset.
[0091] FIG. 8 depicts an embodiment 800 in which the device 102 includes a wearable electronic device 802, illustrated as a “smart watch.” A microphone 810 and the one or more processors 510 including the audio processor 106 are integrated into the wearable electronic device 802. In a particular example, the audio processor 106 is configured to, responsive to user instructions, such as via a graphical user interface at a display screen 804 of the wearable electronic device 802, initiate context-based noise reduction during a voice call and to align the active periods of audio processing with call timing criteria to enable low-power operation (e.g., to support a low power island state) during the voice call. In a particular example, the wearable electronic device 802 includes a haptic device that provides a haptic notification (e.g., vibrates) in response to detection of an incoming call during which the user can initiate context-based noise reduction. For example, the haptic notification can cause a user to look at the display screen 804 of the wearable electronic device 802 to see a displayed notification indicating an incoming call, including a prompt to perform context-based noise reduction during the call with the calling party. The wearable electronic device 802 can thus alert a user of the option to perform context-based noise reduction during a voice call.
[0092] FIG. 9 is an embodiment 900 in which the device 102 includes a wireless speaker and voice activated device 902. The wireless speaker and voice activated device 902 can have wireless network connectivity and is configured to execute an assistant operation. A microphone 910 and the one or more processors 510 including the audio processor 106 are included in the wireless speaker and voice activated device 902. The wireless speaker and voice activated device 902 also includes a speaker 942 and supports use of a wireless headset, illustrated as a pair of in-ear earphones 990, which can optionally be used by a user for participating in voice calls via the wirelessspeaker and voice activated device 902. During operation, in response to receiving a verbal command identified as user speech via the microphone 910 or via wireless signaling from the earphones 990, the wireless speaker and voice activated device 902 can execute assistant operations, such as via execution of a voice activation system (e.g., an integrated assistant application). The assistant operations can include initiating context-based noise reduction during an ongoing voice call or during initiation of a voice call.
[0093] FIG. 10 depicts an embodiment 1000 in which the device 102 corresponds to, or is integrated within, a vehicle 1002, illustrated as a car. The vehicle 1002 includes the one or more processors 510 including the audio processor 106. The vehicle 1002 also includes microphones 1010 positioned to capture utterances of an operator and / or one or more users of the vehicle 1002. User voice activity detection can be performed based on audio signals received from the microphones 1010, including one or more user commands to initiate context-based noise reduction during an ongoing voice call, in response to accepting an incoming voice call, or during initiation of a voice call. For example, when an incoming voice call is detected, the user may be prompted via a display 1046 or via the one or more speakers 1042 if the user would like to initiate context-based noise reduction during the call with the calling party.
[0094] Referring to FIG. 11, a particular embodiment of a method 1100 of performing context-based noise reduction during a voice call is shown. In a particular aspect, one or more operations of the method 1100 are performed by at least one of the audio context detector 132, the noise reducer 134, the audio enhancer 130, the audio processor 106, the modem 150, the device 102, the system 100 of FIG. 1, the configurer 220, the low latency noise suppression engine 202, the high performance noise suppression engine 204, the echo cancellation engine 206 of FIG. 2, the linear EC engine 302, the AINS engine 304, the ENPU 310, the NSPP engine 306, the ECPP engine 308 of FIG. 3, or a combination thereof.
[0095] The method 1100 includes, at block 1102, responsive to transitioning from a low-power state to an active state during a voice call, obtain, at an audio processor, an audio input from an audio source, where the audio input corresponds to the voice call.For example, the audio processor 106 transitions from the low-power state to the active state upon entering the awake period 162 A of the first cycle 158A associated with a voice call with the device 154. The audio processor 106, upon entering the awake period 162 A, obtains the audio input 224 from an audio source. In a particular aspect, the audio input 224 includes the audio input 124 of FIG. 1. The audio input 124 is based on audio data received from the device 154 during the voice call. For example, the audio input 124 represents audio captured by a microphone coupled to the device 154. In another aspect, the audio input 224 corresponds to microphone output received from a microphone during the voice call with the device 154. For example, the microphone is coupled to the device 102 and represents speech of a user of the device 102.
[0096] The method 1100 includes, at block 1104, performing, at the audio processor, audio context detection (ACD) on the audio input to obtain an ACD indicator. For example, the audio context detector 132 performs ACD on the audio input 224 to obtain the ACD indicator 210, as described with reference to FIG. 2.
[0097] The method 1100 also includes, at block 1106, configuring, at the audio processor, a noise reducer based on a value of the ACD indicator. For example, the configurer 220 sends, based on the value of the ACD indicator 210, a command 212 to set the configuration parameters 214 to configure the noise reducer 134, as described with reference to FIG. 2.
[0098] The method 1100 further includes, at block 1108, processing, at the audio processor, the audio input using the configured noise reducer to generate output audio. For example, the audio enhancer 130 uses the configured noise reducer 134 to process the audio input 124 to generate the output audio 142, as described with reference to FIG. 1.
[0099] The method 1100 also includes, at block 1110, after generating the output audio, transitioning from the active state to the low-power state. For example, the audio enhancer 130 generates the output audio 142 during the awake period 162 A of the first cycle 158A associated with the voice call, and transitions from the active state to the low-power state upon exiting the awake period 162 A and entering the low-power period164 A of the first cycle 158 A associated with the voice call. In some aspects, the audio processing period 172 A ends after the awake period 162 A and the audio enhancer 130 transitions from the active state to the low-power state after exiting the audio processing period 172 A during the low-power period 164 A associated with the sleep / wake cycle of the modem 150.
[0100] In some embodiments, the method 1100 also includes transmitting, at a modem, an output signal based on the output audio and corresponding to a voice call. For example, the modem 150 generates the output signal 152 based on the output audio 142 for transmission to the device 154. According to an aspect, duration of audio processing periods 172 associated with the audio processing operations 170 is dynamically adjusted based on a detected audio context so that audio processing is reduced (e.g., not performed) during the low-power periods 164 for less noisy audio, enabling the central sleep manager to initiate transition to the low power island state earlier during the low- power periods 164.
[0101] By configuring the noise reducer 134 based on the detected audio context, the method 1100 enables the audio processor to enter the low-power state earlier for less noisy audio during low-power periods associated with the sleep / wake timing criteria associated with the voice call. As a result, power consumption of the audio processor is reduced for less noisy audio without reducing the quality of noise reduction of more noisy audio.
[0102] The method 1100 of FIG. 11 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a neural processing unit (NPU), a controller, another hardware device, firmware device, or any combination thereof. As an example, the method 1100 of FIG. 11 may be performed by a processor that executes instructions, such as described with reference to FIG. 12.
[0103] Referring to FIG. 12, a block diagram of a particular illustrative embodiment of a device is depicted and generally designated 1200. In various embodiments, the device 1200 may have more or fewer components than illustrated in FIG. 12. In an illustrative embodiment, the device 1200 may correspond to the device 102. In an illustrativeembodiment, the device 1200 may perform one or more operations described with reference to FIGS . 1-11.
[0104] In a particular embodiment, the device 1200 includes a processor 1206 (e.g., a CPU). The device 1200 may include one or more additional processors 1210 (e.g., one or more DSPs, one or more NPUs, or a combination thereof). In a particular aspect, the audio processor 106 of FIG. 1 is included in or corresponds to the processors 1210, the processor 1206, or a combination thereof. The processors 1210 may include a speech and music coder-decoder (CODEC) 1208 that includes a voice coder (“vocoder”) encoder 1236, a vocoder decoder 1238, or a combination thereof. In some embodiments, the speech and music codec 1208 corresponds to, or is included in, the codec 140 of FIG. 1.
[0105] The device 1200 may include a memory 1286 and a CODEC 1234. The memory 1286 may include instructions 1256 that are executable by the one or more additional processors 1210 (or the processor 1206) to implement the functionality described with reference to the audio processor 106, a central sleep manager, a voice timer, or any combination thereof. The device 1200 may include the modem 150 coupled, via a transceiver 1250, to an antenna 1252.
[0106] In a particular aspect, the memory 1286 is configured to store data used or generated by one or more components of the audio processor 106. For example, the memory 1286 is configured to buffer the audio input 224 from an audio source. In a particular aspect, the memory 1286 is configured to buffer the audio input 124 received via the modem 150, the transceiver 1250, and the antenna 1252, from the device 154 of FIG. 1. In a particular aspect, the memory 1286 is configured to buffer microphone output received from the microphone 1220.
[0107] The device 1200 may include a display 1228 coupled to a display controller 1226. One or more speakers 1224 and one or more microphones 1220 may be coupled to the CODEC 1234. In a particular aspect, the audio input 124, the audio input 224, or both, are based on microphone output of the one or more microphones 1220. In a particular aspect, the one or more microphones 1220 include the microphone 630, the microphone 710, the microphone 810, the microphone 910, the microphone 1010, or acombination thereof. The CODEC 1234 may include a digital-to-analog converter (DAC) 1202, an analog-to-digital converter (ADC) 1204, or both. In a particular embodiment, the CODEC 1234 may receive analog signals from the microphone 1220, convert the analog signals to digital signals using the analog-to-digital converter 1204, and provide the digital signals to the speech and music codec 1208. According to an aspect, the digital signals corresponding to the microphone input may be processed by the audio processor 106. In a particular embodiment, the speech and music codec 1208 (e.g., the audio processor 106) may provide digital signals to the CODEC 1234. The CODEC 1234 may convert the digital signals to analog signals using the digital-to- analog converter 1202 and may provide the analog signals to the speaker 1224.
[0108] In a particular embodiment, the device 1200 may be included in a system -in- package or system-on-chip device 1222. In a particular embodiment, the memory 1286, the processor 1206, the processors 1210, the display controller 1226, the CODEC 1234, the transceiver 1250, and the modem 150 are included in the system-in-package or system-on-chip device 1222. In a particular embodiment, an input device 1230 and a power supply 1244 are coupled to the system-in-package or the system-on-chip device 1222. Moreover, in a particular embodiment, as illustrated in FIG. 12, the display 1228, the input device 1230, the speaker 1224, the microphone 1220, the antenna 1252, and the power supply 1244 are external to the system-in-package or the system-on-chip device 1222. In a particular embodiment, each of the display 1228, the input device 1230, the speaker 1224, the microphone 1220, the antenna 1252, and the power supply 1244 may be coupled to a component of the system-in-package or the system-on-chip device 1222, such as an interface or a controller.
[0109] The device 1200 may include a smart speaker, a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a car, a computing device, a communication device, an internet-of-things (loT) device, an extended reality (XR) device, a base station, a mobile device, or any combination thereof.
[0110] In conjunction with the described embodiments, an apparatus includes means for obtaining, from an audio source, an audio input corresponding to a voice call, where the audio input is obtained responsive to transitioning from a low-power state to an active state during the voice call. For example, the means for obtaining the audio input can correspond to the audio context detector 132, the noise reducer 134, the audio enhancer 130, the audio processor 106, the device 102, the processor 1206, the one or more processors 1210, one or more other circuits or components configured to obtain the audio input responsive to transitioning from a low-power state to an active state during a voice call, or any combination thereof.[OHl] The apparatus also include means for performing audio context detection (ACD) on the audio input to obtain an ACD indicator. For example, the means for performing ACD can correspond to the audio context detector 132, the audio enhancer 130, the audio processor 106, the device 102, the processor 1206, the one or more processors 1210, one or more other circuits or components configured to perform ACD, or any combination thereof.
[0112] The apparatus further includes means for configuring a noise reducer based on a value of the ACD indicator. For example, the means for configuring a noise reducer can correspond to the audio context detector 132, the configurer 220, the audio enhancer 130, the audio processor 106, the device 102, the processor 1206, the one or more processors 1210, one or more other circuits or components configured to configure a noise reducer based on a value of an ACD indicator, or any combination thereof.
[0113] The apparatus also includes means for processing the audio input using the configured noise reducer to generate output audio. For example, the means for processing the audio input using the configured noise reducer can correspond to the noise reducer 134, the audio enhancer 130, the audio processor 106, the device 102, the processor 1206, the one or more processors 1210, one or more other circuits or components configured to process audio input using a configured noise reducer to generate output audio, or any combination thereof.
[0114] The apparatus further includes means for transitioning from the active state to the low-power state after generating the output audio. For example, the means for transitioning from the active state to the low-power state after generating the output audio can correspond to the audio enhancer 130, the audio processor 106, the device 102, a central sleep manager, a voice timer, the processor 1206, the one or more processors 1210, one or more other circuits or components configured to transition from the active state to the low-power state after generating the output audio, or any combination thereof.
[0115] In some embodiments, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as the memory 1286) includes instructions (e.g., the instructions 1256) that, when executed by one or more processors (e.g., the audio processor 106, the one or more processors 1210, or the processor 1206), cause the one or more processors to, responsive to transitioning from a low-power state to an active state during a voice call, obtain an audio input (e.g., the audio input 124, the audio input 224, or both) from an audio source (e.g., the device 154 or the microphone 1220), where the audio input corresponds to the voice call. The instructions, when executed by the one or more processors, also cause the one or more processors to perform audio context detection (ACD) on the audio input to obtain an ACD indicator (e.g., the ACD indicator 210). The instructions, when executed by the one or more processors, further cause the one or more processors to configure a noise reducer (e.g., the noise reducer 134) based on a value of the ACD indicator. The instructions, when executed by the one or more processors, also cause the one or more processors to process the audio input using the configured noise reducer to generate output audio (e.g., the output audio 226). The instructions, when executed by the one or more processors, further cause the one or more processors to, after generating the output audio, transition from the active state to the low-power state.
[0116] Particular aspects of the disclosure are described below in sets of interrelated Examples:
[0117] According to Example 1, a device includes: a memory configured to buffer an audio input from an audio source, wherein the audio input corresponds to a voice call;and an audio processor configured to: responsive to a transition from a low-power state to an active state during the voice call: perform audio context detection (ACD) on the audio input to obtain an ACD indicator; configure a noise reducer based on a value of the ACD indicator; and process the audio input using the configured noise reducer to generate output audio; and after generating the output audio, transition from the active state to the low-power state.
[0118] Example 2 includes the device of Example 1, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator to one or more thresholds, configure the noise reducer to use a neural network noise suppression engine.
[0119] Example 3 includes the device of Example 1 or 2, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator to one or more thresholds, configure the noise reducer to bypass a neural network noise suppression engine.
[0120] Example 4 includes the device of any of Examples 1 to 3, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator to one or more thresholds, configure the noise reducer to bypass a neural network noise suppression engine and to use an echo cancellation engine.
[0121] Example 5 includes the device of any of Examples 1 to 4, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator to one or more thresholds, configure the noise reducer to use a low latency noise suppression engine and to bypass a high performance noise suppression engine.
[0122] Example 6 includes the device of any of Examples 1 to 5, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to bypass a neural network noise suppression engine and to bypass an echo cancellation engine.
[0123] Example 7 includes the device of any of Examples 1 to 6, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator toone or more thresholds, configure the noise reducer to restrict noise reduction processing to using an echo cancellation engine.
[0124] Example 8 includes the device of any of Examples 1 to 7, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator to one or more thresholds, disable the noise reducer.
[0125] Example 9 includes the device of any of Examples 1 to 8, wherein the audio processor is configured to set, based on the value of the ACD indicator, configuration parameters to selectively enable one or more of multiple engines of the noise reducer.
[0126] Example 10 includes the device of Example 9, wherein the multiple engines include a low-latency noise suppression engine, a high performance noise suppression engine, and an echo cancellation engine.
[0127] Example 11 includes the device of Example 10, wherein the high performance noise suppression engine includes a neural network noise suppression engine.
[0128] Example 12 includes the device of Example 10 or 11, wherein at least a portion of audio processing of the high performance noise suppression engine is performed at an embedded neural processor unit (ENPU).
[0129] Example 13 includes the device of any of Examples 10 to 12, wherein the low- latency noise suppression engine includes a filter based noise suppression engine.
[0130] Example 14 includes the device of any of Examples 1 to 13, wherein the voice call is over a long-term evolution (LTE) network.
[0131] Example 15 includes the device of any of Examples 1 to 14, wherein the voice call is over at least one of a new radio (NR) network, a fifth generation (5G) wireless network, or a beyond 5G wireless network.
[0132] Example 16 includes the device of any of Examples 1 to 15, further comprising a modem configured to initiate transmission of an output signal based on the output audio.
[0133] Example 17 includes the device of Example 16, wherein transitions between a modem active state and a modem low-power state of the modem are aligned with transitions of the audio processor between the active state and the low-power state to enable synchronized processing using a low power island memory.
[0134] Example 18 includes the device of Example 16 or 17, wherein a modem low- power state of the modem and the low-power state of the audio processor enable transition to a power collapse mode.
[0135] Example 19 includes the device of any of Examples 1 to 18, wherein the voice call is a connected mode discontinuous reception (CDRx) call.
[0136] Example 20 includes the device of any of Examples 1 to 19, further comprising a microphone configured to generate the audio input.
[0137] Example 21 includes the device of any of Examples 1 to 19, further comprising a modem configured to receive the audio input from another device.
[0138] Example 22 includes the device of any of Examples 1 to 21, wherein an integrated circuit includes the audio processor.
[0139] Example 23 includes the device of any of Examples 1 to 22, wherein the audio processor is integrated into a vehicle.
[0140] Example 24 includes the device of any of Examples 1 to 23, wherein the audio processor is integrated into a portable electronic device.
[0141] According to Example 25, a method includes: responsive to transitioning from a low-power state to an active state during a voice call, obtaining, at an audio processor, an audio input from an audio source, wherein the audio input corresponds to the voice call; performing, at the audio processor, audio context detection (ACD) on the audio input to obtain an ACD indicator; configuring, at the audio processor, a noise reducer based on a value of the ACD indicator; processing, at the audio processor, the audio input using the configured noise reducer to generate output audio; and after generating the output audio, transitioning from the active state to the low-power state.
[0142] Example 26 includes the method of Example 25, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to use a neural network noise suppression engine.
[0143] Example 27 includes the method of Example 25 or 26, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to bypass a neural network noise suppression engine and to use an echo cancellation engine.
[0144] Example 28 includes the method of any of Examples 25 to 27, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to bypass a neural network noise suppression engine and to use an echo cancellation engine.
[0145] Example 29 includes the method of any of Examples 25 to 28, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to use a low latency noise suppression engine and to bypass a high performance noise suppression engine.
[0146] Example 30 includes the method of any of Examples 25 to 29, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to bypass a neural network noise suppression engine and to bypass an echo cancellation engine.
[0147] Example 31 includes the method of any of Examples 25 to 30, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to restrict noise reduction processing to using an echo cancellation engine.
[0148] Example 32 includes the method of any of Examples 25 to 31, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is disabled.
[0149] Example 33 includes the method of any of Examples 25 to 32, wherein the configuring the noise reducer includes setting, based on the value of the ACD indicator,configuration parameters to selectively enable one or more of multiple engines of the noise reducer.
[0150] Example 34 includes the method of Example 33, wherein the multiple engines include a low-latency noise suppression engine, a high performance noise suppression engine, and an echo cancellation engine.
[0151] Example 35 includes the method of Example 34, wherein the high performance noise suppression engine includes a neural network noise suppression engine.
[0152] Example 36 includes the method of Example 34 or 35, wherein at least a portion of audio processing of the high performance noise suppression engine is performed at an embedded neural processor unit (ENPU).
[0153] Example 37 includes the method of any of Examples 34 to 36, wherein the low- latency noise suppression engine includes a filter based noise suppression engine.
[0154] Example 38 includes the method of any of Examples 25 to 37, wherein the voice call is over a long-term evolution (LTE) network.
[0155] Example 39 includes the method of any of Examples 25 to 38, wherein the voice call is over at least one of a new radio (NR) network, a fifth generation (29G) wireless network, or a beyond 29G wireless network.
[0156] Example 40 includes the method of any of Examples 25 to 39, further comprising initiating transmission, via a modem, of an output signal based on the output audio.
[0157] Example 41 includes the method of Example 40, wherein transitions between a modem active state and a modem low-power state of the modem are aligned with transitions of the audio processor between the active state and the low-power state to enable synchronized processing using a low power island memory.
[0158] Example 42 includes the method of Example 40 or 41, wherein a modem low- power state of the modem and the low-power state of the audio processor enable transition to a power collapse mode.
[0159] Example 43 includes the method of any of Examples 25 to 42, wherein the voice call is a connected mode discontinuous reception (CDRx) call.
[0160] Example 44 includes the method of any of Examples 25 to 43, further comprising using a microphone to generate the audio input.
[0161] Example 45 includes the method of any of Examples 25 to 43, further comprising receiving the audio input via a modem from another device.
[0162] Example 46 includes the method of any of Examples 25 to 45, wherein an integrated circuit includes the audio processor.
[0163] Example 47 includes the method of any of Examples 25 to 46, wherein the audio processor is integrated into a vehicle.
[0164] Example 48 includes the method of any of Examples 25 to 47, wherein the audio processor is integrated into a portable electronic device.
[0165] According to Example 49, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to: responsive to transitioning from a low-power state to an active state during a voice call, obtain an audio input from an audio source, wherein the audio input corresponds to the voice call; perform audio context detection (ACD) on the audio input to obtain an ACD indicator; configure a noise reducer based on a value of the ACD indicator; process the audio input using the configured noise reducer to generate output audio; and after generating the output audio, transition from the active state to the low- power state.
[0166] Example 50 includes the non-transitory computer-readable medium of Example 49, wherein the instructions, when executed by the one or more processors, cause the one or more processors to set, based on the value of the ACD indicator, configuration parameters to selectively enable one or more of multiple engines of the noise reducer.
[0167] According to Example 51, an apparatus includes: means for obtaining, from an audio source, an audio input corresponding to a voice call, wherein the audio input isobtained responsive to transitioning from a low-power state to an active state during the voice call; means for performing audio context detection (ACD) on the audio input to obtain an ACD indicator; means for configuring a noise reducer based on a value of the ACD indicator; means for processing the audio input using the configured noise reducer to generate output audio; and means for transitioning from the active state to the low- power state after generating the output audio.
[0168] Those of skill would further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or combinations of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or processor executable instructions depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, such implementation decisions are not to be interpreted as causing a departure from the scope of the present disclosure.
[0169] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transient storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor may read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In the alternative, the processor and thestorage medium may reside as discrete components in a computing device or user terminal.
[0170] The previous description of the disclosed aspects is provided to enable a person skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope possible consistent with the principles and novel features as defined by the following claims.
Claims
WHAT IS CLAIMED IS:
1. A device comprising: a memory configured to buffer an audio input from an audio source, wherein the audio input corresponds to a voice call; and an audio processor configured to: responsive to a transition from a low-power state to an active state during the voice call: perform audio context detection (ACD) on the audio input to obtain an ACD indicator; configure a noise reducer based on a value of the ACD indicator; and process the audio input using the configured noise reducer to generate output audio; and after generating the output audio, transition from the active state to the low-power state.
2. The device of claim 1, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator to one or more thresholds, configure the noise reducer to use a neural network noise suppression engine.
3. The device of claim 1, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator to one or more thresholds, configure the noise reducer to bypass a neural network noise suppression engine.
4. The device of claim 1, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator to one or more thresholds, configure the noise reducer to use a low latency noise suppression engine and to bypass a high performance noise suppression engine.
5. The device of claim 1, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator to one or more thresholds, configure the noise reducer to restrict noise reduction processing to using an echo cancellation engine.
6. The device of claim 1, wherein the audio processor is configured to, based on a comparison of the value of the ACD indicator to one or more thresholds, disable the noise reducer.
7. The device of claim 1, wherein the audio processor is configured to set, based on the value of the ACD indicator, configuration parameters to selectively enable one or more of multiple engines of the noise reducer.
8. The device of claim 7, wherein the multiple engines include a low-latency noise suppression engine, a high performance noise suppression engine, and an echo cancellation engine.
9. The device of claim 8, wherein the high performance noise suppression engine includes a neural network noise suppression engine.
10. The device of claim 8, wherein at least a portion of audio processing of the high performance noise suppression engine is performed at an embedded neural processor unit (ENPU).
11. The device of claim 8, wherein the low-latency noise suppression engine includes a filter based noise suppression engine.
12. The device of claim 1, wherein the voice call is over a long-term evolution (LTE) network.
13. The device of claim 1, wherein the voice call is over at least one of a new radio (NR) network, a fifth generation (5G) wireless network, or a beyond 5G wireless network.
14. The device of claim 1, further comprising a modem configured to initiate transmission of an output signal based on the output audio.
15. The device of claim 14, wherein transitions between a modem active state and a modem low-power state of the modem are aligned with transitions of the audio processor between the active state and the low-power state to enable synchronized processing using a low power island memory.
16. The device of claim 14, wherein a modem low-power state of the modem and the low-power state of the audio processor enable transition to a power collapse mode.
17. The device of claim 1, wherein the voice call is a connected mode discontinuous reception (CDRx) call.
18. The device of claim 1, further comprising a microphone configured to generate the audio input.
19. The device of claim 1, further comprising a modem configured to receive the audio input from another device.
20. The device of claim 1, wherein an integrated circuit includes the audio processor.
21. The device of claim 1, wherein the audio processor is integrated into a vehicle.
22. The device of claim 1, wherein the audio processor is integrated into a portable electronic device.
23. A method comprising: responsive to transitioning from a low-power state to an active state during a voice call, obtaining, at an audio processor, an audio input from an audio source, wherein the audio input corresponds to the voice call; performing, at the audio processor, audio context detection (ACD) on the audio input to obtain an ACD indicator; configuring, at the audio processor, a noise reducer based on a value of the ACD indicator; processing, at the audio processor, the audio input using the configured noise reducer to generate output audio; and after generating the output audio, transitioning from the active state to the low- power state.
24. The method of claim 23, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to use a neural network noise suppression engine.
25. The method of claim 23, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to bypass a neural network noise suppression engine and to use an echo cancellation engine.
26. The method of claim 23, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to use a low latency noise suppression engine and to bypass a high performance noise suppression engine.
27. The method of claim 23, wherein, based on a comparison of the value of the ACD indicator to one or more thresholds, the noise reducer is configured to bypass a neural network noise suppression engine and to bypass an echo cancellation engine.
28. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: responsive to transitioning from a low-power state to an active state during a voice call, obtain an audio input from an audio source, wherein the audio input corresponds to the voice call; perform audio context detection (ACD) on the audio input to obtain an ACD indicator; configure a noise reducer based on a value of the ACD indicator; process the audio input using the configured noise reducer to generate output audio; and after generating the output audio, transition from the active state to the low- power state.
29. The non-transitory computer-readable medium of claim 28, wherein the instructions, when executed by the one or more processors, cause the one or moreprocessors to set, based on the value of the ACD indicator, configuration parameters to selectively enable one or more of multiple engines of the noise reducer.
30. An apparatus comprising: means for obtaining, from an audio source, an audio input corresponding to a voice call, wherein the audio input is obtained responsive to transitioning from a low-power state to an active state during the voice call; means for performing audio context detection (ACD) on the audio input to obtain an ACD indicator; means for configuring a noise reducer based on a value of the ACD indicator; means for processing the audio input using the configured noise reducer to generate output audio; and means for transitioning from the active state to the low-power state after generating the output audio.
Citation Information
Patent Citations
Noise suppression
EP1232496B1
Power efficient batch-frame audio decoding apparatus, system and method
US20090070119A1
Fast DRX for DL speech transmission in wireless networks
US20090201892A1
Suppressing noise in an audio signal
US20110081026A1
Real-time assessment of call quality
US20190355377A1