Machine learning-based feedback cancellation

Machine learning-based adaptive feedback cancellation with subband-specific models addresses howling issues in audio devices, enhancing audio quality and speech clarity by separating feedback components effectively.

JP2026517569APending Publication Date: 2026-06-02QUALCOMM INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
QUALCOMM INC
Filing Date
2024-03-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Conventional audio playback devices suffer from feedback artifacts such as howling, which degrade the signal-to-noise ratio and make speech unintelligible, especially in wearable devices like earbuds and headphones.

Method used

Implementing machine learning-based adaptive feedback cancellation using multiple independent models, each optimized for specific frequency subbands, to separate and reduce feedback components while preserving desirable sound components.

Benefits of technology

Enhances audio quality by effectively reducing feedback artifacts and improving speech clarity, with lower computational complexity and resource usage compared to single large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026517569000001_ABST
    Figure 2026517569000001_ABST
Patent Text Reader

Abstract

This disclosure provides systems, methods, and devices for audio signal processing that support feedback cancellation in personal audio amplification systems. In a first aspect, the signal processing method includes receiving an input audio signal, wherein the input audio signal contains desired audio components and feedback components, and reducing the feedback components by applying a machine learning model to the input audio signal to determine an output audio signal. Other aspects and features are also claimed and described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - reference to related applications) This application claims the benefit of U.S. Patent Application No. 18 / 611,494, filed on March 20, 2024, entitled "MACHINE LEARNING - BASED FEEDBACK CANCELLATION", and also claims the benefit of U.S. Provisional Patent Application No. 63 / 611,639, filed on December 18, 2023, entitled "MACHINE LEARNING - BASED FEEDBACK CANCELLATION", and further claims the benefit of U.S. Provisional Patent Application No. 63 / 493,158, filed on March 30, 2023, entitled "LOW - LATENCY NOISE SUPPRESSION", all of which are hereby expressly incorporated by reference in their entirety.

[0002] Aspects of the present disclosure generally relate to audio signal processing, and more specifically, to an amplification system with reduced artifacts. Some features can enable and provide improved audio signal processing, including improved audio quality by reducing the howling sound resulting from feedback when amplifying an audio signal.

Background Art

[0003] An audio playback device is a device that can play one or more audio signals, whether digital or analog. Audio playback can be incorporated into a wide variety of devices. By way of example, audio playback devices can include stand - alone audio devices, mobile phones, cellular or satellite phones, personal digital assistants (PDAs), panels or tablets, game devices, or computing devices.

[0004] One class of audio playback devices is wearable devices (e.g., earbuds, headphones, hearing aids, etc.) that can be used to improve hearing, situational awareness, and / or speech clarity. Generally, such devices apply a relatively simple noise suppression process to remove as much ambient noise as possible. Noise suppression operation may introduce artifacts into the reproduced sound, which reduce the signal-to-noise ratio (SNR) (of speech relative to ambient noise). This can be problematic because desired sounds, such as speech, may be obscured by the introduced artifacts. For some individuals, speech is only easily understandable when the signal-to-noise ratio (of speech relative to ambient noise) exceeds a certain level, and as a result, artifacts can render speech unrecognizable. [Overview of the Initiative]

[0005] The following summarizes several aspects of the disclosure in order to provide a basic understanding of the technology discussed. This summary is not intended to be a comprehensive overview of all conceivable features of the disclosure, nor to identify the main or significant elements of all aspects of the disclosure, nor to specify the scope of any or all aspects of the disclosure. Its sole purpose is to present, in summary form, some concepts of one or more aspects of the disclosure as an introduction to the more detailed explanations that will follow.

[0006] In some embodiments, audio amplification devices may use machine learning (ML)-based adaptive feedback cancellation to remove artifacts such as howling caused by feedback between a speaker and a microphone. In some embodiments, an ML model may be trained to retain speech components and other desirable sound components (such as ambient sounds) while removing undesirable components such as feedback or howling components. That is, an audio signal may contain multiple components from the environment, such as bird sounds, car noise, human speech, and feedback from a speaker. Some components may be desirable to hear (e.g., bird sounds and human speech). Other components may be desirable not to hear (e.g., car noise and feedback). An ML model may be trained to recognize sounds in an audio signal and remove undesirable components of the audio signal.

[0007] ML-based feedback cancellation offers improved performance compared to conventional adaptive filters in conventional feedback cancellation circuits. For example, ML-based feedback cancellation may be more effective in separating feedback components in an audio signal and / or retaining more desirable sonic components when removing feedback components. In some embodiments, ML-based feedback cancellation can compensate for linear and nonlinear components associated with feedback by modeling the nonlinearity of the amplifier circuit, thereby increasing its effectiveness. While ML-based models can be used in audio amplification systems in a similar manner to adaptive filters, ML-based models can solve other problems in feedback cancellation and provide additional functions and advantages in feedback cancellation, which will be described in more detail in the detailed descriptions of some embodiments that follow.

[0008] In one aspect of the present disclosure, a method for signal processing includes receiving an input audio signal, wherein the input audio signal includes desired audio components and feedback components, and reducing the feedback components by applying a machine learning model to the input audio signal to determine an output audio signal. The machine learning model may be trained to generate a feedback cancellation signal that reduces the magnitude or audibility of the feedback components when combined with the input audio signal. Such a machine learning model may be trained using training data including sample microphone signals recorded in the presence of feedback from a loudspeaker, where each set of training data includes a sample microphone signal, as well as a representation of the feedback components present in the sample microphone signal and / or a representation of the desired feedback cancellation signal for the sample microphone signal. The ML model may be trained by loading pre-calculated weights or other parameters of the ML model at startup of the ML model. Alternatively, the ML model may be trained by providing training data and constructing the ML model based on known feedback cancellation signals from the training data.

[0009] In additional aspects of the present disclosure, the apparatus includes at least one processor and memory coupled to the at least one processor. The at least one processor is configured to perform operations including receiving an input audio signal, wherein the input audio signal includes desired audio components and feedback components, and reducing the feedback components by applying a machine learning model to the input audio signal in order to determine an output audio signal.

[0010] In additional aspects of the present disclosure, the apparatus includes means for receiving an input audio signal, wherein the input audio signal includes desired audio components and feedback components; and means for reducing the feedback components by applying a machine learning model to the input audio signal in order to determine an output audio signal.

[0011] In additional aspects of the present disclosure, a non-temporary computer-readable medium stores instructions, and when the instructions are executed by at least one processor, the instructions cause the processor to perform an operation. The operation includes receiving an input audio signal, the input audio signal including desired audio components and feedback components, and reducing the feedback components by applying a machine learning model to the input audio signal to determine an output audio signal.

[0012] The audio signal processing methods described herein may be performed by a signal processing device. Audio signal processing may be applied to audio data captured by one or more microphones of the signal processing device. Audio signal processing devices, devices capable of playing back, recording, and / or processing one or more audio recordings, may be incorporated into a wide variety of devices. For example, audio signal processing devices may include standalone audio devices such as entertainment devices and personal media players, wireless communication devices such as mobile phones, cellular phones, or satellite radio phones, computing devices such as personal digital assistants (PDAs), tablets, gaming devices, webcams, video surveillance cameras, or other devices with audio recording or audio capabilities.

[0013] The audio signal processing techniques described herein may involve devices having a microphone and processing circuitry (e.g., application-specific integrated circuits (ASICs), digital signal processors (DSPs), graphics processing units (GPUs), or central processing units (CPUs)).

[0014] In some embodiments, the device may include a digital signal processor, a processor with specific functions for audio processing (e.g., an application processor). The methods and techniques described herein may be performed entirely by the digital signal processor or processor, or various operations may be divided between the digital signal processor and the processor, and in some embodiments, divided across additional processors. In some embodiments, the methods and techniques disclosed herein may be adapted using input from a neural signal processor (NSP), where one or more parameters of the signal processing are controlled based on the output from a machine learning (ML) model performed by the NSP.

[0015] In additional aspects of this disclosure, devices configured for audio signal processing and / or audio capture are disclosed. The devices include means for recording audio. Exemplary means may include dynamic microphones, condenser microphones, ribbon microphones, carbon microphones, or crystal microphones. Microphones may be interpreted as micro-electromechanical systems (MEMS). These components may be controlled to capture first and / or second recordings, which may correspond to the left and right channels of a recording.

[0016] For any of these types of microphones, the microphone may include analog and / or digital microphones. Analog microphones provide a sensor signal that is tuned or filtered in some embodiments. Analog microphones in digital systems include an external analog-to-digital converter (ADC) for interfacing with digital circuitry. Digital microphones include an ADC and other digital elements to convert the sensor signal into a digital data stream, such as a pulse density modulation (PDM) stream or a pulse code modulation (PCM) stream.

[0017] The embodiments disclosed herein describe the use of machine learning models as a solution to problems involving amplification artifacts such as feedback. While the embodiments of the disclosure described herein illustrate the use of a single machine learning model, embodiments may employ multiple machine learning models to perform the operations described herein. For example, independent machine learning models may be configured to analyze and process each frequency subband of audio data, enabling appropriate processing of the subbands. Each independent machine learning model is trained and optimized to process its respective subband. As an example, the low-frequency subband of an audio segment may correspond to speech, while the high-frequency subband of the same audio segment corresponds to noise. A first machine learning model processes the low-frequency subband audio data to generate a first enhanced subband audio data in which speech is preserved or amplified. A second machine learning model processes the high-frequency subband audio data to generate a second enhanced subband audio data in which noise is reduced. A coupler is used to combine the first enhanced subband audio data and the second enhanced subband audio data to generate enhanced audio data. As another example, a first machine learning model could be trained to preserve speech in low-frequency subband audio, and a second machine learning model could be trained to reduce noise in high-frequency subband audio. This would benefit from lower complexity and higher efficiency than a single machine learning model trained to process a larger frequency band to reduce speech in low-frequency subband audio and noise in high-frequency subband audio.

[0018] A problem with machine learning models for handling larger frequency bands is that some models are better suited to handling certain subbands than others. For example, a long short-term memory (LSTM) based masking network may be better suited to handling low-frequency subband audio, and a convolutional neural network may be better suited to handling high-frequency subband audio. Independent machine learning models can solve this problem by having different model architectures that are better suited to handling each subband. For example, a first machine learning model trained to handle low-frequency subband audio may include an LSTM-based masking network, and a second machine learning model trained to handle high-frequency subband audio may include a convolutional neural network (e.g., U-Net) architecture. In some examples, procedural signal processing may be performed for audio enhancement of a particular subband, or audio enhancement may be bypassed for another subband, or both.

[0019] A single, large, and complex machine learning model may have high resource usage (e.g., compute cycles, memory, etc.) that can limit the types of devices that can support it. This problem can be solved by using independent machine learning models that do not need to be located in the same place. For example, processing of low-frequency subband audio may be performed on a first device containing a first machine learning model, and processing of high-frequency subband audio may be performed on a second device containing a second machine learning model. Thus, at least part of the subband audio processing can be offloaded to another device.

[0020] Reconfiguring a large-scale machine learning model can change how it processes an entire frequency band. Independent machine learning models can solve this problem by being independently configurable. For example, an updated configuration better suited to low-frequency subband audio can be used for the first machine learning model without modifying the second machine learning model. In another example, the second machine learning model can be updated to have a second configuration better suited to high-frequency subband audio. The configurations of independent machine learning models can be obtained from one or more sources, such as other devices, based on the audio context.

[0021] In some cases, different microphones may capture audio with better audio quality in different subbands. For example, a first microphone might be closer to a first sound source (e.g., a speech source), and a second microphone might be closer to a second sound source (e.g., a music source). In this example, first low-frequency subband audio data from the first microphone is selected to be processed using a first machine learning model to generate first enhanced subband audio data in which speech is preserved or emphasized, and second high-frequency subband audio data from the second microphone is selected to be processed using a second machine learning model to generate second enhanced subband audio data in which music is preserved or emphasized.

[0022] In some examples, machine learning models are used to process subband audio data from multiple microphones to generate enhanced subband audio data. A first machine learning model is used to process first low-frequency subband audio data from a first microphone and second low-frequency subband audio data from a second microphone to generate enhanced low-frequency subband audio data. A second machine learning model is used to process first high-frequency subband audio data from a first microphone and second high-frequency subband audio data from a second microphone to generate enhanced high-frequency subband audio data.

[0023] The performance of a machine learning model can be improved by processing audio data from multiple microphones. For example, enhanced low-frequency subband audio data generated by a first machine learning model that processes low-frequency subband audio data from multiple microphones may have enhanced speech and reduced noise compared to enhanced low-frequency subband audio data based on low-frequency subband audio data from a single microphone. As another example, enhanced high-frequency subband audio data generated by a second machine learning model that processes high-frequency subband audio data from multiple microphones may have enhanced speech and reduced noise compared to enhanced high-frequency subband audio data based on high-frequency subband audio data from a single microphone.

[0024] Separate machine learning models trained to handle different subbands may have lower complexity (e.g., fewer network nodes, network layers, etc.) and higher efficiency (e.g., faster processing time, fewer computation cycles, etc.) compared to a single machine learning model trained to handle a larger frequency band that includes the subbands.

[0025] By considering the following descriptions of specific exemplary embodiments in conjunction with the attached figures, other embodiments, features, and implementations will become apparent to those skilled in the art. Features may be discussed in relation to some of the embodiments and figures below, but various embodiments may include one or more of the advantageous features discussed herein. In other words, one or more embodiments may be discussed as having several advantageous features, but one or more of such features may also be used according to various embodiments. Similarly, exemplary embodiments may be discussed below as embodiments of devices, systems, or methods, but exemplary embodiments may be implemented in various devices, systems, and methods.

[0026] This method may be embedded in a computer-readable medium as computer program code containing instructions that cause a processor to perform the steps of this method. In some embodiments, the processor may be part of a mobile device comprising a first network adapter configured to transmit data such as images or video (along with associated or embedded sounds) in recorded data or as streaming data over a first network connection of a plurality of network connections, and a processor coupled to the first network adapter and memory. The processor may trigger the transmission of output image frames described herein over a wireless communication network such as a 5G NR communication network.

[0027] The foregoing has outlined rather broadly the features and technical advantages of examples according to the present disclosure so as to enable a better understanding of the following "Best Mode for Carrying Out the Invention." Additional features and advantages will be described hereinafter. The disclosed concepts and specific examples may be readily utilized as a basis for modifying or designing other structures to accomplish the same purposes of the present disclosure. Such equivalent structures do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein will be better understood from the following description when considered in connection with the accompanying drawings in which each of the figures is provided for purposes of illustration and description only and not as a definition of the limits of the claims.

[0028] While this application describes embodiments and implementations by example to several embodiments, those skilled in the art will understand that additional implementations and use cases may arise in many different configurations and scenarios. The innovations described herein can be realized across many different platform types, devices, systems, forms, sizes, and packaging configurations. For example, embodiments and / or applications may be implemented in integrated chip implementations and other non-modular component-based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, artificial intelligence (AI)-enabled devices, etc.). Some examples may or may not specifically target use cases or applications, but a wide range of combinations of the innovations described may be applicable. Implementations may range from chip-level or modular components to non-modular, non-chip-level implementations, and further to aggregated, distributed, or original equipment manufacturer (OEM) devices or systems incorporating one or more embodiments of the innovations described. In some practical settings, devices incorporating the embodiments and features described may also necessarily include additional components and features for the implementation and practice of the claims and embodiments described. The innovations described herein are intended to be practiced in a wide variety of devices, chip-level components, systems, distributed configurations, end-user devices, and the like, of various sizes, shapes, and structures.

[0029] A further understanding of the nature and advantages of the present disclosure can be realized by referring to the following drawings. In the accompanying drawings, similar components or features may have the same reference labels. Further, various components of the same type can be distinguished by attaching a dash and a second label for distinguishing similar components after the reference label. When only the first reference label is used in this specification, the description is applicable to any one of the similar components having the same first reference label regardless of the second reference label.

Brief Description of the Drawings

[0030] [Figure 1] It is a block diagram of a system-on-chip (SoC) configured to perform signal processing according to one or more aspects of the present disclosure. [Figure 2] It is a block diagram showing an exemplary data flow path for audio signal processing in a multimedia device according to one or more aspects of the present disclosure. [Figure 3] It is a block diagram showing an audio system with machine learning-based feedback cancellation according to some embodiments of the present disclosure. [Figure 4] It is a block diagram showing an audio system with a machine learning-based feedback cancellation application before forward path processing according to some embodiments of the present disclosure. [Figure 5] It is a block diagram showing an audio system with machine learning-based feedback cancellation and other feedback cancellations according to some embodiments of the present disclosure. [Figure 6] It is a block diagram showing an audio system with machine learning-based feedback cancellation based on additional sensors according to some embodiments of the present disclosure. [Figure 7] It is a block diagram showing an audio system with machine learning-based feedback cancellation applied through a time domain filter according to some embodiments of the present disclosure. [Figure 8]This is a flowchart illustrating exemplary methods for processing audio data to perform acoustic source separation, according to some embodiments of the present disclosure. [Figure 9] This is an exemplary block diagram of an exemplary machine learning (ML) model represented by an artificial neural network (ANN) according to some embodiments of the present disclosure. [Figure 10] This diagram shows a headset capable of performing feedback cancellation, as illustrated by some examples of the present disclosure. [Figure 11] Figures of headsets, such as virtual reality, mixed reality, or augmented reality headsets, that are capable of performing feedback cancellation, as illustrated in some examples of the present disclosure. [Figure 12] Figures of augmented reality glasses capable of performing feedback cancellation, as illustrated by some examples of the present disclosure. [Figure 13] Figures of wearable devices capable of performing feedback cancellation, as illustrated by some examples of the present disclosure. [Figure 14] The following are diagrams of earbuds capable of performing feedback cancellation, as illustrated by some examples of the disclosure. [Figure 15] The following are some examples of exemplary embodiments of systems capable of performing machine learning-based audio subband processing, as illustrated by some of the examples in this disclosure. [Figure 16] These are diagrams illustrating exemplary embodiments of system components, some examples of the present disclosure. [Figure 17] These are diagrams illustrating exemplary embodiments of system components, some examples of the present disclosure.

[0031] Similar reference numbers and names in various drawings refer to the same elements. [Modes for carrying out the invention]

[0032] This disclosure provides systems, apparatus, methods, and computer-readable media that support signal processing, including techniques for machine learning (ML)-based adaptive feedback cancellation to remove artifacts such as howling caused by feedback between a speaker and a microphone. In some embodiments, the ML model may be trained to retain speech components and other desirable sound components (such as ambient sounds) while removing other undesirable components such as feedback components. Several embodiments providing different techniques for applying the ML model to an audio signal, which may be used in an amplification system, are disclosed below. Each of the ML models may be trained to acquire processing appropriate to the configuration of its embodiment in order to provide feedback cancellation.

[0033] Certain implementations of the subject matter described herein may be implemented to realize one or more of the following potential advantages or benefits. In some embodiments, the disclosure provides techniques for improving the sound quality of output audio by reducing feedback and resulting artifacts when amplifying an audio signal from a microphone in the sound field of a speaker outputting an amplified audio signal. ML-based feedback cancellation offers improved performance by being more effective in separating feedback components of an audio signal while preserving desirable sound components. In some embodiments, ML-based feedback cancellation can compensate for linear and nonlinear components associated with feedback, thereby increasing its effectiveness.

[0034] With respect to the attached figures referenced herein, the detailed descriptions below are intended to illustrate various embodiments and are not intended to limit the scope of this disclosure. Rather, the detailed descriptions include certain details intended to provide a complete understanding of the subject matter of this disclosure. Those skilled in the art will see that these specific details are not required in all cases, and that in some instances, well-known structures and components are shown in block diagrams for clarity in the description.

[0035] The descriptions of embodiments in this specification include numerous specific details, such as examples of specific components, circuits, and processes, to provide a complete understanding of the disclosure. The term “combined” as used herein means directly connected or connected via one or more intermediary components or circuits. Furthermore, certain technical terms are used in the following descriptions for illustrative purposes to provide a complete understanding of the disclosure. However, it will become clear to those skilled in the art that these specific details may not be necessary to implement the teachings disclosed herein. In other instances, well-known circuits and devices are shown in block diagrams to avoid obscuring the teachings of this disclosure.

[0036] Some parts of the following detailed description are presented in terms of procedures, logical blocks, processes, and other symbolic representations of operations on data bits in computer memory. In this disclosure, procedures, logical blocks, processes, etc., are considered to be self-consistent sequences of steps or instructions that lead to a desired result. These steps require the physical manipulation of physical quantities. Although not always necessary, these quantities usually take the form of electrical or magnetic signals that can be stored, transferred, combined, compared, and otherwise manipulated in a computer system.

[0037] An exemplary device for recording sound and / or processing sound signals using one or more microphones, such as micro-electromechanical system (MEMS) microphones, may include configurations of one, two, three, four, or more microphones at different locations on the device. The exemplary device may include one or more digital signal processors (DSPs), an AI engine, or other suitable circuitry for processing the signals captured by the microphones, such as performing specific noise cancellation operations. One or more digital signal processors (DSPs) may output signals representing sound via a bus for storage in memory, playback by an audio system, and / or further processing by other components (such as an application processor).

[0038] The processing circuit may perform further processing such as encoding, storing, transmitting, or other operations on the audio signal. In some embodiments, the exemplary device may include an audio circuit that includes an audio amplifier (e.g., a Class D amplifier) ​​for driving a converter to reproduce the sound represented by the audio signal. A speaker may be integrated into the device and coupled to the audio amplifier to be driven by the audio amplifier to reproduce sound. Connections may be provided by jacks or other connectors on the device to couple an external converter (e.g., an external speaker or headphones) to an audio amplifier that is driven by the audio circuit to reproduce sound. In some embodiments, the jack may instead be configured to couple to a digital device via a Universal Serial Bus (USB) Type-C (USB-C) connection and may output a digital signal for conversion and amplification by an external device, for example, when some or all of the audio circuit is bypassed.

[0039] Figure 1 shows a block diagram of a system-on-a-chip (SoC) configured to perform signal processing, according to one or more embodiments of the present disclosure. The SoC 100 may include several components coupled to one another via a bus 102, which may be a network-on-a-chip (NoC) or multiple NoCs interconnecting various components. For example, Figure 1 shows several components coupled to bus 102, but some components may be coupled to different buses, with additional buses connecting the different buses to provide a path for communication between the components.

[0040] One exemplary component within the SoC100 is a digital signal processor (DSP) 112 for signal processing. The DSP 112 can process audio signals received from microphones 130A, 130B, and 130C of a microphone array 130 (three are shown, but may include one or any number of microphones). The DSP 112 may include customized hardware to perform a limited set of operations for specific types of data. For example, the DSP may include transistors coupled together to perform operations for streaming data and may use a memory architecture and / or access techniques to fetch multiple data or instructions simultaneously. Such a configuration may enable the DSP 112 to operate for real-time data such as video data, audio data, or modem data in a power-efficient manner.

[0041] The SoC100 also includes a central processing unit (CPU) 104 and memory 106 (e.g., memory for storing processor-readable code, or non-temporary computer-readable medium for storing instructions) for storing instructions 108 that can be executed by the processor of the SoC100. The CPU 104 may be a single central processing unit (CPU) or a CPU cluster including two or more cores, such as core 104A. The CPU 104 may include hardware capable of performing general-purpose arithmetic on many kinds of data, such as hardware capable of executing instructions from advanced RISC machine (ARM®) instruction sets, such as ARMv8 and ARMv9. For example, the CPU 104 may include transistors coupled together to perform operations to support the execution of an operating system and user applications (e.g., camera applications, multimedia applications, game applications, productivity applications, messaging applications, video call applications, audio recording applications, video recording applications). The CPU 104 may execute instructions 108 retrieved from memory 106. In some embodiments, the CPU 104 running the operating system may coordinate the execution of instructions by various components within the SoC100. For example, the CPU 104 can retrieve instruction 108 from memory 106 and execute the instruction on the DSP 112.

[0042] The SoC100 may further include a neural signal processor (NSP) 124 for running machine learning (ML) models related to multimedia applications. The NSP124 may include hardware configured to perform and accelerate convolutional operations involved in the execution of machine learning algorithms. For example, the NSP124 may improve performance when running predictive models such as artificial neural networks (ANNs) (including multilayer feedforward neural networks (MLFFNNs), recurrent neural networks (RNNs), and / or radial basis functions (RBFs)). An ANN run by the NSP124 may have access to default training weights stored in memory 106 to perform operations on user data.

[0043] The SoC100 may be coupled to a display 114 for user interaction. The SoC100 may also include a graphics processing unit (GPU) 126 for rendering images onto the display 114. The images may be coordinated with sound processed and output by an audio processing circuit, for example, when the SoC100 is running a user application for multimedia playback or a game application. In some embodiments, the CPU 104 may perform rendering to the display 114 without the GPU 126. In some embodiments, the GPU 126 may be configured to execute instructions for performing operations unrelated to image rendering, such as processing large datasets in parallel.

[0044] The processing algorithms, techniques, and methods described herein may be executed by at least one processor of the SoC100, which may include execution by all steps on one of the processors (e.g., DSP112, CPU104, NSP124, GPU126) or by execution of steps across one or more combinations of the processors (e.g., DSP112, CPU104, NSP124, GPU126). In some embodiments, at least one of the processors executes instructions to perform various operations described herein, including machine learning-based feedback cancellation. For example, the execution of instructions by CPU104 as part of a multimedia application (e.g., a voice recorder, audio recording, or video recorder) may instruct DSP112 to start or stop capturing audio from one or more microphones 130A-C and digital audio signals processed in NSP124 to reduce the presence of feedback (e.g., howling) in the audio signals. The operation of CPU104 may be based on user input. For example, a voice recorder application running on processor 104 may receive a user command to start voice recording, at which point audio, including one or more channels, is captured and processed for playback and / or storage. Audio processing for determining an “output” or “correction” signal may be applied to one or more segments of audio in the recording sequence, such as the techniques described herein.

[0045] Input / output components may be coupled to the SoC 100 via an input / output (I / O) hub 116. An example of a hub 116 is an interconnection to a peripheral component interconnect express (PCIe) bus. Exemplary components coupled to the hub 116 may be components used to interact with the user, such as a touchscreen interface and / or physical buttons. Some components connected to the hub 116 may also include network interfaces for communicating with other devices, including a wide area network (WAN) adapter (e.g., WAN adapter 152), a local area network (LAN) adapter (e.g., LAN adapter 153), and / or a personal area network (PAN) adapter (e.g., PAN adapter 154). The WAN adapter 152 may be a 4G LTE or 5G NR wireless network adapter. The LAN adapter 153 may be an IEEE 802.11 WiFi wireless network adapter. The PAN adapter 154 may be a Bluetooth wireless network adapter. Each of the WAN adapter 152, LAN adapter 153, and / or PAN adapter 154 may be coupled to an antenna that can be shared by each of the adapters 152, 153, and 154, or to multiple antennas configured for primary and diversity reception and / or configured to receive specific frequency bands. In some embodiments, the WAN adapter 152, LAN adapter 153, and / or PAN adapter 154 may share circuitry such as a portion of a radio frequency front end (RFFE). In some embodiments, data transmitted via the I / O hub 116 may include audio signals processed according to aspects of this disclosure.For example, the processed audio signal can be output to a network media playback device via the WAN adapter 152 or LAN adapter 153, or to a personal audio device (e.g., the user's speakers, headset, or earbuds) via the PAN adapter 154.

[0046] The audio circuit 156, which may be a speaker (either inside or outside the device incorporating the SoC100) or a transducer such as headphones, may be integrated into the SoC100 as a dedicated circuit for coupling the SoC100 to an external speaker 120. The audio circuit 156 may include a coder / decoder (CODEC) function for processing digital audio signals. The audio circuit 156 may further include one or more amplifiers (e.g., Class D amplifiers) for driving transducers coupled to the SoC100 to output sound generated during the execution of applications by the SoC100. The audio signal-related functions described herein may be performed by a combination of the audio circuit 156 of the SoC100 and / or other processors (e.g., CPU104, DSP112, GPU126, NSP124).

[0047] The SoC100 may be coupled to an external device outside the SoC100 package. For example, the SoC100 may be coupled to a power supply 118, such as a battery or adapter, for coupling the SoC100 to an energy source. The signal processing described herein may be adapted to power efficiency to support the operation of the SoC100 from a power supply 118 with limited capacity, such as a battery, and power efficiency may be achieved. For example, the operation may be performed on a portion of the SoC100 configured to perform the operation with the lowest power consumption. As another example, the operation itself may be performed in a manner that reduces the number of computations required to perform the operation, and as a result, the algorithm may be optimized to extend the operating time of the device while powered by a power supply 118 with limited capacity. In some embodiments, the operations described herein may be configured based on the type of power supply 118 that provides energy to the SoC100. For example, a first set of operations may be performed to perform a function when the power supply 118 is a wall adapter. As another example, a second set of operations may be performed to perform a function when the power supply 118 is a battery.

[0048] The SoC100 may also include, or be combined with, additional features or components not shown in Figure 1. While the components are shown integrated as a single SoC100, which may include all components built on a single semiconductor die having a common semiconductor substrate, other configurations of exemplary blocks of different numbers of dies, substrates, and / or packages may be arranged to achieve the same functionality described herein.

[0049] Memory 106 may include a non-transient or non-transitory computer-readable medium that stores computer-executable instructions 108 for performing all or part of one or more operations described in this disclosure. Instructions 108 may include multimedia applications (or other suitable applications such as messaging applications) to be executed by the SoC 100 for recording, processing, or outputting audio signals. Instructions 108 may also include other applications or programs to be executed by the SoC 100, such as an operating system and applications other than those for multimedia processing.

[0050] In addition to instruction 108, memory 106 may also store audio data. The SoC 100 may be coupled to external memory and configured to access memory to write output audio files for later playback or long-term storage. For example, the SoC 100 may be coupled to a flash storage device including NAND memory to store video files containing audio tracks (e.g., MP4 container format files) and / or audio recordings (e.g., MPEG-1 Layer 3 files, also known as MP3 files). A portion of a video or audio file may be transferred to memory 106 for processing by the SoC 100, and the resulting signal after processing may be encoded as a video or audio file in memory 106 for transfer to long-term storage.

[0051] While SoC100 is mentioned in the examples herein for carrying out aspects of the disclosure, some device components may not be shown in Figure 1 to avoid obscuring aspects of the disclosure. In addition, other components, numerous components, or combinations of components may be included in a suitable device for carrying out aspects of the disclosure. Therefore, the disclosure is not limited to the configuration of a particular device or component.

[0052] The SoC in Figure 1 may operate to achieve improved audio recording and / or an improved user experience through higher quality audio playback by applying feedback cancellation to the audio signal based on a machine learning (ML) model. One exemplary method of performing multimedia operations is shown in Figure 2 and described below.

[0053] Figure 2 is a block diagram illustrating an exemplary data flow path for audio signal processing in a multimedia device according to one or more aspects of the present disclosure. The SoC 100 of the system 200 may perform multimedia controls 210, such as part of an operating system or driver, to control the capture of sound from a microphone or other audio source and / or the configuration of the audio processing circuit 156. The audio configuration applied by the multimedia controls 210 to either an output device (e.g., a speaker) or an input device (e.g., a microphone) may include parameters that specify, for example, bit depth, sampling rate, data rate, magnitude, or other parameters.

[0054] The multimedia control 210 may be managed by or provide services to the multimedia application 204. The multimedia application 204 may also run on the SoC 100, which includes one or more processors of the SoC 100. The multimedia application 204 provides the user with accessible settings so that the user can specify individual playback settings or select a profile having corresponding playback settings. The multimedia application 204 may be, for example, a video recording application, a screen sharing application, a virtual conferencing application, an audio playback application, a messaging application, a video communication application, or other applications that process audio data. The multimedia application 204 may include feedback cancellation 206 to improve the quality of audio presented to the user while the multimedia application 204 is running. The feedback cancellation 206 may perform one or more or a combination of the techniques described herein, and in some embodiments, the feedback cancellation 206 may be performed by multiple processing units within the SoC 100 (with a portion performed by the DSP 112 and a portion performed by the NSP 124, etc.).

[0055] One exemplary implementation for feedback cancellation 206 is shown in Figure 3. Figure 3 is a block diagram showing an audio system with machine learning-based feedback cancellation according to several embodiments of the present disclosure. A microphone 130 may receive an audio signal 370 from a source of a desired audio signal. The microphone 130 may also receive feedback 324 from the output audio signal 372 of a speaker 120 (or other transducer in the field of reception of the microphone 130). The microphone 130 generates an audio signal 370 that includes a feedback component corresponding to the feedback 324 and a desired audio component corresponding to the audio signal 304. The audio signal 304 may be processed in a forward path 322, such as by amplifying the audio signal 304. The output of the forward path 322 may be processed using one or more machine learning models 326 to determine an output audio signal 314. The output audio signal 314 may be amplified in an amplifier circuit 336 so that the speaker 120 can reproduce the output audio signal 314 as sound 372. The amplifier circuit 336 could be, for example, a Class D amplifier.

[0056] A machine learning model(s) 326 can be trained to retain a desired input, such as an audio signal 370, while removing artifacts arising from the feedback 324. An exemplary artifact is howling, which arises from amplification in the forward path 322, and the magnitude of the feedback 324 is increased until it is large enough to produce audible howling noise from the speaker 120. The feedback 324 may have linear and nonlinear components, and the machine learning model(s) 326 can be configured to cancel the linear and / or nonlinear components. The feedback 324 can be represented by a transfer function between the speaker 120 and the microphone 130. Conventional techniques for feedback cancellation have involved an adaptive filter configured to estimate this transfer function. However, the adaptive filter was limited to canceling the linear component of the transfer function of the feedback 324.

[0057] The machine learning model(s) 326 may cancel the nonlinear components of the transfer function, and in some embodiments, the linear components. The machine learning model(s) 326 may also be combined in some embodiments with other feedback cancellation processes to further improve the audio quality of the sound 372 output from the speaker 120. Exemplary models of the machine learning model(s) 326 include masking, composite masking, SGN, and its variations including exemplary embodiments of SGN as illustrated and described with reference to Figures 15, 16, and 17, as well as UNet and its variation ShrimpNet.

[0058] In some embodiments, the machine learning model(s)326 may switch between a transparency mode and a noise cancellation mode, where the output signal is generated by processing the audio signal with different coefficients. The operating mode (transparency, noise cancellation, or other) may be selected by the user or determined based on one or more criteria.

[0059] In some embodiments, feedback cancellation is performed by a machine learning model(s) on the output of the forward path 322, as shown in Figure 3. In other embodiments, feedback cancellation may be applied to the audio signal 304 before processing in the forward path 322. An exemplary audio system for pre-amplification feedback cancellation using a machine learning model is shown in Figure 4. Figure 4 is a block diagram showing an audio system with a machine learning-based feedback cancellation application before forward path processing, according to some embodiments of the present disclosure.

[0060] In Figure 4, a machine learning model (one or more) 426, which may be one of the machine learning models described in relation to Figure 3, can modify the audio signal 304 before processing in the forward path 322. The machine learning model 426 can determine a feedback cancellation signal to be combined with the audio signal 304 received from the microphone 130, and the combined signal is processed in the forward path 322. The machine learning model (one or more) 426 can output a feedback cancellation signal based on the audio signal 314 from the output of the forward path 322 being played back by the speaker 120.

[0061] In some embodiments, a machine learning model(s) 326 may be combined with one or more adaptive filters. The combination may be configured such that the adaptive filter, which may be performed by the DSP 112, removes the linear component of the feedback 324, and the machine learning model(s) 326, which may be performed by the NSP 124, removes the nonlinear component of the feedback 324. One exemplary combination of machine learning models with other feedback cancellation is shown in Figure 5. Figure 5 is a block diagram showing audio systems with machine learning-based feedback cancellation and other feedback cancellation according to some embodiments of the present disclosure.

[0062] In Figure 5, a machine learning model(s) 526, which may be configured similarly to the machine learning model(s) 326 in Figure 3, can modify the audio signal after processing in the forward path 322. A feedback canceller 528 may determine a feedback cancellation signal to combine with the audio signal 304 before processing in the forward path 322. The feedback canceller 528 may be an adaptive filter that reduces the linear component of feedback 324 before the forward path 322. The machine learning model(s) 526 may be configured to reduce the nonlinear component of feedback 324 before output to speaker 120. The machine learning model(s) 526 may improve upon the feedback cancellation performed by the audio system in Figure 5. While certain blocks are described as providing linear component cancellation and / or nonlinear component cancellation, each of blocks 526 and 528 may perform either linear component cancellation, nonlinear component cancellation, or a combination of linear and nonlinear component cancellation.

[0063] The machine learning model(s) 526 may be configured to receive feedback from the feedback canceller 528 (as input parameters related to the feedback cancellation signal). For example, the adaptive filter of the feedback canceller 528 may provide the machine learning model(s) 526 with information about the configuration of the adaptive filter that the machine learning model(s) 526 uses as input to estimate the nonlinear component in the output of the forward path 322, and that estimation is used to modify the output of the forward path 322 to reduce artifacts such as howling. In some embodiments, the feedback canceller 528 may provide the machine learning model(s) 526 with an indicator of whether howling has been detected in the audio signal 304, and the indicator may be used to activate the machine learning model(s) 526. In some embodiments, the feedback canceller 528 provides the machine learning model(s) 526 with one or more of the following: the F(q) FIR coefficient of the adaptive filter, the D(q) coefficient of the forward path 322, the open-loop transfer function 1+F(q)D(q), an estimate of the gain margin (e.g., the distance of the open-loop transfer function from the Nyquist point (-1,0)), and / or the proximity of the adaptive filter to an instability point.

[0064] In some embodiments, a machine learning model(s) 626, which may be configured similarly to the machine learning model(s) 526 in Figure 5, may receive information from one or more additional sensors(s) 626. Figure 6 is a block diagram illustrating an audio system with machine learning-based feedback cancellation based on additional sensors, according to some embodiments of the present disclosure. Data from the additional sensors(s) 626 may be input to the machine learning model(s) 626 to improve the feedback cancellation provided by the machine learning model(s) 626. The additional sensors(s) 626 may be located outside the path of feedback 324 or isolated from feedback 324 (e.g., a different sensor modality from the audio sensors), thereby improving the machine learning model(s) 626's ability to distinguish howling resulting from feedback 324 processed in the forward path 322 from howling in the audio signal 370 (e.g., music, bird sounds, sirens, etc.). An example of an additional sensor is a camera that receives optical signals that are not part of feedback 324.

[0065] Figure 7 is a block diagram showing an audio system with machine learning-based feedback cancellation applied via a time-domain filter, according to some embodiments of the present disclosure. This configuration for the machine learning model can reduce latency by performing analysis in the frequency domain in the ML model, with the ML model outputting coefficients of the time-domain filter. In some embodiments, the machine learning model(s) 726 is trained to produce output data that includes the target audio components of the frequency-domain audio data from the forward path 322 and omits or suppresses the non-target audio components of the frequency-domain audio data. Such a model may be called an “inline” model. In some embodiments, the machine learning model(s) 726 is trained to produce output data that includes the non-target audio components of the frequency-domain audio data from the forward path 322 and omits or suppresses the target audio components of the frequency-domain audio data. Such a model may be called a “masking” model. The output data from the masking model may be used directly as a noise suppression output and combined with the output of the forward path 322. The output data from the inline model may be further processed to produce a noise suppression output 128.

[0066] In some embodiments, output coefficients from the ML model may be mask coefficients. These mask coefficients may be input to a time-domain filter and applied to the forward path output. In Figure 7, a machine learning model (one or more) 726, which may be configured similarly to machine learning model 326, receives an audio signal from the forward path 322 to determine the coefficients of the time-domain filter 724. The time-domain filter 724 also receives an output audio signal from the forward path 322 to generate an output audio signal 314 for playback by speaker 120.

[0067] The system 200 in Figure 2 may be configured to perform operations described with reference to embodiments of Figures 3, 4, 5, 6, and / or 7 in order to determine an audio signal. Figure 8 shows a flowchart of an exemplary method for processing audio data to perform acoustic source separation, according to some embodiments of the present disclosure. The operation in Figure 3 may result in an audio signal with improved sonic representation, resulting in an improved user experience. Each of the operations described with reference to Figure 3 may be performed by one or a combination of processors of the SoC 100.

[0068] In block 802, an input audio signal is received. The input audio signal may be received, for example, from a microphone. Alternatively, audio data may be received from a wireless microphone, where the audio data is received via one or more of the WAN adapter 152, LAN adapter 153, and / or PAN adapter 154. Alternatively, audio data may be received from a memory location or network storage location, such as when the audio signal has been previously captured and is now being retrieved from memory 106 and / or from a remote location via one or more of the WAN adapter 152, LAN adapter 153, and / or PAN adapter 154. In some embodiments, the reception (e.g., capture or retrieval) of the audio signal may be initiated by a multimedia application 204 running on the SoC 100. The audio data, including the audio signal, is retrieved in block 802 and may be further processed by the SoC 100 according to the operations described in one or more of the following blocks.

[0069] In block 804, the input audio signal is processed to reduce feedback components by applying a machine learning model to the input audio signal received in block 802. Applying the machine learning model to the input audio signal can be carried out according to the embodiments shown in Figures 3, 4, 5, 6, and / or 7.

[0070] The operations described with reference to blocks 802 and 804 in Figure 8 may be performed on a digital signal processor (DSP), such as the DSP 112 of the SoC 100 shown in Figure 1. However, the operations may alternatively be performed by one or more of the processors in Figure 1, including one or more of the CPU 104, DSP 112, GPU 126, or NSP 124. For example, the CPU 104 may, as part of the operation of block 802, control the recording of input audio signals from the microphone array 130 to memory 106. The DSP 112 and / or NSP 124 may then perform the operations of block 804 on the audio signals stored in memory 106, and the output signals determined by the DSP 112 and / or NSP 124 may then be stored in memory 106, output to audio circuit 156 for playback by speaker 120, and / or transmitted to another device via one or more of the WAN 152, LAN 153, and / or PAN 154. In another example, the processor performing the operations of blocks 802 and / or 804 could be a dedicated logic circuit for performing several operations.

[0071] Certain embodiments and techniques described herein can be implemented, at least in part, using artificial intelligence (AI) programs, such as programs including machine learning (ML) models. For example, a machine learning model (one or more) 326 can implement an AI program as described with reference to Figure 9. An ML model can be trained offline to receive inputs including clear speech, unstable feedback (e.g., howling or other artifacts), and ambient sounds (e.g., sounds similar to howling). An ML model can be trained to output clear speech and ambient sounds without unstable feedback (e.g., feedback artifacts such as howling).

[0072] An exemplary ML model may define computational power for making decisions from input data, where decisions are based on patterns identified in the input data. Computational power may be defined in terms of weights and biases. Weights may represent the relationship between specific input data and specific decisions. Biases may represent a starting point for decisions. An exemplary ML model operating on input data may start with a decision defined by the biases and then modify that decision based on a combination of input data and weights. Decisions from an ML model may be one or more of decisions, predictions, inferences, or values. Decisions, predictions, or inferences may be represented as values ​​output from the ML model. In some embodiments of this disclosure, an ML model may be configured to provide computational power for feedback cancellation in audio signal processing. Such an ML model may be configured with weights and / or biases to perform feedback cancellation by recognizing desirable aspects of audio (e.g., speech and ambient sounds) and reducing or removing undesirable aspects of audio (e.g., feedback artifacts). Therefore, during device operation, the ML model can receive input data (e.g., microphone signals) and make decisions based on weights and / or biases (e.g., it can output an audio signal with reduced feedback artifacts). ML models that can be configured in this way according to embodiments of the present disclosure include supervised ML models and unsupervised ML models, as well as ML models for classification and / or regression.

[0073] The description herein illustrates, as several examples, how one or more tasks / problems in feedback cancellation in audio signal processing can benefit from the application of one or more ML models using ANNs. In some embodiments, other types (one or more) of ML models may be used instead of ANNs. Therefore, unless expressly stated otherwise, the subject matter concerning ML models is not necessarily intended to be limited to ANN solutions. Furthermore, it should be understood that terms such as “AI / ML model,” “ML model,” “trained ML model,” “ANN,” “model,” and “algorithm” are intended to be interchangeable unless specifically noted.

[0074] Figure 9 is an exemplary block diagram of an exemplary machine learning (ML) model represented by an artificial neural network (ANN) 900. The ANN 900 may receive input data 906 which may include one or more bits of data 902, preprocessed data (optional) output from a preprocessor 904, or some combination thereof. Here, data 902 may include, for example, training data, validation data, application-related data, etc., depending on the stage of deployment of the ANN 900. The preprocessor 904 may be included within the ANN 900 in some other implementations. The preprocessor 904 may process all or part of data 902, which may result in, for example, some of the data 902 being modified, replaced, deleted, etc. In some implementations, the preprocessor 904 may add additional data to data 902. In some implementations, the preprocessor 904 may be an ML model such as the ANN.

[0075] ANN 900 includes at least one first layer 908 of artificial neurons 910 to process input data 906 and provide the resulting first layer data to at least a portion of at least one second layer 914 via edge 912. The second layer 914 processes the data received via edge 912 and provides the output data of the second layer to at least a portion of at least one third layer 918 via edge 916. The third layer 918 processes the data received via edge 916 and provides the third layer output data to at least a portion of a final layer 922, which includes one or more neurons, via edge 920 to provide output data 924. All or part of the output data 924 may be further processed in several ways by an (optional) post-processor 926. Thus, in a particular example, ANN 900 may provide output data 928 based on the output data 924, post-processed data output from the post-processor 926, or several combinations thereof. The post-processor 926 may be contained within the ANN 900 in some other implementations. The post-processor 926 may process all or part of the output data 924, for example, so that the output data 928 differs from the output data 924 at least partially as a result of the data being modified, replaced, deleted, etc. In some implementations, the post-processor 926 may be configured to append additional data to the output data 924. In this example, the second layer 914 and the third layer 918 represent intermediate or hidden layers that may be arranged in a hierarchical structure or other similar structure. Although not explicitly shown, one or more further intermediate layers may exist between the second layer 914 and the third layer 918. In some implementations, the post-processor 926 may be an ML model such as an ANN.

[0076] The ANN 900 or other ML models, along with memory and applicable instructions within it, can be implemented in various types of processing circuits. For example, general-purpose hardware circuits such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs) may be employed to realize the model. One or more tensor processing units (TPUs), neural processing units (NPUs), or other dedicated processors, and / or field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., may also be employed, or alternatively. In some implementations, the ML model can be implemented by NPUs and / or TPUs integrated into a system-on-a-chip (SoC) along with other components such as one or more CPUs and / or GPUs. The SoC includes several components manufactured on a shared semiconductor substrate. The NPU and / or TPU may be controlled by one or more CPUs by constructing an ML model implemented by the NPU and / or TPU using weights and / or biases, providing specific training data to the ML model to construct it, and / or providing input data to the ML model to obtain decisions. One or more CPUs may also be configured to receive decisions and perform specific actions based on the decisions made by the ML model.

[0077] The signal processing modes described in Figures 3, 4, 5, 6, and / or 7 may be incorporated into exemplary devices such as the exemplary devices in Figures 10, 11, 12, 13, and 14. Examples of such wearable devices include, but are not limited to, a headset device as further described with reference to Figure 10, a virtual reality, mixed reality, or augmented reality headset as described with reference to Figure 11, augmented reality glasses as described with reference to Figure 12, a hearing aid device as described with reference to Figure 13, or earbuds as described with reference to Figure 14.

[0078] Figure 10 shows an implementation form 1000 of a headset device 1002, which may include an SoC 100 configured according to one of Figures 2, 3, 4, 5, 6, 7, and / or 8. The headset device 1002 includes one or more microphones 130 and one or more speakers 120. In the example shown in Figure 10, microphone 130A is primarily positioned to detect speech from the person wearing the headset device 1002, and microphone 130B is positioned to detect speech from another person or ambient sounds such as other sounds. Components of the SoC 100, including feedback cancellation 206, are integrated into the headset device 1002 and are shown using dashed lines to indicate components that are not generally visible to the user of the headset device 1002.

[0079] In a specific example of operation, the microphone 130B may detect sounds in the environment around the headset device 1002 and generate audio data representing those sounds. The audio data may be provided to the SoC 100, which may process the audio data in the time domain to generate a noise-suppressed output signal, and may process the audio data in the frequency domain to generate (or update) time-domain filter coefficients to be applied to the subsequently received audio data. In this example, the time-domain filter coefficients determined via frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to the audio data received in the time domain, little to no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the headset device 1002 can provide high-quality, low-latency noise suppression.

[0080] Figure 11 shows one implementation form 1100, which includes a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality headset device 1102. The headset device 1102 includes one or more microphones 130 and one or more speakers 120. Additionally, components of the SoC 100, including a feedback canceller 206, are integrated into the headset device 1102. In a particular example of operation, one or more microphones 130 may detect sounds in the environment around the headset device 1102 and generate audio data representing those sounds. The audio data may be provided to the SoC 100, which may process the audio data in the time domain to generate a noise-suppressed output signal, and may process the audio data in the frequency domain to generate (or update) time-domain filter coefficients to be applied to subsequently received audio data. In this example, the time-domain filter coefficients determined via frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to receive audio data in the time domain, using such time-domain filter coefficients to process the audio data results in little to no latency. Therefore, the headset device 1102 can provide high-quality, low-latency noise suppression.

[0081] Figure 12 shows one implementation form 1200 corresponding to augmented reality or mixed reality glasses 1202. The glasses 1202 include a holographic projection unit 1204 configured to project visual data onto the surface of a lens 1206 or to reflect visual data from the surface of the lens 1206 onto the wearer's retina. The glasses 1202 also include a microphone(s) 130, a speaker(s) 120, and an SoC 100.

[0082] In a specific example of operation, a microphone(s) 130 may detect sounds in the environment around the glasses 1202 and generate audio data representing those sounds. The audio data may be provided to a feedback canceller 206, which may process the audio data in the time domain to generate a noise-suppressed output signal, and may process the audio data in the frequency domain to generate (or update) time-domain filter coefficients to be applied to the audio data received thereafter. In this example, the time-domain filter coefficients determined via frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to receive the audio data in the time domain, little to no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the glasses 1202 may provide high-quality, low-latency noise suppression.

[0083] In some implementations, the holographic projection unit 1204 may display information related to sound detected by the microphone(s) 130. For example, the holographic projection unit 1204 may display a notification indicating that sound has been detected. In another example, the holographic projection unit 1204 may display a notification indicating a detected audio event. For example, the notification may be superimposed on the user's field of view at a specific location that coincides with the location of the sound source associated with the audio event.

[0084] Figure 13 is a diagram of a wearable device capable of performing low-latency noise suppression, according to some examples of the present disclosure. In the example shown in Figure 13, the wearable device is a hearing aid device 1302. The hearing aid device 1302 includes a microphone (one or more) 130, a speaker (one or more) 120, and a SoC 100 including a feedback canceller 206. In the example shown in Figure 13, the hearing aid device 1302 is shown to be configured to be worn behind the user's ear and to include a portion 1308 configured to extend over the ear and a portion 1306 that is worn in or near the user's ear canal. In other examples, the hearing aid device 1302 may have a different configuration or form factor. For example, the hearing aid device 1302 may be an in-ear device that does not include a portion 1304 configured to be worn behind the ear and a portion 1308 configured to extend over the ear.

[0085] In a specific example of the operation of the hearing aid device 1302, a microphone(s) 130 may detect sounds in the environment surrounding the hearing aid device 1302 and generate audio data representing those sounds. The audio data may be provided to an audio component 940, which may process the audio data in the time domain to generate a noise-suppressed output signal, and may process the audio data in the frequency domain to generate (or update) time-domain filter coefficients to be applied to the audio data subsequently received. In this example, the time-domain filter coefficients determined via frequency-domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to receive the audio data in the time domain, little to no latency is added by using such time-domain filter coefficients to process the audio data. Thus, the hearing aid device 1302 may provide high-quality low-latency noise suppression.

[0086] Figure 14 shows one implementation configuration 1400, which includes a portable electronic device corresponding to one or more earbuds 1404 (e.g., a first earbud 1406, a second earbud 1402, or both). Although earbud 1406 is described, it should be understood that this technology may be applicable to other in-ear or over-ear audio devices.

[0087] In the example shown in Figure 14, the first earbud 1402 includes a first microphone 1410A, such as a high-signal-to-noise microphone, positioned to capture the voice of the wearer of the first earbud 1402; one or more other microphones, indicated as microphone(s) 1412A, configured to detect ambient sounds and spatially distributed to support beamforming; an "inner" microphone 1414A positioned close to the wearer's ear canal (for example, to assist with active noise cancellation); and a self-speaking microphone 1416A, such as a bone conduction microphone, configured to convert sound vibrations of the wearer's ossicles or skull into audio signals.

[0088] The second earbud 1404 may be configured in substantially the same manner as the first earbud 1402. For example, the second earbud may include a microphone 1410B positioned to capture the voice of the wearer of the second earbud 1404, one or more other microphones 1412B configured to detect ambient sounds and spatially distributed to support beamforming, an "inner" microphone 1414B, and a self-speaking microphone 1416B.

[0089] In some implementations, the earbuds 1402, 1404 are configured to automatically switch between various operating modes, such as a pass-through mode in which ambient sound is processed by the audio component 940 for output via speaker(s) 120, and a playback mode in which non-ambient sound (e.g., streaming audio corresponding to telephone conversations, media playback, video games, etc.) is played through speaker(s) 120. In other implementations, the earbuds 1402, 1404 may support fewer modes, or one or more other modes in place of or in addition to the described modes. In an exemplary example of operation in pass-through mode, one or more of the microphone(s) 130 (e.g., microphone(s) 1412A, 1412B) may detect sounds in the environment around the earbuds 1402, 1404 and generate audio data representing those sounds.

[0090] Audio data may be provided to a feedback canceller 206, which may process the audio data in the time domain using a machine learning model to generate a noise-suppressed output signal, and may process the audio data in the frequency domain to generate (or update) time-domain filter coefficients to be applied to the subsequently received audio data. In this example, the time-domain filter coefficients determined via frequency domain processing provide high-quality noise suppression, and since the time-domain filter coefficients are applied to the received audio data in the time domain, little to no latency is added by using such time-domain filter coefficients to process the audio data. Thus, earbuds 1402, 1404 can provide high-quality, low-latency noise suppression.

[0091] Referring to Figure 15, a diagram of an exemplary embodiment of a system 660 capable of performing machine learning-based audio subband processing is shown. In a particular embodiment, the systems in Figures 1 to 14 may include one or more components of system 660. The audio frequency divider 142 generates subband audio data for the audio data. For example, the audio frequency divider 142 processes audio data 617B to generate subband audio data 618AA associated with a first frequency subband and subband audio data 618AB associated with a second frequency subband. Subband audio data 618AA represents the first frequency subband of the audio captured by the microphone. Subband audio data 618AB represents the second frequency subband of the audio captured by the microphone.

[0092] Furthermore, the audio frequency divider 142 processes the audio data 627B to generate subband audio data 628AA associated with a first frequency subband and subband audio data 628AB associated with a second frequency subband. Subband audio data 628AA represents the first frequency subband of the audio captured by one of the microphones. Subband audio data 628AB represents the second frequency subband of the audio captured by one of the microphones.

[0093] One or more audio subband enhancers 144 process a set of subband audio data associated with a corresponding frequency subband to generate enhanced audio data for that frequency subband. For example, audio subband enhancer 144A processes subband audio data for a first frequency subband to generate enhanced audio data 135A. Exemplarily, audio subband enhancer 144A processes subband audio data 618AA and subband audio data 628AA to generate enhanced audio data 135A. As another example, audio subband enhancer 144B processes subband audio data 618AB and subband audio data 628AB to generate enhanced audio data 135B.

[0094] Referring to Figure 16, a diagram 700 of exemplary embodiments of the system components is shown. In certain embodiments, the diagram 700 includes an example of an exemplary optional embodiment of the highlighted subband audio generator 140 and an example of an exemplary optional embodiment of the coupler 148.

[0095] The audio subband enhancer 144 includes multiple machine learning models (e.g., LSTMs) associated with each subband. For example, LSTM704A coupled with LSTM706A and LSTM708A corresponds to audio subband enhancer 144A associated with the first frequency subband. As another example, LSTM704B coupled with LSTM706B and LSTM708B corresponds to audio subband enhancer 144B associated with the second frequency subband. As yet another example, LSTM704C coupled with LSTM706C and LSTM708C corresponds to audio subband enhancer 144C associated with the third frequency subband. In an additional example, LSTM704D coupled with LSTM706D and LSTM708D corresponds to audio subband enhancer 144D associated with the fourth frequency subband. The enhanced subband audio generator 140, which includes an audio subband enhancer 144 corresponding to four frequency subbands, is provided as an illustrative example. It should be understood that in other examples, the enhanced subband audio generator 140 may include an audio subband enhancer 144 associated with fewer than four frequency subbands or more than four frequency subbands.

[0096] Coupler 148 includes a coupling layer 748A coupled to the fully coupled layer 750A. Coupler 148 also includes a coupling layer 748B coupled to the fully coupled layer 750B. Audio subband enhancer 144A processes audio data representing a first frequency subband (e.g., subband audio data 118A, subband audio data 128A, one or more additional sets of subband audio data, or a combination thereof) to generate outputs provided to each of the LSTM 706A and LSTM 708A of audio subband enhancer 144A.

[0097] The audio frequency divider 142 processes the audio data 117 to generate subband audio data 118A, subband audio data 118B, subband audio data 118C, and subband audio data 118D, corresponding to the first frequency subband, second frequency subband, third frequency subband, and fourth frequency subband, respectively. The audio frequency divider 142 processes the audio data 127 to generate subband audio data 128A, subband audio data 128B, subband audio data 128C, and subband audio data 128D, corresponding to the first frequency subband, second frequency subband, third frequency subband, and fourth frequency subband, respectively.

[0098] The audio frequency divider 142 provides subband audio data for the frequency subbands to the corresponding LSTM704. For example, the audio frequency divider 142 provides subband audio data 118A, subband audio data 128A, or both, to LSTM704A. As another example, the audio frequency divider 142 provides subband audio data 118B, subband audio data 128B, or both, to LSTM704B. The output of LSTM704 is provided to the corresponding LSTM706, the corresponding LSTM708, or both. For example, the output of LSTM704A is provided to LSTM706A, LSTM708A, or both.

[0099] The output of LSTM706 is provided to the concatenation layer 748A, and the output of LSTM708 is provided to the concatenation layer 748B. For example, the output of LSTM706A is provided to the concatenation layer 748A, and the output of LSTM708A is provided to the concatenation layer 748B. In certain embodiments, the output of LSTM706A, the output of LSTM708A, or both, corresponds to the emphasized subband audio data 136A.

[0100] The coupling layer 748A couples the outputs of LSTM706A, LSTM706B, LSTM706C, LSTM706D, one or more additional LSTMs, or combinations thereof, to generate a first coupled audio data representing a frequency band. In one example, the frequency band includes a first frequency subband, a second frequency subband, a third frequency subband, a fourth frequency subband, one or more additional frequency subbands, or combinations thereof. The first coupled audio data is processed by the fully coupled layer 750A. The coupler 148 applies a sigmoid function 752 to the output of the fully coupled layer 750A to generate a mask value 764. For example, the output of the fully coupled layer 750A includes a first count value (e.g., 257 integer values). By applying the sigmoid function 752 to the output of the fully coupled layer 750A, the first mask count value 764 (e.g., 257 mask values) is generated. In certain optional embodiments, the mask value is either 0 or 1.

[0101] The coupler 148 applies a delay 740 to the audio data 127 to generate delayed audio data 762. The coupler 148 includes a multiplier 754 that applies a mask value 764 to the delayed audio data 762 to generate masked audio data 766. For example, if the delayed audio data 762 contains a first count value (e.g., 257 values), applying the mask value 764 to the delayed audio data 762 includes applying the first mask value to the first value of the delayed audio data 762 to generate the first value of the masked audio data 766. In a particular optional embodiment, if the first mask value is 0, the first value of the masked audio data 766 is 0. Alternatively, if the first mask value is 1, the first value of the masked audio data 766 is the same as the first value of the delayed audio data 762. Thus, the mask value 764 allows a selected value of the delayed audio data 762 to be included in the masked audio data 766.

[0102] The coupling layer 748B couples the outputs of LSTM708A, LSTM708B, LSTM708C, LSTM708D, one or more additional LSTMs, or combinations thereof, to generate a second coupled audio data representing a frequency band. The second coupled audio data is processed by the fully coupled layer 750B to generate audio data 768. The coupler 148 generates enhanced audio data 135 based on the combination of masked audio data 766 and audio data 768.

[0103] In certain optional embodiments, the model architecture of the audio subband enhancer 144 is based on subband enhancer data 346. For example, each LSTM of audio subband enhancer 144A includes two hidden layers, each LSTM of audio subband enhancer 144B includes two hidden layers, each LSTM of audio subband enhancer 144C includes four hidden layers, and each LSTM of audio subband enhancer 144D includes four hidden layers. In certain embodiments, the audio subband enhancer 144 and coupler 148 correspond to SGNs. In certain embodiments, the audio subband enhancer 144 includes multiple LSTMs for generating enhanced audio data for each subband, which are smaller as a group than a single LSTM configured to generate enhanced audio data for the entire frequency band.

[0104] It should be understood that applying a delay 740 to audio data 127 to generate delayed audio data 762 is provided as an illustrative example. In another example, a delay 740 may be applied to audio data 117, audio data 127, or a combination thereof to generate delayed audio data 762.

[0105] Referring to Figure 17, a diagram 800 of exemplary embodiments of the operation of the system components is shown. In certain embodiments, the diagram 800 includes an example of an exemplary optional implementation of the highlighted subband audio generator 140 and an example of an exemplary optional implementation of the coupler 148.

[0106] In certain optional embodiments, one or more of the audio subband enhancers 144 are configured to perform procedural signal processing. For example, the audio subband enhancer 144 includes an audio subband enhancer 144E configured to process audio data of a fifth frequency subband using procedural signal processing to generate enhanced subband audio data 136E of a fifth frequency subband.

[0107] In an exemplary example, the audio frequency divider 142 processes audio data 117 to generate subband audio data 118E for a fifth frequency subband, in addition to generating subband audio data 118A, subband audio data 118B, subband audio data 118C, and subband audio data 118D. The audio frequency divider 142 processes audio data 127 to generate subband audio data 118A for a fifth frequency subband, in addition to generating subband audio data 128E, subband audio data 118B, subband audio data 118C, and subband audio data 118D. The audio frequency divider 142 that generates audio data associated with five frequency subbands is provided as an exemplary example; in other examples, the audio frequency divider 142 may generate audio data associated with fewer than five or more frequency subbands.

[0108] The audio subband enhancer 144E applies procedural signal processing to the subband audio data 118E, subband audio data 128E, or a combination thereof, to generate enhanced subband audio data 136E. In optional embodiments, the audio subband enhancer 144E applies procedural signal processing based on voice activity information 810 from one or more of the audio subband enhancers 144A-D. In certain embodiments, a fifth frequency subband (e.g., 8-16 kHz) corresponds to a higher frequency range, and subband SGN processing (using a generative network such as an LSTM) is bypassed for the higher frequency range because speech in the higher frequency range appears similar to noise to the generative network. In some optional embodiments, the audio subband enhancer 144E includes a machine learning model other than a generative network.

[0109] The coupler 148 generates audio data 864 based on the combination of the masked audio data 766 and audio data 768. The audio data 864 is for a specific frequency subband (e.g., a first frequency subband, a second frequency subband, a third frequency subband, and a fourth frequency subband). The coupler 148 includes a coupling layer 812 that couples the audio data 864 with the enhanced subband audio data 136E to generate enhanced audio data 135. The enhanced audio data 135 is for a specific frequency band (e.g., a fifth frequency subband and a fifth frequency subband).

[0110] In one or more embodiments, the techniques for supporting signal processing may include additional embodiments, such as any single embodiment or any combination of embodiments, as described below or in relation to one or more other processes or devices described elsewhere in this specification. In a first embodiment, supporting signal processing may include an apparatus configured for feedback reduction. The apparatus is further configured to perform feedback reduction using a trained ML model that determines an output audio signal by combining the output of the machine learning (ML) model with the microphone signal in order to reduce feedback components in the microphone signal.

[0111] In addition, the device may implement or operate according to one or more embodiments, as described below. In some implementations, the device includes a wireless device such as a UE. In some implementations, the device includes a remote server, such as a cloud-based computing solution, which receives image data for processing to determine an output image frame. In some implementations, the device may include at least one processor and memory coupled to that processor. The processor may be configured to perform the operations described herein with respect to the device. In some other implementations, the device may include a non-temporary computer-readable medium recording program code, which may be executable by a computer to cause the computer to perform the operations described herein with respect to the device. In some implementations, the device may include one or more means configured to perform the operations described herein. In some implementations, a wireless communication method may include one or more operations described herein with respect to the device.

[0112] In a second embodiment, in combination with the first embodiment, the apparatus is further configured to receive an input audio signal, determine an output audio signal by applying a machine learning model to the input audio signal, wherein the input audio signal includes desired audio components and feedback components, and the machine learning model is configured to reduce the feedback components.

[0113] In a third embodiment, the machine learning model is configured, in combination with one or more of the first or second embodiments, to retain desired components and remove feedback components.

[0114] In a fourth embodiment, in combination with one or more of the first to third embodiments, the apparatus further includes an amplification circuit coupled to one or more processors and configured to drive a converter from an output audio signal.

[0115] In the fifth embodiment, in combination with one or more of the first to fourth embodiments, one or more processors are configured to reduce feedback components by determining an output audio signal in combination with an input audio signal and a cancellation signal generated by a machine learning model, wherein the machine learning model is configured to generate a cancellation signal to cancel the nonlinearity created by an amplification circuit that amplifies the output audio signal.

[0116] In the sixth aspect, in combination with one or more of the first to fifth aspects, one or more processors are further configured to determine a feedback cancellation signal to reduce the linear component of the feedback component of the input audio signal, and to combine the feedback cancellation signal with the input audio signal before determining the output audio signal by applying a machine learning model.

[0117] In the seventh embodiment, in combination with one or more of the first to sixth embodiments, the apparatus further comprises an amplification circuit coupled to one or more processors and configured to amplify the output audio signal to drive a converter from the output audio signal, wherein a machine learning model is configured to reduce feedback components by reducing the nonlinearity of the amplification circuit.

[0118] In the eighth aspect, in combination with one or more of the first to seventh aspects, the apparatus further comprises an additional amplification circuit coupled to one or more processors, configured to amplify the input audio signal after combining a feedback cancellation signal with the input audio signal and before reducing the feedback component by applying a machine learning model, wherein the machine learning model is configured to reduce the feedback component by reducing the nonlinearity of the additional amplification circuit.

[0119] In the ninth aspect, in combination with one or more of the first to eighth aspects, the machine learning model is configured to reduce the feedback component based on parameters related to the feedback cancellation signal.

[0120] In the tenth embodiment, in combination with one or more of the first to ninth embodiments, one or more processors include a digital signal processor configured to determine a feedback cancellation signal and output parameters related to the feedback cancellation signal, and a neural signal processor configured to run a machine learning model based on the parameters related to the feedback cancellation signal.

[0121] In the eleventh embodiment, in combination with one or more of the first to tenth embodiments, the machine learning model is configured to reduce the feedback component based on input parameters corresponding to sensor inputs that do not correlate with the feedback component.

[0122] In the twelfth embodiment, in combination with one or more of the first to eleventh embodiments, the machine learning model is configured to reduce one or more artifacts originating from the amplification circuit without reducing other feedback in the input audio signal.

[0123] In the 13th aspect, in combination with one or more of the first to 12th aspects, one or more processors are configured to reduce feedback components by applying a machine learning model which includes applying a time-domain filter to the input audio signal after amplification, wherein the time-domain filter is configured based on the machine learning model.

[0124] In the 14th embodiment, in combination with one or more of the first to 13th embodiments, the apparatus further comprises: a first microphone coupled to one or more processors, from which an input audio signal is received; and a converter coupled to one or more processors, configured to reproduce an output audio signal.

[0125] In the 15th aspect, in combination with one or more of the first to 14th aspects, the method includes receiving an input audio signal, wherein the input audio signal includes a desired audio component and a feedback component, and reducing the feedback component by applying a machine learning model to the input audio signal to determine an output audio signal.

[0126] In the sixteenth embodiment, the machine learning model is configured, in combination with one or more of the first to fifteenth embodiments, to retain a desired component and remove a feedback component.

[0127] In the 17th aspect, in combination with one or more of the first to 16th aspects, the method further includes: amplifying an output audio signal for output to a converter; reducing feedback components; and determining the output audio signal by combining the input audio signal with a cancellation signal generated by a machine learning model before amplifying the output audio signal, wherein the machine learning model is configured to generate a cancellation signal to cancel out any nonlinearity created by amplifying the output audio signal.

[0128] In the 18th aspect, in combination with one or more of the first to 17th aspects, the method further includes determining a feedback cancellation signal to reduce the linear component of the feedback component of an input audio signal; combining the feedback cancellation signal with the input audio signal before applying a machine learning model to reduce the feedback component; and amplifying the output audio signal to drive a converter from the output audio signal, wherein the machine learning model is configured to reduce the feedback component by reducing the nonlinearity of amplifying the output audio signal.

[0129] In the 19th aspect, in combination with one or more of the first to 18th aspects, the method further comprises amplifying the input audio signal after combining it with a feedback cancellation signal and before reducing the feedback component by applying a machine learning model, wherein the machine learning model is configured to reduce the feedback component by reducing the nonlinearity of amplifying the input audio signal.

[0130] In the 20th embodiment, in combination with one or more of the first to 19 embodiments, the machine learning model is configured to reduce the feedback component based on input parameters related to the feedback cancellation signal.

[0131] In the 21st embodiment, in combination with one or more of the first to 20 embodiments, the machine learning model is configured to reduce the feedback component based on input parameters corresponding to sensor inputs that are not correlated with the feedback component.

[0132] In the 22nd aspect, in combination with one or more of the 1st to 21st aspects, amplification results in one or more artifacts arising from feedback components in the input audio signal, one or more artifacts including howling, and the machine learning model is configured to reduce one or more artifacts arising from amplification without reducing other howling in the input audio signal.

[0133] In the 23rd aspect, in combination with one or more of the first to 22nd aspects, the method further includes amplifying an input audio signal, wherein the amplification results in one or more artifacts arising from feedback components in the input audio signal, and a machine learning model is configured to reduce one or more artifacts.

[0134] In the diagrams, a single block may be described as performing one or more functions. The one or more functions performed by that block may be implemented in a single component or across multiple components, and / or using hardware, software, or a combination of hardware and software. To clearly demonstrate this hardware-software compatibility, various exemplary components, blocks, modules, circuits, and steps are briefly described below in relation to their functionality. Whether such functionality is implemented in hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for specific applications, but such decisions on implementation should not be construed as a departure from the scope of this disclosure. Furthermore, exemplary devices may include components other than those shown, including well-known components such as processors and memory.

[0135] Unless otherwise stated, as will be evident from the following discussions, discussions throughout this application that use terms such as “access,” “receive,” “send,” “use,” “select,” “determine,” “normalize,” “multiply,” “average,” “monitor,” “compare,” “apply,” “update,” “measure,” “derive,” “solve,” and “generate” refer to the actions and processes of a computer system or similar electronic computing device that manipulate data represented as physical (electronic) quantities in the registers and memory of a computer system to convert that data into other data similarly represented as physical quantities in the registers, memory, or other such information storage devices, transmission devices, or display devices of a computer system. The use of different terms to refer to actions or processes of a computer system does not necessarily indicate different behaviors. For example, “determining” data may mean “generating” data. Another example is that “determining” data may mean “retrieving” data.

[0136] The terms “device” and “apparatus” are not limited to one physical object or a specific number of physical objects (such as one smartphone, one camera controller, or one processing system). As used herein, a device may be any electronic device having one or more components capable of implementing at least some parts of this disclosure. While the descriptions and examples herein use the term “device” to illustrate various aspects of this disclosure, the term “device” is not limited to a specific configuration, type, or number of objects. As used herein, an apparatus may include a device that performs the operation described, or a part of such a device.

[0137] Some components in a device or apparatus described as “means for accessing,” “means for receiving,” “means for sending,” “means for using,” “means for selecting,” “means for deciding,” “means for normalizing,” “means for multiplying,” or other similarly named terms referring to one or more operations on data such as image data, may refer to processing circuits configured to perform the listed functions via hardware, software, or a combination of software-configured hardware (e.g., application-specific integrated circuits (ASICs), digital signal processors (DSPs), graphics processing units (GPUs), central processing units (CPUs), computer vision processors (CVPs), or neural signal processors (NSPs)).

[0138] Those skilled in the art will understand that information and signals may be represented using any of a variety of different techniques and methods. For example, data, instructions, commands, information, signals, bits, symbols, and chips which may be mentioned throughout the above description may be represented by voltage, electric current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof.

[0139] With respect to the diagrams referenced above, the components, functional blocks, and modules described herein include, in particular, processors, electronic devices, hardware devices, electronic components, logic circuits, memory, software code, firmware code, or any combination thereof. Software, whether called software, firmware, middleware, microcode, hardware description language, or otherwise, should be broadly interpreted in the examples to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, and / or functions. In addition, the features discussed herein may be implemented via dedicated processor circuits, via executable instructions, or a combination thereof.

[0140] Those skilled in the art will know that one or more blocks (or actions) described with reference to Figure 3 can be combined with one or more blocks (or actions) described with reference to another of the figures. For example, one or more blocks (or actions) in Figure 3 can be combined with one or more blocks (or actions) in Figure 1 or Figure 2. As another example, one or more blocks associated with Figures 4-9 can be combined with one or more blocks (or actions) associated with Figures 1-3. As yet another example, one or more blocks associated with Figures 15-17 can be combined with one or more blocks (or actions) associated with Figures 3-9. As yet another example, one or more blocks associated with Figures 10-14 can be combined with one or more blocks (or actions) associated with Figures 1-2 and / or Figures 3-9.

[0141] Those skilled in the art will further understand that various exemplary logic blocks, modules, circuits, and algorithmic steps described herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly demonstrate this hardware-software compatibility, various exemplary components, blocks, modules, circuits, and steps have been outlined above in relation to their functionality. Whether such functionality is implemented as hardware or as software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for specific applications, but such decisions on implementation forms should not be construed as causing a departure from the scope of this disclosure. Those skilled in the art will also readily recognize that the order or combination of components, methods, or interactions described herein are merely examples, and that components, methods, or interactions of various aspects of this disclosure may be combined or implemented in ways other than those shown and described herein.

[0142] The various exemplary logics, logic blocks, modules, circuits, and algorithmic processes described in relation to the implementations disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. Hardware-software compatibility is generally described in terms of functionality, as shown in the various exemplary components, blocks, modules, circuits, and processes described above. Whether such functionality is implemented in hardware or software depends on the specific application and the design constraints imposed on the overall system.

[0143] In one or more embodiments, the described operations may be implemented in hardware, digital electronic circuitry, computer software, firmware, or any combination thereof, including the structures disclosed herein and their structural equivalents. Implementations of the subject matter described herein may also be implemented as one or more computer programs, which are one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing device or for controlling the operation of a data processing device.

[0144] The operation of methods or algorithms disclosed herein may be performed in a processor-executable software module that resides on a computer-readable medium and may be commercially available as a computer program product as software. Computer-readable medium includes both computer storage medium and communication medium, which may include any medium that may enable the transfer of computer programs from one location to another. Storage medium may be any available medium that can be accessed by a computer. Such computer-readable medium may include, but not exclusively, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that may be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection may also be appropriately referred to as computer-readable medium. As used herein, the terms "disk" and "disc" include compact discs (CDs), laser discs, optical discs, digital multipurpose discs (DVDs), floppy disks, and Blu-ray discs. A disk typically reproduces data magnetically, while a disc reproduces data optically using a laser. Any combination of the above should also be included within the scope of computer-readable media.

[0145] Various modifications of the implementations described herein may be readily apparent to those skilled in the art, and the general principles defined herein may be applied to several other implementations without departing from the spirit or scope of this disclosure. Therefore, the claims are not intended to be limited to the implementations shown herein, but should be given the broadest scope consistent with this disclosure, the principles disclosed herein, and the novel features.

[0146] In addition, antonyms such as "top" and "bottom," or "front" and "back," or "upper" and "lower," or "front" and "rear," or "left" and "right" may be used to facilitate the description of a figure and indicate a relative position corresponding to the orientation of the figure on a properly oriented page, and may not reflect the proper orientation of any device being implemented. This will be easily understood by those skilled in the art.

[0147] Some features described herein in the context of separate implementations may also be realized in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be realized separately or in any suitable partial combination in multiple implementations. Furthermore, even if features have been described above as functioning in a particular combination and are initially claimed as such, one or more features from the claimed combination may, in some cases, be removed from that combination, and the claimed combination may cover partial combinations or variations of partial combinations.

[0148] Similarly, while actions are shown in a specific order in the diagram, this should not be understood as requiring that such actions be performed in a specific or sequential order, or that all shown actions be performed, in order to achieve the desired result. Furthermore, the diagram may schematically represent one or more exemplary processes in the form of a flow chart. However, other actions not shown may be incorporated into the schematicly represented exemplary process. For example, one or more additional actions may be performed before, after, simultaneously with, or between any of the shown actions. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the implementation forms described above should not be understood as requiring such separation in all implementation forms, and it should be understood that the described program components and systems may generally be integrated together within a single software product or packaged within multiple software products. In addition, several other implementation forms fall within the scope of the following claims. In some cases, the actions enumerated in the claims may be performed in a different order and still achieve the desired result.

[0149] As used herein, including in the claims, the term “or” means that, when used in a list of two or more items, any one of the listed items may be taken alone, or any combination of two or more of the listed items may be taken. For example, if a composition is described as containing component A, B, or C, the composition may include A only, B only, C only, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B, and C. Also, as used herein, including in the claims, “or” means a disjunctive list, for example, that the list “at least one of A, B, or C” means any of these in the case of A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C), or any combination thereof.

[0150] As a person skilled in the art would understand, the term “substantially” broadly defines (and includes) what is specified; for example, substantially 90 degrees includes 90 degrees, and substantially parallel includes parallel, but not necessarily the whole. In any disclosed implementation, the term “substantially” may be replaced by “within [percentage] of” what is specified, where the percentage includes 0.1 percent, 1 percent, 5 percent, or 10 percent.

[0151] The above description in this disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to this disclosure will be readily apparent to a person skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the embodiments and designs described herein, but should be given the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. It is a device, Memory for storing processor-readable code, One or more processors coupled to the memory, The system includes, and the one or more processors, An input audio signal, wherein the input audio signal includes a desired audio component and a feedback component, and the input audio signal is received. An apparatus configured to determine an output audio signal by applying a machine learning model to the input audio signal, wherein the machine learning model is configured to reduce the feedback component.

2. The apparatus according to claim 1, wherein the machine learning model is configured to retain a desired component.

3. The circuit further comprises an amplification circuit coupled to one or more of the aforementioned processors and configured to drive a converter from the output audio signal, The one or more processors are configured to reduce the feedback component by combining the input audio signal and the cancellation signal generated by the machine learning model to determine the output audio signal. The apparatus according to claim 1, wherein the machine learning model is configured to generate the cancellation signal in order to cancel the nonlinearity created by the amplification circuit amplifying the output audio signal.

4. The aforementioned one or more processors In order to reduce the linear component of the feedback component of the input audio signal, a feedback cancellation signal is determined. Before determining the output audio signal by applying the machine learning model, the feedback cancellation signal is combined with the input audio signal. It is further structured in the following way: The aforementioned device The system further comprises an amplification circuit coupled to one or more of the processors and configured to amplify the output audio signal to drive a converter from the output audio signal, The apparatus according to claim 1, wherein the machine learning model is configured to reduce the feedback component by reducing the nonlinearity of the amplification circuit.

5. The system further comprises an additional amplification circuit coupled to one or more of the processors, configured to amplify the input audio signal after combining the feedback cancellation signal with the input audio signal and before reducing the feedback component by applying the machine learning model, The apparatus according to claim 4, wherein the machine learning model is configured to reduce the feedback component by reducing the nonlinearity of the additional amplification circuit.

6. The apparatus according to claim 4, wherein the machine learning model is configured to reduce the feedback component based on parameters related to the feedback cancellation signal.

7. The aforementioned one or more processors A digital signal processor configured to determine the feedback cancellation signal and output the parameters related to the feedback cancellation signal, A neural signal processor configured to execute the machine learning model based on the parameters related to the feedback cancellation signal, The apparatus according to claim 6, including the apparatus described in claim 6.

8. The apparatus according to claim 4, wherein the machine learning model is configured to reduce the feedback component based on input parameters corresponding to inputs from sensors that do not correlate with the feedback component.

9. The apparatus according to claim 8, wherein the machine learning model is configured to reduce one or more artifacts originating from the amplification circuit without reducing other feedback in the input audio signal.

10. The apparatus according to claim 9, wherein one or more processors are configured to reduce the feedback component by applying the machine learning model, which includes amplifying the input audio signal and then applying a time-domain filter to the input audio signal, and the time-domain filter is configured based on the machine learning model.

11. A first microphone coupled to one or more processors, wherein the input audio signal is received from the first microphone, A converter coupled to one or more of the aforementioned processors, configured to reproduce the output audio signal, The apparatus according to claim 1, further comprising the following:

12. It is a method, The system receives an input audio signal, wherein the input audio signal includes desired audio components and feedback components. In order to determine the output audio signal, the feedback component is reduced by applying a machine learning model to the input audio signal, Methods that include...

13. The method according to claim 12, wherein the machine learning model is configured to retain a desired component.

14. To amplify the output audio signal for output to the converter, It further includes, Reducing the feedback component includes determining the output audio signal by combining the input audio signal with a cancellation signal generated by the machine learning model before amplifying the output audio signal. The method according to claim 12, wherein the machine learning model is configured to generate the cancellation signal in order to cancel the nonlinearity created by amplifying the output audio signal.

15. In order to reduce the linear component of the feedback component of the input audio signal, a feedback cancellation signal is determined, Before reducing the feedback component by applying the machine learning model, the feedback cancellation signal is combined with the input audio signal. To drive the converter from the output audio signal, the output audio signal is amplified, It further includes, The method according to claim 12, wherein the machine learning model is configured to reduce the feedback component by reducing the nonlinearity of amplifying the output audio signal.

16. After combining the feedback cancellation signal with the input audio signal, and before reducing the feedback component by applying the machine learning model, amplify the input audio signal. It further includes, The method according to claim 15, wherein the machine learning model is configured to reduce the feedback component by reducing the nonlinearity of amplifying the input audio signal.

17. The method according to claim 15, wherein the machine learning model is configured to reduce the feedback component based on input parameters related to the feedback cancellation signal.

18. The method according to claim 15, wherein the machine learning model is configured to reduce the feedback component based on input parameters corresponding to inputs from sensors that do not correlate with the feedback component.

19. The amplification results in one or more artifacts arising from the feedback component in the input audio signal, and the one or more artifacts include howling. The method according to claim 18, wherein the machine learning model is configured to reduce the one or more artifacts resulting from the amplification without reducing other feedback in the input audio signal.

20. The input audio signal, which amplifies the input audio signal, causing one or more artifacts arising from the feedback component within the input audio signal. It further includes, The method according to claim 12, wherein the machine learning model is configured to reduce the one or more artifacts.