Audio device with efficient neural network processing and related methods

By adopting a combined solution of dynamic neural networks and exit modules in small devices, the problem of computing power limitations is solved, and efficient neural network processing is achieved, reducing computing costs and extending battery life.

CN120201347APending Publication Date: 2025-06-24GN HEARING AS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411894030.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-20
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Computational power limitations become a challenge when using deep neural networks (DNNs) in small devices, especially on devices with limited processing performance and battery life.

Method used

Dynamic neural network (DyNN) is used to reduce computational costs by adjusting the configuration and parameters of the model at runtime. The specific implementation includes using a neural network with a dynamic model layer in the audio device, combining an exit module to determine the quality of the intermediate layer output, and based on this, whether to exit the neural network processing in advance.

Benefits of technology

Through the use of dynamic neural networks, efficient neural network processing on small devices is achieved, reducing computing costs and battery consumption, extending battery life, while maintaining performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201347A_ABST
    Figure CN120201347A_ABST
Patent Text Reader

Abstract

The invention relates to an audio device with efficient neural network processing and related methods. An audio device comprising: an audio enhancement module comprising a first neural network having a first model layer, the first model layer comprising a first input layer, a plurality of first intermediate layers, and a first output layer; and a first exit module; wherein the audio enhancement module is configured to process an audio input signal using a first neural network to provide an audio output signal, and wherein the at least one first intermediate layer has an exit likelihood to provide an intermediate layer output, and wherein the first exit module is configured to determine whether the intermediate layer output satisfies a first criterion, the first criterion is indicative of performance, quality and / or efficiency of the intermediate layer output, and wherein the audio device is configured to determine an audio output signal based on the intermediate layer output in accordance with the intermediate layer output satisfying the first criterion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of audio devices and methods performed by audio devices, and more particularly to audio devices and related methods for implementing efficient neural network processing. Background Art

[0002] For example, signal processing using deep neural networks (DNNs) and other types of neural networks is rapidly becoming an integral part of electronic devices (e.g., audio devices) as it can solve problems that were previously unsolvable using traditional methods. As neural networks continue to evolve and their use increases, e.g., in audio signal processing, the reliability of their predictions becomes increasingly important. DNNs can learn complex patterns in large amounts of data and make accurate predictions based on this knowledge. DNNs have proven to be effective in tasks such as speech recognition, speaker identification, and speech synthesis.

[0003] A key advantage of DNNs in speech processing is the ability to learn hierarchical representations of speech signals, capturing both low-level acoustic features and high-level semantic information. This enables DNNs to perform well even in noisy environments where traditional speech processing methods struggle. Summary of the Invention

[0004] However, despite their impressive performance, there are some limitations to using DNNs in small devices. A major limitation is computational power. DNNs are complex models with many parameters that require a large amount of computational resources to train and use. This can be a challenge for small devices that typically have limited processing power and battery life.

[0005] Dynamic neural networks (DyNNs), also known as dynamic models, can help reduce the computational cost of running DNNs in small devices by adjusting the configuration (e.g., the architecture, computational graph, path, and / or route) and parameters of the model at runtime based on the input data. Thus, dynamic neural networks (DyNNs) can have higher computational efficiency than traditional static neural networks.

[0006] Therefore, there is a need for audio devices with efficient neural network processing and methods performed by audio devices that can mitigate, alleviate, or solve existing drawbacks and can improve neural network processing efficiency, thereby achieving lower computational costs and battery savings.

[0007] An audio device is disclosed. The audio device may be configured to act as a receiver device and / or a transmitter device. The audio device may include a memory, an interface, and one or more processors. Optionally, the audio device includes one or more output transducers (e.g., one or more speakers) and one or more input transducers (e.g., one or more microphones). In one or more examples or embodiments, the one or more processors are configured to obtain audio data, e.g., an audio input signal. In other words, the audio device may be configured to obtain audio data, e.g., an audio input signal, using one or more processors and / or via the interface.

[0008] The audio device includes an audio enhancement module, the audio enhancement module includes a first neural network having a first model layer, the first model layer includes a first input layer, a plurality of first intermediate layers, and a first output layer. The audio device includes a first dropout module. The audio device (e.g., the audio enhancement module) is configured to process an audio input signal using the first neural network to provide an audio output signal. Optionally, at least one of the first intermediate layers has a dropout probability to provide an intermediate layer output. Optionally, the audio device (e.g., the first dropout module) is configured to determine whether the intermediate layer output meets a first criterion. The first criterion may indicate the quality of the intermediate layer output. Depending on the intermediate layer output meeting the first criterion, the audio device may be configured to determine the audio output signal, e.g., based on the intermediate layer output.

[0009] A method performed by an audio device is disclosed. The method may be used to implement efficient neural network processing, wherein the audio device includes: an audio enhancement module, the audio enhancement module includes a first neural network having a first model layer, the first model layer includes a first input layer, a plurality of first intermediate model layers, and a first output layer; and a first dropout module. The method includes, e.g., processing an audio input signal using the first neural network to provide an audio output signal. Optionally, at least one of the first intermediate layers has a dropout probability to provide an intermediate layer output. The method includes, e.g., using the first dropout module to determine whether the intermediate layer output meets a first criterion, e.g., wherein the first criterion indicates the quality of the intermediate layer output. The method includes determining the audio output signal depending on the intermediate layer output meeting the first criterion, e.g., based on the intermediate layer output.

[0010] The present disclosure provides an audio device and related methods having improved processing efficiency when using a neural network (e.g., having improved processing efficiency when using a dynamic neural network (e.g., using a conditional computing neural network)). The present disclosure can reduce the scope of the neural network used to perform a specific task (e.g., reduce a portion of the neural network used to perform a specific task). For example, the present disclosure allows a specific task to be performed with a neural network by using only a subset of the layers of the neural network, e.g., without having to execute or complete all layers of the neural network. Further, the present disclosure provides an audio device that reduces battery consumption and thus increases battery life while performing more efficient processing. In other words, for the same performance, the present disclosure allows a reduction in computational cost and an increase in battery life. The present disclosure provides a more general audio device having dynamically adjustable neural network processing depending on one or more parameters (e.g., output quality, parameters of the audio device (e.g., capabilities), and / or user preferences). In other words, the present disclosure allows the computational cost of processing the neural network to be adjusted based on the size and / or capabilities of the audio device. For example, the present disclosure can allow the architecture and parameters of a neural network model to be adjusted at runtime based on the input data and output of the neural network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features and advantages of the present disclosure will become apparent to those skilled in the art from the following detailed description of examples of the present disclosure with reference to the accompanying drawings, in which:

[0012] Figure 1 An exemplary audio device according to the present disclosure is schematically illustrated;

[0013] Figure 2A and Figure 2B A flowchart showing an exemplary method according to the present disclosure is shown;

[0014] Figure 3 An exemplary audio device according to the present disclosure is schematically illustrated, in which the techniques disclosed herein are applied;

[0015] Figure 4 An exemplary audio device according to the present disclosure is schematically illustrated, in which the first training technique disclosed herein is applied;

[0016] Figure 5 An exemplary audio device according to the present disclosure is schematically illustrated, in which the second training technique disclosed herein is applied; and

[0017] Figure 6 An exemplary audio device according to the present disclosure is schematically illustrated, in which the third training technique disclosed herein is applied. DETAILED DESCRIPTION

[0018] Various examples and details will be described below with reference to the related drawings. It should be noted that the drawings may or may not be drawn to scale, and elements having similar structures or functions in all the drawings are denoted by the same reference numerals. It should also be noted that the drawings are only intended to facilitate the description of the examples. The drawings are not intended to be an exhaustive description of the present disclosure or a limitation on the scope of the present disclosure. In addition, the examples shown need not have all the aspects or advantages shown. Aspects or advantages described in connection with a particular example are not necessarily limited to that example and may be practiced in any other example, even if not so stated or not so explicitly described.

[0019] These figures are schematic and simplified for clarity, and they only show the details that contribute to the understanding of the present disclosure, while other details have been omitted. Throughout the text, the same reference numerals are used for the same or corresponding parts.

[0020] An audio device is disclosed. The audio device may be configured to act as a receiver device and / or a transmitter device. In other words, the audio device is configured to receive an input signal, e.g., audio data, from an audio device configured to act as a transmitter device, or vice versa. The audio device as disclosed herein may include one or more interfaces, one or more audio speakers, one or more microphones (e.g., including a first microphone), one or more processors, and one or more memories. The one or more interfaces may include one or more of the following: a wireless interface, a wireless transceiver, an antenna, an antenna interface, a microphone interface, and a speaker interface.

[0021] In addition, the audio device may include one or more microphones, e.g., a first microphone, an optional second microphone, an optional third microphone, and an optional fourth microphone. The audio device may include one or more audio speakers, e.g., an audio receiver, e.g., a speaker.

[0022] An audio device can be regarded as an audio device configured to obtain audio data (e.g., an input signal, e.g., an audio input signal), output an audio signal, and process the input signal (e.g., an audio input signal). The audio device can be regarded as or include headphones, a speakerphone, a hearing aid, and / or a video bar. The audio device can be regarded as, for example, a conference audio device, e.g., the conference audio device is configured to be used by one party (e.g., one or more users at the proximal end) to communicate with one or more other parties (e.g., one or more users at the distal end). The audio device configured to act as a receiver device can also be configured to act as a transmitter device when transmitting the output signal back to the distal end. The receiver audio device and the transmitter audio device can thus switch between being a receiver audio device and a transmitter audio device. The audio device can be regarded as an intelligent audio device. The audio device can be used for meetings and / or encounters between two or more parties that are far apart from each other. The audio device can be used by one or more users in the vicinity (also called the proximal end) where the audio device is located. The audio device can be configured to, for example, use an audio speaker and output an audio device output at the receiver end based on the input signal. The audio device output can be regarded as an audio output signal, which is the output of the audio speaker at the proximal end where the audio device and the user of the audio device are located.

[0023] The audio device can be a single audio device. The audio device can be regarded as multiple interconnected audio devices, e.g., a system, e.g., an audio device system. The system can include one or more users.

[0024] In one or more example audio devices, the interface includes a wireless transceiver (also called a radio transceiver) and an antenna for wirelessly transmitting and receiving input signals (e.g., audio signals), e.g., an antenna for wirelessly transmitting the output signal and / or wirelessly receiving a wireless input signal. The audio device can be configured to wirelessly communicate with one or more electronic devices (e.g., another audio device, a smart phone, a tablet computer, a computer, and / or a smart watch). The audio device optionally includes an antenna for converting one or more wireless input audio signals into an antenna output signal. The audio device system and / or the audio device can be configured to wirelessly communicate via a wireless communication system (e.g., a short-range wireless communication system, e.g., Wi-Fi, Bluetooth, Zigbee, IEEE 802.11, IEEE 802.15, infrared, etc.).

[0025] An audio device system and / or an audio device may be configured for wireless communication via a wireless communication system, such as a 3GPP system, such as a 3GPP system supporting one or more of the following: New Radio (NR), Narrowband IoT (NB-IoT), and Long-Term Evolution-Machine Type Communication Enhanced (LTE-M) millimeter-wave communication (e.g., millimeter-wave communication in a licensed band, e.g., device-to-device millimeter-wave communication in a licensed band).

[0026] In one or more example audio device systems and / or audio devices, the interface of the audio device includes one or more of the following: a Bluetooth interface, a Bluetooth Low Energy interface, and a magnetic induction interface. For example, the interface of the audio device may include a Bluetooth antenna and / or a magnetic interference antenna.

[0027] In one or more example audio devices, the interface may include a connector for wired communication via a connector (e.g., by using a cable). The connector may connect one or more microphones to the audio device. The connector may connect the audio device to an electronic device, e.g., for a wired connection. The connector may be regarded as an electrical connector, e.g., a physical connector for connecting the audio device to another device via a wire.

[0028] One or more interfaces may be or include a wireless interface, e.g., a transmitter and / or a receiver, and / or a wired interface, e.g., a connector for physical coupling. For example, an audio device may have an input interface configured to receive data (e.g., a microphone input signal). In one or more example audio devices, the audio device may be used in all form factors in all types of environments, e.g., for headphones and / or video conferencing devices. For example, the audio device may not have specific microphone placement requirements. In one or more example audio devices, the audio device may include an external microphone.

[0029] The audio device includes an audio enhancement module, which includes a first neural network having a first model layer. The first model layer includes a first input layer, a plurality of first intermediate model layers, and a first output layer. The audio enhancement module may be regarded as a module configured to operate according to the first neural network. The audio enhancement module (e.g., the first neural network) may be used to process audio data, e.g., process an audio input signal. The audio enhancement module may be configured to process the audio input signal by using an ML-based method and / or a signal processing-based method, e.g., by using the first neural network. It can be understood that the audio enhancement module may be regarded as and / or represented as an audio model module. The technical terms audio enhancement module and audio model module may be used interchangeably. In one or more examples or embodiments, the audio enhancement module may alternatively or additionally be regarded as a noise suppression module and / or a denoising module.

[0030] Multiple first intermediate layers can be considered as hidden layers (e.g., hidden features). The multiple first intermediate layers can include a first primary intermediate layer, a first secondary intermediate layer, a first tertiary intermediate layer, etc. The first neural network can be configured to operate according to a model (e.g., a machine learning model, e.g., a first machine learning model). The model mentioned herein (e.g., the first machine learning model) can be regarded as a model and / or scheme and / or mechanism and / or method configured to process an audio input signal based on the layer outputs of the first neural network (e.g., based on the intermediate layer outputs and / or a previous model). In one or more examples or embodiments, the first neural network is a dynamic neural network (DyNN). The first neural network can be configured to operate according to a first dynamic model. The dynamic model can be regarded as a model capable of adjusting its architecture and parameters at runtime.

[0031] In one or more example audio devices, the model mentioned herein can be stored on a non-transitory storage medium (e.g., on the memory of the audio device). The model can be stored on the non-transitory storage medium of the audio device configured to execute the model. In one or more example audio devices, the model can include model data and / or computer-readable instructions (e.g., based on the audio input signal, the features of the audio input signal, the audio device parameters, and / or the intermediate layer outputs disclosed herein). The audio device can use the model data and / or the computer-readable instructions. The audio device can use the model (e.g., the model data and / or the computer-readable instructions) to process audio data, e.g., the audio input signal.

[0032] In one or more examples or embodiments, an audio device is configured to obtain audio data, such as an audio input signal, for example, using one or more processors and / or via an interface. In one or more example audio devices, the audio device may be configured to obtain an input signal, such as an audio input signal, from a transmitter device. In one or more example audio devices, the audio device is configured to obtain audio data from a remote end (e.g., a remote party or user). For example, a processor may be configured to obtain audio data (e.g., an audio input signal) via one or more microphones of the audio device (e.g., microphones associated with and / or included in the audio device). The audio data may include and / or be based on one or more audio signals obtained by the audio device. In other words, the transmitter device may be regarded as an audio device of the remote end. The audio data may be regarded as data including audio. In one or more embodiments or examples, the audio input signal has been signal-processed at the transmitter device, such as encoded, compressed, and / or enhanced. The audio data may indicate an audio signal generated by a user at the remote end. In other words, the audio data may indicate speech, such as speech from the remote transmitter device. The audio data may be based on and / or regarded as an output signal of the transmitter device, such as an output signal of a signal processor of the transmitter device. Obtaining the audio data may include retrieving and / or receiving the audio data. When obtained from one or more microphones of the audio device (e.g., a first microphone and / or a second microphone), the audio data (e.g., an audio input signal) may be based on an input signal from a proximal end (e.g., speech). The audio data may be based on the input signal, such as based on a first microphone input signal, a second microphone input signal, and / or a transceiver input signal.

[0033] The audio device is configured to process an audio input signal, such as audio data, for example, using one or more processors to provide an audio output signal using a first neural network. Processing the audio input signal to provide the audio output signal may include performing one or more audio processing steps on the audio input signal. For example, processing the audio input signal to provide the audio output signal may include, for example, using a signal processor to perform noise reduction on the audio input signal, such as background noise reduction, for example, to provide a denoised audio output signal. Other examples may include processing the audio input signal to provide the audio output signal, which may include, for example, using a first neural network to perform filtering on the audio input signal to provide a filtered audio output signal and / or a speech enhancement task on the audio input signal. In addition, processing the audio input signal to provide the audio output signal may include performing compression of the audio input signal. The signal processor may include an audio enhancement module and perform processing according to the audio enhancement module to provide the audio output signal, such as echo control, dereverberation, noise reduction, and / or beamforming.

[0034] In one or more examples or embodiments, at least one first intermediate layer has an exit possibility to provide an intermediate layer output. In other words, at least one of the plurality of first intermediate layers is configured to provide an exit possibility in the processing of an audio input signal by an audio enhancement module to provide an intermediate layer output. The intermediate layer output can be regarded as the result of the exit possibility. The exit possibility can be regarded as the possibility that an audio device (e.g., an audio enhancement module) exits during the processing of an audio input signal and generates an intermediate layer output based on the first intermediate layer. In other words, the exit possibility can be regarded as the possibility of exiting the processing of an audio input signal using a first neural network before the first output layer, e.g., the possibility of early exit. For example, the exit possibility can be regarded as at least one first intermediate layer being configured to provide a useful intermediate layer output for determining an audio output signal for which a given audio processing task has been performed. In other words, at least one first intermediate layer can be configured to provide an intermediate layer output that is configured to perform an expected audio processing task, e.g., denoising, echo suppression, and / or dereverberation of an audio input signal.

[0035] In one or more examples or embodiments, at least some of the first intermediate layers have an exit possibility. For example, at least two, at least three, at least five, at least ten of the first intermediate layers have an exit possibility. In one or more examples or embodiments, each first intermediate layer has an exit possibility. In one or more examples or embodiments, the first intermediate layers having an exit possibility are evenly distributed or partitioned. The even distribution or partitioning can be regarded as each first intermediate layer including or requiring substantially the same amount of processing, and / or each first intermediate layer providing substantially the same amount of progress, e.g., processing progress. In other words, the even distribution or partitioning can be regarded as the processing of an audio input signal being evenly distributed or partitioned among the first intermediate layers. One or more first intermediate layers can include a first primary intermediate layer, a first secondary intermediate layer, a first tertiary intermediate layer, etc. One or more first intermediate layers can be regarded as being located between a first input layer and a first output layer.

[0036] The audio device includes a first exit module. The first exit module can be regarded as a module configured to evaluate one or more features of the output of a model layer (e.g., the first model layer). In other words, the first exit module can be configured to determine whether the output of the model layer (e.g., the first model layer) meets one or more criteria (e.g., performance, quality, and / or efficiency criteria, such as the first criterion disclosed herein) or whether more processing is required. One or more criteria can include, for example, or be based on one or more of the following: a voice quality criterion, a latency criterion, a voice intelligibility criterion, an SNR criterion, a learning criterion (e.g., learned by the first exit module), and a multi-dimensional threshold (e.g., a threshold that considers multiple parameters (e.g., performance, quality, and / or efficiency) simultaneously). In one or more examples or embodiments, the first neural network can be regarded as or include a recurrent neural network, e.g., having the same weights in each layer. By having the recurrent neural network have the same weights in each layer, the first exit module can more easily evaluate the output of the model layer, e.g., evaluate whether the output meets the criteria, because the output of each layer can represent a similar transformation (which may be easier to compare and / or evaluate considering the common criteria and / or thresholds).

[0037] The first exit module can be regarded as an early exit estimator module and / or a quality assessment module. The terms first exit module, early exit estimator module, and quality assessment module can be used interchangeably. The first exit module can be regarded as a gating module, e.g., the gating module is configured to compare the intermediate layer output with a threshold (e.g., the first threshold), and / or compare the intermediate layer output combined with the audio input signal with a threshold (e.g., the first threshold). It can be understood that the first exit module can be included in the audio enhancement module or form a part of the audio enhancement module.

[0038] The first exit module is configured to determine whether the intermediate layer output meets the first criterion, where the first criterion indicates the performance, quality, and / or efficiency of the intermediate layer output.

[0039] The first criterion can be regarded as a criterion indicating when the intermediate layer output indicates satisfactory processing performance, quality, and / or efficiency. In other words, when the first criterion is met, the intermediate layer output indicates satisfactory processing performance, quality, and / or efficiency for a given task to be performed by the first neural network. The first criterion may include a first threshold. In one or more example audio devices, in accordance with the intermediate layer output meeting the first criterion, one or more processors are configured to determine an audio output signal based on the intermediate layer output. It can be understood that the first criterion is met when the intermediate layer output is higher than or equal to the first threshold. The first criterion may, for example, include or be based on one or more of the following: a voice quality criterion, a latency criterion, a speech intelligibility criterion, an SNR criterion, a learning criterion (e.g., learned by the first exit module), and a multi-dimensional threshold (e.g., a threshold considering multiple parameters simultaneously, such as performance, quality, and / or efficiency). The first criterion can be regarded as a learning criterion indicating certain characteristics of the intermediate layer output. The learning criterion can be based on voting and / or a score threshold to determine whether to continue the processing of the first neural network (e.g., stay in the first neural network) or leave or exit the processing of the first neural network, where the highest vote or score compared to the score threshold determines further processing. For example, the first criterion can indicate that no further processing is required, e.g., the intermediate layer output has met a certain condition (e.g., a quality condition) and / or cannot be significantly improved. The first threshold can be adaptive, e.g., based on the desired voice quality, audio device parameters, and / or audio input signal characteristics.

[0040] The first exit module can determine whether the estimated signal-to-noise ratio SNR of the intermediate layer output and / or the predicted mean opinion score MOS of the intermediate layer output meet the first criterion. In other words, the first criterion can be based on the SNR and / or MOS. For example, the first threshold can include or be based on the SNR and / or MOS.

[0041] In one or more examples or embodiments, the first exit module may determine whether one or more parameters of the intermediate layer output meet a first criterion. For example, the first exit module may determine whether parameters based on the Perceptual Evaluation of Speech Quality (PESQ) and / or the Deep Neural Network - based Speech Enhancement (DNSMOS) meet the first criterion. However, PESQ and DNSMOS may require samples in the range of seconds (e.g., audio samples of several seconds). Therefore, using these parameters to evaluate whether the intermediate layer output meets the first criterion may take too much time. In one or more examples or embodiments, the first exit module may determine whether the parameter is based on a parameter / feature that only requires signal samples in the range of milliseconds to estimate the quality of the intermediate layer output. For example, the first exit module may determine whether the intermediate layer output meets the first criterion by determining whether the non - intrusive performance, quality, and / or efficiency score of the intermediate layer output meets the first criterion. The non - intrusive performance, quality, and / or efficiency score may include, for example, the SNR score. It can be understood that the first exit module may determine the non - intrusive performance, quality, and / or efficiency score by directly extracting tags from the intermediate layer output without applying the intermediate layer output to the audio input signal. The parameters of the intermediate layer output may include, for example, speech quality metrics such as the Virtual Speech Quality Objective Listener (ViSQOL) and / or the Perceptual Objective Listening Quality Analysis (POLQA). Another example of the parameters of the intermediate layer output may be, for example, the energy consumption of the processing of a certain intermediate layer vs. the quality and / or performance of the intermediate layer output.

[0042] In one or more examples or embodiments, the intermediate layer output may include a mask applied in processing an audio input signal to provide an audio output signal. For example, the intermediate layer output may include a mask for performing a specific task of audio processing, such as removing noise from the audio input signal to provide a denoised audio output signal. For example, the intermediate layer output may include a set of filtering parameters that form a mask to be applied to the audio input signal to filter it. For example, the intermediate layer output may include an estimated spectrogram that indicates how much each frequency component should be attenuated or amplified each time.

[0043] Based on the intermediate layer output meeting the first criterion, the audio device is configured to determine an audio output signal based on the intermediate layer output. The audio device may be configured to apply the intermediate layer output in the processing of the audio input signal to provide the audio output signal. For example, the audio device may be configured to apply a mask from the intermediate layer output to the audio input signal. When it is determined that the intermediate layer output meets the first criterion, the audio device is configured to perform an early exit from the processing of the first neural network and exit the first neural network at the first intermediate layer that provides the intermediate layer output. In other words, when the intermediate layer output is determined / evaluated to meet specific performance, quality, and / or efficiency requirements for the audio input signal processing, the audio device is configured to stop or interrupt the processing using the first neural network and exit the first neural network at the first intermediate layer that provides the intermediate layer output.

[0044] When it is determined that the intermediate layer output does not meet the first criterion, the audio device is configured to avoid performing an early exit from the processing of the first neural network and avoid exiting the first neural network at the first intermediate layer that provides the intermediate layer output. In other words, when the intermediate layer output is determined / evaluated not to meet specific performance, quality, and / or efficiency requirements for the audio input signal processing, the audio device is configured to continue or proceed with the processing using the first neural network, e.g., move to the next first intermediate layer or the output layer.

[0045] It can be understood that when the intermediate layer output is below the first threshold, the first criterion is not met.

[0046] In one or more examples or embodiments, the first exit module is configured to output the result of determining whether the intermediate layer output meets the first criterion. In one or more examples or embodiments, the audio enhancement module is configured to process the audio input signal based on the result of determining whether the intermediate layer output meets the first criterion.

[0047] Thus, the present disclosure can reduce the scope of the neural network used to perform a specific task (e.g., reduce the portion of the first neural network used to perform a specific task). For example, the present disclosure allows performing a specific task with the first neural network by using only a subset of the layers of the first neural network (e.g., not having to execute or complete all layers of the first neural network, i.e., by allowing early exit). In other words, based on the input data, certain parts of the first neural network can be activated or deactivated, for example, by the first exit module and / or the second exit module. For example, based on the determination of whether the first output layer output meets the first criterion, certain parts of the first neural network can be activated or deactivated, for example, by the first exit module and / or the second exit module. For example, in a speech enhancement task, the first neural network can use only a subset of its layers for speech that is quiet or has no background noise and activate additional layers for speech in a noisy environment. In this way, the first neural network can process the input data using fewer resources, resulting in a lower computational cost.

[0048] In one or more examples or embodiments, the audio device includes a second exit module configured to obtain one or more features of an audio input signal that include a first feature. The second exit module may be considered a module of the audio device, e.g., a module of one or more processors of the audio device, that may be configured to preprocess the audio input signal before the audio input signal is processed by the audio device, e.g., before the audio input signal is processed by an audio enhancement module. Obtaining one or more features of the audio input signal may include extracting and / or determining one or more features from the audio input signal, e.g., one or more audio features. The second exit module may be configured to predict which layer (e.g., which first model layer of a first neural network) will be the best layer to exit when processing the audio input signal. The prediction of the layer (e.g., the first prediction layer, the second prediction layer, and / or the third prediction layer disclosed herein) may be based on one or more features (e.g., the quality of the audio input signal), a prediction of the output quality (e.g., based on the output quality of the first prediction layer), the performance of the output of the layer (e.g., based on the performance of the output of the first prediction layer), user preferences, and / or one or more audio device parameters (e.g., based on the capabilities of the audio device). The second exit module may be considered a preprocessing module. The terms second exit module and preprocessing module may be used interchangeably.

[0049] In one or more examples or embodiments, the audio device (e.g., the second exit module) is configured to process the audio input signal to provide one or more audio parameters indicative of one or more features or characteristics of the audio input signal. One or more features of the audio input signal (e.g., the first feature) may be considered audio parameters. One or more features may be considered quality parameters of the audio input signal (e.g., indicative of the audio quality of the audio data). One or more features may, for example, indicate one or more features of the audio input signal, e.g., one or more of bit rate, sample rate, dynamic range, frequency response, distortion, noise level, stereo imaging, compression artifacts, interference, abnormal noise, and voice artifacts. One or more features may indicate features including one or more of signal-to-noise ratio, confidence probability map, quality representation, and mean opinion score. The confidence probability map (time-frequency T-F map) may indicate the confidence of the denoised signal, e.g., the reliability of the gain time-frequency T-F map. The mean opinion score may be considered a predicted mean opinion score, e.g., a predicted mean opinion score quality prediction. For example, the mean opinion score may be determined based on an intrusive method (e.g., by comparing the audio input signal with a reference signal, e.g., a reference audio signal). Alternatively or additionally, the mean opinion score may be determined based on a non-intrusive method, e.g., by performing a blind prediction, e.g., using a pre-trained neural network dedicated to MOS score and / or alternative score estimation.

[0050] In one or more example audio devices, one or more features include direct reverberation ratio (DRR), coherent diffusion ratio (CDR), spatial noise coherence, room impulse response, noise / voice / interference level / direction, and an audio copy of the audio input signal. The direct reverberation ratio (DRR), coherent diffusion ratio (CDR), spatial noise coherence, room impulse response, and noise / voice / interference level / direction may be associated with the room and / or location where the transmitter and / or the audio device is located.

[0051] In one or more examples or embodiments, the second exit module is configured to predict, based on the first feature, from which first prediction layer of the first model layer to exit when processing the audio input signal. In other words, the second exit module may be configured to predict, based on one or more features, from which first prediction layer of the first model layer to exit. It can be understood that the first prediction layer may be selected from one or more first intermediate layers or output layers. For example, the second exit module may be configured to predict from which intermediate layer of one or more first intermediate layers to exit. The second exit module may be configured to predict the first prediction layer before the audio enhancement module has started processing the audio input signal.

[0052] In one or more examples or embodiments, the second exit module is configured to output to the audio enhancement module the prediction result of from which first prediction layer of the first model layer to exit when processing the audio input signal. In one or more examples or embodiments, the audio enhancement module is configured to process the audio input signal based on the prediction result of from which first prediction layer of the first model layer is going to exit when processing the audio input signal. The result may, for example, indicate and / or include the predicted layer to exit, such as the first prediction layer, the second prediction layer, and / or the third prediction layer, and / or indicate and / or include the predicted layer output, such as the first prediction layer output, the second prediction layer output, and / or the third prediction layer output. It can be understood that the second exit module and the first exit module may have a synergistic effect. The first exit module may, for example, use the predicted layer predicted by the second exit module and avoid processing one or more first intermediate layers before the predicted layer. This would be advantageous for reducing the resources used to process the audio input signal.

[0053] In one or more examples or embodiments, the second exit module includes a third neural network. In one or more examples or embodiments, predicting from which first prediction layer of the first model layer to exit when processing the audio input signal based on the first feature includes predicting from which first prediction layer of the first model layer to exit when processing the audio input signal using the third neural network.

[0054] The second exit module can be configured to predict, based on a first feature, from which first prediction layer of a first model layer when processing an audio input signal by using an ML-based method (e.g., by using a third neural network). The third neural network can include a third model layer, and the third model layer includes a third input layer, a plurality of third intermediate layers, and a third output layer.

[0055] The plurality of third intermediate layers can be considered hidden layers (e.g., hidden features). The plurality of third intermediate layers can include a third primary intermediate layer, a third secondary intermediate layer, a third tertiary intermediate layer, etc. The third neural network can be configured to operate according to a model (e.g., a machine learning model, e.g., a second machine learning model). The model mentioned herein (e.g., the second machine learning model) can be regarded as a model and / or scheme and / or mechanism and / or method configured to predict, based on a first feature, from which first prediction layer of a first model layer when processing an audio input signal. In one or more examples or embodiments, the third neural network is a deep neural network DNN.

[0056] In one or more examples or embodiments, the first exit module includes a second neural network. In one or more examples or embodiments, determining whether an intermediate layer output meets a first criterion includes using the second neural network to determine whether the intermediate layer output meets the first criterion.

[0057] The first exit module (e.g., the second neural network) can be configured to determine whether an intermediate layer output meets a first criterion. The first exit module can be configured to determine whether an intermediate layer output meets a first criterion by using an ML-based method and / or a signal processing-based method, e.g., by using the second neural network. The second neural network can include a second model layer, and the second model layer includes a second input layer, a plurality of second intermediate layers, and a second output layer.

[0058] The plurality of second intermediate layers can be considered hidden layers (e.g., hidden features). The plurality of second intermediate layers can include a second primary intermediate layer, a second secondary intermediate layer, a second tertiary intermediate layer, etc. The second neural network can be configured to operate according to a model (e.g., a machine learning model, e.g., a second machine learning model). The model mentioned herein (e.g., the second machine learning model) can be regarded as a model and / or scheme and / or mechanism and / or method configured to determine whether an intermediate layer output meets a first criterion. In one or more examples or embodiments, the second neural network is a deep neural network DNN.

[0059] In one or more examples or embodiments, the first prediction layer is an intermediate layer among one or more first intermediate layers. In one or more examples or embodiments, the first prediction layer is different from the one or more first intermediate layers. For example, the first prediction layer is the first output layer. In one or more examples or embodiments, the first prediction layer may be the first output layer.

[0060] In one or more examples or embodiments, the first prediction layer is configured to provide a first prediction layer output. The first prediction layer may provide an exit likelihood as mentioned herein. The first prediction layer output may be regarded as the result of the exit likelihood provided by the first prediction layer. In one or more examples or embodiments, the first prediction layer output may include a mask that is applied in the processing of the audio input signal to provide an audio output signal. For example, the first prediction layer output may include a mask for performing a specific task of audio processing, such as removing noise from the audio input signal to provide a denoised audio output signal. For example, the first prediction layer output may include a set of filtering parameters that constitute a mask to be applied to the audio input signal to filter it.

[0061] In one or more examples or embodiments, the audio device is configured to determine the audio output signal based on the first prediction layer output. In other words, the audio device may not need to determine whether the first prediction layer output meets a first criterion to evaluate whether the audio output signal should be determined based on the first prediction layer. The prediction of the first prediction layer can thus further reduce the required processing because the audio enhancement module can continue to process the audio input signal until the first prediction layer without having to evaluate the performance, quality, and / or efficiency of the intermediate layer outputs. It can be understood that in some embodiments, the prediction aspect of the first prediction layer and the aspect of determining whether the intermediate layer output meets the first criterion may be independent aspects or may be combined aspects. In one or more examples or embodiments, the first prediction layer may be an accurate layer for when to exit. Or, the first prediction layer may be the layer at which it starts to determine whether the layer meets the first criterion. For example, the second exit module may predict that it may not make sense to exit the intermediate layer before the first prediction layer. In one or more examples or embodiments, the audio enhancement module is configured to determine the audio input signal based on the first prediction layer output, for example, based on the result from the second exit module.

[0062] In one or more examples or embodiments, the first exit module and / or the second exit module may be trained during operation (e.g., while running). For example, the second exit module may be configured to occasionally perform a verification (e.g., a sanity check) to verify that the predicted first prediction layer is the best prediction. It will be appreciated that the first exit module may provide verification of the prediction of the second exit module by verifying the intermediate layer output. The second exit module may, for example, predict an exit at an earlier or later layer (e.g., stage) to verify whether the earlier or later layer is better than the first prediction layer. The first exit module may, for example, exit the first neural network at a later intermediate layer to evaluate the intermediate layer output of the later intermediate layer. The first exit module may thereby determine whether the intermediate layer output of the later intermediate layer provides a significant improvement in terms of the processing efficiency, quality, and / or performance of the audio input signal compared to the intermediate layer output that meets the first criterion. In other words, the first exit module may evaluate whether an exit layer that is later than the intermediate layer that provides the intermediate layer output that meets the first criterion can provide results that are significantly improved compared to the results of the previous intermediate layer output. In this case, best may be understood as having a better quality VS performance / required resources ratio. For example, if processing by exiting at an earlier layer saves resources, but the output of that layer has nearly the same quality, the second exit module may evaluate that earlier layer as better. When being trained during operation, the feedback to the first exit module and / or the second exit module may be provided as an increase or decrease in a threshold and / or a weighting (e.g., a correction) of the result. For example, for the first exit module, training during operation may allow for an evaluation of whether there may be further significant improvements in the processing of the first neural network (e.g., by adjusting the first criterion, e.g., the threshold, and potentially exiting at a different layer), e.g., improvements in performance, efficiency, and / or quality. For example, for the second exit module, training during operation may allow for an evaluation of an exit layer different from the predicted exit layer, e.g., whether adjusting the predicted weights will also provide suitable results.

[0063] In one or more examples or embodiments, the first exit module is configured to determine, based on an audio input signal and an intermediate layer output, an intermediate layer of a first model layer where the processing performance, quality, and / or efficiency of the audio input signal converges. In other words, the first exit module can be configured to determine at which intermediate layer of the first neural network a point of diminishing return is reached. The point of diminishing return and / or the convergence point can be considered as the point at which each additional input (e.g., additional processing of a further intermediate layer) gives a slower improvement in the output. In other words, the point of diminishing return and / or the convergence point can be the point at which further processing of the audio input signal will give a slower output improvement than the previous layer. Determining when the processing performance, quality, and / or efficiency of the audio input signal converges can include determining the improvement in the processing of the audio input signal using the output of a layer of the first neural network and comparing that output with one or more previous outputs. If it is determined that the improvement in the processing of the audio input signal using the output of the layer is low, then that layer can be predicted as the intermediate layer where the processing performance, quality, and / or efficiency of the audio input signal converges. For example, the first exit module can determine that the output of an intermediate layer provides an improvement compared to the previous output of a previous intermediate layer that is less than a certain improvement threshold. When this occurs, the first exit module can determine the intermediate layer where the processing performance, quality, and / or efficiency of the audio input signal converges (e.g., where diminishing returns occur). The determination of the intermediate layer where the processing performance, quality, and / or efficiency of the audio input signal converges can be used as feedback to adjust a first criterion (e.g., a threshold) so as to further improve the processing performance, quality, and / or efficiency of the audio input signal by exiting at the optimal intermediate layer.

[0064] In one or more examples or embodiments, the second exit module is configured to obtain one or more audio device parameters including a first audio device parameter. Obtaining one or more audio device parameters can include determining one or more audio device parameters based on one or more. Audio device parameters can be considered as parameters indicating characteristics of an audio device. Audio device parameters can be considered as parameters indicating one or more of the following: the capabilities of the audio device, the state of the audio device, and the type of the audio device. The capabilities of the audio device can be, for example, the processing capabilities of the audio device. The state of the audio device can be, for example, the battery state of the audio device and / or the power state of the audio device. The type of the audio device can be, for example, the model of the audio device, e.g., a headphone model or a speakerphone model. The second exit module can thus consider the audio device itself and its characteristics when predicting the exit layer.

[0065] In one or more examples or embodiments, the first audio device parameter is a power parameter, a battery parameter, and / or a processing capability parameter. For example, the power parameter of the audio device, the battery parameter of the audio device, and / or the processing capability parameter of the audio device. In other words, one or more audio device parameters may include one or more of the following items: a power parameter, a battery parameter, and a processing capability parameter. The power parameter may be regarded as the processing power parameter and / or the electrical power of the audio device. The battery parameter may be regarded as the battery status of the audio device, the available battery power, and / or the battery size of the audio device. The processing capability parameter may be regarded as the processing capability of the audio device. It can be understood that the power parameter, the battery parameter, and the processing capability parameter may be interrelated. For example, the processing capability parameter may depend on the power parameter and / or the battery parameter.

[0066] In one or more examples or embodiments, the second exit module is configured to predict, based on the first audio device parameter, from which second prediction layer of the first model layer to exit when processing an audio input signal.

[0067] The second exit module may be configured to predict which layer (e.g., which first model layer of the first neural network) will be the best layer to exit when processing an audio input signal. The prediction of the prediction layer (e.g., the first prediction layer, the second prediction layer, and / or the third prediction layer disclosed herein) may be based on one or more features (e.g., based on the quality of the audio input signal), the prediction of the output quality (e.g., based on the output quality of the first prediction layer), the performance of the output of the layer (e.g., based on the performance of the output of the first prediction layer), user preferences, and / or one or more audio device parameters (e.g., based on the capabilities of the audio device).

[0068] In other words, the second exit module can be configured to predict which second prediction layer to exit from the first model layer based on one or more audio device parameters. It can be understood that the second prediction layer can be selected from one or more first intermediate layers or output layers. For example, the second exit module can be configured to predict which intermediate layer to exit from among one or more first intermediate layers. The second exit module can be configured to predict the second prediction layer before the audio enhancement module has started processing the audio input signal. The second prediction layer can be an intermediate layer among one or more first intermediate layers. In one or more examples or embodiments, the second prediction layer is different from one or more first intermediate layers. For example, the second prediction layer is the first output layer. In one or more examples or embodiments, the second prediction layer is the same layer as the first prediction layer and / or the third prediction layer disclosed herein. In one or more examples or embodiments, the second prediction layer is different from the first prediction layer. It can be understood that when the first prediction layer, the second prediction layer, and / or the third prediction layer are different, the audio device can be configured to evaluate which one of the first prediction layer and the second prediction layer is optimal in terms of one or more features (e.g., based on the quality of the audio input signal), prediction of output quality (e.g., based on the output quality of the first prediction layer), output performance of the layer (e.g., based on the output performance of the first prediction layer), and / or one or more audio device parameters (e.g., based on the capabilities of the audio device).

[0069] It can be understood that in some embodiments, the aspect of predicting the first prediction layer, the second prediction layer, the third prediction layer and / or the aspect of determining whether the intermediate layer output meets the first criterion can be an independent aspect or can be a combined aspect. In one or more examples or embodiments, the first prediction layer can be the accurate layer of when to exit. Or, the first prediction layer can be the layer that starts to determine whether the layer meets the first criterion. For example, it may not make sense for the second exit module to predict the exit of an intermediate layer before the first prediction layer.

[0070] In one or more examples or embodiments, the second exit module is configured to predict which second prediction layer of the first neural network to exit during the processing of the audio input signal based on power parameters, battery parameters, and / or processing capability parameters.

[0071] In one or more examples or embodiments, the second prediction layer is configured to provide a second prediction layer output.

[0072] The second prediction layer can provide the exit probability mentioned in this article. The second prediction layer output can be regarded as the result of the exit probability provided by the second prediction layer. In one or more examples or embodiments, the second prediction layer output can include a mask, which is applied in the processing of the audio input signal to provide an audio output signal. For example, the second prediction layer output can include a mask for performing a specific task of audio processing, such as a mask for removing noise from the audio input signal to provide a denoised audio output signal. For example, the second prediction layer output can include a set of filtering parameters, and this set of filtering parameters constitutes a mask to be applied to the audio input signal to filter it.

[0073] In one or more examples or embodiments, the audio device is configured to determine the audio output signal based on the second prediction layer output. The description regarding determining the audio output signal based on the first prediction layer output can also be applied to the description of determining the audio output signal based on the second prediction layer output. In one or more examples or embodiments, the audio enhancement module is configured to determine the audio input signal based on the second prediction layer output, for example, based on the result from the second exit module.

[0074] In one or more examples or embodiments, the second exit module is configured to predict, based on the audio input signal, which third prediction layer of the first model layer the processing performance, quality, and / or efficiency of the audio input signal converges to. In other words, the second exit module can be configured to predict at which third prediction layer of the first neural network the point of diminishing returns is reached. The point of diminishing returns and / or the convergence point can be regarded as the point at which each additional input gives a slower improvement in the output. In other words, the point of diminishing returns and / or the convergence point can be the point at which further processing of the audio input signal will give a slower output improvement than the previous layer. Predicting when the processing performance, quality, and / or efficiency of the audio input signal converges can include predicting the improvement potential of processing the audio input signal using the output of the layers of the first neural network. If the predicted improvement potential of processing the audio input signal using the output of a layer is low, then that layer can be predicted as the third prediction layer. For example, the second exit module can predict that even if the audio enhancement module continues to process all the layers of the first neural network, the first neural network cannot achieve more than 90% improvement. Then, the second exit module can predict that the layer achieving an improvement between 80 - 90% is the third prediction layer at which the processing performance, quality, and / or efficiency of the audio input signal converges. In one or more examples or embodiments, the third prediction layer is configured to provide a third prediction layer output.

[0075] In one or more examples or embodiments, the audio device is configured to determine an audio output signal based on a third prediction layer. In other words, determining the audio output signal based on the third prediction layer may include determining the audio output signal based on the prediction of the layer where the processing converges. In other words, the audio output signal may not be determined directly based on the third prediction layer (e.g., the third prediction layer output), but the third prediction layer may be used to determine from which layer (or layer output) the audio output signal is determined. In one or more examples or embodiments, the audio enhancement module is configured to determine an audio input signal based on the third prediction layer, e.g., based on the result from the second exit module.

[0076] In one or more examples or embodiments, the audio device is configured to determine the audio output signal based on the output of the layer before the third prediction layer. It can be understood that if layer M is the third prediction layer, the audio device (e.g., the audio enhancement module) may be configured to exit the first neural network at, for example, layer M-1, M-2, or M-5. In one or more examples or embodiments, the audio enhancement module is configured to determine the audio input signal based on the output of the layer before the third prediction layer, e.g., based on the result from the second exit module.

[0077] In one or more examples or embodiments, determining the audio output signal based on the third prediction layer (e.g., based on the prediction of the layer where the processing converges) includes determining from which layer of the first neural network to exit when processing the audio input signal based on the third prediction layer. In some embodiments, the prediction layer to be exited may be considered different from the third prediction layer, as also shown in the above examples.

[0078] In one or more examples or embodiments, the third prediction layer is configured to provide a third prediction layer output. In one or more examples or embodiments, the audio device is configured to determine the audio output signal based on the third prediction layer output. The third prediction layer output may be regarded as the result of the exit possibility provided by the third prediction layer. In one or more examples or embodiments, the third prediction layer output may include a mask that is applied in the processing of the audio input signal to provide the audio output signal. For example, the third prediction layer output may include a mask for performing a specific task of audio processing, e.g., a mask for removing noise from the audio input signal to provide a denoised audio output signal. For example, the third prediction layer output may include a set of filtering parameters that form a mask to be applied to the audio input signal to filter it.

[0079] The description regarding determining the audio output signal based on the first prediction layer output can also be applied to the description of determining the audio output signal based on the third prediction layer output.

[0080] In one or more examples or embodiments, an audio device (e.g., a second exit module) may be configured to use a first neural network to predict the minimum layer at which an audio enhancement module may exit processing in order to provide an acceptable audio output signal. In other words, the audio device (e.g., the second exit module) may be configured to use the first neural network to predict the minimum layer at which the audio enhancement module may exit processing in order to have at least to some extent performed a particular audio processing task.

[0081] As described herein, based on what may be considered a "function of..." and / or "used as an input to...", e.g., the audio output signal may be a function of an intermediate layer output, a first prediction layer output, a second prediction layer output, and / or a third prediction layer output. The intermediate layer output, the first prediction layer output, the second prediction layer output, and / or the third prediction layer output may be used as inputs for determining the audio output signal.

[0082] In one or more examples or embodiments, an audio device (e.g., a second exit module) is configured to determine an uncertainty parameter of a prediction layer based on an audio input signal, e.g., an uncertainty parameter of a first prediction layer, a second prediction layer, and / or a third prediction layer.

[0083] For example, the audio device may be configured to determine the uncertainty parameter based on one or more features of audio data (e.g., the audio input signal), as disclosed herein. In one or more examples or embodiments, the audio device may be configured to determine the uncertainty parameter based on one or more features, user preferences, and / or one or more audio parameters, as disclosed herein. In other words, the audio device is configured to determine an uncertainty parameter that indicates the uncertainty of estimating the prediction layer. The uncertainty parameter may indicate an estimate of the quality of the processing result of the audio input signal based on the output of a given layer, e.g., an estimate of the quality of the audio output signal. For example, the uncertainty parameter may indicate an estimate of the quality of the processing result of the audio enhancement module on the audio input signal, as disclosed herein. Based on the output of a given prediction layer, the uncertainty parameter may be considered a prediction of the processing quality of the audio device on the audio input signal. The uncertainty parameter may be considered an estimate of the confidence in the processing quality of the audio device when using the audio enhancement module on a given input (e.g., a given audio input signal).

[0084] In one or more example audio devices, one or more processors include a digital signal processor. In one or more examples or embodiments, the audio enhancement module disclosed herein forms part of the digital signal processor.

[0085] In one or more example audio devices, the models mentioned herein may be stored on a non-transitory storage medium (e.g., on the memory of the audio device). The model may be stored on the non-transitory storage medium of the audio device configured to execute the model. In one or more example audio devices, the model may include model data and / or computer-readable instructions (e.g., based on an audio input signal and / or layer outputs, such as intermediate layer outputs, first prediction layer, second prediction layer, and / or third prediction layer, as disclosed herein). The audio device may use the model data and / or computer-readable instructions. The audio device may use the model (e.g., model data and / or computer-readable instructions) to process audio data, e.g., an audio input signal.

[0086] For example, a digital signal processor may include a noise reducer and / or an echo controller, e.g., a deep noise suppression DNS noise reducer, which is configured to operate according to a neural network NN (e.g., DNN).

[0087] In one or more example audio devices, one or more processors are configured to output an audio output signal via an interface, e.g., an audio output. In other words, the audio device may be configured to output the audio output signal to a remote end, e.g., via a wired and / or wireless interface, and / or output the audio output signal at the proximal end of the audio device itself via one or more speakers (e.g., receivers).

[0088] A method of operating an audio device (e.g., an audio device configured to act as a receiver device) is disclosed. The method includes obtaining audio data, e.g., via an interface and / or using one or more processors of the audio device. The method includes processing the audio data, e.g., using one or more processors of the audio device, to provide an audio output. The method includes determining an uncertainty parameter based on the audio data, e.g., using one or more processors of the audio device. The method includes controlling the processing of the audio data, e.g., using one or more processors of the audio device, to provide an audio output based on the uncertainty parameter.

[0089] In one or more examples or embodiments, the method includes, for example, using one or more processors of an audio device to process audio data to provide one or more audio parameters indicative of one or more characteristics of the audio data. In one or more examples or embodiments, the method includes, for example, using one or more processors of an audio device to map the one or more audio parameters to a first latent space of a first neural network to provide mapping parameters indicative of whether the one or more audio parameters belong to a training manifold of the first latent space. In one or more examples or embodiments, the method includes, for example, using one or more processors of an audio device to determine an uncertainty parameter based on the mapping parameters, the uncertainty parameter indicative of uncertainty in processing quality.

[0090] It should be understood that the description of the features of the audio device also applies to the corresponding features in the method of operating the audio device disclosed herein, and vice versa.

[0091] A computer-implemented method for training a first neural network disclosed herein is provided. The method includes obtaining an audio data set including one or more audio signals. The method includes training the first neural network based on the audio data set to perform an audio processing task at at least one first intermediate layer of the first neural network to provide a first intermediate layer with an exit probability to provide an intermediate layer output. In other words, the method may include training the audio enhancement module disclosed herein to perform an audio processing task at each predefined latent exit stage (e.g., exit probability).

[0092] In one or more examples or embodiments, the method includes training a second exit module disclosed herein to predict one or more of the first prediction layer, the second prediction layer, and the third prediction layer disclosed herein. In other words, the method may include training the second exit module to pick the best exit for the first neural network (e.g., for the audio enhancement module). It can be understood that the cost function for training the second exit module may be a combination of audio quality VS computation (e.g., computational cost).

[0093] Figure 1 An example audio device according to the present disclosure is schematically illustrated, e.g., audio device 10. Audio device 10 may be regarded as an audio communication device. Audio device 10 may be regarded as a communication device for performing calls (e.g., audio and / or video calls). Audio device 10 may be regarded as an audio device for implementing efficient neural network processing, e.g., for implementing efficient dynamic neural network processing.

[0094] The audio device 10 can be configured to act as a receiver device and / or a transmitter device. In other words, the audio device 10 can be configured to receive an input signal from another audio device that is configured to act as a transmitter device and / or is configured to transmit an output signal to other audio devices. The audio device 10 includes an interface and a memory (not shown). Optionally, the audio device 10 includes an audio speaker 10D and one or more microphones, for example, a first microphone 10E1 and a second microphone 10E2. Optionally, the audio device 10 includes one or more transceivers, for example, a first wireless transceiver 10F. The audio device 10 can be regarded as an audio device configured to obtain, output, and process audio signals. The audio device 10 can be regarded as a conference audio device. For example, the conference audio device is configured to be used by one party (e.g., one or more users at the proximal end) to communicate with one or more other parties (e.g., one or more users at the distal end). The audio device 10 can be regarded as an intelligent audio device. The audio device 10 can be used for communication, conferencing, and / or meetings between two or more parties that are far apart from each other. The audio device 10 can be used by one or more users in the vicinity (also referred to as the proximal end) where the audio device 10 is located. In this example, the receiver side can be regarded as the proximal end, and the transmitter side can be regarded as the distal end.

[0095] The audio device 10 includes one or more processors 10C. The one or more processors 10C can be configured to obtain audio data, e.g., an audio input signal. The audio device 10 can be configured to obtain, e.g., using the one or more processors 10C and / or via the input interface 10B, a first microphone input signal 50 from the first microphone 10E1, a second microphone input signal 72 from the second microphone 10E2, and / or a transceiver input signal 74 from the first transceiver. The first microphone input signal 50, the second microphone input signal 72, and / or the transceiver input signal 74 can be regarded as input signals. In one or more example audio devices, the transceiver input signal 74 can be obtained via the first transceiver interface 16 (e.g., as a transceiver interface input signal) and forwarded as a transceiver interface output 76 to the input interface 10B. In one or more example audio devices, the audio device 10 is configured to obtain an input signal (e.g., the transceiver input signal 74) from a remote end (e.g., a remote party or user). It can be understood that the input signal includes audio. In one or more embodiments or examples, the input signal (e.g., the transceiver input signal 74) has been signal processed, e.g., encoded, compressed, and / or enhanced. The transceiver input signal 74 can indicate an audio signal generated by a user at the remote end. In other words, the transceiver input signal 74 can indicate speech, e.g., speech from a remote transmitter device. When obtained from the first microphone 10E1 and / or the second microphone 10E2, the input signal can indicate audio from the proximal end, e.g., speech. The audio data can be based on the input signal, e.g., based on the first microphone input signal 50, the second microphone input signal 72, and / or the transceiver input signal 74. The input interface 10B can be configured to provide an input interface output 52 based on the first microphone input signal 50, the second microphone input signal 72, and / or the transceiver input signal 74. The audio data (e.g., the audio input signal 53) can be based on the input interface output 52.

[0096] The audio device 10 is configured to process audio data (e.g., audio input signal 53) using, for example, one or more processors 10C to provide an audio output. In other words, the audio device 10 can be configured to process the first microphone input signal 50, the second microphone input signal 72, the transceiver input signal 74, and / or the input interface output 52 to provide an audio output. The audio device 10 can include an output interface 10A that is configured to output an audio output, e.g., an audio output signal 58. For example, the audio device 10 can be configured to output the audio output signal to an audio speaker 10D via the output interface 10A as an audio speaker input 78. The audio speaker 10D can be configured to output an audio output based on the audio speaker input 78, e.g., output an audio output proximally. For example, the audio device 10 can be configured to output the audio output signal to a second wireless transceiver 10G via the output interface 10A as a second transceiver input 80. The second wireless transceiver 10G can be configured to output an audio output based on the second transceiver input 80, e.g., output an audio output signal distally. In one or more example audio devices, the second transceiver input 80 (e.g., as a second transceiver interface input signal) can be output via a second transceiver interface 18 and forwarded as a second transceiver output 82 to a second wireless transceiver interface 10G. It will be appreciated that the audio output can be based on and / or include the audio speaker input 78 and / or the second transceiver input 80.

[0097] The audio device 10 (e.g., one or more processors 10C) includes an audio enhancement module 13 that includes a first neural network having a first model layer that includes a first input layer, a plurality of first intermediate layers, and a first output layer.

[0098] In one or more example audio devices, one or more processors 10C include a digital signal processor. In one or more examples or embodiments, the audio enhancement module 13 disclosed herein includes a digital signal processor.

[0099] In one or more examples or embodiments, the audio device 10 (e.g., one or more processors 10C) includes a first dropout module 14.

[0100] The audio enhancement module 13 is configured to process an audio input signal 53 using a first neural network to provide an audio output signal 58. In one or more examples or embodiments, at least one first intermediate layer has an exit possibility to provide an intermediate layer output. The audio enhancement module 13 may be configured to output 55 the intermediate layer output 56 to a first exit module 14. In one or more examples or embodiments, the first exit module 14 is configured to determine whether the intermediate layer output 56 meets a first criterion, where the first criterion indicates the performance, quality, and / or efficiency of the intermediate layer output 56. In one or more examples or embodiments, according to the intermediate layer output 56 meeting the first criterion, the audio device 10 is configured to determine, for example using the audio enhancement module 13, the audio output signal 58 based on the intermediate layer output 56.

[0101] In one or more examples or embodiments, the first exit module 14 is configured to output 57 the result of determining whether the intermediate layer output 56 meets the first criterion to the audio enhancement module 13. In one or more examples or embodiments, the audio enhancement module 13 is configured to process the audio input signal 53 based on the result 57 of determining whether the intermediate layer output 56 meets the first criterion. In one or more examples or embodiments, the first exit module 14 may be configured to output 57 the intermediate layer output, for example, to a digital signal processor (not shown) of the audio device 10. In one or more examples or embodiments, the first exit module 14 may be included by the audio enhancement module 13 or form a part of the audio enhancement module 13.

[0102] In one or more examples or embodiments, the audio device 10 includes a second exit module 12 that is configured to obtain one or more features of the audio input signal 53 that include a first feature. In one or more examples or embodiments, the second exit module 12 is configured to obtain the audio input signal 53, for example, via an input interface 10B.

[0103] In one or more examples or embodiments, the second exit module 12 is configured to predict, based on a first feature, from which first prediction layer of a first model layer to exit when processing an audio input signal 53 (e.g., when the audio enhancement module 13 processes the audio input signal 53). In one or more examples or embodiments, the second exit module 12 is configured to output to the audio enhancement module 13 a prediction result 54 of from which first prediction layer of the first model layer to exit when processing the audio input signal 53. In one or more examples or embodiments, the audio enhancement module 13 is configured to process the audio input signal 53 based on the prediction result 54 of from which first prediction layer of the first model layer to exit when processing the audio input signal 53. The result may, for example, indicate and / or include the prediction layer to exit, e.g., the first prediction layer, the second prediction layer, and / or the third prediction layer, and / or indicate and / or include the prediction layer output, e.g., the first prediction layer output, the second prediction layer output, and / or the third prediction layer output.

[0104] In one or more examples or embodiments, the first prediction layer is an intermediate layer among one or more first intermediate layers.

[0105] In one or more examples or embodiments, the first prediction layer is configured to provide a first prediction layer output.

[0106] In one or more examples or embodiments, the audio device 10 is configured to determine an audio output signal 58 based on the first prediction layer output.

[0107] In one or more examples or embodiments, the first exit module 14 includes a second neural network. In one or more examples or embodiments, determining whether the intermediate layer output 56 meets a first criterion includes using the second neural network to determine whether the intermediate layer output 56 meets the first criterion. In one or more examples or embodiments, the audio enhancement module 13 is configured to output the intermediate layer output 56 to the second exit module 12.

[0108] In one or more examples or embodiments, the second exit module 12 is configured to obtain one or more audio device parameters including a first audio device parameter. The second exit module 12 may be configured to obtain the one or more audio device parameters from the audio device 10, e.g., from a memory (not shown) of the audio device 10. In one or more examples or embodiments, the first audio device parameter is a power parameter, a battery parameter, and / or a processing capability parameter, e.g., the power parameter of the audio device, the battery parameter of the audio device, and / or the processing capability parameter of the audio device. In one or more examples or embodiments, the second exit module 12 is configured to predict, based on the first audio device parameter, from which second prediction layer of the first model layer to exit when processing the audio input signal 53.

[0109] In one or more examples or embodiments, the second exit module 12 is configured to predict which second prediction layer of the first neural network to exit when processing the audio input signal 53 based on power parameters, battery parameters, and / or processing capability parameters. This prediction can be indicated and / or included in the result 54 of the audio enhancement module 13.

[0110] In one or more examples or embodiments, the second prediction layer is configured to provide a second prediction layer output. The second prediction layer output can be indicated and / or included in the result 54.

[0111] In one or more examples or embodiments, the audio device 10 is configured to determine the audio output signal 58 based on the second prediction layer output. In one or more examples or embodiments, the audio enhancement module 13 is configured to determine the audio input signal 58 based on the second prediction layer output (e.g., based on the result 54 from the second exit module 12).

[0112] In one or more examples or embodiments, the second exit module 12 is configured to predict, based on the audio input signal 53, which third prediction layer of the first model layer the processing performance, quality, and / or efficiency of the audio input signal 53 converges to when the processing of the audio input signal 53 by the audio enhancement module 13 converges.

[0113] In one or more examples or embodiments, the audio device 10 is configured to determine the audio output signal 58 based on the third prediction layer. In one or more examples or embodiments, the audio enhancement module 13 is configured to determine the audio input signal 58 based on the third prediction layer, e.g., based on the result 54 from the second exit module 12.

[0114] In one or more examples or embodiments, the audio device 10 is configured to determine the audio output signal 53 based on the output of the layer before the third prediction layer. In one or more examples or embodiments, the audio enhancement module 13 is configured to determine the audio input signal 58 based on the output of the layer before the third prediction layer, e.g., based on the result 54 from the second exit module 12.

[0115] In one or more examples or embodiments, determining the audio output signal 53 based on the third prediction layer (e.g., based on the prediction) includes determining which layer of the first neural network to exit when processing the audio input signal 53 based on the third prediction layer.

[0116] In one or more examples or embodiments, a third prediction layer is configured to provide a third prediction layer output. In one or more examples or embodiments, the audio device 10 is configured to determine an audio output signal based on the third prediction layer output. In one or more examples or embodiments, the audio enhancement module 13 is configured to determine an audio input signal 58 based on the third prediction layer output, e.g., based on the result 54 from the second exit module 12.

[0117] The audio device 10 may be configured to perform Figure 2A and Figure 2B any method disclosed in

[0118] The operation of the audio device 10 may be embodied in the form of executable logic routines (e.g., lines of code, software programs, etc.) stored on a non-transitory computer-readable medium (e.g., a memory) and executed by one or more processors 10C.

[0119] In addition, the operation of the audio device 10 may be considered as a method that the audio device 10 is configured to perform. In addition, although the described functions and operations may be implemented in software, such functions may also be performed via dedicated hardware or firmware or some combination of hardware, firmware, and / or software.

[0120] The memory of the audio device may be one or more of a buffer, flash memory, hard disk drive, removable media, volatile memory, non-volatile memory, random access memory (RAM), or other suitable devices. In a typical arrangement, the memory may include non-volatile memory for long-term data storage and volatile memory that serves as the system memory of the processor 10C. The memory may exchange data with the processor 10C via a data bus. There may also be control lines and address buses ( Figure 1 not shown in

[0121] between the memory and the processor 10C). The memory is considered a non-transitory computer-readable medium.

[0122] Figure 2A and Figure 2B show a flowchart of an example method (e.g., method 100).

[0123] Method 100 performed by an audio device is disclosed, e.g., a method of operating an audio device as disclosed herein. Method 100 can be used to implement efficient neural network processing. In one or more examples or embodiments, the audio device includes an audio enhancement module, which includes a first neural network having a first model layer, the first model layer including a first input layer, a plurality of first intermediate model layers, and a first output layer; and a first dropout module. Method 100 includes processing S106 an audio input signal using the first neural network to provide an audio output signal. At least one of the first intermediate layers has a dropout probability to provide an intermediate layer output.

[0124] In one or more examples or embodiments, method 100 includes determining S108 whether the intermediate layer output satisfies a first criterion using the first dropout module. In one or more examples or embodiments, the first criterion indicates the performance, quality, and / or efficiency of the intermediate layer output.

[0125] In one or more example methods, determining S108 whether the intermediate layer output satisfies the first criterion includes determining S108A whether the intermediate layer output satisfies the first criterion using a second neural network.

[0126] In one or more examples or embodiments, method 100 includes determining S110 the audio output signal based on the intermediate layer output according to the intermediate layer output satisfying the first criterion.

[0127] In one or more examples or embodiments, method 100 includes continuing S109 processing the next layer of the first neural network, e.g., the next intermediate layer, according to the intermediate layer output not satisfying the first criterion.

[0128] In one or more example methods, method 100 includes obtaining S102 one or more features of the audio input signal including a first feature using a second dropout module.

[0129] In one or more example methods, method 100 includes predicting S104 from which first prediction layer of the first model layer to dropout when processing S106 the audio input signal based on the first feature.

[0130] In one or more example methods, the first prediction layer is configured to provide a first prediction layer output. In one or more examples or embodiments, method 100 includes determining the first prediction layer output based on the first prediction layer.

[0131] In one or more example methods, method 100 includes determining S110A the audio output signal based on the first prediction layer output.

[0132] In one or more example methods, method 100 includes obtaining, using a second exit module, one or more audio device parameters including a first audio device parameter at S103.

[0133] In one or more example methods, method 100 includes predicting, using a second exit module and based on the first audio device parameter, at S105 which second prediction layer to exit from in the first model layer when processing an S106 audio input signal.

[0134] In one or more example methods, the second prediction layer is configured to provide a second prediction layer output.

[0135] In one or more example methods, method 100 includes determining an S110B audio output signal based on the second prediction layer output.

[0136] In one or more example methods, predicting which second prediction layer to exit from in the first model layer at S105 includes predicting at S105A which second prediction layer to exit from in the first neural network when processing the audio input signal based on a power parameter, a battery parameter, and / or a processing capability parameter.

[0137] In one or more example methods, method 100 includes predicting, using a second exit module and based on the audio input signal, at S107 at which third prediction layer the processing performance, quality, and / or efficiency of the audio input signal converges in the first model layer.

[0138] In one or more example methods, method 100 includes determining an S110C audio output signal based on the third prediction layer.

[0139] In one or more example methods, determining the S110C audio output signal includes determining at S110C1 which layer to exit from in the first model layer when processing the audio input signal based on the third prediction layer.

[0140] In one or more example methods, the third prediction layer is configured to provide a third prediction layer output.

[0141] In one or more example methods, method 100 includes determining an S110D audio output signal based on the third prediction layer output.

[0142] In one or more example methods, method 100 includes determining an S110E audio output signal based on the output of the layer before the third prediction layer.

[0143] In one or more example methods, the first exit module includes a second neural network.

[0144] Figure 3Schematically illustrated is an example audio device according to the present disclosure, e.g., audio device 10, in which the techniques disclosed herein are applied.

[0145] Optionally, audio device 10 includes one or more microphones 10E. Optionally, audio device 10 includes one or more transceivers, e.g., first wireless transceiver 10F.

[0146] Audio device 10 includes one or more processors 10C. One or more processors 10C may be configured to obtain audio data, e.g., audio input signal 53.

[0147] Audio device 10 is configured to process audio data, e.g., audio input signal 53, using one or more processors 10C, for example, to provide an audio output, e.g., audio output signal 58.

[0148] Audio device 10 (e.g., one or more processors 10C) includes an audio enhancement module 13, which includes a first neural network 20 having a first model layer, and the first model layer includes a first input layer (not shown), a plurality of first intermediate layers (20A, 20B), and a first output layer 20C. In one or more examples or embodiments, the first input layer may be intermediate layer 20A. The first neural network 20 (e.g., the first model layer) may include a first three - level intermediate layer, a first four - level intermediate layer, a first five - level intermediate layer, etc.

[0149] In one or more example audio devices, one or more processors 10C include a digital signal processor. In one or more examples or embodiments, the audio enhancement module 13 disclosed herein includes a digital signal processor.

[0150] In one or more examples or embodiments, audio device 10 (e.g., one or more processors 10C) includes a first dropout module 14.

[0151] Audio enhancement module 13 is configured to process audio input signal 53 using the first neural network 20 to provide an audio output signal 58. In one or more examples or embodiments, at least one first intermediate layer has a dropout probability to provide an intermediate layer output 56. In Figure 3In the example, the first primary intermediate layer 20A has an exit possibility to provide a first primary intermediate layer output 56. The first primary intermediate layer output 56 can be regarded as the intermediate layer output disclosed herein. The audio enhancement module 13 can be configured to output the intermediate layer output 56 (e.g., the first primary intermediate layer output 56) to the first exit module 14. In one or more examples or embodiments, the first exit module 14 is configured to determine whether the intermediate layer output 56 (e.g., the first primary intermediate layer output 56) meets a first criterion, where the first criterion indicates the performance, quality, and / or efficiency of the intermediate layer output 56 (e.g., the first primary intermediate layer output 56). In one or more examples or embodiments, according to the intermediate layer output 56 meeting the first criterion, the audio device 10 is configured to determine, for example using the audio enhancement module 13, an audio output signal 58 based on the intermediate layer output 56.

[0152] In one or more examples or embodiments, the first exit module 14 is configured to output a result 57 of determining whether the intermediate layer output 56 meets the first criterion to the audio enhancement module 13. In one or more examples or embodiments, the audio enhancement module 13 is configured to process the audio input signal 53 based on the result 57 of determining whether the intermediate layer output 56 meets the first criterion. For example, according to the intermediate layer output 56 meeting the first criterion, the first exit module 14 can be configured to confirm to the audio enhancement module 13 in the result 57 that the intermediate layer output 56 can be used to process the audio input signal 53. The result 57 can include, for example, instructions on how to process the audio input signal 53. For example, according to the intermediate layer output 56 not meeting the first criterion, the first exit module 14 can be configured to instruct the audio enhancement module 13 to proceed to the next layer of the first neural network 20, for example, to the first secondary intermediate layer 20B of the first neural network 20. The result 57 can include, for example, instructions to continue with the first secondary intermediate layer 20B of the first neural network 20. In one or more examples or embodiments, according to the intermediate layer output 56 not meeting the first criterion, the first exit module 14 can be configured to instruct the audio enhancement module 13 to advance 60 to the next layer of the first neural network 20, for example, to advance 62 to the output layer 20C of the first neural network 20. In one or more examples or embodiments, the first exit module 14 can be configured to output 57 the intermediate layer output, for example, output 57 the intermediate layer output to a digital signal processor (not shown) of the audio device 10.

[0153] In one or more examples or embodiments, the audio device 10 includes a second exit module 12, which is configured to obtain one or more features of the audio input signal 53 including a first feature.

[0154] In one or more examples or embodiments, the second exit module 12 is configured to predict, based on a first feature, from which first prediction layer of a first model layer to exit when processing an audio input signal 53, e.g., when the audio enhancement module 13 processes the audio input signal 53. In one or more examples or embodiments, the second exit module 12 is configured to output, to the audio enhancement module 13, a prediction result 54 of from which first prediction layer of a first model layer to exit when processing the audio input signal 53. In one or more examples or embodiments, the audio enhancement module 13 is configured to process the audio input signal 53 based on the prediction result 54 of from which first prediction layer of a first model layer to exit when processing the audio input signal 53. The result may, for example, indicate and / or include the prediction layer from which to exit, e.g., a first prediction layer, a second prediction layer, and / or a third prediction layer, and / or indicate and / or include a prediction layer output, e.g., a first prediction layer output, a second prediction layer output, and / or a third prediction layer output. It will be appreciated that the first prediction layer, the second prediction layer, and / or the third prediction layer may, for example, be a first primary intermediate layer 20A, a first secondary intermediate layer 20B, and / or a first output layer 20C.

[0155] In one or more examples or embodiments, the first prediction layer is an intermediate layer among one or more first intermediate layers, e.g., a first primary intermediate layer 20A and / or a first secondary intermediate layer 20B.

[0156] In one or more examples or embodiments, the first prediction layer is configured to provide a first prediction layer output.

[0157] In one or more examples or embodiments, the audio device 10 is configured to determine an audio output signal 58 based on the first prediction layer output. In one or more examples or embodiments, the first prediction layer output may be an intermediate layer output 56, e.g., a first primary intermediate layer output 56.

[0158] In one or more examples or embodiments, the first exit module 14 includes a second neural network. In one or more examples or embodiments, determining whether the intermediate layer output 56 meets a first criterion includes using the second neural network to determine whether the intermediate layer output 56 meets the first criterion. In one or more examples or embodiments, the audio enhancement module 13 is configured to output the intermediate layer output 56 to the second exit module 12.

[0159] In one or more examples or embodiments, the second exit module 12 is configured to obtain one or more audio device parameters including a first audio device parameter. The second exit module 12 may be configured to obtain one or more audio device parameters from the audio device 10, e.g., from a memory (not shown) of the audio device 10. In one or more examples or embodiments, the first audio device parameter is a power parameter, a battery parameter, and / or a processing capability parameter, e.g., the power parameter of the audio device, the battery parameter of the audio device, and / or the processing capability parameter of the audio device 10. In one or more examples or embodiments, the second exit module 12 is configured to predict from which second prediction layer of the first model layer to exit when processing the audio input signal 53 based on the first audio device parameter.

[0160] In one or more examples or embodiments, the second exit module 12 is configured to predict from which second prediction layer of the first neural network 20 to exit when processing the audio input signal 53 based on the power parameter, the battery parameter, and / or the processing capability parameter. The prediction may be indicated and / or included in the result 54 of the audio enhancement module 13.

[0161] In one or more examples or embodiments, the second prediction layer is configured to provide a second prediction layer output. The second prediction layer output may be indicated and / or included in the result 54.

[0162] In one or more examples or embodiments, the audio device 10 is configured to determine the audio output signal 58 based on the second prediction layer output. In one or more examples or embodiments, the audio enhancement module 13 is configured to determine the audio output signal 58 based on the second prediction layer output, e.g., based on the result 54 from the second exit module 12.

[0163] In one or more examples or embodiments, the second exit module 12 is configured to predict, e.g., when the processing of the audio input signal 53 by the audio enhancement module 13 converges, at which third prediction layer of the first model layer the processing performance, quality, and / or efficiency of the audio input signal 53 converges based on the audio input signal 53.

[0164] In one or more examples or embodiments, the audio device 10 is configured to determine the audio output signal 58 based on the third prediction layer. In one or more examples or embodiments, the audio enhancement module 13 is configured to determine the audio output signal 58 based on the third prediction layer, e.g., based on the result 54 from the second exit module 12.

[0165] In one or more examples or embodiments, the audio device 10 is configured to determine the audio output signal 58 based on the output of the layers before the third prediction layer. In one or more examples or embodiments, the audio enhancement module 13 is configured to determine the audio output signal 58 based on the output of the layers before the third prediction layer, e.g., based on the result 54 from the second exit module 12.

[0166] In one or more examples or embodiments, determining the audio output signal 58 based on the third prediction layer (e.g., based on a prediction) includes determining from which layer of the first neural network 20 to exit when processing the audio input signal 53 based on the third prediction layer.

[0167] In one or more examples or embodiments, the third prediction layer is configured to provide a third prediction layer output. In one or more examples or embodiments, the audio device 10 is configured to determine the audio output signal based on the third prediction layer output. In one or more examples or embodiments, the audio enhancement module 13 is configured to determine the audio input signal 58 based on the third prediction layer output, e.g., based on the result 54 from the second exit module 12.

[0168] Figure 4 An example audio device according to the present disclosure is schematically shown, e.g., the audio device 10, in which the techniques disclosed herein are applied. Figure 4 A first example of a computer-implemented method for training the first neural network disclosed herein is shown, e.g., the first example of the computer-implemented method disclosed herein. Figure 4 May show Figure 3 the training of the audio device 10, e.g., Figure 3 the training of the audio enhancement module 13 and the first neural network.

[0169] The audio device 10 (e.g., one or more processors 10C) includes an audio enhancement module 13, which includes a first neural network 20 having a first model layer. The first model layer includes a first input layer (not shown), a plurality of first intermediate layers (20A, 20B), and a first output layer 20C. In one or more examples or embodiments, the first input layer may be an intermediate layer 20A. The first neural network 20 (e.g., the first model layer) may include a first three-level intermediate layer, a first four-level intermediate layer, a first five-level intermediate layer, etc.

[0170] The audio enhancement module 13 is configured to use the first neural network 20 to process the audio input signal 53 to provide an audio output signal. In one or more examples or embodiments, at least one of the first intermediate layers has an exit possibility to provide an intermediate layer output. In Figure 4In the example, the first primary intermediate layer 20A provides a first primary intermediate layer output 56 (which can become an exit probability when the first neural network has been trained), the first secondary intermediate layer 20B provides a first secondary intermediate layer output 64 (which can become an exit probability when the first neural network has been trained), and the first output layer 20C provides a first output layer output 68 (which can become an exit probability when the first neural network has been trained). Arrows 60 and 62 illustrate the intermediate layers between the first secondary intermediate layer 20B and the first output layer 20C.

[0171] A method for training a first neural network includes obtaining an audio data set 53 including one or more audio signals. The one or more audio signals 53 can be obtained via one or more microphones 10E and / or via a first wireless transceiver 10F. In one or more examples or embodiments, the one or more audio signals can be obtained from a database and / or a memory of the audio device 10. The method includes training the first neural network based on the audio data set (e.g., based on the one or more audio signals 53) to perform an audio processing task at at least one first intermediate layer of the first neural network, the at least one first intermediate layer providing a first intermediate layer with an exit probability to provide an intermediate layer output. In Figure 4 the example, the method includes training the first neural network based on the audio data set (e.g., based on the one or more audio signals 53) to perform an audio processing task at each first intermediate layer of the first neural network, each of the first intermediate layers providing a first intermediate layer with an exit probability to provide an intermediate layer output. In other words, the method can include training the audio enhancement module disclosed herein to perform an audio processing task at each predefined potential exit stage (e.g., exit probability). Figure 4 The example of illustrates a first example of a computer-implemented method for training the first neural network disclosed herein (i.e., global training of the first neural network). Global training can be regarded as end-to-end training of the entire dynamic model. In other words, all layers and their parameters are updated simultaneously during the training process. This training method may be advantageous for tasks with relatively simple relationships between inputs and outputs. As in Figure 4 can be observed, the layer outputs of the first model layer of the first neural network are output together to determine a loss function 66. The loss function 66 is determined based on the output from the first model layer and based on a clean audio signal 70. Thus, the first neural network can be globally trained by repeating this process.

[0172] Figure 5 Schematically illustrates an example audio device according to the present disclosure, e.g., the audio device 10, in which the techniques disclosed herein are applied. Figure 5Shows a second example of a computer-implemented method for training the first neural network disclosed herein. For example, a second example of the computer-implemented method disclosed herein. Figure 5 May show Figure 3 training of the audio device 10, for example, Figure 3 training of the audio enhancement module 13 and the first neural network.

[0173] Figure 5 An example shows a first example of a computer-implemented method for training the first neural network disclosed herein (i.e., layer-by-layer training of the first neural network). Layer-by-layer training can be regarded as training a dynamic model one layer at a time. In other words, each layer is trained based on the activations of the previous layer, and the parameters of that layer are updated before moving to the next layer. This training method may be advantageous for tasks where the relationship between input and output is more complex, the computational resources available for training are limited, and / or new samples are expected to be significantly different from existing inputs. This method improves the generalization ability of the model. Layer-by-layer training can also be used to train dynamic models with different structures, where the number of neurons or connections in each layer can vary dynamically based on the input data. This helps to optimize the use of computational resources and improve the accuracy of the model.

[0174] The method for training the first neural network includes obtaining an audio data set 53 including one or more audio signals. The one or more audio signals 53 may be obtained via one or more microphones 10E and / or via the first wireless transceiver 10F. In one or more examples or embodiments, the one or more audio signals may be obtained from a database and / or memory of the audio device 10. In Figure 5 the training method, the audio signal 53 is input for the training of each layer. In Figure 5 the training method, one layer is trained each time.

[0175] The method includes training the first neural network based on the audio data set (e.g., based on the one or more audio signals 53) to perform an audio processing task at at least one first intermediate layer of the first neural network, the at least one first intermediate layer providing a first intermediate layer with an exit possibility to provide an intermediate layer output. In Figure 5 an example, the method includes training the first neural network based on the audio data set (e.g., based on the one or more audio signals 53) to perform an audio processing task at each first intermediate layer of the first neural network, each of the first intermediate providing layers being a first intermediate layer with an exit possibility to provide an intermediate layer output. In other words, the method may include training the audio enhancement module disclosed herein to perform an audio processing task at each predefined potential exit stage (e.g., exit possibility). Figure 5The example of [[ID=]] shows a second example of a computer-implemented method for training the first neural network disclosed herein, namely, the layer-by-layer training of the first neural network.

[0176] In Figure 5 the example of [[ID=]], the first primary intermediate layer 20A is trained in the first step S1. The first primary intermediate layer 20A provides the first primary intermediate layer output 56. As Figure 5 can be observed, the first primary intermediate layer output 56 is used to determine the first loss function 66A based on the first primary intermediate layer output 56 and the first pure audio signal 70A.

[0177] In the second step S2, the audio signal 53 is input into the frozen first primary intermediate layer 20A1, that is, where the parameters of the first primary intermediate layer 20A are frozen or fixed. Then, the output of the frozen first primary intermediate layer 20A1 is input into the first secondary intermediate layer 20B. The first secondary intermediate layer 20B provides the first secondary intermediate layer output 64. Then, the first secondary intermediate layer output 64 is used to determine the second loss function 66B based on the first secondary intermediate layer output 64 and the second pure audio signal 70B.

[0178] In the nth step SN, the audio signal 53 is input into the frozen first primary intermediate layer 20A1, that is, where the parameters of the first primary intermediate layer 20A are frozen or fixed. Then, the output of the frozen first primary intermediate layer 20A1 is input into the frozen first secondary intermediate layer 20B1, which then outputs to the next layer until the first output layer 20C of the first neural network (e.g., the nth layer of the first neural network). Finally, the first output layer output 68 is used to determine the nth loss function 66C based on the first output layer output 68 and the nth pure audio signal 70C. The first neural network can thus be trained layer by layer by repeating this process.

[0179] Figure 6 Schematically shows the training of an example audio device 10 according to the present disclosure, e.g., the training of one or more processors 10C disclosed herein, where the third training technique disclosed herein is applied.

[0180] Figure 6 Shows a third example of a computer-implemented method for training the second exit module 12 disclosed herein, e.g., a third example of the computer-implemented method disclosed herein. For example, Figure 6 a third example of a computer-implemented method for training the third neural network disclosed herein can be shown. Figure 6 A training of Figure 3 the audio device 10 can be shown, e.g., Figure 3 the training of the second exit module 12 (e.g., the third neural network) of

[0181] In one or more examples or embodiments, the second exit module 12 is configured to predict, based on the first audio device parameter, from which second prediction layer of the first model layer to exit when processing an audio input signal. The second exit module 12 may be configured to predict which layer (e.g., which first model layer of the first neural network (e.g., 20A, 20B, 20C, 20D)) will be the best layer to exit when processing the audio input signal. The prediction of the prediction layer (e.g., the first prediction layer, the second prediction layer, and / or the third prediction layer disclosed herein) may be based on one or more features (e.g., based on the quality of the audio input signal 53), the prediction of the output quality (e.g., based on the output quality of the first prediction layer), the performance of the output of the layer (e.g., based on the performance of the output of the first prediction layer), user preferences, and / or one or more audio device parameters (e.g., based on the capabilities of the audio device).

[0182] In other words, the second exit module 12 may be configured to predict from which second prediction layer of the first model layer to exit based on one or more audio device parameters disclosed herein. It can be understood that the second prediction layer may be selected from one or more first intermediate layers or output layers. For example, the second exit module 12 may be configured to predict from which intermediate layer among one or more first intermediate layers to exit. The second exit module 12 may be configured to predict the second prediction layer before the audio enhancement module 13 has started processing the audio input signal 53. The second prediction layer may be an intermediate layer among one or more first intermediate layers. In one or more examples or embodiments, the second prediction layer is different from one or more first intermediate layers.

[0183] In one or more examples or embodiments, the audio device 10 (e.g., the second exit module 12) may be configured to use the first neural network 20 to predict the minimum layer at which the audio enhancement module 13 may exit processing in order to provide an acceptable audio output signal. In other words, the audio device 10 (e.g., the second exit module 12) may be configured to use the first neural network 20 to predict the minimum layer at which the audio enhancement module 13 may exit processing in order to have at least to some extent performed a specific audio processing task.

[0184] In one or more examples or embodiments, the second exit module 12 is configured to predict which third prediction layer of the first model layer (20A, 20B, 20C, 20D) the processing performance, quality, and / or efficiency of the audio input signal 53 converges to based on the audio input signal 53. In other words, the second exit module 12 can be configured to predict at which third prediction layer of the first neural network 20 the point of diminishing returns is reached. The point of diminishing returns and / or the convergence point can be regarded as the point at which each additional input gives a slower improvement in the output. In other words, the point of diminishing returns and / or the convergence point can be the point at which further processing of the audio input signal will give a slower output improvement than the previous layer. Predicting when the processing performance, quality, and / or efficiency of the audio input signal converges can include predicting the improvement potential of the processing of the audio input signal using the output of the layers of the first neural network 20. If the predicted improvement potential of the processing of the audio input signal using the output of the layer is low, then that layer can be predicted as the third prediction layer. For example, the second exit module 12 can predict that even if the audio enhancement module continues to process all the layers of the first neural network 20, the first neural network 20 cannot achieve more than 90% improvement. Then, the second exit module 12 can predict that the layer achieving the improvement between 80% and 90% is the third prediction layer at which the processing performance, quality, and / or efficiency of the audio input signal converges. In one or more examples or embodiments, the third prediction layer is configured to provide a third prediction layer output.

[0185] Figure 6 The third training technique shown can be used to train the prediction of the second exit module 12.

[0186] The method for training the second exit module includes obtaining an audio data set 53 including one or more audio signals. The one or more audio signals 53 can be obtained via one or more microphones 10E and / or via the first wireless transceiver 10F. In one or more examples or embodiments, the one or more audio signals can be obtained from the database and / or memory of the audio device 10. In Figure 6 the training method, the audio signal 53 is input into the audio enhancement module 13, for example, the first neural network 20. The audio enhancement module 13 can then output an intermediate layer output for each first intermediate layer of the first neural network 20. For example, the first primary intermediate layer 20A outputs the first primary intermediate layer output 56, the first secondary intermediate layer 20B outputs the first secondary intermediate layer output 64, the first tertiary intermediate layer 20C outputs the first tertiary intermediate layer output 68, and the first quaternary intermediate layer 20D outputs the first quaternary intermediate layer output 69. Then, the intermediate layer outputs 56, 64, 68, 69 are evaluated, for example, the performance, quality, and / or efficiency of the intermediate layer outputs are evaluated. For example, the intermediate layer outputs 56, 64, 68, 69 are input into the first exit module disclosed herein for evaluating their performance, quality, and / or efficiency.

[0187] One or more audio signals 53 can also be directly evaluated by bypassing the audio enhancement module 13. For example, one or more audio signals 53 can also be directly input into the first exit module 14, for example, bypassing the audio enhancement module 13. This can allow comparison of the processing of the audio enhancement module 13 with the unprocessed audio signals 53.

[0188] The audio device 10 (e.g., one or more processors 10C and / or the first exit module 14) is then configured to evaluate the performance, quality, and / or efficiency of the intermediate layer outputs by determining a speech quality point SQP for each intermediate layer output.

[0189] In addition, the audio device 10 (e.g., one or more processors 10C and / or the first exit module 14) is then configured to determine a mean opinion score MOS for each intermediate layer output (e.g., each SQP). For example, MOS 0 = 1 is the score of the unprocessed audio input signal 53, MOS 1 = 2 is the score of the first primary intermediate layer output 56, MOS 2 = 3 is the score of the first secondary intermediate layer output 64, MOS 3 = 3.4 is the score of the first tertiary intermediate layer output 68, and MOS 4 = 3.5 is the score of the first quaternary intermediate layer output 69.

[0190] The audio device 10 (e.g., one or more processors 10C and / or the first exit module 14) is then configured to determine the difference (e.g., increment) between the scores. For example, given the score of the unprocessed audio input signal 53, the increment of the score of the first primary intermediate layer output 56 can be regarded as +1. Therefore, the first primary intermediate layer 20A can be assigned the index 20A + 1.

[0191] Given the score of the first primary intermediate layer output 56, the increment of the score of the first secondary intermediate layer output 64 can be regarded as +1. Therefore, the first secondary intermediate layer 20B can be assigned the index 20B + 1.

[0192] Given the score of the first secondary intermediate layer output 64, the increment of the score of the first tertiary intermediate layer output 68 can be regarded as +0.4. Therefore, the first tertiary intermediate layer 20C can be assigned the index 20C + 0.4.

[0193] Given the score of the first tertiary intermediate layer output 68, the increment of the score of the first quaternary intermediate layer output 69 can be regarded as +0.1. Therefore, the first quaternary intermediate layer 20D can be assigned the index 20D + 0.1.

[0194] The audio device 10 (e.g., one or more processors 10C and / or the first exit module 14) is then configured to determine 82 whether the fraction of the intermediate layer output meets a criterion, e.g., the first criterion disclosed herein. The criterion can include, for example, an incremental threshold to be met. In this example, the incremental threshold to be met can be set to 0.3. The audio device 10 (e.g., one or more processors 10C and / or the first exit module 14) can then be configured to select the intermediate layer output whose increment meets the incremental threshold of 0.3. Thus, the first tertiary intermediate layer output 68 is selected. As described herein, this can allow finding the point of diminishing returns, e.g., the convergence point. Thus, the index of the first intermediate layer (in this case the first tertiary intermediate layer 20C) corresponding to the increment that meets the incremental threshold is selected as the best first intermediate layer to exit. In one or more examples or embodiments, the audio device 10 (e.g., one or more processors 10C and / or the first exit module 14) is then configured to transform 84 the result using one-hot encoding. For example, since the first tertiary intermediate layer 20C has been determined as the best exit layer, it can be encoded to read 3[0, 0, 1, 0]. The transformed result can input the result of the intermediate layer output evaluation into the loss function 66 for training the second exit module 12. In other words, the result 3[0, 0, 1, 0] can be read by the second exit module 12 as the probability that the intermediate layer output is the best first intermediate layer to exit. According to this example, the probability will read: p(20A)=0, p(20B)=0, p(20C)=1, p(20D)=0. In other words, the first tertiary intermediate layer 20C is determined as the best first intermediate layer to exit. It can be understood that the loss function 66 for training the second exit module 12 can be regarded as a cross-entropy loss function.

[0195] In one or more examples or embodiments, the second exit module 12 can be trained during operation (e.g., during runtime). For example, the second exit module 12 can be configured to occasionally perform a verification, e.g., a sanity check, to verify that the predicted first prediction layer is the best prediction. It can be understood that the first exit module 14 can provide verification of the prediction of the second exit module 12 by verifying the intermediate layer output.

[0196] Examples of the audio device, system, and method according to the present disclosure are set forth in the following items:

[0197] Item 1. An audio device, comprising:

[0198] An audio enhancement module, comprising a first neural network having a first model layer, the first model layer including a first input layer, a plurality of first intermediate layers, and a first output layer; and

[0199] A first exit module;

[0200] Wherein, the audio enhancement module is configured to process an audio input signal using a first neural network to provide an audio output signal, and wherein at least one first intermediate layer has an exit possibility to provide an intermediate layer output, and wherein a first exit module is configured to determine whether the intermediate layer output meets a first criterion, where the first criterion indicates the performance, quality, and / or efficiency of the intermediate layer output, and wherein, based on the intermediate layer output meeting the first criterion, the audio device is configured to determine the audio output signal based on the intermediate layer output.

[0201] Item 2. The audio device according to item 1, wherein the audio device includes a second exit module configured to obtain one or more features of the audio input signal including a first feature, and wherein the second exit module is configured to predict, based on the first feature, from which first prediction layer of a first model layer to exit when processing the audio input signal, where the first prediction layer is configured to provide a first prediction layer output, and wherein the audio device is configured to determine the audio output signal based on the first prediction layer output.

[0202] Item 3. The audio device according to item 2, wherein the first prediction layer is an intermediate layer among one or more first intermediate layers.

[0203] Item 4. The audio device according to any one of items 2-3, wherein the second exit module is configured to obtain one or more audio device parameters including a first audio device parameter, and wherein the second exit module is configured to predict, based on the first audio device parameter, from which second prediction layer of the first model layer to exit when processing the audio input signal, where the second prediction layer is configured to provide a second prediction layer output, and wherein the audio device is configured to determine the audio output signal based on the second prediction layer output.

[0204] Item 5. The audio device according to item 4, wherein the first audio device parameter is a power parameter, a battery parameter, and / or a processing capacity parameter, and wherein the second exit module is configured to predict, based on the power parameter, the battery parameter, and / or the processing capacity parameter, from which second prediction layer of the first neural network to exit when processing the audio input signal.

[0205] Item 6. The audio device according to any one of items 2-5, wherein the second exit module is configured to predict, based on the audio input signal, at which third prediction layer of the first model layer the processing performance, quality, and / or efficiency of the audio input signal converges, and wherein the audio device is configured to determine the audio output signal based on the third prediction layer.

[0206] Item 7. The audio device according to item 6, wherein determining the audio output signal based on the prediction includes determining, based on the third prediction layer, from which layer of the first model layer to exit when processing the audio input signal.

[0207] Item 8. The audio device according to any one of Items 6 - 7, wherein the third prediction layer is configured to provide an output of the third prediction layer, and wherein the audio device is configured to determine an audio output signal based on the output of the third prediction layer.

[0208] Item 9. The audio device according to any one of Items 6 - 8, wherein the audio device is configured to determine an audio output signal based on the output of the layer before the third prediction layer.

[0209] Item 10. The audio device according to any one of Items 2 - 9, wherein the first exit module includes a second neural network, and wherein determining whether the intermediate layer output meets the first criterion includes using the second neural network to determine whether the intermediate layer output meets the first criterion.

[0210] Item 11. A method for implementing efficient neural network processing performed by an audio device, wherein the audio device includes an audio enhancement module, the audio enhancement module includes a first neural network having a first model layer, the first model layer includes a first input layer, a plurality of first intermediate model layers, and a first output layer; and a first exit module, and wherein the method includes:

[0211] Processing (S106) an audio input signal using the first neural network to provide an audio output signal, wherein at least one of the first intermediate layers has an exit possibility to provide an intermediate layer output,

[0212] Using the first exit module to determine (S108) whether the intermediate layer output meets a first criterion, wherein the first criterion indicates the performance, quality, and / or efficiency of the intermediate layer output, and

[0213] Based on the intermediate layer output meeting the first criterion, determining (S110) the audio output signal based on the intermediate layer output.

[0214] Item 12. The method according to Item 11, the method includes:

[0215] Obtaining (S102) one or more features including the first feature of the audio input signal using a second exit module,

[0216] Based on the first feature, predicting (S104) from which first prediction layer of the first model layer to exit when processing (S106) the audio input signal, wherein the first prediction layer is configured to provide an output of the first prediction layer, and

[0217] Determining (S110A) the audio output signal based on the output of the first prediction layer.

[0218] Item 13. The method according to Item 12, the method includes:

[0219] Obtain (S103), using a second exit module, one or more audio device parameters including first audio device parameters,

[0220] Predict (S105), using the second exit module and based on the first audio device parameters, from which second prediction layer of a first model layer to exit when processing (S106) an audio input signal, wherein the second prediction layer is configured to provide a second prediction layer output, and

[0221] Determine (S110B) an audio output signal based on the second prediction layer output.

[0222] Item 14. The method according to item 13, wherein predicting (S105) from which second prediction layer of a first model layer to exit includes predicting (S105A) from which second prediction layer of a first neural network to exit when processing an audio input signal based on power parameters, battery parameters, and / or processing capability parameters

[0223] Item 15. The method according to any one of items 12 - 14, the method comprising:

[0224] Predict (S107), using the second exit module and based on the audio input signal, at which third prediction layer of a first model layer the processing performance, quality, and / or efficiency of the audio input signal converges, and

[0225] Determine (S110C) an audio output signal based on the third prediction layer.

[0226] Item 16. The method according to item 15, wherein determining (S110C) an audio output signal includes determining (S110C1) from which layer of a first model layer to exit when processing an audio input signal based on the third prediction layer.

[0227] Item 17. The method according to any one of items 15 - 16, wherein the third prediction layer is configured to provide a third prediction layer output, and wherein the method includes determining (S110D) an audio output signal based on the third prediction layer output.

[0228] Item 18. The method according to any one of items 15 - 17, the method comprising:

[0229] Determine (S110E) an audio output signal based on the output of a layer before the third prediction layer.

[0230] Item 19. The method according to any one of items 11 - 18, wherein the first exit module includes a second neural network, and wherein determining (S108) whether an intermediate layer output meets a first criterion includes using the second neural network to determine (S108A) whether the intermediate layer output meets the first criterion.

[0231] Item 20. A computer-implemented method for training a first neural network in any one of Items 1-10, wherein the method comprises:

[0232] Obtaining an audio data set comprising one or more audio signals; and

[0233] Based on the audio data set, training the first neural network to perform an audio processing task at at least one first intermediate layer of the first neural network to provide a first intermediate layer with an exit probability to provide an intermediate layer output.

[0234] Item 21. An audio device, comprising:

[0235] An audio enhancement module comprising a first neural network having a first model layer, the first model layer comprising a first input layer, a plurality of first intermediate layers, and a first output layer;

[0236] A second exit module configured to obtain one or more features of the audio input signal comprising a first feature,

[0237] wherein the audio enhancement module is configured to process the audio input signal using the first neural network to provide an audio output signal, and wherein at least one first intermediate layer has an exit probability to provide an intermediate layer output, and wherein the second exit module is configured to predict, based on the first feature, from which first prediction layer of the first model layer to exit when processing the audio input signal, wherein the first prediction layer is configured to provide a first prediction layer output, and wherein the audio device is configured to determine the audio output signal based on the first prediction layer output.

[0238] Item 22. An audio device, comprising:

[0239] An audio enhancement module comprising a first neural network having a first model layer, the first model layer comprising a first input layer, a plurality of first intermediate layers, and a first output layer;

[0240] A second exit module, the second exit module being configured to obtain one or more audio device parameters comprising a first audio device parameter, and wherein the second exit module is configured to predict, based on the first audio device parameter, from which second prediction layer of the first model layer to exit when processing the audio input signal, wherein the second prediction layer is configured to provide a second prediction layer output, and wherein the audio device is configured to determine the audio output signal based on the second prediction layer output.

[0241] Item 23. An audio device, comprising:

[0242] An audio enhancement module comprising a first neural network having a first model layer, the first model layer comprising a first input layer, a plurality of first intermediate layers, and a first output layer;

[0243] A second exit module, the second exit module being configured to predict which third prediction layer of the audio input signal's processing performance, quality, and / or efficiency converges to a first model layer, and wherein the audio device is configured to determine an audio output signal based on the third prediction layer.

[0244] The use of the terms "first", "second", "third", and "fourth", "primary", "secondary", "tertiary", etc. does not imply any particular order, but includes these terms for identifying individual elements. In addition, the use of the terms "first", "second", "third", and "fourth", "primary", "secondary", "tertiary", etc. does not denote any order or importance, but the terms "first", "second", "third", and "fourth", "primary", "secondary", "tertiary", etc. are used to distinguish one element from another. Note that the words "first", "second", "third", and "fourth", "primary", "secondary", "tertiary", etc. are used here and elsewhere for labeling purposes only and are not intended to indicate any particular spatial or temporal ordering. In addition, the labeling of a first element does not imply the existence of a second element, and vice versa.

[0245] It can be understood that the accompanying drawings include some circuits or operations shown in solid lines and some circuits, components, features, or operations shown in dashed lines. The circuits or operations included in the solid lines are the circuits, components, features, or operations included in the broadest example. The circuits, components, features, or operations included in the dashed lines are examples that can be included in or are a part of the circuits, components, features, or operations of the solid line example, or other circuits, components, features, or operations that can be adopted in addition to the circuits, components, features, or operations of the solid line example. It should be understood that these operations do not need to be performed in the order presented. In addition, it should be understood that not all operations need to be performed. The example operations can be performed in any order and in any combination. It should be understood that these operations do not need to be performed in the order presented. The circuits, components, features, or operations included in the dashed lines can be considered optional.

[0246] Other operations not described herein can be incorporated into the example operations. For example, one or more additional operations can be performed before, after, simultaneously with, or between any of the described operations.

[0247] Certain features discussed above as separate implementations may also be implemented in combination as a single implementation. Conversely, features described as a single implementation may also be implemented in multiple implementations separately or in any suitable subcombination. Furthermore, although features may be described above as acting in certain combinations, in some cases one or more features in a claimed combination may be removed from that combination, and the combination may be claimed as any subcombination or variation of any subcombination.

[0248] It should be noted that the word "comprising" does not necessarily exclude the presence of other elements or steps than those listed.

[0249] It should be noted that the word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements.

[0250] It should be noted that the term "indication" may be viewed as "associated with", "involving", "describing", "characterizing" and / or "defining". The terms "indication", "associated with", "involving", "describing", "characterizing" and "defining" may be used interchangeably. The term "indication" may be viewed as indicating a relationship. For example, weight data indicating a weight may include one or more weight parameters.

[0251] It should be noted that the word "based on" may be viewed as "according to" and / or "derived from". The terms "based on" and "according to" may be used interchangeably. For example, a parameter determined "based on" a data set may be viewed as a parameter determined "according to" the data set. In other words, a parameter may be the output of one or more functions that take a data set as input.

[0252] A function may represent a relationship between input and output, such as a mathematical relationship, a database relationship, a hardware relationship, a logical relationship, and / or other suitable relationship.

[0253] It should also be noted that any reference signs do not limit the scope of the claims, that examples may be implemented at least partially by means of hardware and software, and that several "means", "units" or "devices" may be represented by the same item of hardware.

[0254] Although features have been shown and described, it should be understood that these features are not intended to limit the claimed disclosure, and it will be apparent to those skilled in the art that various changes and modifications may be made without departing from the scope of the claimed disclosure. Therefore, the specification and drawings are to be regarded as illustrative rather than restrictive. The claimed disclosure is intended to cover all alternatives, modifications and equivalents.

Claims

1. An audio device, comprising: An audio enhancement module comprising a first neural network having a first model layer, wherein the first model layer comprises a first input layer, a plurality of first intermediate layers, and a first output layer; and First exit module; wherein the audio enhancement module is configured to process an audio input signal using the first neural network to provide an audio output signal, and wherein at least one first intermediate layer has an exit possibility to provide an intermediate layer output, and wherein the first exit module is configured to determine whether the intermediate layer output satisfies a first criterion, wherein the first criterion indicates performance, quality and / or efficiency of the intermediate layer output, and wherein, based on the intermediate layer output satisfying the first criterion, the audio device is configured to determine the audio output signal based on the intermediate layer output.

2. The audio device according to claim 1, wherein The audio device includes a second exit module, which is configured to obtain one or more features of the audio input signal including a first feature, and wherein the second exit module is configured to predict which first prediction layer of the first model layer to exit from when processing the audio input signal based on the first feature, wherein the first prediction layer is configured to provide a first prediction layer output, and wherein the audio device is configured to determine the audio output signal based on the first prediction layer output.

3. The audio device according to claim 2, wherein: The first prediction layer is an intermediate layer of the one or more first intermediate layers.

4. The audio device according to any one of claims 2 to 3, wherein: The second exit module is configured to obtain one or more audio device parameters including a first audio device parameter, and wherein the second exit module is configured to predict, based on the first audio device parameter, which second prediction layer of the first model layer to exit from when processing the audio input signal, wherein the second prediction layer is configured to provide a second prediction layer output, and wherein the audio device is configured to determine the audio output signal based on the second prediction layer output.

5. The audio device according to claim 4, wherein: The first audio device parameter is a power parameter, a battery parameter and / or a processing capability parameter, and wherein the second exit module is configured to predict which second prediction layer of the first neural network to exit from when processing the audio input signal based on the power parameter, the battery parameter and / or the processing capability parameter.

6. The audio device according to any one of claims 2 to 5, wherein: The second exit module is configured to predict, based on the audio input signal, which third prediction layer of the first model layer the processing performance, quality and / or efficiency of the audio input signal converges to, and wherein the audio device is configured to determine the audio output signal based on the third prediction layer.

7. The audio device according to claim 6, wherein: Determining the audio output signal based on the prediction includes determining, based on the third prediction layer, which layer of the first model layers to exit from when processing the audio input signal.

8. The audio device according to any one of claims 6 to 7, wherein: The third prediction layer is configured to provide a third prediction layer output, and wherein the audio device is configured to determine the audio output signal based on the third prediction layer output.

9. The audio device according to any one of claims 6 to 8, wherein: The audio device is configured to determine the audio output signal based on an output of a layer preceding the third prediction layer.

10. The audio device according to any one of claims 2 to 9, wherein: The first exit module includes a second neural network, and wherein determining whether the intermediate layer output meets the first criterion includes using the second neural network to determine whether the intermediate layer output meets the first criterion.

11. A method for implementing efficient neural network processing performed by an audio device, wherein: The audio device comprises: an audio enhancement module, the audio enhancement module comprising a first neural network having a first model layer, the first model layer comprising a first input layer, a plurality of first intermediate model layers and a first output layer; and a first exit module, wherein the method comprises: processing (S106) an audio input signal using the first neural network to provide an audio output signal, wherein at least one first intermediate layer has a dropout possibility to provide an intermediate layer output, determining (S108) using the first exit module whether the intermediate layer output meets a first criterion, wherein the first criterion indicates performance, quality and / or efficiency of the intermediate layer output, and According to the intermediate layer output satisfying the first criterion, the audio output signal is determined ( S110 ) based on the intermediate layer output.

12. The method according to claim 11, comprising: obtaining (S102) one or more features of the audio input signal including the first feature using a second exit module, Based on the first feature, predicting (S104) from which first prediction layer of the first model layer to exit when processing (S106) the audio input signal, wherein the first prediction layer is configured to provide a first prediction layer output, and The audio output signal is determined (S110A) based on the first prediction layer output.

13. The method according to claim 12, comprising: using the second exit module to obtain (S103) one or more audio device parameters including the first audio device parameter, using the second dropout module and based on the first audio device parameter, predicting (S105) which second prediction layer of the first model layer to drop out from when processing (S106) the audio input signal, wherein the second prediction layer is configured to provide a second prediction layer output, and The audio output signal is determined ( S110B) based on the second prediction layer output.

14. The method according to claim 13, wherein: Predicting (S105) which second prediction layer of the first model layer to exit from comprises predicting (S105A) which second prediction layer of the first neural network to exit from when processing the audio input signal based on the power parameter, the battery parameter and / or the processing capability parameter.

15. A computer-implemented method for training the first neural network of any one of claims 1 to 10, wherein: The method comprises: obtaining an audio data set comprising one or more audio signals; and Based on the audio data set, the first neural network is trained to perform an audio processing task in at least one first intermediate layer of the first neural network to provide a first intermediate layer with an exit possibility to provide an intermediate layer output.