Method of processing audio data in audio device by using neural network
By adjusting the activation function of neural network nodes, the problems of high computing costs and difficult to adjust in the existing technology are solved, and efficient audio processing and sound quality improvement are achieved.
Patent Information
- Application Number
- CN202411883002.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art uses neural networks to process audio data, the calculation cost is high and difficult to adjust to meet specific user needs.
By adjusting the activation function of the neural network node, keeping the neural network topology unchanged to achieve efficient computational adjustment. The method includes obtaining audio data and input, adjusting the activation function based on the input, and processing the audio data using the adjusted neural network.
It realizes that without changing the neural network topology, the computational efficiency and sound quality of audio processing are improved while reducing noise.
Smart Images

Figure CN120183373A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computer-implemented method for processing audio data in an audio device using a neural network, the neural network being defined by its topology, the topology including its number of layers and nodes, where each node has an activation function. More specifically, the present disclosure relates to a method including obtaining audio data and obtaining an input (e.g., from a user). Background Art
[0002] In the field of audio processing, the use of neural networks is well known. For example, it is well known to use neural networks to improve speech clarity. Another example is the use of neural networks to reduce noise. One reason for using neural networks in such applications is that several parameters need to be considered, and there are linear and non-linear dependencies between different parameters.
[0003] Although neural networks have proven successful in different aspects of audio processing, they also require a certain computational cost. For example, if a neural network is to be used in a wide variety of sound environments, a wide variety of devices, etc., the neural network must be extended to handle all possible combinations. The disadvantage of this increased complexity is an increase in computational cost, i.e., an increase in the number of processors required for the operations.
[0004] Another problem with using neural networks is that it is very difficult, and usually impossible, to use a neural network to adjust or regulate an audio device for audio processing. Since neural networks used to solve complex problems are not designed for adjustment, the only option for adjusting a neural network is usually to retrain the neural network with a different set of training data.
[0005] For the above reasons, there is a need to find a balance between the advantages of using neural networks to solve non-linear multi-parameter problems and keeping the computational cost at a reasonable level, while ensuring that an audio processing device using a neural network can be adjusted to meet specific user needs.
[0006] In the literature of neural networks, activation functions can be used to provide non-linearity, which plays an important role in the underlying network decisions in real-world applications. Although many well-established activation functions, such as ReLU, Sigmoid, linear, etc., have been frequently applied over the years, there has recently been a renewed interest in the research community in activation functions that can improve the performance of neural networks.
[0007] A survey on available activation functions can be found in "Activation Functions in Deep Learning: A Comprehensive Survey and Benchmark" by SR Dubey et al. in Neurocomputing, while more modern trainable activation functions can be found in "A survey on modern trainable activation functions" by A. Apicella et al. in Neural Networks 138 (2021) 14 - 32.
[0008] However, there is a need for an improved computer - implemented method for processing audio data in an audio device by using a neural network. Summary of the Invention
[0009] A computer - implemented method for processing audio data in an audio device using a neural network is disclosed. The neural network is defined by its topology, which includes the number of its layers and nodes, where each node has an activation function.
[0010] The method includes obtaining first audio data.
[0011] The method includes obtaining an input, where the input includes one or more of the following: an input from an audio engineer who adjusts the audio device; an input defining preferences from a user of the audio device; a hearing diagram of a user of the audio device; and device characteristics of the audio device.
[0012] The method includes adjusting the activation function of one or more nodes of the neural network based on the input while maintaining the topology of the neural network, thereby allowing the neural network to be adjusted in a computationally efficient manner.
[0013] The method includes processing the first audio data into processed audio data by using the neural network with the adjusted activation function.
[0014] The method includes outputting the processed audio data.
[0015] The processed audio data can improve the sound quality and can reduce noise in the audio signal. The method can be applied, for example, to deep noise suppression, deep echo suppression, bandwidth expansion, etc.
[0016] The method can be executed when a user actually uses the audio device, for example, when making a call using headphones, performing hearing compensation using a hearing aid, participating in an audio conference using a speaker, etc.
[0017] The method can also be performed during the setup / regulation of an audio device, for example, by an audio engineer to adjust the audio device to be used, or by a user to set up the audio device before use, such as setting preferences like gain, volume, etc.
[0018] An audio device can be a device for capturing and / or processing and / or outputting audio data (such as audio signals, sounds, etc.). Sound can be captured by one or more input transducers (such as microphones). Sound can be processed by a processing unit. Sound can be output by output transducers (such as speakers, loudspeakers, receivers, etc.). The audio device can be headphones, a hands-free phone, a hearing device, a hearing aid, a smartphone, etc.
[0019] Audio data can be obtained by a microphone of the audio device. Audio data can also be obtained by a microphone of other devices outside the audio device, such as a remote headset. Therefore, the audio data may originate from the surrounding environment of the audio device and / or may originate from other locations outside the audio device. Therefore, the audio data may originate from many different sources, and thus the audio data does not need to be received by the microphone of the audio device. For example, the audio data may come from a remote device, such as a remote caller in a phone call. The audio data can be a sound signal.
[0020] The method includes obtaining an input, where the input includes one or more of the following: an input from an audio engineer who adjusts the audio device; an input from a user of the audio device who defines preferences; a hearing profile of the user of the audio device; and device characteristics of the audio device. Therefore, the input can be an individual's input, such as the user of the audio device, or an employee who manufactures the audio device, such as an audio engineer who adjusts the audio device before use. The input can alternatively and / or additionally be in the form of a hearing profile that defines the user's hearing loss. The input can alternatively and / or additionally be the device characteristics of the audio device, such as the device characteristics of headphones, hands-free phones, hearing aids, etc.
[0021] For example, in the case where the input is a user input from an audio engineer, the engineer can select a specific function or a similar function from a code library.
[0022] For example, in the case where the input is a user input that defines preferences, the user can undergo one or more hearing tests, where the user rates audio clips and selects an activation function based on the ratings. For example, the user can listen to the same audio clip processed with different activation functions and rate them, and then select the audio clip with the highest rating and the associated activation function.
[0023] For example, in the case where the input is a hearing profile, the activation function can be constrained based on the user's hearing profile.
[0024] For example, in the case where the input is device characteristics, the device characteristics can be determined according to the device type of the audio device. For example, if the audio device is a speaker, the device characteristics may be specific to the speaker; if the audio device is headphones, the device characteristics may be specific to the headphones; if the audio device is a hearing aid, the device characteristics may be specific to the hearing aid, etc.
[0025] Therefore, obtaining the input can be done before using the audio device, i.e., when setting / adjusting the audio device. For example, obtaining the input from an audio engineer who adjusts the audio device can be done before use; obtaining the input from a user who defines preferences for the audio device can be done before or during use; obtaining the input as the user's audiogram for the audio device can be done before use; and obtaining the input as the device characteristics of the audio device can also be done before use.
[0026] Therefore, obtaining the input can be done before or during the use of the audio device, and the above examples are all before use, but obtaining the input that defines preferences from the user of the audio device can also be done during use, where if the audio output does not meet the requirements, the user can change the settings.
[0027] Artificial neural networks are used to process audio data. A neural network is defined by its topology, which includes the number of layers and nodes of the neural network, where each node has an activation function. Artificial neural networks can be used to solve artificial intelligence (AI) problems; they model the connections of biological neurons as weights between nodes. Positive weights reflect excitatory connections, while negative values represent inhibitory connections. All inputs are modified by the weights and added together. This activity is called a linear combination. Finally, the activation function controls the amplitude of the output. For example, the acceptable output range is usually between 0 and 1, or it can be between -1 and 1. These artificial networks can be used for predictive modeling, adaptive control, and applications that can be trained via a dataset. Generally, nodes or neurons are aggregated into layers. Different layers can perform different transformations on their inputs. Signals are transmitted from the first layer (input layer) to the last layer (output layer), possibly after traversing these layers multiple times.
[0028] The method includes adjusting the activation function of one or more nodes of a neural network based on an input while maintaining the topology of the neural network, thereby allowing the neural network to be adjusted in a computationally efficient manner. Thus, the activation function is adjusted based on the input, such as selecting, changing, updating, etc. Thus, the "default" activation function can be changed to another activation function, or the "default" activation function can be adjusted, such as by adjusting some parameters, values, etc. of the activation function. The aim is to find a suitable activation function according to the situation, and then the existing activation function of the neural network can be replaced with the suitable activation function, thereby attempting to enhance the performance of the neural network without retraining the neural network or changing the topology of the neural network. In some cases, the activation function may be associated only with the last node in the output layer of the neural network. Thus, the neural network can process different acoustic environments, different device characteristics, different user preferences, etc. in a computationally efficient manner.
[0029] In some embodiments, adjusting the activation function of one or more nodes of a neural network while maintaining the neural network topology includes: selecting an activation function from a codebook including a plurality of activation functions based on an input, and associating the selected activation function with one or more nodes of the neural network. Associating the selected activation function with one or more nodes of the neural network can be understood as inserting the selected activation function into one or more nodes of the neural network.
[0030] In some embodiments, adjusting the activation function of one or more nodes of a neural network while maintaining the neural network topology includes: determining one or more parameters related to the activation function based on an input. For example, the activation function can be a parameterized function, such as a parameterized ReLU function:
[0031] f(y) = max(0, y) + αmin(0, y)
[0032] where y is the weighted sum input to the node, and α is a parameter that can be determined based on the input. For example, this parameter can be determined based on the input provided by an audio engineer. Although ReLU is specifically mentioned, other parameterized functions (such as a parameterized sigmoid or similar functions) are equally feasible.
[0033] In some embodiments, adjusting the activation function of one or more nodes of a neural network while maintaining the topology of the neural network includes: synthesizing an activation function based on an input.
[0034] The method includes processing first audio data into processed audio data using a neural network having an adjusted activation function. Thus, audio data will be processed using the neural network, and due to the adjusted and preferably optimized activation function, the processed audio data can be improved.
[0035] The method includes outputting processed audio data.
[0036] The advantages of the present method and audio device are that it can modify the achievable performance of the neural network during runtime without modifying the neural network topology. By keeping the neural network topology, additional conditions can be added. In other words, instead of allowing the neural network to add nodes and / or layers, the topology is set. In other words, the neural network is not only trained to process audio data into an output signal using training data, but also processes in this way for the selected topology. The effect of setting or selecting the topology is that the computational cost can be predicted more accurately. This in turn can adjust the hardware accordingly. Adjusting the hardware in turn can improve the computational efficiency.
[0037] The advantage is that this can be achieved by changing the activation function to a function learned in the use case scenario encountered during the inference process. For example, during the certification test of an audio device product, this effect can only be seen by changing the activation function to a function learned from the ROM (Read-Only Memory), so the runtime performance can be adjusted according to the end user's intention or certification requirements.
[0038] The advantage is that there is no need to modify the neural network topology.
[0039] The advantages of the present method and audio device are that signal processing can be improved, the target signal can be improved, noise in the audio signal can be reduced, and so on.
[0040] The advantage is that there is no need to change the neural network or add the number of layers of the neural network.
[0041] The advantage is that less audio data processing may be required, and thus less power consumption (e.g., less battery usage) may be required.
[0042] According to one aspect, an audio device is disclosed, including a processor and a memory, wherein the audio device is configured to execute the methods disclosed above and below.
[0043] In some embodiments, the method includes:
[0044] - Obtaining second audio data;
[0045] - Determining one or more eigenvalue based on the second audio data, wherein the one or more eigenvalue are related to the acoustic environment; and
[0046] - Adjusting the activation function of one or more nodes of the neural network based on the input and based on the determined one or more eigenvalue, while keeping the neural network topology,
[0047] Thereby allowing the neural network to be adjusted in a computationally efficient manner.
[0048] The second audio data can be an audio signal. The first audio data and the second audio data can be the same or different audio data. When adjusting / setting the audio device, the first audio data and the second audio data can be the same. When using the audio device, for example, during inference, the first audio data and the second audio data can be different.
[0049] The second audio data can be received through a microphone, such as the microphone of the audio device, or the microphone of another device (such as a remote device).
[0050] One or more eigenvalue(s) is / are determined based on the second audio data.
[0051] The eigenvalue(s) is / are related to the acoustic environment from which the second audio data is obtained. The acoustic environment can be a location, a place, such as outdoors, such as indoors, such as outdoors under windy conditions, such as a quiet indoors. The acoustic environment can be a noisy restaurant, a noisy office space, a room with echo, a reverberant room. The acoustic environment can be defined by the signal-to-noise ratio. The acoustic environment can be a car in which the user is making a call. The acoustic environment can be the place where the microphone is capturing the second audio data, and / or the location where the user or the audio device is located, etc. The acoustic environment may be important for determining the noise level, the speech level, the echo level, etc. A sound event detection module can be implemented in the audio device to determine when there is a sound in the environment and what kind of sound it is.
[0052] The eigenvalue(s) can be the sound level, the noise level, the speech level, the amplitude, the frequency, etc. of the audio data.
[0053] Based on the input and based on the determined one or more eigenvalue(s), the activation function is adjusted to provide improved sound processing.
[0054] In some embodiments, an activation function is selected from a code library that includes multiple activation functions, and the code library optionally further includes coefficients and / or boundary coefficients associated with the activation function.
[0055] Therefore, different from other neural network applications in the field of audio processing, after the neural network training is completed, a specific activation function can be adjusted (i.e., changed or replaced). By setting (i.e., fixing) the topology of the neural network, different activation function packages associated with different adjustment options can be used to create the code library.
[0056] In some embodiments, the activation function code library includes at least one or more of the following activation functions:
[0057] - ReLU (Rectified Linear Unit);
[0058] - tanh (Hyperbolic Tangent Function);
[0059] - sigmoid;
[0060] - linear; and / or
[0061] - identity.
[0062] In some embodiments,
[0063] - when the input includes an input from an audio engineer, the audio engineer selects a specific activation function from the code library to test the specific activation function;
[0064] - when the input includes an input from user-defined preferences, the user undergoes one or more hearing tests, wherein the user rates audio clips and selects a specific activation function based on the ratings; and / or
[0065] - when the input includes the audiogram of a user of an audio device, the activation function is constrained based on the user's hearing profile.
[0066] When the input includes an input from a user with defined preferences, the user undergoes one or more hearing tests, wherein the user rates audio clips and selects a specific activation function based on the ratings. For example, the user may listen to and rate the same audio clip processed with different activation functions and then select the audio clip with the highest rating and the associated activation function.
[0067] In some embodiments, the first audio data or one or more eigenvalue(s) based on the first audio data is / are provided as a further input to the activation function code library to further guide the selection of the activation function.
[0068] For example, if the first audio data is a reverberant sound signal, it may help guide the selection of the activation function if the audio engineer has defined that a specific activation function should be used for reverberant signals or if the user has shown a preference for a specific activation function for reverberant signals.
[0069] In some embodiments, the activation function is adjusted based on the first audio data or one or more eigenvalue(s) based on the first audio data.
[0070] For example, if the first audio data is a reverberant sound signal, if the audio engineer has defined that a specific activation function should be used for the reverberant signal, or if the user has shown a preference for a specific activation function for the reverberant signal, it can help guide the selection of the activation function. Based on the input provided by the audio engineer or the user of the audio device, other characteristics of the first audio signal (such as noise or echo) can also be associated with different activation functions. In some embodiments, the method includes determining one or more nodes of the neural network to be adjusted based on one or more eigenvalue. The one or more eigenvalue can be based on the first audio data or the second audio data.
[0071] In some embodiments, the method includes:
[0072] In response to obtaining the first audio data, determining whether the activation function provides an undesired processed audio data output,
[0073] If so, determining one or more nodes of the activation function to be adjusted, adjusting the activation function of the nodes, and processing the audio data,
[0074] Otherwise, processing the audio data without adjusting the neural network and outputting the processed audio data.
[0075] Thus, this embodiment can explain the decision of when it is appropriate to adjust the activation function. If the activation function set during setup / regulation does not provide satisfactory results, the activation function can be adjusted. For example, a specific sound environment may not have been tested during the setup / regulation of the audio device, such as a room with very strong reverberation, and then the method can determine whether the activation function that has been set is suitable for this specific sound environment. For example, the specific device characteristics of the audio device can provide whether the activation function provides a desired or undesired processed audio data output, and if the output is undesired, the activation function is adjusted.
[0076] In some embodiments, the activation function is only associated with the nodes of the output layer of the neural network. This is an advantage because this may be the most influential, as the output layer of the neural network, i.e., the last layer, may be directly related to the mask, so it is sufficient to only use the output layer. This can save processing power by not performing unnecessary calculations.
[0077] In other cases, the activation function may be associated with more than just the nodes of the output layer of the neural network. One or more activation functions may be associated with the nodes of the input layer of the neural network. One or more activation functions may be associated with one or more intermediate layer nodes of the neural network.
[0078] In some embodiments, an activation function is adjusted at least in part based on first audio data received by a microphone of an audio device. Thus, the adjustment of the activation function can be based on the first audio data. The first audio data is received by the microphone of the audio device. The adjustment of the activation function can also be based on other things, such as an input, such as a user input as described above. It has been disclosed above that the activation function is adjusted based on an input. However, as defined in this embodiment, the activation function may not be determined entirely based on the input and may be determined based on the first audio data.
[0079] In some embodiments, one or more eigenvalue sets include a first set related to a sound environment, and / or a second set related to device characteristics of the audio device, and / or a third set related to user preferences of user settings of the audio device, wherein the eigenvalue set used to determine one or more nodes to be adjusted is the first set, the second set, the third set, or any combination thereof.
[0080] For example, the first sound environment can be a reverberation chamber having a first set of specific eigenvalue sets, while the second sound environment can be an anechoic chamber having another first set of other specific eigenvalue sets. Echo, reverberation, and voice activity can be examples of the first set.
[0081] For example, the first device characteristic can be a speakerphone having a second set of specific eigenvalue sets, and the second device characteristic can be a headphone chamber having another second set of other specific eigenvalue sets, and the third device characteristic can be a hearing aid having yet another second set of other specific eigenvalue sets.
[0082] For example, the first user preference can be a specific gain setting having a third set of specific eigenvalue sets, while the second user preference can be a specific frequency range having another third set of other specific eigenvalue sets.
[0083] In some embodiments, one or more eigenvalue sets include a fourth set related to data transmission effects, such as signal-to-noise ratio, wherein the eigenvalue set used to determine one or more nodes to be adjusted is the first set, the second set, the third set, the fourth set, or any combination thereof.
[0084] In some embodiments, one or more eigenvalue sets include a fifth set related to the first audio data and / or a sixth set related to the second audio data, wherein the eigenvalue set used to determine one or more nodes to be adjusted is the first set, the second set, the third set, the fourth set, the fifth set, the sixth set, or any combination thereof.
[0085] In some embodiments, a first set of coefficients is defined / determined for each of a first set of eigenvalue, and / or a second set of coefficients is defined / determined for each of a second set of eigenvalue, and / or a third set of coefficients is defined / determined for each of a third set of eigenvalue, and / or a fourth set of coefficients is defined / determined for each of a fourth set of eigenvalue, and / or a fifth set of coefficients is defined / determined for each of a fifth set of eigenvalue, and / or a sixth set of coefficients is defined / determined for each of a sixth set of eigenvalue.
[0086] In an embodiment, the audio device is configured to be worn by a user. The audio device may be arranged at, on, above, in, within the ear canal, behind, and / or within the concha of the user's ear, i.e., the audio device is configured to be worn within, on, above, and / or at the user's ear. The user may wear two audio devices, one for each ear. The two audio devices may be connected, e.g., wirelessly and / or wired, e.g., a binaural hearing aid system.
[0087] The audio device may be an audible device, such as a headset, a headset with microphone, earbuds, earplugs, a hearing aid, a personal sound amplification product (PSAP), an over-the-counter (OTC) hearing device, a hearing protection device, a general-purpose hearing device, a custom hearing device, or other head-mounted hearing devices. The hearing device may include prescription and non-prescription devices.
[0088] The audio device may adopt various housing styles or form factors. Some of the form factors are behind-the-ear (BTE) hearing aids, receiver-in-canal (RIC) hearing aids, receiver-in-ear (RIE) hearing aids, or microphone-and-receiver-in-ear (MaRIE) hearing aids. Some of the form factors are earplugs, supra-aural headphones, or ear-hook headphones. Those skilled in the art are well aware of different types of audio / hearing devices and different options for arranging the audio / hearing device within, on, above, and / or near the ear of the audio / hearing device wearer.
[0089] In an embodiment, the audio device may include one or more input transducers. The one or more input transducers may include one or more microphones. The one or more input transducers may include one or more vibration sensors configured to detect bone vibrations. The one or more input transducers may be configured to convert an acoustic signal into a first electrical input signal. The first electrical input signal may be an analog signal. The first electrical input signal may be a digital signal. The one or more input transducers may be coupled to one or more analog-to-digital converters configured to convert the analog first input signal into a digital first input signal.
[0090] In an embodiment, the audio device may include one or more antennas configured for wireless communication. The one or more antennas may include electrical antennas. The electrical antennas may be configured to communicate wirelessly at a first frequency. The first frequency may be higher than 800 MHz, preferably with a wavelength between 900 MHz and 6 GHz. The first frequency may be from 902 MHz to 928 MHz. The first frequency may be from 2.4 GHz to 2.5 GHz. The first frequency may be from 5.725 GHz to 5.875 GHz. The one or more antennas may include magnetic antennas. The magnetic antennas may include magnetic cores. The magnetic antennas may include coils. The coils may be wound around the magnetic cores. The magnetic antennas may be configured to communicate wirelessly at a second frequency. The second frequency may be lower than 100 MHz. The second frequency may be between 9 MHz and 15 MHz.
[0091] In an embodiment, the audio device may include one or more wireless communication units. The one or more wireless communication units may include one or more wireless receivers, one or more wireless transmitters, one or more transmitter-receiver pairs, and / or one or more transceivers. At least one of the one or more wireless communication units may be coupled to the one or more antennas. The wireless communication units may be configured to convert a wireless signal received by at least one of the one or more antennas into a second electrical input signal. The audio device may be configured for wired / wireless audio communication, e.g., enabling a user to listen to media such as music or radio and / or enabling the user to make a phone call.
[0092] In an embodiment, the wireless signal may originate from one or more external sources and / or external devices such as a spouse microphone device, a wireless audio transmitter, a smart computer, and / or a distributed microphone array associated with a wireless transmitter. The wireless input signal may originate from another audio device, e.g., as part of a binaural hearing system and / or from one or more accessory devices such as a smartphone and / or a smartwatch.
[0093] In an embodiment, the audio device may include a processing unit. The processing unit may be configured to process a first electrical input signal and / or a second electrical input signal. The processing may include compensating for the user's hearing loss, i.e., applying frequency-dependent gain to the input signal according to the user's frequency-dependent hearing impairment. The processing may include performing feedback cancellation, beamforming, tinnitus reduction / masking, noise reduction, noise cancellation, speech recognition, bass adjustment, treble adjustment, and / or processing of user input. The processing unit may be a processor, an integrated circuit, an application, a functional module, etc. The processing unit may be implemented in a signal processing chip or a printed circuit board (PCB). The processing unit may be configured to provide a first electrical output signal based on the processing of the first electrical input signal and / or the second electrical input signal. The processing unit may be configured to provide a second electrical output signal. The second electrical output signal may be based on the processing of the first electrical input signal and / or the second electrical input signal.
[0094] In an embodiment, the audio device may include an output transducer. The output transducer may be coupled to the processing unit. The output transducer may be a receiver. It should be noted that in this case, the receiver may be a speaker, and the wireless receiver may be a device configured to process wireless signals. The receiver may be configured to convert the first electrical output signal into an acoustic output signal. The output transducer may be coupled to the processing unit via a magnetic antenna. The output transducer may be included in the ITE unit or the earphone of the audio device, such as an in-the-ear receiver (RIE) unit or an in-the-ear microphone and receiver (MaRIE) unit. One or more input transducers may be included in the ITE unit or the earphone.
[0095] In an embodiment, the wireless communication unit may be configured to convert the second electrical output signal into a wireless output signal. The wireless output signal may include synchronization data. The wireless communication unit may be configured to transmit the wireless output signal through at least one of one or more antennas.
[0096] In an embodiment, the audio device may include a digital-to-analog converter configured to convert the first electrical output signal, the second electrical output signal, and / or the wireless output signal into an analog signal.
[0097] In an embodiment, the audio device may include a power supply. The power supply may include a battery providing a first voltage. The battery may be a rechargeable battery. The battery may be a replaceable battery. The power supply may include a power management unit. The power management unit may be configured to convert the first voltage into a second voltage. The power supply may include a charging coil. The charging coil may be provided by a magnetic antenna.
[0098] In an embodiment, the audio device may include a memory, including memory in volatile and non-volatile forms.
[0099] The present invention relates to different aspects, including the methods and audio devices described above and below, as well as corresponding method and device components, each aspect giving rise to one or more benefits and advantages associated with the first-mentioned aspect, and each aspect having one or more embodiments corresponding to the embodiments associated with the first-mentioned aspect and / or the embodiments disclosed in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Those skilled in the art will readily appreciate the above and other features and advantages from the following detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0101] Figure 1 A method of the prior art is schematically illustrated.
[0102] Figure 2 An exemplary computer-implemented method 200 of processing audio data in an audio device using a neural network is schematically illustrated.
[0103] Figure 3 An exemplary computer-implemented method 300 of processing audio data in an audio device using a neural network is schematically illustrated.
[0104] Figure 4 An exemplary activation function code library 402 is schematically illustrated.
[0105] Figure 5 An exemplary neural network 502 defined by its topology is schematically illustrated, the topology including the number of layers 504 and the number of nodes 506 of the exemplary neural network, where each node 506 has an activation function.
[0106] Figure 6 An exemplary node 606 of a neural network 602 and the activation function f(z) 608 of the node 606 are schematically illustrated. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0107] Various embodiments will be described below with reference to the accompanying drawings. Like reference numerals always denote like elements. Therefore, like elements will not be described in detail for each drawing description. It should also be noted that the drawings are only for facilitating the description of the embodiments. They are not intended as an exhaustive description of the claimed invention or a limitation on the scope of the claimed invention. Moreover, the illustrated embodiments need not have all the aspects or advantages shown. Aspects or advantages described in connection with a particular embodiment are not necessarily limited to that embodiment and may be practiced in any other embodiment, even if not so shown or not so explicitly described.
[0108] Figure 1An existing technology method is schematically shown. The existing technology method 100 shows an adjustment process performed in the prior art. In the prior art adjustment process, an input signal 102 (such as an audio signal) is provided to a trained neural network 103. The trained neural network 103 outputs a first output signal 104 (such as a processed audio signal). Then the first output signal 104 is provided to a post-processor 106, and a user input 108 is also provided to the post-processor 106, such as the user input 108 provided by an audio engineer. Therefore, the audio engineer uses the post-processor 106 to adjust the processing pipeline so that it meets the specifications stipulated by, for example, Microsoft Teams TM or similar audio / video conferencing software. After the post-processing of the signal, a second output signal 110 is output, for example, output by a speaker in an audio device.
[0109] Figure 2 An exemplary computer-implemented method 200 for processing audio data in an audio device using a neural network is schematically shown. The neural network is defined by its topology, which includes the number of its layers and nodes, where each node has an activation function.
[0110] Method 200 includes obtaining 202 first audio data.
[0111] Method 200 includes obtaining 204 an input, where the input includes one or more of the following: an input from an audio engineer who adjusts the audio device; an input from a user of the audio device defining preferences; a hearing diagram of a user of the audio device; and device characteristics of the audio device.
[0112] Method 200 includes adjusting 206 the activation function of one or more nodes of the neural network based on the input while maintaining the topology of the neural network, thereby allowing the neural network to be adjusted in a computationally efficient manner.
[0113] Method 200 includes processing 208 the first audio data into processed audio data by using the neural network with the adjusted activation function.
[0114] Method 200 includes outputting 210 the processed audio data.
[0115] By maintaining the topology of the neural network, additional conditions are added. In other words, instead of allowing the neural network to add nodes and / or layers, the topology is set. In other words, the neural network is not only trained to process audio data into an output signal using training data, but also to perform such processing for the selected topology. The effect of setting or selecting the topology is that the computational cost can be predicted more accurately. This in turn can adjust the hardware accordingly. Adjusting the hardware can in turn improve the computational efficiency.
[0116] Figure 3 Exemplary computer - implemented method 300 for processing audio data in an audio device using a neural network is schematically shown. The neural network is defined by its topology, which includes the number of layers and nodes of the neural network, where each node has an activation function.
[0117] Method 300 includes obtaining 304 an input, where the input includes one or more of the following: input from an audio engineer who adjusts the audio device; input defining preferences from a user of the audio device; a hearing profile of a user of the audio device; and device characteristics of the audio device. The input (e.g., user input) is provided to an activation function code library 306. Based on the input, an activation function is selected in the code library. The selected activation function is passed 308 to a trained neural network 309, where the selected activation function is subsequently associated with one or more predetermined nodes of the trained neural network.
[0118] Method 300 includes obtaining 310 first audio data, e.g., an audio input signal. The audio data, e.g., the input signal, is provided to the trained neural network 309, and then the neural network processes the audio data using the adjusted activation function and provides an output signal 312.
[0119] In the case where the input is from an audio engineer, the engineer can pick a specific function or a similar function from the code library. In the case where the input is a user preference, the user can undergo one or more hearing tests, where the user rates audio clips and selects an activation function based on the ratings. For example, the user can listen to the same audio clip processed with different activation functions, rate them, and then select the audio clip with the highest rating and the associated activation function. If the input is a hearing profile, the activation function can be constrained based on the user's hearing profile. The input can also be related to device details.
[0120] Thus, an activation function suitable for the scenario or the expectations of the user or audio engineer is learned / selected.
[0121] An activation function can be selected in the code library, which can include multiple activation functions, and the code library can also optionally include coefficients and / or boundary coefficients associated with the activation functions.
[0122] The activation function code library can include at least one or more of the following activation functions:
[0123] - ReLU;
[0124] - tanh;
[0125] - sigmoid;
[0126] - linear; and / or
[0127] - identity.
[0128] Figure 3 It is also schematically shown that the first audio data 310 or one or more eigenvalue(s) based on the first audio data can be provided as further input to the activation function code library 306 to further guide the selection of the activation function (314).
[0129] As shown, an audio signal (e.g., an input signal or a parameter of the input signal) can also be provided as input to the code library to further guide the selection of the activation function. For example, if the input signal is a reverberant signal, it can help guide the selection of the activation function if the audio engineer has defined that this activation function should be used for the reverberant signal, or if the user has shown a preference for the activation function of the reverberant signal.
[0130] As Figure 3 shown, different from other neural network applications in the field of audio processing, after the neural network training is completed, a specific activation function can be adjusted (i.e., changed or replaced). By setting (i.e., fixing) the topology of the neural network, different activation function packages related to different adjustment options can be used to make the code library.
[0131] Figure 4 An exemplary activation function code library 402 is schematically shown. As shown, the content of the code library can be different types of activation functions, such as ReLU, tanh, sigmoid, linear, and / or identity. The code library can also include coefficients related to the activation function, boundary coefficients, or similar functions.
[0132] Figure 5 An exemplary neural network 502 defined by its topology is schematically shown. The topology includes the number of layers 504 and the number of nodes 506 of the exemplary neural network, where each node 506 has an activation function (see Figure 6 ). The input layer 504' is shown on the left. The output layer 504” is shown on the right. Between the input layer 504' and the output layer 504”, one or more other layers 504 are shown. It can be understood that the number of layers 504 can be arbitrary. In some cases, the activation function can be associated only with the nodes 506” of the output layer 504” of the neural network 502.
[0133] Figure 6 An exemplary node 606 of the neural network 602 is schematically shown. The activation function f(z) 608 of the node 606 in the artificial neural network 602 is a function that calculates the output a of the node 606 based on the inputs x1, x2, x3, etc. of the node 606 and the weights w1, w2, w3, etc. on each input.
[0134] Therefore, Figure 5 and Figure 6Artificially neural networks 502, 602 for processing audio data are schematically shown. A neural network is defined by its topology, which includes the number of layers 504 and the number of nodes 506, 606 of the neural network, where each node has an activation function 608. Artificial neural networks can be used to solve artificial intelligence (AI) problems; they model the connections of biological neurons as weights between nodes. Positive weights reflect excitatory connections, while negative values indicate inhibitory connections. All inputs are modified by weights and added together. This activity is called a linear combination. Finally, the activation function controls the amplitude of the output. For example, the acceptable output range is typically between 0 and 1, or it can be between -1 and 1. These artificial networks can be used for predictive modeling, adaptive control, and applications that can be trained through a dataset. Generally, nodes or neurons are aggregated into layers. Different layers can perform different transformations on their inputs. Signals are transmitted from the first layer (input layer) to the last layer (output layer), possibly after traversing these layers multiple times.
[0135] Although specific features have been shown and described, it should be understood that these features are not intended to limit the claimed invention, and those skilled in the art will appreciate that various changes and modifications can be made without departing from the scope of the claimed invention. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive. The claimed invention is intended to cover all alternatives, modifications, and equivalents.
[0136] Entry:
[0137] 1. A computer-implemented method for processing audio data in an audio device using a neural network, the neural network being defined by the topology of the neural network, the topology including the number of layers of the neural network and the number of nodes of the neural network, where each node has an activation function, the method comprising:
[0138] - Obtaining first audio data,
[0139] - Based on the input, adjusting the activation function of one or more nodes of the neural network while maintaining the topology of the neural network, thereby allowing the neural network to be adjusted in a computationally efficient manner,
[0140] - Using the neural network with the adjusted activation function to process the first audio data into processed audio data, and
[0141] - Outputting the processed audio data.
[0142] 2. The method according to entry 1, wherein the method comprises:
[0143] - Obtaining second audio data;
[0144] - Determine one or more eigenvalue based on the second audio data, wherein the one or more eigenvalue are related to the acoustic environment; and
[0145] - Based on the input and based on the determined one or more eigenvalue, adjust the activation function of the one or more nodes of the neural network while maintaining the topology of the neural network, thereby allowing the neural network to be adjusted in a computationally efficient manner.
[0146] 3. The method according to any one of item 1 or 2, wherein the activation function is selected from a code library including a plurality of activation functions, and the code library optionally further includes coefficients and / or boundary coefficients associated with the activation function.
[0147] 4. The method according to any one of the preceding items, wherein the activation function code library includes at least one or more of the following activation functions:
[0148] - ReLU;
[0149] - tanh;
[0150] - sigmoid;
[0151] - Linear; and / or
[0152] - Identity.
[0153] 5. The method according to any one of the preceding items, wherein:
[0154] - When the input includes an input from an audio engineer, the audio engineer selects a specific activation function from the code library to test the specific activation function;
[0155] - When the input includes an input from user-defined preferences, the user undergoes one or more hearing tests, wherein the user rates audio clips and selects a specific activation function based on the ratings; and / or
[0156] - When the input includes the audiogram of the user for the audio device, constrain the activation function based on the user's hearing profile.
[0157] 6. The method according to items 3 to 5, wherein the first audio data or the one or more eigenvalue based on the first audio data is provided as a further input to the activation function code library to further guide the selection of the activation function.
[0158] 7. The method according to any one of items 2 to 6, wherein the method includes:
[0159] - Determine the one or more nodes of the neural network to be adjusted based on the one or more eigenvalue.
[0160] 8. The method according to any one of the preceding items, comprising:
[0161] In response to obtaining the first audio data, determine whether the activation function provides an undesired processed audio data output,
[0162] If so, determine the one or more nodes of the activation function to be adjusted, and adjust the activation function of the nodes, and process the audio data,
[0163] Otherwise, process the audio data without adjusting the neural network, and output the processed audio data.
[0164] 9. The method according to any one of the preceding items, wherein the activation function is only associated with the nodes of the output layer of the neural network.
[0165] 10. The method according to any one of the preceding items, wherein the activation function is adjusted at least partially based on the first audio data received via the microphone of the audio device.
[0166] 11. The method according to any one of the preceding items, wherein the one or more eigenvalue include a first set related to the sound environment, and / or a second set related to the device characteristics of the audio device, and / or a third set related to the user preferences of the user settings of the audio device, wherein the eigenvalue for determining the one or more nodes to be adjusted is the first set, the second set, the third set or any combination thereof.
[0167] 12. The method according to the preceding item, wherein the one or more eigenvalue include a fourth set related to the data transmission effect, such as signal-to-noise ratio, wherein the eigenvalue for determining the one or more nodes to be adjusted is the first set, the second set, the third set, the fourth set or any combination thereof.
[0168] 13. The method according to the preceding item, wherein the one or more eigenvalue include a fifth set related to the first audio data, and / or a sixth set related to the second audio data, wherein the eigenvalue for determining the one or more nodes to be adjusted is the first set, the second set, the third set, the fourth set, the fifth set, the sixth set or any combination thereof.
[0169] 14. The method according to item 11, 12 or 13, wherein a first set of coefficients is defined / determined for each of the first set of eigenvalue, and / or wherein a second set of coefficients is defined / determined for each of the second set of eigenvalue, and / or wherein a third set of coefficients is defined / determined for each of the third set of eigenvalue, and / or wherein a fourth set of coefficients is defined / determined for each of the fourth set of eigenvalue, and / or wherein a fifth set of coefficients is defined / determined for each of the fifth set of eigenvalue, and / or wherein a sixth set of coefficients is defined / determined for each of the sixth set of eigenvalue.
[0170] 15. The method according to any one of the preceding items, wherein adjusting the activation function of the one or more nodes of the neural network while maintaining the topology of the neural network includes: selecting the activation function from a codebook including a plurality of activation functions based on the input, and associating the selected activation function with the one or more nodes of the neural network.
[0171] 16. The method according to the preceding item, wherein associating the selected activation function with the one or more nodes of the neural network includes: inserting the selected activation function into the one or more nodes of the neural network.
[0172] 17. The method according to any one of the preceding items, wherein adjusting the activation function of the one or more nodes of the neural network while maintaining the topology of the neural network includes: determining one or more parameters associated with the activation function based on the input.
[0173] 18. The method according to the preceding item, wherein the activation function is a parameterized function, such as a parameterized ReLU function:
[0174] f(y) = max(0, y) + αmin(0, y)
[0175] where y is the weighted sum input to the node, and α is a parameter determined based on the input.
[0176] 19. The method according to any one of the preceding items, wherein adjusting the activation function of the one or more nodes of the neural network while maintaining the topology of the neural network includes: synthesizing an activation function based on the input.
[0177] 20. An audio device, comprising a processor and a memory, wherein the audio device is configured to execute the method according to any one of items 1 to 19.
[0178] List of reference numerals
[0179] 100 Prior art method
[0180] 102 Signal
[0181] 103 Trained neural network
[0182] 104 First output signal
[0183] 106 Post - processor
[0184] 108 User input
[0185] 110 Second output signal
[0186] 200 Computer - implemented method
[0187] 202 Obtain first audio data
[0188] 204 Obtain input
[0189] 206 Adjust the activation function of one or more nodes of the neural network
[0190] 208 Process the first audio data
[0191] 210 Output the processed audio data
[0192] 300 Computer - implemented method
[0193] 304 Obtain input
[0194] 306 Activation function code library
[0195] 308 Pass the selected activation function
[0196] 309 Trained neural network
[0197] 310 Obtain first audio data
[0198] 312 Output signal
[0199] 314 Provide the first audio data as a further input to the activation function code library
[0200] 402 Activation function code library
[0201] 502 Neural network
[0202] 504 Layer
[0203] 504' Input layer
[0204] 504” Output layer
[0205] 506 Node
[0206] Nodes in the "506" output layer
[0207] 602 Neural network
[0208] 606 Node
[0209] 608 Activation function f(z)
[0210] x1, x2, x3 inputs
[0211] w1, w2, w3 weights
[0212] a Output
Claims
1. A computer-implemented method for processing audio data in an audio device by using a neural network, the neural network being defined by a topology of the neural network, the topology comprising a number of layers of the neural network and a number of nodes of the neural network, wherein: Each node has an activation function, and the method comprises: - obtaining first audio data, - based on the input, adjusting an activation function of one or more nodes of the neural network while maintaining the topology of the neural network, thereby allowing the neural network to be adjusted in a computationally efficient manner, - processing the first audio data into processed audio data by using the neural network with the adjusted activation function, and - outputting the processed audio data.
2. The method according to claim 1, wherein: The method comprises: - obtaining second audio data; - based on the second audio data, determining one or more feature values, wherein the one or more feature values are related to the sound environment; and - Based on the input and based on the determined one or more feature values, adjusting an activation function of one or more nodes of the neural network while maintaining a topology of the neural network, thereby allowing the neural network to be adjusted in a computationally efficient manner.
3. The method according to any one of claims 1 or 2, wherein: The activation function is selected from a code base comprising a plurality of activation functions, and the code base optionally also comprises coefficients and / or boundary coefficients associated with the activation function.
4. A method according to any one of the preceding claims, wherein: The activation function code library includes at least one or more of the following activation functions: -ReLU; -tanh; -sigmoid; - linear; and / or -Constant.
5. A method according to any one of the preceding claims, wherein: - when the input comprises an input from an audio engineer, the audio engineer selects a specific activation function from a code base to test the specific activation function; - when the input comprises input from a user defining preferences, the user undergoes one or more hearing tests, wherein the user rates audio clips and based on the ratings, a particular activation function is selected; and / or - when the input comprises an audiogram for a user of the audio device, constraining the activation function based on a hearing profile of the user.
6. The method according to claims 3 to 5, wherein: The first audio data or one or more feature values based on the first audio data are provided as further input to an activation function code library to further guide the selection of an activation function.
7. The method according to any one of claims 2 to 6, wherein: The method comprises: - Based on the one or more feature values, determining one or more nodes of the neural network to be adjusted.
8. A method according to any one of the preceding claims, comprising: In response to obtaining the first audio data, determining whether the activation function provides an undesirable processed audio data output, If yes, determining one or more nodes of the activation function to be adjusted, adjusting the activation function of the node, and processing the audio data, Otherwise, the audio data is processed without adjusting the neural network, and the processed audio data is output.
9. A method according to any one of the preceding claims, wherein: The activation function is only associated with the nodes of the output layer of the neural network.
10. A method according to any one of the preceding claims, wherein: The activation function is adjusted based at least in part on the first audio data received via a microphone of the audio device.
11. A method according to any one of the preceding claims, wherein: One or more characteristic values include a first group related to the sound environment, and / or a second group related to device characteristics of the audio device, and / or a third group related to user preferences set by a user of the audio device, wherein the characteristic values used to determine the one or more nodes to be adjusted are the first group, the second group, the third group or any combination thereof.
12. The method according to the preceding claim, wherein: The one or more characteristic values include a fourth group related to data transmission effect, such as signal-to-noise ratio, wherein the characteristic values used to determine the one or more nodes to be adjusted are the first group, the second group, the third group, the fourth group or any combination thereof.
13. A method according to the preceding claim, wherein: One or more characteristic values include a fifth group related to the first audio data and / or a sixth group related to the second audio data, wherein the characteristic values used to determine the one or more nodes to be adjusted are the first group, the second group, the third group, the fourth group, the fifth group, the sixth group or any combination thereof.
14. The method according to claim 11, 12 or 13, wherein: A first set of coefficients is defined / determined for each of the first set of eigenvalues, and / or wherein, A second set of coefficients is defined / determined for each of the second set of eigenvalues, and / or wherein a third set of coefficients is defined / determined for each of the third set of eigenvalues, and / or wherein a fourth set of coefficients is defined / determined for each of the fourth set of eigenvalues, and / or wherein a fifth set of coefficients is defined / determined for each of the fifth set of eigenvalues, and / or wherein a sixth set of coefficients is defined / determined for each of the sixth set of eigenvalues.
15. An audio device comprising a processor and a memory, wherein: The audio device is configured to perform the method according to any one of claims 1 to 14.