Method for encoding an audio signal using a pulsed neural network and vehicle

By equalizing neuron firing probabilities in a spiking neural network to maximize entropy, the audio signal encoding method improves energy efficiency in battery-electric vehicles, addressing the inefficiency of conventional signal processing.

DE102024003973B3Active Publication Date: 2026-03-12MERCEDES BENZ GROUP AG
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Conventional voice assistants in vehicles rely on energy-inefficient analog or digital signal processing, which is problematic for battery-electric vehicles as it affects the vehicle's range.

Method used

An audio signal encoding method using a spiking neural network (SNN) that adjusts firing thresholds to equalize the firing probability of neurons across channels, maximizing Shannon entropy and optimizing neuron populations to reduce the number of neurons required, thereby increasing energy efficiency.

Benefits of technology

The method enhances energy efficiency by reducing the number of neurons by up to 50% while maintaining information content, thus increasing the range of battery-electric vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for encoding an audio signal (1) by a pulsed neural network (2), wherein the audio signal (1) is divided into at least two channels (1.1, 1.2), each comprising a different frequency band of the audio signal (1), wherein each channel (1.1, 1.2) is assigned its own neuron population (2.1, 2.2) for reading the frequencies (f) assigned to the channel (1.1, 1.2) and outputting a spike train (3), wherein the neurons (4) of the neuron populations (2.1, 2.2) fire to generate a pulse peak (5) in the respective spike train (4) as soon as a signal intensity in the read frequencies exceeds a neuron-specific firing threshold (6). The method according to the invention is characterized in that for at least one channel (1.1, 1.2) the firing thresholds (6) of the neurons (4) are changed until the firing probability (7) of all neurons (4) of the neuron population (2) assigned to the channel (1.1, 1.2) is1, 2.2) during the processing of the audio signal (1) is of equal size within a fixed tolerance limit.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for encoding an audio signal by a pulsed neural network of the type defined in more detail in the preamble of claim 1, and to a vehicle with a computing unit for carrying out the method.

[0002] Speech-based human-machine interfaces have proven their worth in everyday life. Such an interface can also be called a voice assistant. A user can issue a voice command, which is recorded by a microphone and converted into text through appropriate signal processing. This text is then processed using computational linguistics methods to identify semantic content and the user's intention. Artificial intelligence can be used for this purpose. The voice assistant can be used, for example, to retrieve information from external sources such as the internet or to control functions. For instance, it can ask for the time or weather, create an entry on a shopping list or an appointment in a calendar, and so on.External systems such as smart home devices can also be integrated, allowing, for example, smart lighting to be switched on and off via voice command. Such voice assistants can be integrated into a wide variety of devices, such as smartphones, laptops, desktop computers, vehicles, and the like. Voice assistants are particularly beneficial when integrated into vehicles, as this enables interaction with computer systems without requiring the driver to take their eyes off the road or use their hands to input commands. This increases road safety.

[0003] Conventional voice assistants typically use analog or digital signal processing methods, which are effective but not particularly energy-efficient. This is especially problematic for battery-electric vehicles, as energy consumption can affect the vehicle's range.

[0004] A spiking neural network (SNN), also known as a pulsed neural network, is a special type of artificial neural network. While classical artificial neural networks operate all underlying neurons while processing input data, the neurons in a spiking neural network only activate when they receive an input signal. This makes spiking neural networks more energy-efficient than classical neural networks. A spiking neural network can include one or more input neurons. A time-dependent signal is applied to the input neurons. Each neuron is assigned a threshold value for its membrane potential, which is referred to as the firing threshold during registration. The magnitude of the input signal at a given time causes an increase in the neuron's membrane potential.If the firing threshold is exceeded, the neuron "fires," meaning it outputs a pulse. This pulse is referred to as a pulse peak during registration. The pulse peaks recorded over time are aggregated into a so-called spike train. The spike train is propagated by the pulsed neural network and processed by downstream neurons. The pulsed neural network comprises one or more output neurons, which also output a spike train as a result.

[0005] It is known to encode audio signals using pulsed neural networks. For example, US 2019 / 0394568 A1 discloses an auditory signal processor using pulsed neural networks and stimulus reconstruction with a top-down attention control mechanism. The audio signal is divided into several frequency bands using a filter bank, each of which is fed as input to a separate neuron population of the pulsed neural network. For each of these channels, the pulsed neural network outputs a separate spike train to encode the audio signal. This involves assigning sound sources to spatial regions. The encoded audio signal can be decoded using stimulus reconstruction to reconstruct the original audio signal. The original audio signal is divided into left and right channels, i.e., it is a stereo recording.The signal processor disclosed in the publication can, for example, be integrated into hearing devices such as headphones, hearing aids, or cochlear implants. The underlying signal processing algorithm can be used for speech recognition. In particular, this allows for more reliable speech recognition in noisy environments.

[0006] Furthermore, the document WU, J. [et al.]: “A Spiking Neural Network Framework for Robust Sound Classification.” In: Frontiers in Neuroscience, Vol. 12, 2018, Article 836, pp. 1–17. – ISSN 1662-453X presents a method for classifying audio signals using a neural network. The document DIEHL, PU [et al.]: “Fastclassifying, high-accuracy spiking deep networks through weight and threshold balancing.” In: 2015 International Joint Conference on Neural Networks (IJCNN), 2015, pp. 1–8. – ISSN 2161-4407 describes a method for balancing weights and thresholds in spike networks.

[0007] The present invention is based on the objective of providing an improved method for encoding an audio signal.

[0008] According to the invention, this problem is solved by a method for encoding an audio signal using a pulsed neural network with the features of claim 1. Advantageous embodiments and further developments, as well as a vehicle with a computing unit for carrying out the method, are described in the dependent claims.

[0009] A generic method for encoding an audio signal using a pulsed neural network, wherein the audio signal is divided into at least two channels, each comprising a different frequency band of the audio signal, wherein each channel is assigned its own neuron population for reading the frequencies assigned to the channel and outputting a spike train, wherein the neurons of the neuron populations fire to generate a pulse peak in the respective spike train as soon as a signal intensity in the read frequencies exceeds a neuron-specific firing threshold, is further developed according to the invention in that for at least one channel the firing thresholds of the neurons are changed until the firing probability of all neurons of the neuron population assigned to the channel is equal within a defined tolerance limit during the processing of the audio signal.

[0010] The inventive method is based on the understanding that the information content of the resulting spike trains can be maximized by maximizing the Shannon entropy of the respective spike train. Shannon entropy is a measure of the uncertainty or information content of a system. Shannon entropy is explained in more detail later in the figures using mathematical equations. The Shannon entropy of a respective spike train depends on the probability that individual neurons in the channel's neuron population fire. Maximum Shannon entropy is achieved when the firing probability of all neurons in the neuron population assigned to the channel is equal. In other words, the firing probability for each neuron within a channel is thus uniformly distributed.

[0011] To find the desired firing thresholds of the neurons, an optimization problem is solved. Established optimization algorithms can be used for this purpose. For example, a specific audio signal can be repeatedly processed, with the respective firing thresholds of the neurons being changed each time. The firing probability of the neurons is monitored from iteration to iteration, allowing the system to determine when a corresponding uniform distribution of firing probabilities across the neuron population is achieved.

[0012] The defined tolerance threshold, within which the firing probability of the neurons is considered equal, can, for example, correspond to a maximum deviation of 10%, 5%, 1%, or even smaller values. Ideally, the firing probability of all neurons should be exactly the same.

[0013] An advantageous further development of the method according to the invention provides that the firing thresholds of the neuronal populations of all channels are adjusted. Thus, the information content can be increased not only proportionally for the encoded audio signal, but for the entire audio signal. The total entropy across all channels is obtained by summing the individual entropies of the respective channels, since the channels are independent of each other.

[0014] According to a first embodiment of the method according to the invention, it is possible for the individual firing thresholds of the neurons of the neuron populations assigned to different channels to differ across the channels. However, the firing probability is the same for each channel considered individually.

[0015] According to a second possible embodiment of the method according to the invention, the firing probability of all neurons in the pulsed neural network is essentially the same. This means that even across multiple channels, the neurons in the neuron population exhibit essentially the same firing probability.

[0016] According to a further advantageous embodiment of the method according to the invention, it is further provided that the firing thresholds of the neurons are modified such that the global average firing rate of at least the neuron population assigned to the respective channel remains constant or decreases. Here, too, the average firing rate can be kept constant or decreased not only on a channel-specific basis, but also across all channels for the entire pulsed neural network.

[0017] A possible optimization algorithm for finding suitable threshold values ​​λ could look like this: 1. Define the target average firing rate µ in Hz for all neurons across all channels. 2. Determining an initial distribution of the frequency bands across the channels, as well as the total number of pulse peaks resulting from the target average rate of fire. These can be differentiated into so-called onset and offset spikes. 3. Calculate the number of neurons to be provided by dividing the total number of pulse peaks by the target average firing rate. By differentiating between onset and offset spikes, an individual number of onset and offset neurons can be calculated. The respective numbers must be whole numbers. 4. Creating vectors m on and m off with evenly distributed points between 0 and 1. 5. Determining fire threshold values, in particular onset fire threshold values ​​and offset fire threshold values ​​λcon and λ coff , by estimating m-quantiles from the estimated distributions of the onset and offset neurons.

[0018] A first embodiment of the method according to the invention further provides that the number of neurons, at least in the neuron population assigned to the respective channel, is reduced during the adjustment of the firing thresholds. As already mentioned, maximizing the Shannon entropy assigned to the neuron populations increases the information content of the output spike trains. This can be used either to increase the information content while keeping the number of neurons constant, or to decrease the number of neurons while keeping the information content constant. Reducing the number of neurons lowers the energy consumption of the pulsed neural network and thus also of the method for encoding the audio signal. In an embodiment of the method according to the invention in a battery-electric vehicle, this in turn lowers the vehicle's energy consumption and thus increases its range.This allows the efficiency of the disclosed coding method to be increased. In experiments, a reduction in the number of neurons of up to 50% was achieved, while maintaining the same information content.

[0019] According to a further embodiment of the method according to the invention, it is also provided that the distribution of the audio signal frequencies across the channels is changed during the processing of the audio signal. This allows the method according to the invention to be flexibly adapted to different applications. Changing the frequencies assigned to the respective channels affects the respective firing thresholds of the neurons in order to maintain the even distribution of firing probabilities. Thus, if the assignment of frequencies to the channels is changed, the firing thresholds of the neurons that lead to maximum information content also change.

[0020] An advantageous further development of the method according to the invention provides that the spike train output by the pulsed neural network is fed as an input parameter to a trained artificial neural network, wherein the artificial neural network outputs a loss function, wherein an error gradient for the loss function is calculated, and wherein the assignment of the frequencies of the audio signal to the channels is adjusted by backward propagation of the error gradient, taking the error gradient into account. Using the gradient signals generated in this way, the firing thresholds of the neurons can thus be dynamically changed in order to optimally adapt the neuron populations to the respective coding task.

[0021] Preferably, the artificial neural network solves a classification or regression task based on the input spike train. For the classification task, the loss function could be, for example, the so-called cross-entropy loss, and for a regression task, the mean squared error. The artificial neural network outputs this loss function during the processing of the input data, specifically during the processing of the classification or regression task. This loss function determines the error gradient, which can be used to adjust the weights of the artificial neural network. According to the invention, however, not only the weights of the neurons of the trained artificial neural network can be adjusted, but also the assignment of the respective frequencies to the channels.This not only maximizes global entropy but also the system's performance in specific tasks.

[0022] It is conceivable that the input parameter could be not only an audio signal, but also, for example, a video signal. While the audio signal consists of various frequencies to which a volume or level is assigned over time, such a video signal consists of several color channels to which a brightness value is assigned over time. These color channels can, in turn, be spatially resolved, for example, for individual pixels of an image or image regions. A corresponding video signal can also be processed in the same way as the audio signal, whereby the spatially resolved color channels are assigned to individual neuron populations.

[0023] The method according to the invention can be used to improve object classification in images of a video stream. This increases neuronal activity for those image segments that contain relevant image content, while reducing neuronal activity for less important image areas.

[0024] A further advantageous embodiment of the method according to the invention provides that the firing thresholds of the neurons are adaptively adjusted during processing of the audio signal, depending on the audio signal. This enables dynamic fidelity through adaptive firing threshold shifting. Thus, fixed firing thresholds are not used during audio signal processing; instead, the firing thresholds of the neurons can be modified during processing. In other words, the underlying system can learn in real time which parts of the audio signal are particularly relevant and devote more attention, and therefore higher fidelity, to these parts. This improves performance in solving the underlying task. This can be done in combination with, or separately from, reassigning frequencies to channels.

[0025] Similar to what was already described in the context of additionally applying a trained artificial neural network, particularly for solving a classification or regression task to assign the frequencies of the audio signal to the channels, another artificial neural network can be used to identify relevant frequencies of the audio signal. This artificial neural network then identifies neurons within neuron populations for which the respective firing threshold should be lowered or raised. The system thus learns to react dynamically to changes in the audio signal and to allocate neural resources efficiently.

[0026] According to a further advantageous embodiment of the method according to the invention, the audio signal includes speech, and the firing threshold of neurons that process frequencies contained in the speech is lowered. This approach can be combined, in particular, with the adaptive adjustment of the firing thresholds described above. For example, the artificial neural network can identify those frequency components of the audio signal to which speech is assigned. This increases the performance of speech recognition algorithms. Frequency bands that are less important for speech recognition can be encoded with a lower resolution, which further increases energy efficiency.

[0027] The advantages of dynamic fidelity lie in improved attention control, enabling the automatic detection and higher-precision encoding of relevant signal areas. Furthermore, resources can be used more efficiently. Less relevant parts of the signal are encoded at a lower resolution, increasing energy efficiency. Additionally, learning-based optimization is possible. By incorporating gradient signals from downstream classification tasks, the system continuously learns which parts of the audio signal should receive more or less attention to maximize performance.

[0028] For example, a suitable speech classifier can learn to recognize certain speech sounds or words, with the loss function of the classifier being propagated back as a gradient signal.

[0029] In a vehicle comprising a computing unit, the computing unit is configured, according to the invention, to execute a method described above. This allows for particularly resource-saving and thus energy-efficient coding of information in the vehicle. The vehicle can be any road vehicle such as a car, truck, van, bus, or the like. Generally, it could also be a rail vehicle, watercraft, or aircraft. The vehicle is preferably battery-electric powered. Thus, the vehicle's range can be increased by implementing the method according to the invention.

[0030] An advantageous embodiment of the vehicle according to the invention provides that it includes a microphone, wherein the processing unit is further configured to process an audio signal recorded by the microphone using a method described above, particularly in the context of a voice assistant. As already explained in the course of the method, by adjusting the firing thresholds of the neurons to one another, the information content for each channel can be increased or the number of neurons reduced. The performance of a voice assistant provided in the vehicle can thereby be increased and energy consumption reduced.

[0031] Further advantageous embodiments of the inventive method for encoding an audio signal by a pulsed neural network also result from the exemplary embodiments which are described in more detail below with reference to the figures.

[0032] This shows: Fig. 1 a schematic representation of the processing pipeline of an audio signal encoding process according to the invention; Fig. 2 a schematic representation of the operating principle of a neuron of a pulsed neural network; Fig. 3 a schematic representation of various equations; and Fig. 4 a schematic flowchart of the inventive method for encoding the audio signal.

[0033] Using a method according to the invention, a Fig. The audio signal 1 shown in Figure 1 is encoded using a pulsed neural network 2. The method is executed by a processing unit, for example, a processing unit integrated into a vehicle. The method according to the invention is particularly preferred for use in speech processing, especially speech recognition. In this case, the audio signal 1 was generated, in particular, by a microphone 8. The audio signal 1 is divided into at least two channels 1.1 and 1.2. This is done, for example, using a filter bank 9. The filter bank 9 can represent a so-called cochlear model. Several frequencies of the audio signal 1 are assigned to each channel 1.1, 1.2. In particular, bandpass filtering is performed so that different contiguous frequency bands are supplied to each channel 1.1, 1.2. The pulsed neural network 2 then processes the frequency bands assigned to channels 1.1, 1.2. For this purpose, each channel 1.1, 1.2 is assigned a frequency band 1.1.2. A separate neuron population 2.1, 2.2 is assigned. In . Fig. The neuron populations 2.1 and 2.2 are shown in Figure 1 as purely exemplary representations. Thus, neuron populations 2.1 and 2.2 can comprise a suitable number of input neurons, intermediate neurons, and output neurons. One of the neurons 4 is in Fig. 1 is shown as an example with a reference symbol. As output, the respective neuron populations 2.1 and 2.2 each generate a spike train 3. The entirety of the spike trains 3 then represents the encoded audio signal 1.

[0034] According to the invention, for at least one channel 1.1, 1.2 respectively, Fig. The two fire thresholds shown in the diagram were changed by 6 of the neurons until the values ​​in the diagram were reached. Fig. The firing probability shown for all neurons of the neuron population 2.1, 2.2 assigned to channel 1.1, 1.2 during the processing of audio signal 1 is equal within a defined tolerance limit. This results in the Shannon entropy 10 or H being... c The respective spike train 3, and thus the encoded audio signal 1, is enlarged. This increases the information content, ultimately allowing the number of neurons 4 in the respective neuron populations 2.1, 2.2 to be reduced while maintaining the same information content. This reduces energy consumption during the execution of the method according to the invention. This can be used to lower the energy consumption of a vehicle executing the corresponding method and thus to increase its range.

[0035] The firing thresholds 6 of neuron populations 2.1 and 2.2, as well as the assignment of the respective frequencies f to channels 1.1 and 1.2, can be adjusted during the processing of the audio signal 1. For this purpose, the encoded audio signal 1, i.e., the aforementioned spike trains 3, is fed as an input parameter to a trained artificial neural network (ANN) downstream in the processing pipeline. For example, the ANN can solve a classification or regression task. In doing so, the ANN outputs a loss function for which an error gradient is determined. This error gradient is calculated in step 101 and, as indicated by arrow 102, can be propagated back in the system to adjust the assignment of the frequency ranges to channels 1.1 and 1.2.The adjustment of the firing thresholds 6 can be adaptive, taking into account the audio signal 1, so that the firing thresholds 6 change continuously during time t. It is also conceivable to use static firing thresholds 6.

[0036] The functional principle of each neuron 4 of the pulsed neural network 2 is explained using the following: Fig. Figure 2 illustrates this. Three diagrams are shown, depicting the volume I or level for different frequencies f of the audio signal 1 at different times t1, t2, and t3. The firing threshold 6 is indicated in each diagram. It is conceivable that there is a separate neuron 4 for each frequency f. However, it is also conceivable that one neuron 4 is assigned to several frequencies f and thus to a frequency band. If the volume I of a particular frequency f exceeds the firing threshold 6, the neuron 4 outputs a pulse peak 5; that is, the neuron 4 fires. The variable P here indicates the potential. In the Fig. In the embodiment shown in Figure 2, this is the case at time t2.

[0037] Fig. Figure 3 shows in subfigure (1) the Shannon entropy 10 of a spike train 3 for channel c. Here, p represents c iThe firing probability 7. The firing probability 7 corresponds to the probability that the i-th neuron 4 fires in channel c.

[0038] The total entropy 11 over all channels C is obtained by summing the entropies of the individual channels. This is shown using the in Fig. The equation shown in 3(2) illustrates this. To maximize the information content, it is necessary to determine the probabilities p. c i For each neuron 4 within each channel 1.1, 1.2 are equally distributed, that is, p c i = 1 / N.

[0039] If the firing probabilities 7 within each channel 1.1, 1.2 are equally distributed, the entropy H simplifies c for each channel c to the in Fig. 3(3) shown equation. The total entropy 11 is then obtained by the Fig. Equation 3(4) shown.

[0040] It follows that the entropy, and thus the information content, is maximized when the firing probability 7 for all neurons 4 within a channel 1.1, 1.2 is uniformly distributed. This means that by appropriately adjusting the firing thresholds 6, usually denoted by the symbol λ, maximum entropy can be achieved in each channel 1.1, 1.2. Optimizing the firing thresholds 6 aims to achieve this uniform distribution to ensure the information content per channel 1.1, 1.2 without exceeding the global average firing rate, typically denoted by the symbol µ.

[0041] As previously described, the distribution of frequencies f across channels 1.1 and 1.2 can be modified based on a backpropagated error gradient of the artificial neural network (ANN). For the firing thresholds 6 of neurons 4, this results in the following: Fig. Equation 5 is shown. Here, λ corresponds to the fire threshold value 6 at a given time t or t+1. η corresponds to the learning rate and the fraction to the gradient 12 of the loss function with respect to the fire threshold values ​​6.

[0042] The in Fig. Equation 3(6) illustrates the influence on the estimated distribution of frequencies f across channels 1.1 and 1.2. It is assumed that the distribution of frequencies f in channel c is described by a density function 13 as a function of x and λ. To account for the influence of the gradient signal, the distribution is fitted by the gradient. The new distribution is obtained by the equation shown in Fig. Equation 3(6) shows how the gradient descent changes the fire thresholds 6 and λ, and thus the underlying frequency distribution f. This change, in turn, affects the population coding.

[0043] The goal is to maximize the information content, i.e., the entropy, of neurons 4. Therefore, gradient descent is not only aimed at minimizing the loss function L, but also at fitting the distribution to maximize entropy. The entropy H with respect to the distribution p(x; λ) is determined by the in Fig. Equation 3(7) shown is adapted. By adjusting the fire thresholds 6 taking into account the gradient signal, the distribution of frequencies f is influenced in such a way that the entropy H is maximized and at the same time the task to be solved by the artificial neural network KNN is optimized.

[0044] Fig.Figure 4 illustrates a flowchart of the method according to the invention. In step 401, the method starts. The audio signal 1 is recorded or provided. In step 402, the audio signal 1 is divided into the respective channels 1.1 and 1.2. The audio signal 1 can also be divided into more than two channels. For example, one channel is provided for low frequency bands, one channel for mid-frequency bands, and one channel for high frequency bands. For example, the lowest frequency can be 50 Hz and the highest frequency 20 kHz.

[0045] In step 403, the frequencies assigned to each channel 1.1, 1.2 are processed by the neurons 4 of the respective neuron populations 2.1, 2.2. If it is necessary to reassign the frequencies f to the channels 1.1, 1.2, a new distribution can be estimated in step 404.

[0046] In step 405, it is checked whether a corresponding gradient signal is present. If this is not the case, the firing thresholds 6 are calculated statically in step 406. This maximizes the entropies in each channel 1.1, 1.2 in step 407. In step 408, the respective spike train 3 is output for each channel 1.1, 1.2.

[0047] If, however, a gradient signal is present, the fire threshold values ​​6 are adjusted by the gradient in step 409. Finally, the procedure ends in step 410.

Claims

[1] Method for encoding an audio signal (1) by a pulsed neural network (2), wherein the audio signal (1) is divided into at least two channels (1.1, 1.2), each comprising a different frequency band of the audio signal (1), wherein each channel (1.1, 1.2) is assigned its own neuron population (2.1, 2.2) for reading the frequencies (f) assigned to the channel (1.1, 1.2) and outputting a spike train (3), wherein the neurons (4) of the neuron populations (2.1, 2.2) fire to generate a pulse spike (5) in the respective spike train (4) as soon as a signal intensity in the read frequencies exceeds a neuron-specific firing threshold (6), characterized by, that for at least one channel (1.1, 1.2) the firing thresholds (6) of the neurons (4) are changed until the firing probability (7) of all neurons (4) of the neuron population (2.1, 2.2) assigned to the channel (1.1, 1.2) during the processing of the audio signal (1) is equal within a specified tolerance limit, wherein - the number of neurons, at least in the neuron population (2.1, 2.2) assigned to the respective channel (1.1, 1.2), is reduced as part of the adjustment of the firing thresholds (6); and / or - the distribution of the frequencies of the audio signal (1) across the channels (1.1, 1.2) is changed during the processing of the audio signal (1). [2] Method according to claim 1, characterized by , that the firing thresholds (6) of the neuron populations (2.1, 2.2) of all channels (1.1, 1.2) are adjusted. [3] Method according to claim 1 or 2, characterized by, that the firing thresholds (6) of the neurons (4) are changed such that the global average firing rate of at least the neuron population (2.1, 2.2) assigned to the respective channel (1.1, 1.2) remains constant or decreases. [4] Method according to claim 1, characterized by , that if the distribution of the frequencies of the audio signal (1) onto the channels (1.1, 1.2) is changed during the processing of the audio signal (1), the spike train (3) output by the pulsed neural network (2) is fed to a trained artificial neural network (ANN) as an input parameter, wherein the artificial neural network (ANN) outputs a loss function, wherein an error gradient for the loss function is calculated, and wherein the assignment of the frequencies (f) of the audio signal (1) to the channels (1.1, 1.2) is adjusted taking into account the error gradient by backward propagation of the error gradient. [5] Method according to claim 4, characterized by, that the artificial neural network (ANN) solves a classification task or a regression task based on the input spike train (3). [6] Method according to any one of claims 1 to 5, characterized by , that the firing thresholds (6) of the neurons (4) are adaptively adjusted depending on the audio signal (1) during the processing of the audio signal (1). [7] Method according to any one of claims 1 to 6, characterized by , that the audio signal (1) includes speech, whereby the firing threshold (6) of neurons (4) that process frequencies contained in the speech is lowered. [8] Vehicle comprising a computing unit, characterized by that the computing unit is configured to execute a method according to any one of claims 1 to 7. [9] Vehicle according to claim 8, further comprising a microphone (8), characterized by, that the computing unit is configured to process an audio signal recorded by the microphone (8) using a method according to one of claims 1 to 7, in particular in the course of a voice assistant.

Citation Information

Patent Citations

  • Auditory signal processor using spiking neural network and stimulus reconstruction with top-down attention control

    US20190394568A1