Data efficient and personalized audio scene classifier adaptation
By extracting and transmitting the statistical aggregation of feature vectors from hearing aids to a server for deep neural network training, the problem of insufficient audio scene classification adaptation in existing technologies is solved, realizing personalized audio scene classification and privacy protection, and improving the listening experience of hearing aids.
Patent Information
- Application Number
- CN202480020904.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-08-18
- Filing Date
- 2024-02-21
- Publication Date
- 2025-11-11
AI Technical Summary
Existing audio scene classification methods cannot be effectively adapted to the specific environment of an individual and require a large amount of data manipulation. They are usually limited by the storage and processing capabilities of hearing aids, and privacy protection requirements restrict data transmission.
By extracting feature vectors from hearing aids and generating statistical aggregates, which are then transmitted to a server for extrapolation, deep neural networks are used to train and update classifier parameters, enabling adaptive adaptation of audio scene classification while ensuring data privacy protection.
Under limited computing and storage conditions, personalized audio scene classification was achieved, which improved the listening experience of hearing aids, reduced data transmission, and met privacy protection requirements.
Smart Images

Figure CN120937392A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Application 63 / 447,487, filed February 22, 2023, and European Patent Application EP23192168.5, filed August 18, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This invention relates to audio scene classifier adaptation, and more specifically, to neural network scene classification adaptation in hearing aids, such as hearing aids and cochlear implants, using a server, such as a cloud-based server. Background Technology
[0003] Normal ear Figure 1 As shown in the prior art, sound is transmitted via the outer ear 101 to the tympanic membrane 102, which moves the ossicles (malleus, incus, and stapes) of the middle ear 103, causing the oval and round window openings of the cochlea 104 to vibrate. The cochlea 104 is a long, narrow duct that spirals about two and a half turns around its axis. It includes a superior canal called the scala vestibulae and a inferior canal called the scala tympani, which are connected by the cochlear duct. The cochlea 104 forms an upright spiral cone with a central axis called the cochlear axis, where the spiral ganglion cells of the auditory nerve 113 reside. In response to the received sound transmitted by the middle ear 103, the fluid-filled cochlea 104 acts as a transducer to generate electrical impulses, which are transmitted to the cochlear nerve 113 and ultimately to the brain.
[0004] Hearing loss occurs when there is a problem with the ability to transduce external sounds into meaningful action potentials along the neural matrix of the cochlea 104. To improve impaired hearing, hearing aids have been developed. For example, when the damage is related to the functioning of the middle ear 103, conventional hearing aids can be used to provide mechanical stimulation to the auditory system in the form of amplified sound. Alternatively, when the damage is related to the cochlea 104, a cochlear implant with implanted stimulating electrodes can electrically stimulate the auditory nerve tissue using small currents delivered through multiple electrode contacts distributed along the electrodes.
[0005] Figure 1 Some components of a typical cochlear implant system are also shown, including an external microphone that provides audio signal input to an external signal processor 111, in which various signal processing schemes can be implemented. The processed signal is converted into a digital data format, such as a sequence of data frames, for transmission to the implant 108. In addition to receiving the processed audio information, the implant 108 performs additional signal processing, such as error correction, impulse generation, etc., and generates stimulation patterns (based on the extracted audio information) that are transmitted via electrode leads 109 to the implantable electrode array 110.
[0006] Typically, the electrode array 110 includes a plurality of electrode contacts 112 on its surface, which provide selective stimulation of the cochlea 104. Depending on the context, the electrode contacts 112 are also referred to as electrode channels. In modern cochlear implants, a relatively small number of electrode channels are each associated with a relatively wide frequency band, and each electrode contact 112 addresses a group of neurons with a charged electrical stimulation pulse, the charge originating from the instantaneous amplitude of the signal envelope within that frequency band.
[0007] In some stimulation signal encoding strategies, stimulation pulses are applied at a constant rate across all electrode channels, while in others, they are applied at a channel-specific rate. Various specific signal processing schemes can be implemented to generate electrical stimulation signals. Signal processing methods well-known in the field of cochlear implants include sequential interleaved sampling (CIS), channel-specific sampling sequences (CSSS) (as described in U.S. Patent No. 6,348,070, which is incorporated herein by reference in its entirety), fine structure processing (FSP) strategies, spectral peaking (SPEAK), and compressed analog processing (CA).
[0008] In the CIS strategy, the signal processor uses only the bandpass signal envelope for further processing; that is, these envelopes contain all stimulus information. For each electrode channel, the signal envelope is represented as a biphasic pulse sequence with a constant repetition rate. A characteristic of CIS is that the stimulation rate is equal for all electrode channels and is independent of the center frequency of each channel. The intention is that the pulse repetition rate is not a temporal cue for the patient (i.e., it should be high enough that the patient does not perceive a tone with a frequency equal to the pulse repetition rate). The pulse repetition rate is typically chosen to be greater than twice the bandwidth of the envelope signal (based on the Nyquist theorem).
[0009] In a CIS system, stimulation pulses are applied in a strictly non-overlapping sequence. Therefore, as is typical of CIS, only one electrode channel is active at a time, and the total stimulation rate is relatively high. For example, assuming a total stimulation rate of 18 kpps and a 12-channel filter bank, the stimulation rate per channel is 1.5 kpps. Such a per-channel stimulation rate is generally sufficient for a adequate temporal representation of the envelope signal. The maximum total stimulation rate is limited by the minimum phase duration per pulse. The phase duration cannot be arbitrarily short because the shorter the pulse, the higher the current amplitude must be to elicit an action potential in the neuron, and the current amplitude is limited for various practical reasons. For a total stimulation rate of 18 kpps, the phase duration is 27 μs, which is close to the lower limit.
[0010] Med-El's Fine Structure Processing (FSP) strategy uses CIS in higher frequency channels and fine structure information present in the bandpass signal in lower frequency, top-level electrode channels. In the FSP electrode channels, zero-crossings of the bandpass-filtered time signal are tracked, and at each negative-to-positive zero-crossing, a Channel-Specific Sampling Sequence (CSSS) is initiated. Typically, the CSSS sequence is applied to up to three of the top electrode channels, covering a frequency range up to 200 or 330 Hz. The FSP is deployed in Hochmair I, Nopp P, Jolly C, and Schmidt M. Further description can be found in H. Garnham, C. Anderson, I. MED-EL Cochlear, *Implants: State of the Art and a Glimpse into the Future*, *Trends in Amplification*, vol. 10, 201-219, 2006, which is incorporated herein by reference in its entirety. The FS4 coding strategy differs from FSP in that up to four apical channels can utilize their fine-structure information. In FS4-p, stimulus pulse sequences can be delivered in parallel on any two of the four FSP electrode channels. Utilizing both FSP and FS4 coding strategies, the fine-structure information is the instantaneous frequency information of a given electrode channel, which can provide users with improved auditory experience, better speech comprehension, and enhanced perceived audio quality. See, for example, U.S. Patent 7,561,709; Lorens et al., “Fine structure processing improves speech perception as well as objective and subjective benefits in pediatric MED-EL COMBI 40+ users.” International Journal of Pediatric Otorhinolaryngology 74.12 (2010):1372-1378; and Vermeire et al., “Better speech recognition in noise with the fine structure processing coding strategy.” ORL 72.6 (2010):305-311; all of which are incorporated herein by reference in their entirety.
[0011] Many cochlear implant coding strategies use the so-called n-of-m method, where only a few (n) electrode channels with the maximum amplitude are stimulated in a given sampling timeframe. If the amplitude of a particular electrode channel remains higher than that of other channels in a given timeframe, that channel is selected for the entire timeframe. Subsequently, the number of electrode channels available for encoding information is reduced by one, which leads to clustering of stimulation pulses. Therefore, fewer electrode channels are available to encode important temporal and spectral characteristics of the sound signal (e.g., speech initiation).
[0012] In addition to the specific processing and encoding methods discussed above, different specific pulse stimulation patterns may exist to deliver stimulation pulses with specific electrodes—i.e., monopolar, bipolar, tripolar, multipolar, and phased array stimulation. Furthermore, different stimulation pulse shapes exist—i.e., biphasic, symmetrical triphasic, asymmetrical triphasic pulses, or asymmetrical pulse shapes. These various pulse stimulation patterns and pulse shapes each offer different benefits; for example, higher tone selectivity, smaller electrical threshold, higher electrical dynamic range, fewer unwanted side effects (e.g., facial nerve stimulation), etc.
[0013] Fine-structured coding strategies such as FSP and FS4 use zero-crossing of the bandpass signal to initiate a channel-specific sampled sequence (CSSS) pulse sequence to be delivered to the corresponding electrode contact. In the absence of other spectral components, zero-crossing robustly reflects the dominant instantaneous frequency. However, problems can arise in the presence of higher harmonics and noise. See, for example, WO 2010 / 085477 and Gerhard, David, Pitch extraction and fundamental frequency: History and current techniques, Regina: Department of Computer Science, University of Regina, 2003; both are incorporated herein by reference in their entirety.
[0014] Cochlear implant users, as well as general hearing aid users, are frequently exposed to a wide variety of sound environments. These environments include, for example, noisy environments, verbal conversations, a (quiet) living room, meetings, music, restaurants, cars, streets, offices, stadiums, and combinations thereof. Various techniques are known and used to classify a user's sound environment, such as Bayesian classifiers, Hidden Markov Models (HMMs), and Gaussian Mixture Models (GMMs). Based on the classified sound environment, hearing aid devices can apply parameter settings suitable for that specific sound environment, thereby improving the user's listening experience.
[0015] Figure 2A functional schematic diagram of a conventional signal processing system using Audio Scene Classification (ASC) in generating stimulus signals for a hearing implant is shown. An external signal processor 200 includes an audio scene classifier 201 that utilizes a neural network 203 configured to generate an audio scene classification output based on scene classification parameters. A processor 202 is configured to process the audio input signal and the output of the audio scene classifier 201 to generate a stimulus signal for a pulse generator 204, which is then provided to the hearing implant 205 for the patient to perceive sound. See, for example, U.S. Patent Publication 2021 / 0174824, which is incorporated herein by reference in its entirety.
[0016] Conventional methods for Audio Scene Classification (ASC) have drawbacks in that they are either unsuitable for the specific individual's environment, require expansion to previously unseen signal categories, and / or often require significant data manipulation, which typically necessitates data protection due to privacy concerns. Cochlear implant systems and other hearing aids generally have limited storage and processing capabilities. Furthermore, cochlear implant users face a wide variety of acoustic environments, which may be referred to below as audio or acoustic scenes or environments. Summary of the Invention
[0017] Embodiments of the present invention lead to a system and method adapted for audio scene classification in a hearing aid device. The method includes receiving an audio input signal at the hearing aid device; extracting at least one feature vector from the audio input signal; then processing the at least one feature vector at the hearing aid device, including using a first classifier to generate an audio scene classification output; the hearing aid device generating at least one stimulus signal based on the audio input signal and the audio scene classification output; furthermore, the hearing aid device generating a statistical aggregation of the at least one feature vector and providing the statistical aggregation of the at least one feature vector to a server; the server generating new feature vectors from the received statistical aggregation of the at least one feature vector, the generation being based on extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling from this distribution to generate extrapolated feature vectors for input to a second classifier to generate a second audio scene classification output; the server retraining the first classifier based on the output of the second audio scene classification output to generate updated parameters for the first classifier; the server providing these updated parameters to the hearing aid device, which then updates the first classifier using the updated parameters.
[0018] According to relevant embodiments of the present invention, the statistical aggregation of the at least one feature vector provided to the server can occur at random intervals, user-controlled intervals, or periodic intervals (e.g., hourly, daily, or monthly). The statistical aggregation of the at least one feature vector can also occur upon encountering a new acoustic environment. The statistical aggregation of the at least one feature vector may include standard deviation and / or mean and / or covariance matrix and / or covariance and / or higher-order statistical moments and / or average energy and / or audio scene classification output from the first classifier. The statistical aggregation of the feature vector can comply with the General Data Protection Regulation (GDPR). Training at the server can be at least partially based on extrapolated feature vectors. Extending the statistical aggregation of the at least one feature vector to a statistical distribution and sampling from this distribution to generate extrapolated feature vectors can be based on a mixture model, such as a Gaussian mixture model (GMM).
[0019] According to further related embodiments of the present invention, the hearing aid device may be a hearing aid, a middle ear implant, a bone conduction implant, or a cochlear implant. The audio scene classification output may be, by way of example and not limitation, one or more of the following audio scenes: living room, meeting room, restaurant, car, office, sports, or a combination thereof. The second classifier may be a deep neural network. The first classifier may include a neural network, a linear discriminant analysis (LDA) classifier, and / or a support vector machine (SVM). The feature vector may include Mel-frequency cepstral coefficients (MFCC), convolutional neural network (CNN) features, amplitude modulation spectrum (AMS) features, and / or low-level features such as zero-crossing rate, spectral statistics, and / or timbre.
[0020] According to another embodiment of the present invention, a hearing aid system for generating stimulus signals includes a hearing aid device having a signal processor. The signal processor is configured to: extract at least one feature vector from an audio input signal received by the hearing aid device; process the feature vector using a first classifier to generate an audio scene classification output; generate at least one stimulus signal based on the audio input signal and the audio scene classification output of the hearing aid device; generate a statistical aggregation of the feature vector; and transmit the statistical aggregation. The hearing aid system further includes a server. The server is configured to: receive the statistical aggregation of the at least one feature vector; process the statistical aggregation of the at least one feature vector, the processing of the statistical aggregation of the at least one feature vector including expanding and extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling from this distribution to generate a feature vector for extrapolation to a second classifier to generate a second audio scene classification output; train a copy of the first classifier based on the second audio scene classification output to generate updated parameters for the first classifier; and transmit the updated parameters to the hearing aid device. The signal processor is configured to update the first classifier based on the updated parameters.
[0021] According to relevant embodiments of the present invention, the statistical aggregation of the at least one feature vector provided to the server can occur at random intervals, user-controlled intervals, or periodic intervals (e.g., hourly, daily, or monthly). The statistical aggregation of the at least one feature vector can also occur upon encountering a new acoustic environment. The statistical aggregation of the at least one feature vector may include standard deviation and / or mean and / or covariance matrix and / or higher-order statistical moments and / or average energy and / or audio scene classification output from the first classifier. The statistical aggregation of the at least one feature vector may comply with the General Data Protection Regulation (GDPR). Training a copy of the first classifier at the server may additionally include, at least in part, extrapolated feature vectors. Extending the statistical aggregation of the at least one feature vector to an extrapolated statistical distribution and sampling to produce the extrapolated feature vector may include a mixture model, such as a Gaussian mixture model (GMM).
[0022] According to a further related embodiment of the present invention, the hearing aid device may be a hearing aid, a middle ear implant, or a cochlear implant. The audio scene classification output may be, by way of example and not limitation, one or more of the following audio scenes: living room, meeting room, restaurant, car, office, sports, or a combination thereof. The second classifier may be a deep neural network. The first classifier may include a neural network, a linear discriminant analysis (LDA) classifier, and / or a support vector machine (SVM). The feature vector may include Mel-frequency cepstral coefficients (MFCC), convolutional neural network (CNN) features, amplitude modulation spectrum (AMS) features, and / or low-level features such as zero-crossing rate, spectral statistics, and / or timbre.
[0023] According to another embodiment of the present invention, a hearing aid device for generating stimulus signals includes a signal processor. The signal processor is configured to: extract at least one feature vector from an audio input signal received by the hearing aid device; process the feature vector using a first neural network to generate an audio scene classification output for the hearing aid device; generate at least one stimulus signal based on the audio input signal and the audio scene classification output for the hearing aid device; generate a statistical aggregation of the at least one feature vector; provide the statistical aggregation of the at least one feature vector to a server; receive updated parameters for the first neural network from the server; and update the first neural network with the updated parameters.
[0024] According to relevant embodiments of the present invention, the signal processor can be configured to provide the statistical aggregation of the at least one feature vector at periodic intervals. The signal processor can be configured to provide the statistical aggregation of the at least one feature vector when a new acoustic environment is encountered. The statistical aggregation of the at least one feature vector may include standard deviation and / or mean and / or covariance and / or higher-order statistical moments and / or average energy and / or audio scene classification output from the first neural network. The hearing aid device may be a hearing aid, a middle ear implant, or a cochlear implant.
[0025] According to a further embodiment of the present invention, a server is provided for updating a first classifier of a hearing aid device, the first classifier generating an audio scene classification output of the hearing aid device based on at least one feature vector characterizing an audio input signal received at the hearing aid device. The server includes a server application configured to: receive a statistical aggregation of the at least one feature vector from the hearing aid device; process the statistical aggregation of the at least one feature vector, the processing of the statistical aggregation of the at least one feature vector including expanding and extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling from this distribution to generate a feature vector for extrapolation to a second classifier to generate a second audio scene classification output; train a copy of the first classifier based on the second audio scene classification output to generate parameters for updating the first classifier; and provide the updated parameters for the first classifier to the hearing aid device.
[0026] The server can be configured to utilize a mixture model, such as a Gaussian mixture model (GMM), when extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling to generate the extrapolated feature vector. The second classifier can be a deep neural network.
[0027] According to another embodiment of the present invention, a method for generating stimulus signals in a hearing aid device is provided. The method includes extracting at least one feature vector from an audio input signal received by the hearing aid device. Processing the feature vector using a first neural network at the hearing aid device to generate an audio scene classification output. Generating at least one stimulus signal at the hearing aid device based on the audio input signal and the audio scene classification output of the hearing aid device. Generating a statistical aggregation of the at least one feature vector at the hearing aid device and providing it to a server. Receiving updated parameters for the first neural network from the server by the hearing aid device. Updating the first neural network at the hearing aid device using the updated parameters.
[0028] According to relevant embodiments of the present invention, the hearing aid device may be a hearing aid, a middle ear implant, or a cochlear implant. The statistical aggregation providing the at least one feature vector may occur at periodic intervals. The statistical aggregation providing the at least one feature vector may occur upon determining that a new acoustic environment has been encountered. The statistical data may include standard deviation and / or mean.
[0029] According to another embodiment of the present invention, a method for updating a first neural network for a hearing aid device is provided. The first neural network generates an audio scene classification output for the hearing aid device based on at least one feature vector characterizing an audio input signal received at the hearing aid device. The method includes: receiving, at a server, a statistical aggregation of the at least one feature vector from the hearing aid device; processing, at the server, the statistical aggregation of the at least one feature vector, the processing of the statistical aggregation of the at least one feature vector including expanding and extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling from this distribution to generate an extrapolated feature vector for input to a second neural network to generate a second audio scene classification output; training, at the server, a copy of the first neural network based on the second audio scene classification output to generate parameters for updating the first classifier; and providing the updated parameters for the first classifier to the hearing aid device by the server.
[0030] Extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling from this distribution to generate the extrapolated feature vector may include using a mixture model, such as a Gaussian mixture model (GMM). The second neural network may be a deep neural network. Attached Figure Description
[0031] The foregoing features of the embodiments will be more readily understood by referring to the following detailed description taken in conjunction with the accompanying drawings, in which:
[0032] Figure 1 (Prior art) illustrates the anatomy of a typical human ear and components in a cochlear implant system;
[0033] Figure 2 (Prior art) illustrates a functional schematic of a conventional signal processing system that uses Audio Scene Classification (ASC) when generating stimulus signals for a hearing implant in a user's body; and
[0034] Figure 3 A functional schematic diagram of a hearing aid system with adaptive audio scene classification according to an embodiment of the present invention is shown. Detailed Implementation
[0035] In illustrative embodiments of the invention, systems and methods are provided for adapting audio scene classifiers to devices with, for example, limited computing and / or storage capabilities. For example, scene-based audio analysis / classification in hearing aids (e.g., cochlear implants) can be adapted to the user's existing environment. This is advantageously accomplished within the hearing aid with minimal computing and memory requirements, without direct user interaction. To address the issue of limited storage and processing capabilities, an external server is employed, and data communicated between the hearing aid and the server is aggregated, resulting in a low data rate between the device and the server and ensuring privacy. Details are provided below.
[0036] Figure 3 A functional schematic diagram of a hearing aid system with adaptive audio scene classification according to an embodiment of the present invention is shown. The system includes a hearing aid device 301. The hearing aid device 301 may be, but is not limited to, a hearing aid and / or a hearing implant. The hearing implant may be, but is not limited to, a cochlear implant, wherein the electrodes of a multi-channel electrode array are positioned such that they are, for example, spatially separated within the cochlea. The cochlear implant may be partially implanted and includes, but is not limited to, an external speech / signal processor, a microphone and / or a coil, and an implantable stimulator and / or electrode array. In other embodiments, the cochlear implant may be a fully implantable cochlear implant. In a further embodiment, the multi-channel electrodes may be associated with a brainstem implant, such as an auditory brainstem implant (ABI).
[0037] The hearing aid device 301 includes a signal processor 305. The signal processor 305 is configured to perform various preprocessing strategies / functions on the audio signal input 350, as well as processing steps according to embodiments of the present invention, as described above without limitation. Additionally, the signal processor 305 includes a feature extraction module 309 and a classification module 311.
[0038] The feature extraction module 309 generates a set of audio features based on the audio input 304, which may include, for example, one or more feature vectors. The feature vectors may include, for example, Mel-frequency cepstral coefficients (MFCC), convolutional neural network (CNN) features, or amplitude modulation spectrum (AMS) features. These features may be fixed and predefined (MFCC), or trained on a specific dataset to achieve maximum performance and minimum computational cost (CNN or AMS). The feature vectors may include a set of feature vectors, for example, where each feature vector is a collection generated based on temporally continuous audio inputs 304. The temporally continuous audio inputs 304 used to generate the feature vectors may at least partially overlap. In one example, the feature vectors may include 5 feature vectors, and each feature vector is generated based on audio inputs 304 lasting 2 seconds. The feature vectors in any of the foregoing examples include sufficient information from the audio signal to be reconstructed from the statistical aggregation of the at least one feature vector 317 by extrapolating an artificial feature vector (point) representing the audio signal (point) based on statistical aggregation of the at least one feature vector 317 and a given distribution model. The artificial feature vectors generated in this way statistically represent the audio signals that a user typically hears in a user-specific acoustic environment and can be used to train neural networks, but they are not audio signals themselves. The sequence of generated artificial feature vectors may not represent a sequence of feature vectors derived from environmental audio signals, but only resembles them in statistical properties and distribution.
[0039] The at least one feature vector is then passed to a classification module 311, which may be based on a linear discriminant classifier, a support vector machine (SVM), a (deep) neural network, or a combination of these or other classification techniques. In one specific embodiment, a data-driven AMS feature extractor may be combined with a (first) deep neural network classification module to output an audio scene classification output 313.
[0040] After determining the audio scene classification output 313, the hearing aid device 301 can then apply parameter settings suitable for that specific sound environment, thereby improving the user's listening experience. Illustratively, the audio input signal 304 and the output of the audio scene classifier 313 are used by the hearing aid device 301 to generate a stimulus signal, which can then be provided to the hearing implant for the patient to perceive sound.
[0041] This configuration of the feature extraction module 309 and the classification module 313 (which includes, for example, a first classifier) can be advantageously trained and fully adapted to one or more acoustic environments. In an illustrative embodiment of the invention, the classification module 311 is adapted after the system is delivered to a user of a cochlear implant (or other hearing aid) to take into account the user-specific acoustic environment encountered. For example, a user may be exposed to different environments and situations throughout the day. These different environments may include a living room (e.g., listening to the radio, watching television, etc.), driving to work (e.g., a car, the street, etc.), an office (e.g., conversation, keyboard sounds, etc.), and / or physical activity (e.g., jogging in a park, a stadium).
[0042] The extracted feature vectors can be statistically aggregated 315, but not limited to, the audio scene classification output or statistical moments from the audio scene classifier 313, such as mean, covariance, average energy, and higher-order statistical moments, and then stored in the internal memory of the hearing aid device 301 (e.g., a cochlear implant), thereby obtaining the statistical aggregation of the at least one feature vector 317. In a preferred embodiment, the statistical aggregation of the at least one feature vector 317 restarts when the audio scene classifier 313 detects a change in the encountered acoustic environment. For example, if the audio scene classifier 313 detects a change in the user's existing acoustic environment, and its output indicates the changed, dominant audio scene, this will trigger the restart of the statistical aggregation of the at least one feature vector 317. This has the advantage that the feature vectors of the statistical aggregation 315 originate from the feature vectors 317 of the same audio scene, which can be understood as being in the same region of the feature vector space.
[0043] Changes in an audio scene or environment can be detected based on a comparison of statistical aggregations of feature vectors or AMS from short and long time spans, where the long time span includes the short time span. The statistical aggregation can be an average value (μ). short ,μ long ), variance or standard deviation or covariance matrix (Σ) short ,Σ long Then changes in the audio scene can be detected, for example, when the difference between statistical aggregations from long and short time spans exceeds a certain threshold, or when the value calculated from said statistical aggregation exceeds a certain threshold. Such a calculated value can be the Mahalanobis distance and is calculated according to the following formula: Other methods for statistical aggregation of the at least one feature vector 317 that do not require detection of scene changes can also be utilized. For example, those described in “Sketching for large-scale learning of mixture models”, Nicolas Keriven, Theses at the Université de Rennes, 2017, are incorporated herein by reference. If such a method for statistical aggregation of the at least one feature vector 317 is used, the statistical aggregation of the at least one feature vector 317 does not need to be restarted.
[0044] At predefined intervals, these statistical aggregations are transmitted to and received by application 306 on server 303 via a wireless or wired connection, whereby the statistical aggregations 315 can be used to determine the fit of cochlear implant (or other hearing aid) classification module 311. The predefined intervals can be, but are not limited to, hourly, daily, or monthly. Alternatively, instead of predefined intervals, the statistical aggregations 315 can be provided to server 303 when certain conditions are met. For example, if the user desires it, or adaptively depending on the training status of classification module 311. For example, the statistical aggregations 315 can be provided to server 303 more frequently during the initial training and fitting of classification module 311, or when the user's specific acoustic environment has changed (e.g., when the user is on vacation), while during regular use, the statistical aggregations 315 can be provided to server 303 less frequently.
[0045] Sending statistical aggregations of the feature vectors to server 303, instead of the entire feature vectors themselves, advantageously reduces the vector size—making subsequent communication to server 303 data-efficient. Only statistical data of the feature vectors, such as standard deviation and mean, can be sent to the server. Since no "raw" data is sent to the server, this also means that the data on the server is anonymized or compliant with the General Data Protection Regulation (GDPR).
[0046] In various embodiments, server 303 performs processing, such as statistical modeling 319, on the received statistical aggregation of the at least one feature vector 317. For example, server 303 may be configured to reconstruct the received statistical aggregation of the at least one feature vector 317 based on the statistical aggregation of the at least one feature vector 317 and a given distribution model extrapolation point. The given distribution model may be, but is not limited to, a Gaussian or heavy-tailed distribution model, a mixture model such as a Gaussian mixture model (GMM), or a neural network-based model. A Gaussian mixture model (GMM) can be expressed by the following formula: Where x represents the feature vector, p(x) represents the probability of a specific feature vector x, and μ i Σ represents the average eigenvector. i Let α represent the covariance matrix. i This represents the weight of each of the N components (dimensions). Let represent the probability density function of a normal distribution, and let the weight α be... i It can be normalized, that is Statistical sampling can be performed on any of these models to generate a specific instance of an acoustic feature set 321 representing audio features of the captured acoustic environment from the user.
[0047] These artificially generated training samples can then be fed into and labeled by a precise, cloud-based or server-based classifier 323 at server 303. In an illustrative embodiment, the classifier 323 may include a deep neural network. This deep neural network may have millions of parameters and is advantageously more powerful than the acoustic scene classifier 311 located in the hearing aid device 301 (which has limited resources due to space and power constraints).
[0048] The output of the cloud- or server-based classifier 323 (i.e., the second audio scene classification output) can be a labeled dataset, which is then combined with an existing training foundation to train a copy of the first classifier 325 embedded at server 303. Subsequently, one or more modules of the embedded copy of the first classifier 325 are retrained, and the updated classifier parameters 327 are transmitted back to the user's hearing aid device 301. The signal processor 305 is configured to update the first classifier 311 based on the updated parameters 327 received from server 303.
[0049] After retraining, the user's audio classifier 311 improves its ability to classify sounds within previously captured user-specific contexts (e.g., living room, car, office, and sports). As mentioned above, this process can be repeated whenever a specific new environment is encountered, or at a predetermined time. Changes in the user's environment can be monitored using additional data from the positioning system, or by comparing the data aggregated in the hearing aid device 301 with data available at a cloud-based or server-based server 303.
[0050] Embodiments of the present invention can be implemented in part in any conventional computer programming language. For example, preferred embodiments can be implemented in a procedural programming language (e.g., "C") or an object-oriented programming language (e.g., "C++", Python). Alternative embodiments of the invention can be implemented as pre-programmed hardware elements, other related components, or a combination of hardware and software components.
[0051] The embodiments can be implemented in part as a computer program product for a computer system. Such implementation may include a set of computer instructions fixed on a tangible medium (e.g., a computer-readable medium, such as a floppy disk, CD-ROM, ROM, or fixed disk) or transferable to the computer system via a medium through a modem or other interface device (e.g., a communication adapter connected to a network). The medium may be a tangible medium (e.g., an optical or analog communication line) or a medium implemented using wireless technology (e.g., microwave, infrared, or other transmission technologies). The set of computer instructions embodies all or part of the functionality previously described herein with respect to the system. Those skilled in the art will understand that such computer instructions can be written in various programming languages for use with many computer architectures or operating systems. Furthermore, such instructions can be stored in any storage device, such as semiconductor, magnetic, optical, or other storage devices, and can be transmitted using any communication technology (e.g., optical, infrared, microwave, or other transmission technologies). Such computer program products are intended to be distributed as removable media along with accompanying printed or electronic documentation (e.g., shrink wrapping software), pre-installed on computer systems (e.g., on system ROM or a fixed disk), or distributed from servers or electronic bulletin boards via networks (e.g., the Internet or the World Wide Web). Of course, some embodiments of the invention can be implemented as a combination of both software (e.g., computer program products) and hardware. Other embodiments of the invention are implemented entirely as hardware or entirely as software (e.g., computer program products).
[0052] Although various exemplary embodiments of the present invention have been disclosed, it will be apparent to those skilled in the art that various changes and modifications can be made to achieve some of the advantages of the present invention without departing from the true scope of the invention.
Claims
1. A method for audio scene classification adapted to hearing aids, characterized in that, The method includes: The hearing aid receives an audio input signal. At the hearing aid device, at least one feature vector is extracted from the audio input signal; At the hearing aid device, the at least one feature vector is processed, and the processing at the hearing aid device includes using a first classifier to generate an audio scene classification output; At the hearing aid device, at least one stimulus signal is generated based on the audio input signal and the audio scene classification output; At the hearing aid device, a statistical aggregation of the at least one feature vector is generated; The hearing aid device provides the server with the statistical aggregation of the at least one feature vector; At the server, the statistical aggregation of the at least one feature vector is processed, and the processing at the server includes expanding and extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling from the distribution to generate an extrapolated feature vector for input to a second classifier to generate a second audio scene classification output; At the server, a copy of the first classifier is trained based on the second audio scene classification output to generate parameters for updating the first classifier; The server provides the updated parameters to the hearing aid device; and At the hearing aid device, the first classifier is updated with the updated parameters.
2. The method according to claim 1, characterized in that, The statistical aggregation of the feature vectors is provided at periodic intervals.
3. The method according to claim 2, characterized in that, The periodic interval is either daily or monthly.
4. The method according to claim 1, characterized in that, The statistical aggregation of the at least one feature vector is provided when a new acoustic environment is encountered.
5. The method according to claim 1, characterized in that, The hearing aid device is one of a hearing aid, a middle ear implant, a bone conduction implant, and a cochlear implant.
6. The method according to claim 1, characterized in that, The statistical aggregation of the at least one feature vector includes standard deviation and / or mean and / or covariance and / or higher-order statistical moments and / or average energy and / or audio scene classification output from the first classifier.
7. The method according to claim 1, characterized in that, The statistical aggregation of the feature vectors complies with the General Data Protection Regulation (GDPR).
8. The method according to claim 1, characterized in that, The training at the server is also based, at least in part, on extrapolated feature vectors.
9. The method according to claim 1, characterized in that, The statistical aggregation extension extrapolation of the at least one feature vector to a statistical distribution and sampling from the distribution to produce the extrapolated feature vector is based on a mixture model, such as a Gaussian mixture model (GMM).
10. The method according to claim 1, characterized in that, The audio scene classification output is an audio scene selected from a group of audio scenes consisting of living room, meeting room, restaurant, car, office, sports, and combinations thereof.
11. The method according to claim 1, characterized in that, The second classifier is a deep neural network.
12. The method according to claim 1, characterized in that, The first classifier includes a linear discriminant analysis (LDA) classifier and / or a support vector machine (SVM) and / or a neural network.
13. The method according to claim 1, characterized in that, The feature vectors include Mel frequency cepstral coefficients (MFCC), convolutional neural network (CNN) features, or amplitude modulated spectrum (AMS) features.
14. A hearing aid system for generating stimulus signals, characterized in that, include: Hearing aid device, including a signal processor, the signal processor being configured to: Extract at least one feature vector from the audio input signal received by the hearing aid device; The first neural network is used to process the at least one feature vector to generate an audio scene classification output; At least one stimulus signal is generated based on the audio input signal and the audio scene classification output of the hearing aid device; Statistical aggregation of the at least one feature vector is generated; as well as Transmit the statistical aggregation; as well as The server is configured as follows: The statistical aggregation received from the at least one feature vector; The processing of the statistical aggregation of the at least one feature vector includes expanding and extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling from the distribution to generate an extrapolated feature vector for input to a second neural network to generate a second audio scene classification output; A copy of the first classifier is trained based on the second audio scene classification output to generate parameters for updating the first classifier; as well as The updated parameters are transmitted to the hearing aid device. The signal processor is configured to update the first classifier based on the updated parameters.
15. The hearing aid system according to claim 14, characterized in that, The hearing aid device is one of a hearing aid, a middle ear implant, and a cochlear implant.
16. The hearing aid system according to claim 14, characterized in that, The signal processor is configured to provide the statistical aggregation of the at least one feature vector at periodic intervals.
17. The hearing aid system according to claim 14, characterized in that, The signal processor is configured to provide the statistical aggregation of the at least one feature vector when a new acoustic environment is encountered.
18. The hearing aid system according to claim 14, characterized in that, The statistical aggregation of the at least one feature vector includes standard deviation and / or mean and / or covariance and / or higher-order statistical moments and / or average energy and / or audio scene classification output from the first classifier.
19. The hearing aid system according to claim 14, characterized in that, The statistical aggregation of the at least one feature vector complies with the General Data Protection Regulation (GDPR).
20. The hearing aid system according to claim 14, characterized in that, The server is configured to train a copy of the first classifier, at least in part, based on extrapolated feature vectors.
21. The hearing aid system according to claim 14, characterized in that, The server is configured to utilize a mixture model, such as a Gaussian mixture model (GMM), when extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling to generate the extrapolated feature vector.
22. The hearing aid system according to claim 14, characterized in that, The audio scene classification output is an audio scene selected from a group of audio scenes consisting of living room, meeting room, restaurant, car, office, sports, and combinations thereof.
23. The hearing aid system according to claim 14, characterized in that, The second neural network is a deep neural network.
24. The hearing aid system according to claim 14, characterized in that, The first neural network includes a linear discriminant analysis (LDA) classifier and / or a support vector machine (SVM).
25. The hearing aid system according to claim 14, characterized in that, The feature vectors include Mel frequency cepstral coefficients (MFCC), convolutional neural network (CNN) features, or amplitude modulated spectrum (AMS) features.
26. A hearing aid device for generating stimulus signals, characterized in that, The hearing aid device includes: The signal processor is configured to: Extract at least one feature vector from the audio input signal received by the hearing aid device; The first neural network is used to process the at least one feature vector to generate an audio scene classification output; At least one stimulus signal is generated based on the audio input signal and the audio scene classification output of the hearing aid device; Statistical aggregation of the at least one feature vector is generated; Provide the server with statistical aggregation of the at least one feature vector; Receive parameters from the server for updating the first classifier; and Update the first classifier with the updated parameters.
27. The hearing aid device according to claim 26, characterized in that, The hearing aid device is one of a hearing aid, a middle ear implant, and a cochlear implant.
28. The hearing aid device according to claim 26, characterized in that, The signal processor is configured to provide the statistical aggregation of the at least one feature vector at periodic intervals.
29. The hearing aid device according to claim 26, characterized in that, The signal processor is configured to provide the statistical aggregation of the at least one feature vector when a new acoustic environment is encountered.
30. The hearing aid device according to claim 26, characterized in that, The statistical data includes standard deviation and / or mean and / or covariance and / or higher-order statistical moments and / or average energy and / or audio scene classification output from the first neural network.
31. A server for updating a first classifier of a hearing aid device, characterized in that, The first classifier generates an audio scene classification output for the hearing aid based on a feature vector characterizing the audio input signal received at the hearing aid, and the server includes: The server application is configured as follows: Statistical aggregation of the at least one feature vector received from the hearing aid device; The processing of the statistical aggregation of the feature vectors includes expanding and extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling from the distribution to generate an extrapolated feature vector for input to a second neural network to generate a second audio scene classification output; A copy of the first classifier is trained based on the second audio scene classification output to generate parameters for updating the first classifier; and The hearing aid device is provided with the updated parameters for the first classifier.
32. The server according to claim 31, characterized in that, The server is configured to utilize a mixture model, such as a Gaussian mixture model (GMM), when extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling to generate the extrapolated feature vector.
33. The server according to claim 31, characterized in that, The second classifier is a deep neural network.
34. A method for generating stimulation signals in a hearing aid device, characterized in that, The method includes: At the hearing aid device, at least one feature vector is extracted from the audio input signal received by the hearing aid device; at the hearing aid device, the at least one feature vector is processed using a first classifier to generate an audio scene classification output; At the hearing aid device, at least one stimulus signal is generated based on the audio input signal and the audio scene classification output of the hearing aid device; At the hearing aid device, a statistical aggregation of the at least one feature vector is generated; Statistical aggregation of the at least one feature vector is provided from the hearing aid device to the server; The hearing aid device receives parameters from the server for updating the first classifier; and At the hearing aid device, the first classifier is updated with the updated parameters.
35. The method according to claim 34, characterized in that, The hearing aid device is one of a hearing aid, a middle ear implant, and a cochlear implant.
36. The method according to claim 34, characterized in that, The statistical aggregation of the at least one feature vector is provided at periodic intervals.
37. The method according to claim 34, characterized in that, The statistical aggregation of the feature vectors is provided when a new acoustic environment is encountered.
38. The method according to claim 34, characterized in that, The statistical data includes standard deviation and / or mean.
39. A method for updating a first neural network in a hearing aid device, characterized in that, The first classifier generates an audio scene classification output for the hearing aid based on at least one feature vector characterizing the audio input signal received at the hearing aid, the method comprising: At the server, statistical aggregation of the at least one feature vector is received from the hearing aid device; At the server, the statistical aggregation of the at least one feature vector is processed, the processing of the statistical aggregation of the at least one feature vector includes expanding the statistical aggregation of the at least one feature vector to an extrapolation to a statistical distribution and sampling from the distribution to generate an extrapolation feature vector for input to a classifier to generate a second audio scene classification output; At the server, a copy of the first neural network is trained based on the second audio scene classification output to generate parameters for updating the first classifier; and The server provides the hearing aid device with the updated parameters for the first classifier.
40. The method according to claim 39, characterized in that, Extrapolating the statistical aggregation of the at least one feature vector to a statistical distribution and sampling from the distribution to produce the extrapolated feature vector includes using a mixture model, such as a Gaussian mixture model (GMM).
41. The method according to claim 39, characterized in that, The second classifier is a deep neural network.
Citation Information
Patent Citations
Neural Network Audio Scene Classifier for Hearing Implants
US20210174824A1
Magnetic-interference-free surgical prostheses
US6348070B1
Modulation depth enhancement for tone perception
US7561709B2
High accuracy tonotopic and periodic coding with enhanced harmonic resolution
WO2010085477A1