System and method for in-ear microphone speech signal processing
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- EERS GLOBAL TECH
- Filing Date
- 2024-06-13
- Publication Date
- 2026-04-22
AI Technical Summary
Existing techniques for improving the signal-to-noise ratio of speech signals captured by in-ear microphones in noisy settings are complex and expensive, often limiting the frequency bandwidth and reducing signal quality and intelligibility.
A method and system for speech signal processing that involves obtaining and decomposing in-ear microphone speech signals into frequency bands, determining time-varying corrective gains based on feature differences with reference signals, and applying these gains to enhance the signal quality through a data-driven model, including pre-processing, voice activity detection, and feature extraction using techniques like FFT and machine learning models.
This approach improves the quality of in-ear microphone speech signals by adapting to various speech content in real-time, effectively matching air-conducted speech and enhancing signal quality in noisy environments.
Smart Images

Figure CA2024050800_19122024_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR IN-EAR MICROPHONE SPEECH SIGNAL PROCESSINGCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims the benefit of United States Provisional Patent Application No. 63 / 472,808 filed on June 13, 2023, the contents of which are hereby incorporated by reference.FIELD
[0002] The improvements generally relate to the field of in-ear microphone speech signal processing, and more specifically to in-ear microphone speech signal processing using bandwidth extension.BACKGROUND
[0003] In order to improve the signal-to-noise ratio (SNR) of speech signals captured in noisy settings, existing techniques capture speech through bone and tissue vibrations. This may be achieved using microphones placed inside an occluded ear (referred to as “Inner-Ear Microphones” or IEM) or through bone conduction sensors placed on the cranium. However, the frequency bandwidth of the signals captured using these techniques is typically limited, which in turn reduces signal quality and intelligibility. Although several techniques have been developed for the bandwidth extension of bone and tissue conducted (BTC) speech, and particularly of signals captured by an IEM in noisy conditions, these techniques generally prove complex and expensive.
[0004] Therefore, there is room for improvement.SUMMARY
[0005] In accordance with one aspect, there is provided a method for speech signal processing. The method comprises, at a computing device, obtaining a first in-ear microphone (IEM) speech signal and a reference speech signal, decomposing the first IEM speech signal into a plurality of first IEM speech signal components, and the reference speech signal into a plurality of reference speech signal components, the first IEM speech signal components and the reference speech signal components spanning a plurality of frequency bands, determining, for each frequency band, a plurality of time-varying ground truth corrective gains based on adifference between corresponding ones of a first plurality of features associated with the first IEM speech signal components and the reference speech signal components, the first plurality of features comprising at least one of first temporal features and first spectral features, determining a second plurality of time-varying features of the first IEM speech signal, the second plurality of features comprising at least one of second temporal features and second spectral features, configuring at least one model to predict, from the second plurality of features, a plurality of time-varying corrective gains that estimate the plurality of ground truth corrective gains, obtaining a second IEM speech signal from an lEM-equipped device, decomposing the second IEM speech signal into a plurality of second IEM speech signal components spanning the plurality of frequency bands, determining a third plurality of timevarying features of the second IEM speech signal, the third plurality of features matching the second plurality of features, executing the at least one model for predicting the plurality of corrective gains based on the third plurality of features, applying the plurality of corrective gains to the second IEM speech signal to obtain a processed IEM speech signal, and outputting the processed IEM speech signal.
[0006] In at least embodiment in accordance with any previous / other embodiment described herein, the method further comprises pre-processing each of the first IEM speech signal, the reference speech signal, and the second IEM speech signal prior to the decomposing.
[0007] In at least embodiment in accordance with any previous / other embodiment described herein, the method further comprises applying, prior to decomposing the second IEM speech signal, voice activity detection (VAD) to the second IEM speech signal to detect a presence or an absence of voice activity in the second IEM speech signal, where the plurality of corrective gains are applied to speech frames obtained based on applying the VAD to the second IEM speech signal, the speech frames indicative of the presence of voice activity in the second IEM speech signal.
[0008] In at least embodiment in accordance with any previous / other embodiment described herein, determining each of the second plurality of features and the third plurality of features comprises determining at least one of a pitch, a harmonic ratio, an energy, a transience, a zero-crossing rate, a spectral skewness, a spectral entropy, a spectral centroid, a spectral kurtosis, a spectral roll-off point, a spectral slope, a spectral crest, a spectral decrease, a spectral flux, a spectral flatness, a spectral spread, and a plurality of Fast Fourier Transform(FFT) bins obtained by applying a FFT to windowed frames of a respective one of the first IEM speech signal and the second IEM speech signal.
[0009] In at least embodiment in accordance with any previous / other embodiment described herein, the second plurality of features and the third plurality of features are determined based on at least one of: a respective one of the first IEM speech signal components and the second IEM speech signal components, and a full bandwidth of a respective one of the first IEM speech signal and the second IEM speech signal.
[0010] In at least embodiment in accordance with any previous / other embodiment described herein, decomposing each of the first IEM speech signal and the second IEM speech signal comprises applying a filter to a respective one of the first IEM speech signal and the second IEM speech signal, the filter spanning a frequency spectrum including the plurality of frequency bands.
[0011] In at least embodiment in accordance with any previous / other embodiment described herein, the plurality of time-varying ground truth corrective gains are determined based on at least one of an energy difference between the first IEM speech signal components and the reference speech signal components, an amplitude difference between the first IEM speech signal components and the reference speech signal components, an amplitude and phase difference between the first IEM speech signal components and the reference speech signal components, and an energy difference between consecutive frames of the first IEM speech signal components and the reference speech signal components.
[0012] In at least embodiment in accordance with any previous / other embodiment described herein, each of the first IEM speech signal and the reference speech signal is obtained from a dataset generated based on at least one multi-subject recording.
[0013] In at least embodiment in accordance with any previous / other embodiment described herein, each of the first IEM speech signal and the reference speech signal is obtained from at least one field recording, the first IEM speech signal and the reference speech signal being respectively captured by an IEM and an OEM in response to the IEM and the OEM detecting disturbance-free user speech.
[0014] In at least embodiment in accordance with any previous / other embodiment described herein, the method further comprises fine-tuning the at least one model based on a differencebetween corresponding features of the second IEM speech signal and the reference speech signal, where each of the second IEM speech signal and the reference speech signal is obtained from at least one field recording, the second IEM speech signal and the reference speech signal being respectively captured by an IEM and an OEM in response to the IEM and the OEM detecting disturbance-free user speech.
[0015] In at least embodiment in accordance with any previous / other embodiment described herein, configuring the at least one model comprises configuring a plurality of models, and executing the at least one model comprises selecting a given model among the plurality of models based on the third plurality of features and executing the given model.
[0016] In at least embodiment in accordance with any previous / other embodiment described herein, configuring the at least one model comprises one of training at least one machine learning model and defining at least one empirical model based on a set of empirical rules.
[0017] In accordance with another aspect, there is provided a system for speech signal processing. The system comprises a processing unit, and a non-transitory computer-readable medium having stored thereon program instructions executable by the processing unit for obtaining a first in-ear microphone (IEM) speech signal and a reference speech signal, decomposing the first IEM speech signal into a plurality of first IEM speech signal components, and the reference speech signal into a plurality of reference speech signal components, the first IEM speech signal components and the reference speech signal components spanning a plurality of frequency bands, determining, for each frequency band, a plurality of time-varying ground truth corrective gains based on a difference between corresponding ones of a first plurality of features associated with the first IEM speech signal components and the reference speech signal components, the first plurality of features comprising at least one of first temporal features and first spectral features, determining a second plurality of time-varying features of the first IEM speech signal, the second plurality of features comprising at least one of second temporal features and second spectral features, configuring at least one model to predict, from the second plurality of features, a plurality of time-varying corrective gains that estimate the plurality of ground truth corrective gains, obtaining a second IEM speech signal from an lEM-equipped device, decomposing the second IEM speech signal into a plurality of second IEM speech signal components spanning the plurality of frequency bands, determining a third plurality of time-varying features of the second IEM speech signal, the third plurality offeatures matching the second plurality of features, executing the at least one model for predicting the plurality of corrective gains based on the third plurality of features, applying the plurality of corrective gains to the second IEM speech signal to obtain a processed IEM speech signal, and outputting the processed IEM speech signal.
[0018] In at least embodiment in accordance with any previous / other embodiment described herein, the program instructions are further executable by the processing unit for applying, prior to decomposing the second IEM speech signal, voice activity detection (VAD) to the second IEM speech signal to detect a presence or an absence of voice activity in the second IEM speech signal, and the program instructions are further executable by the processing unit for applying the plurality of corrective gains to speech frames obtained based on the applying the VAD to the second IEM speech signal, the speech frames indicative of the presence of voice activity in the second IEM speech signal.
[0019] In at least embodiment in accordance with any previous / other embodiment described herein, the program instructions are executable by the processing unit for determining each of the second plurality of features and the third plurality of features comprising determining at least one of a pitch, a harmonic ratio, an energy, a transience, a zero-crossing rate, a spectral skewness, a spectral entropy, a spectral centroid, a spectral kurtosis, a spectral roll-off point, a spectral slope, a spectral crest, a spectral decrease, a spectral flux, a spectral flatness, a spectral spread, and a plurality of Fast Fourier Transform (FFT) bins obtained by applying a FFT to windowed frames of a respective one of the first IEM speech signal and the second IEM speech signal.
[0020] In at least embodiment in accordance with any previous / other embodiment described herein, the program instructions are executable by the processing unit for determining the plurality of time-varying ground truth corrective gains based on at least one of an energy difference between the first IEM speech signal components and the reference speech signal components, an amplitude difference between the first IEM speech signal components and the reference speech signal components, an amplitude and phase difference between the first IEM speech signal components and the reference speech signal components, and an energy difference between consecutive frames of the first IEM speech signal components and the reference speech signal components.
[0021] In at least embodiment in accordance with any previous / other embodiment described herein, each of the first IEM speech signal and the reference speech signal is obtained from one of: a dataset generated based on at least one multi-subject recording, and at least one field recording, with the first IEM speech signal and the reference speech signal being respectively captured during the at least one field recording by an IEM and an OEM in response to the IEM and the OEM detecting disturbance-free user speech.
[0022] In at least embodiment in accordance with any previous / other embodiment described herein, the program instructions are further executable by the processing unit for fine-tuning the at least one model based on a difference between corresponding features of the second IEM speech signal and the reference speech signal, where each of the second IEM speech signal and the reference speech signal is obtained from at least one field recording, the second IEM speech signal and the reference speech signal being respectively captured by an IEM and an OEM in response to the IEM and the OEM detecting disturbance-free user speech.
[0023] In at least embodiment in accordance with any previous / other embodiment described herein, the program instructions are executable by the processing unit for configuring the at least one model comprising configuring a plurality of models, and the program are instructions executable by the processing unit for executing the at least one model comprising selecting a given model among the plurality of models based on the third plurality of features and executing the given model.
[0024] In accordance with another aspect, there is provided a non-transitory computer readable medium having stored thereon program code executable by at least one processor for obtaining a first in-ear microphone (IEM) speech signal and a reference speech signal, decomposing the first IEM speech signal into a plurality of first IEM speech signal components, and the reference speech signal into a plurality of reference speech signal components, the first IEM speech signal components and the reference speech signal components spanning a plurality of frequency bands, determining, for each frequency band, a plurality of time-varying ground truth corrective gains based on a difference between corresponding ones of a first plurality of features associated with the first IEM speech signal components and the reference speech signal components, the first plurality of features comprising at least one of first temporal features and first spectral features, determining a second plurality of time-varying features of the first IEM speech signal, the second plurality of features comprising at least oneof second temporal features and second spectral features, configuring at least one model to predict, from the second plurality of features, a plurality of time-varying corrective gains that estimate the plurality of ground truth corrective gains, obtaining a second IEM speech signal from an lEM-equipped device, decomposing the second IEM speech signal into a plurality of second IEM speech signal components spanning the plurality of frequency bands, determining a third plurality of time-varying features of the second IEM speech signal, the third plurality of features matching the second plurality of features, executing the at least one model for predicting the plurality of corrective gains based on the third plurality of features, applying the plurality of corrective gains to the second IEM speech signal to obtain a processed IEM speech signal, and outputting the processed IEM speech signal.
[0025] Many further features and combinations thereof concerning embodiments described herein will appear to those skilled in the art following a reading of the instant disclosure.DESCRIPTION OF THE FIGURES
[0026] In the figures,
[0027] Fig. 1A is a schematic diagram of an example system for detecting the voice of a user, in accordance with one embodiment;
[0028] Fig. 1 B is a block diagram of the signal processing unit of Fig. 1A, in accordance with one embodiment;
[0029] Fig. 2A is a plot of individual IEM speech signals and a static filter implemented to pre- process the IEM speech signals, in accordance with one embodiment;
[0030] Fig. 2B is a plot of the static filter of Fig. 2A and a matched filter derived based on the static filter, in accordance with one embodiment;
[0031] Fig. 3 is a plot illustrating voice-activity detection (VAD)-based denoising applied to a noisy IEM speech signal, in accordance with one embodiment;
[0032] Fig. 4A is a block diagram of the features extraction module of Fig. 1 B, in accordance with one embodiment;
[0033] Fig. 4B illustrates a brute force regression feature technique implemented by the features extraction module of Fig. 4A, in accordance with one embodiment;
[0034] Fig. 4C illustrates a ranking technique implemented by the features extraction module of Fig. 4A, in accordance with one embodiment;
[0035] Fig. 5A is a plot illustrating the magnitude response of the cross-over filter of Fig. 1 B, in accordance with one embodiment;
[0036] Fig. 5B is a plot illustrating the output of the cross-over filter of Fig. 1 B, in accordance with one embodiment;
[0037] Fig. 6 are plots illustrating filtering model fits across four frequency bands for the filtering model of Fig. 1 B, in accordance with one embodiment;
[0038] Fig. 7A illustrates the variation over time of corrective gains for four frequency bands, in accordance with one embodiment;
[0039] Fig. 7B illustrates band-wise corrective gains applied to an IEM speech signal for four frequency bands, in accordance with one embodiment;
[0040] Fig. 8 illustrates an example neural network architecture for use with the system of Fig. 1A, in accordance with one embodiment;
[0041] Figs. 9A and 9B are flowcharts of a method for speech signal processing, in accordance with one embodiment; and
[0042] Fig. 10 is a block diagram illustrating an example computing device, in accordance with one embodiment.
[0043] It will be noted that throughout the appended drawings, like features are identified by like reference numerals.DETAILED DESCRIPTION
[0044] Described herein are systems and methods for in-ear microphone (IEM) speech signal processing. The systems and methods described herein may be used to improve the quality of IEM speech signals, particularly in noisy environments, via bandwidth extension. For this purpose and as will be described further below, a data-driven model is used in which features are extracted from an IEM speech signal and used to predict band-wise corrective gains across time frames. The time-varying corrective gains are used to filter the IEM speech signal and achieve spectral characteristics of natural speech for various speech content across time.In this manner, it becomes possible to truly match air-conducted speech by adapting to the various speech content in real-time.
[0045] Fig. 1A shows an example of a system 10 for detecting the voice of a user, in accordance with one embodiment. The system 10 comprises an in-ear microphone (IEM) 12 that is in fluid communication with an outer ear canal 14 of an ear 16 of the user. The ear 16 is occluded from an environment outside thereof via an lEM-equipped device, e.g. an intra- aural device 18 which comprises the IEM 12. In operation, the IEM 12 picks up speech generated from bone and tissue conduction of the user and generates a speech signal (referred to herein as an “IEM speech signal”). The IEM speech signal comprises several speech fragments (referred to herein as “frames”) having a given duration. In one embodiment, the IEM 12 is as described in U.S. Patent No. 10,783,904 filed on November 6, 2018, the entire contents of which are incorporated herein by reference. It should however be understood that other embodiments may apply. Although an intra-aural device 18 is shown in Fig. 1A, any suitable type of hearing protection device that provides the required occlusion and enables the IEM 12 to capture a signal from inside the occlusion may apply. For example, an earmuff, an extra-aural device, an over the ear device, or the like, may apply.
[0046] The system 10 further comprises a signal processing unit 100 that is operatively connected to the IEM 12 to receive the IEM speech signal therefrom. As will be described further below, the signal processing unit 100 is configured to process the IEM speech signal to improve the quality thereof. Although illustrated as separate (i.e. located away) from the IEM 12, it should be understood that the signal processing unit 100 may be integral to (i.e. embedded into) the IEM 12.
[0047] Fig. 1 B shows an example of the signal processing unit 100, in accordance with one embodiment. The signal processing unit 100 comprises a plurality of interconnected components, namely a pre-processing module 102, a features extraction module 104, a crossover filter 106, a data-driven filtering model 108, an interpolation / smoothing module 1 10, a multiplier 112, and a summation block 114.
[0048] The IEM speech frames captured by the IEM (reference 12 in Fig. 1 A) are sent to the pre-processing module 102 where the frames are pre-processed. As will be described further below, the pre-processing performed by the pre-processing module 102 may comprise, but isnot limited to, dynamic compression / expansion, static filtering, and / or voice-activity detection (VAD)-based denoising. For example, adaptive filtering-based denoising of the IEM signal can be performed using additional microphones (not shown), such as outer ear microphones (OEMs) or in-ear accelerometers. As will be described further below, a static filter may also be applied to the speech signal (obtained by estimating a transfer function using the IEM signal and a reference signal). As will also be described further below, a voice activity detector (VAD) may be used to obtain only speech frames and the decisions of the VAD may be used for further denoising and model training. Further denoising of the IEM signal may be used to eliminate any residual noise. Other examples of pre-processing include, but are not limited, automatic gain control, transient boosting, harmonic excitations, transience enhancers, amplitude modulation, bandwidth extension, side-chaining, noise-gating, and other modulation, time-based, spectral and dynamic processing methods.
[0049] Referring to Figs. 2A and 2B in addition to Fig. 1 B, in one embodiment, the preprocessing module 102 comprises a static filter used to pre-process the IEM speech frames. Using the static filter, an overall (i.e. general) correction can be applied to the IEM speech frames. In one embodiment, the static filter is a finite impulse response (FIR) filter. The static filter is derived based on the mean spectral difference between a reference signal (e.g., captured using a reference microphone) and the full-bandwidth IEM speech signals of a population of subjects. In one embodiment, the static filter is derived using the pwelch method in which the power spectral density of the IEM speech signal is computed (using Welch’s method) after the IEM speech signal is divided into windowed segments of specific length and overlap factor. It should however be understood that other suitable techniques for deriving the means spectral difference, and accordingly deriving the static filter, may apply.
[0050] Fig. 2A illustrates a plot 200 showing individual IEM speech signals (illustrated as curves 202i, 2022, 202s, 2024, and 202s) obtained from five (5) subject speakers. Plot 200 also shows the static filter (illustrated as curve 204) obtained by computing the mean spectral difference between a reference signal (not shown) and the IEM speech signals 202i, 2022, 202s, 2O , and 2025. In one embodiment, in order to achieve cost-efficient filtering targeted for hardware platforms, the static filter may be matched (e.g., using a curve fitting process) with a cascade of second-order infinite impulse response (HR) filters (referred to herein as “cascaded biquad filters”) to generate a resulting matched HR filter. An optimization step maybe performed to match the cascaded biquad filters to the static filter. Fig. 2B illustrates a plot 210 of the static filter (curve 204) and the resulting matching filter (curve 206).
[0051] Referring now to Fig. 3 in addition to Fig. 1 B, in one embodiment, the pre-processing module 102 pre-processes the IEM speech frames using a technique referred to herein as a “VAD-aware denoising”. In particular, VAD is used to indicate the presence or absence of voice activity in the IEM speech signal (see plot 302 in Fig. 3). It thus becomes possible to ensure that the systems and methods described herein act only when the IEM speech signal contains voice. Using the VAD-aware denoising process, any residual noise post adaptive filtering-based denoising of the IEM speech signal using OEMs can therefore be reduced while ensuring that the speech content in the IEM speech signal remains unaffected. Implementation of VAD by the pre-processing module 102 also ensures that the filtering model 108 is bypassed during silences and that the time-varying corrective gains are applied only on speech frames, thus avoiding unwanted amplification of noise signals.
[0052] In one embodiment, the VAD-aware denoising step is based on Wiener filtering which is based on adapting and learning noise statistics of non-speech frames followed by application of the denoising on the subsequent noisy speech frames. IEM speech is analyzed in a frame-wise manner in which noise statistics are derived from accumulated non-speech (or “VAD-OFF”) frames. At the onset of a speech (or “VAD-ON”) frame, the noise statistics derived from the latest frame of N samples, typically 0.3 seconds worth of samples (see plot 304, bars 306), are analyzed to denoise the VAD-ON frame samples, until the next VAD-OFF. The resulting denoised signal is illustrated in plot 308 of Fig. 3.
[0053] In some embodiments, several VAD techniques may be implemented and optimized using a principle based on tracking a ratio of low and high band energy. When the ratio is above a given threshold, the speech frame is determined to be a VAD-ON frame. The threshold may be adjusted using a 1 -euro filter to account for varying SNR of user speech with respect to noise floor and residual ambient noise. An effective VAD is an additional layer which works with a certain persistence / inertia to avoid momentary switching of VAD on / off decisions which could result in poor denoising, choppy speech and affect perception as well as listening comfort.
[0054] Referring back to Fig. 1 B, the output of the pre-processing module 102 is a pre- processed IEM speech signal that is fed to the features extraction module 104. The pre- processed IEM speech signal is further separated (or decomposed) to generate filtered version of the pre-processed IEM speech signal, as will be described further below. This may be achieved using the cross-over filter 106. For this purpose, the output of the pre-processing module 102 (i.e. the pre-processed IEM speech signal) is also fed to the cross-over filter 106, which is configured to operate in parallel with the features extraction module 104.
[0055] In parallel with the cross-over filter 106, the features extraction module 104 determines one or more temporal and spectral features from the pre-processed IEM speech signal. Fig. 4A shows a block diagram of the features extraction module 104, in accordance with one embodiment. In operation, an input audio signal stream (labelled “audioIn” in Fig. 4A) is fed into a circular buffer 402. In one embodiment, the features extraction module 104 is configured to determine the features band-wise and for the full-bandwidth of the IEM speech signal, such that the input audio signal stream may comprise full bandwidth and band-wise signals. The circular buffer 402 is then configured to capture (i.e. buffer) the pre-processed IEM speech signal in a frame-wise manner (e.g., using Short Time Fourier Transform, STFT) with a chosen window and hop size (i.e. the number of samples between each successive FFT window). A Fast Fourier Transform (FFT) block 404 is then used to apply FFT on the windowed frames and FFT bins are output. In the embodiment of Fig. 4A, four (4) FFT bins, labelled “linearSpectrum”, “melSpectrum”, “barkSpectrum”, and “erbSpectrum”, are output. It should be understood that this is for illustrative purposes and any suitable number of FFT bins may apply. The length of the FFT (labelled “FFTLength”) may also vary depending on the application. The FFT bins are then grouped into bands based on a spectral descriptor input (labelled “SpectralDescriptorlnput” in Fig. 4A) and features are computed.
[0056] In some embodiments and as previously noted, the features extraction module 104 may be configured to compute the full-bandwidth features (e.g., energy, spectral entropy, etc.) and the band-wise features (e.g., energy for a given frequency band, spectral entropy for a given frequency band, etc.). The features extraction module 104 may determine the bandwise features based on the full bandwidth signal received as input (e.g., as audioIn), with the FFT block 404 being used to group the FFT bins to obtain band-wise features. A single FFT block may thus be used. Alternatively, the features extraction module 104 may determine theband-wise features based on band-wise signals (i.e. bandlimited audio for each band) received as input (e.g., as audioIn) and FFT may be carried out by one or more of the FFT block 404 for each frequency band in order to obtain the band-wise features. The features are then concatenated (using concatenate block 406) and fed to the filtering model 108 for the latter to determine time-varying corrective gains.
[0057] In one embodiment, the features are extracted in a frame-wise manner with optimized window size, choice of windowing, hop size, FFT size, and grouping of FFT bins. The features may be warped (non-linearscaling) and normalized to achieve optimum model fits. The feature set could also include memory of past feature time frames.
[0058] It should be understood that models implemented by the signal processing unit 100, as described herein, are tuned and / or trained for optimum prediction of the time-varying corrective gains. The choice of models controls the dynamic processing performed by the signal processing unit 100. Examples of models include, but are not limited to, regression, support vector machine (SVM), tree-based models, neural networks, Long short-term memory (LSTM), and other suitable machine learning (ML) and / or deep learning (DL) architectures.
[0059] Examples of features determined by the features extraction module 104 include, but are not limited to, pitch, harmonic ratio, energy, transience, zero-crossing rate, spectral skewness, spectral entropy, spectral centroid, spectral kurtosis, spectral roll-off point, spectral slope, spectral crest, spectral decrease, spectral flux, spectral flatness, and spectral spread. For example, the features extraction module 104 may be configured to determine the pitch by performing a frequency domain decomposition of the pre-processed IEM speech signal at a given time. The features extraction module 104 may also be configured to determine the energy of the IEM speech signal (i.e. determine whether, at a given time, the pre-processed IEM speech signal has an energy level above or below a predetermined threshold). The features extraction module 104 may also be configured to evaluate the decomposition of the pre-processed IEM speech signal (e.g., assess whether the signal’s frequency emphasis is uniform or low) in order to determine spectral skewness. To determine spectral entropy, the features extraction module 104 may be configured to classify the pre-processed IEM speech signal as voiced or unvoiced based on the amount of time that the time domain signal goes above and below zero. In the illustrated embodiment, the features extraction module 104 isconfigured to determine pitch and harmonic ratio directly from the time domain samples rather than from the FFT bins.
[0060] It should however be understood that, while Fig. 4A provides an illustration of one possible embodiment of the features extraction module 104, other embodiments may apply. In particular, in some embodiments, the features determined by the features extraction module 104 may comprise the IEM samples. In other embodiments, the features may comprise FFT bin magnitudes. This is illustrated in the example of Fig. 8 (described further below), in which the input to the first neural network layer may be IEM FFT bins (e.g., 128 bins) or IEM samples (e.g., 128 samples).
[0061] Referring now to Figs. 4B and 4C, one or more feature selection strategies may be implemented by the features extraction module 104 in orderto select a subset of features from the entire set of features that is determined in the manner described above with reference to Fig. 4A. In some embodiments, it may indeed be desirable to determine the time-varying corrective gains based on a subset of features only. The number of features in the subset may vary depending on the application. The model complexity may be reduced, and robustness improved, using hyperparameter tuning, smoothing of features and ground truth, pausing training during no-speech time frames, regularization, feature selection strategies, and data augmentation.
[0062] In one embodiment, a stepwise regression technique may be performed in which the filtering model 108 is fit with each feature and the one with the least p-value (which determines the significance of a given model compared with a null model) or the highest R-square value (which is a measure of how well the given model explains the data associated therewith) is chosen for the next iteration. A second feature is then used along with the chosen feature and the combination of features with the least p-value (or the highest R-square) is chosen for the subsequent iteration. The iterations are carried out until an acceptable R-square value (e.g., 0.7 or more) is achieved with the least number of features. In one example, the proposed stepwise regression technique allowed to reduce the number of features from 158 features to 15 features, while achieving acceptable model accuracy (e.g., maintaining an R-square of 0.7 or more).
[0063] In some embodiments, a second round of feature selection may be performed by implementing a brute force regression technique in which all combinations of features from the short-listed set of features obtained using the stepwise regression technique is determined to attain an acceptable train-test model fit and R-square loss. This is illustrated in the plot 410 of Fig. 4B, which shows the model fit for a first frequency band (subplot 412), a second frequency band (subplot 414), a third frequency band (subplot 416), and a fourth frequency band (subplot 418). For each frequency band, the x-axis represents the model / combination index, the y-axis represents the train-test model fit, and the R-square loss is marked with a cross.
[0064] In another embodiment illustrated in Fig. 4C, feature selection may be based on removing correlated features from the feature set. This helps remove redundant features and reduce model complexity and processing power spent of feature computations. A correlation map 420 is generated based on all the features. Among feature pairs that have high correlation (i.e., a correlation threshold typically greater than 0.8), the feature that is least correlated with the corrective gain values are removed. The correlation threshold can be adjusted to control the number of features to be used in the model. This control of the number of features allows adjusting or balancing the model performance, complexity and processing power needed for the corrective gain prediction model.
[0065] In some embodiments, in addition to or as an alternative to the feature selection techniques described above, techniques including, but not limited to, regularization, lasso regression, and rule-based approaches, may be performed to reduce the number of features that is output by the features extraction module 104 to the filtering model 108, thus reducing model overfitting and complexity.
[0066] Referring now to Figs. 5A and 5B in addition to Fig. 1 B, operation of the cross-over filter 106 will now be described. As previously noted, in one embodiment, the IEM speech signal is decomposed by generating multiple filtered versions of the IEM speech signal. One of more filters, such as the filter 106, may be applied to the IEM speech signal to produce the filtered versions thereof. The properties (e.g., cut-off frequency, gain, Q-factor, and the like) of the filter(s) as in 106 may be fixed or time-varying. The IEM speech signal may therefore be decomposed into a plurality of frequency bands, in either the time domain or in the frequency domain. The gain of each frequency band may be varied overtime such that the filter 106 mayeffectively result in a time-varying filter. In one embodiment, the filter 106 splits the IEM speech signal into four (4) frequency bands to generate four (4) filtered versions of the IEM speech signal labelled IEM_1 , IEM_2, IEM_3, and IEM_4 in Fig. 1 B. It should however be understood that any other suitable number of frequency bands and filtered IEM speech signals may apply. Furthermore, depending on the application, the systems and methods described herein may apply to a single frequency band or filtered IEM speech signal, or to more than one frequency band or filtered IEM speech signal.
[0067] The plot 510 of Fig. 5A illustrates the magnitude response of the cross-over filter 106. As can be seen from Fig. 5A, the filter 106 spans a frequency spectrum including a first frequency range or band (curve 502), a second frequency range (curve 504), a third frequency range (curve 506), and a fourth frequency range (curve 508). The first frequency band comprises low-range frequencies (i.e. below a cutoff-frequency fd) and has a bandwidth BWi. The second frequency band comprises mid-range frequencies (i.e. between cutofffrequencies fd and fc2) and has a bandwidth BW2. The third frequency band comprises midrange frequencies (i.e. between cutoff-frequencies fc2and fa) and has a bandwidth BW3. The fourth frequency band comprises high-range frequencies (i.e. above cutoff-frequency fC3) and has a bandwidth BW4.
[0068] The gains of the different frequency bands 502, 504, 506, 508 are determined by the filtering model 108 described herein below. The bandwidths BW1, BW2, BW3, BW4 and the cutoff frequencies fd, fc2, fc3, are determined by an optimization procedure which is implemented to achieve optimum matching between the quantized bins and the FIR filter corrections for various speech content. Indeed, the IEM signal is decomposed into one or more frequency bands (i.e. into one or more filtered IEM signals) using cross-over filtering with optimized filter orders, cut-off frequencies, and bandwidths selected to achieve all-pass filter characteristics (i.e., to achieve unity gain across the frequency spectrum, thus avoiding ripple effects around cut-off frequencies of two consecutive filter bands 502, 504, 506, 508) and optimum spectral characteristics for the processed output IEM signal. The frequency bands are chosen to maximize similarity of the processed IEM speech to the reference speech.
[0069] Fig. 5B illustrates a plot 510 showing the IEM speech signal (curve 512) provided as an input to the cross-over filter 106 and which is then processed by the filter 106 to output different filtered versions (also referred to herein as “filtered signals” or “signal components”)of the IEM speech signal. In the illustrated example, a first filtered signal (curve 514) containing low-range frequencies, a second filtered signal (516) containing mid-range frequencies, and a third filtered signal (518) containing high-range frequencies are output.
[0070] Referring back to Fig. 1 B, the filtering model 108 performs a linear combination of the features received from the features extraction module 104 in order to predict time-varying corrective gains (CG) to be applied for each of the filtered IEM speech signals output by the cross-over filter 106. As used herein, the term “corrective gain” refers to the gain that is applied to each filtered IEM speech signal output from the IEM cross-over filter 106, in order for the summed filtered IEM speech signals (i.e. the processed IEM speech) to have similar temporal and spectral characteristics as air-conducted natural speech. The features received from the features extraction module 104 are used by the filtering model 108 to determine the values to which the gains of filter 106 should be set for each frequency band in order to obtain the best results (i.e. improve the quality of the IEM speech signal). In the illustrated embodiment, the IEM speech signal is split into four (4) signal components IEM_1 , IEM_2, IEM_3, and IEM_4 such that the filtering model 108 outputs four (4) time-varying corrective gains (labelled CG1 , CG2, CG3, and CG4 in Fig. 1 B) to be respectively applied (after smoothing thereof) to each IEM signal component IEM_1 , IEM_2, IEM_3, IEM_4.
[0071] The filtering model 108 may be established in any suitable manner. In one embodiment, the filtering model 108 is trained (i.e. its weights determined) or otherwise established based on a dataset generated from at least one multi-subject database recording session. During the recording session, a subject wears a hearing protection device (HPD) featuring an IEM in an audiometric booth. A fit test is performed to ensure an acceptable passive attenuation rating. A reference microphone is then placed in front of the subject’s mouth and text is presented visually on a screen. The subject is prompted to narrate the text in a natural tone, which is recorded by the reference microphone and results in the generation of a reference speech signal. In one embodiment, the text that is presented to the subject is selected such that the speech content which is recorded involves consonant-vowel-consonant words (with all combinations from the alphabet) and sentences that have a balanced set of various speech sounds. The text is also selected such that the words in the text allow to localize and conduct statistical analysis as well as filtering model generation across multiplerenditions of specific consonants and vowels. A database is further augmented with various noise conditions and varying SNR levels to account for model sensitivity and robustness.
[0072] In another embodiment, the filtering model 108 may be learnt, entirely or partially, through field recording(s), using the OEM signal as a reference (or ‘REF’) signal. In this case, both the OEM (or reference) signal and the IEM speech signal used to establish the filtering model 108 may be obtained from the field recording(s). The field recording(s) may be captured at appropriate moments, when each of the IEM (reference 12 in Fig. 1A) and OEM mostly captures clean user speech, also referred to herein as “disturbance-free user speech”. For example, the field recording(s) may be captured in response to the IEM and OEM detecting the absence of ambient noise and the presence of user speech. Thus, an example of an appropriate moment is when there is no ambient noise but there is user speech and high coherence between IEM and REF speech. Therefore, in one embodiment, rather than being trained offline (e.g., on a dataset obtained through multi-subject measurements as described herein above) prior to application in the field, the filtering model 108 may be trained opportunistically using the OEM (e.g., on a single observation).
[0073] In some embodiments, the already-established filtering model 108 may also be finetuned (e.g., its weights updated) on the field, based on new observations between the IEM and OEM. In this case, a subsequent IEM speech signal may be recorded using the IEM 12 and the OEM signal captured by the OEM may be used as a reference signal. The filtering model 108 may then be fine-tuned based on a difference between corresponding features of the reference speech signal and the subsequent IEM speech signal.
[0074] In another embodiment, the IEM signal and the reference (REF) signal may be synthesized signals rather than actual recordings. An example of a synthesized dataset involves using available samples of natural speech from existing databases (as if the signals were captured by a reference microphone) and using a model that transforms REF speech to IEM speech. This model could be synthesized for a target earpiece (e.g., the intra-aural device 18 of Fig. 1A) which would have specific acoustical properties.
[0075] In yet another embodiment, rather than using a standard model based on an entire recorded dataset which would be used as is or fine-tuned for the user and / or user conditions (as described herein above), there could exist several pre-trained models which are trainedfor specific classes and / or groups of subjects based on voice characteristics or features. For example, specific models may be pre-trained, with some model(s) being trained for low- pitched voice subjects and other model(s) being trained for high-pitched voice subjects. The appropriate model may then be chosen for use as the filtering model 108 based on a calibration step during which one or more relevant features (e.g., the pitch feature) are evaluated to select the appropriate model among the plurality of pre-trained models. Alternatively, the appropriate model may be obtained by adjusting a given pre-trained model as the user speaks. In one example, the calibration step may be based on the IEM, REF, or IEM-REF features, such as frequency response, or similar temporal / spectral features with target thresholds and / or target properties. In another example, the calibration step may be based on a classifier model which would determine the best model to apply. This classifier model may be based on physical signal properties and / or features, or a data-driven model may be used. Alternatively, the calibration step may be carried out by using a pair of earpieces (e.g., a pair of intra-aural devices 18 of Fig. 1A), with one earpiece placed near the user’s mouth to act as a reference microphone. The calibration step may also be performed using the OEM on the fitted earpiece acting as a reference microphone.
[0076] In yet another embodiment, specific models may be chosen across time for various types of speech sounds with an appropriate detector based on IEM and reference microphone features. An example of such a classification may involve two models of which one is trained for voiced / tonal speech sounds, and the other model is trained for unvoiced / fricative speech sounds which could be detected by a voiced / unvoiced (or VUV) decision algorithm. The model would then be applied in the corresponding time frame. Another classification example involves having separate models for consonants and vowels. Yet another classification may be based on the phonetics-based concept of manner of articulation, place of articulation, and voicing of speech sounds.
[0077] Although reference is made herein to the filtering model 108 being trained based on a dataset, it should be understood that the filtering model 108 may be established in any other suitable manner. For example, in some embodiments, the filtering model 108 may be an empirical model in which an empirical set of rules is established and used to determine the corrective gains to be applied for each filtered IEM speech signal (and accordingly for each frequency band). For example, such an empirical model may define one or more equations tobe used to determine the corrective gains, with the equations being tunable to achieve a desired output.
[0078] In one embodiment, the filtering model 108 is configured to determine a time-varying ground truth corrective gain for each frequency band in a frame-wise manner. This may be achieved by computing the difference between corresponding features (e.g., temporal and / or spectral features) of the reference (‘REF’) signal and of the IEM speech signal, for each frequency band. In one embodiment, computing the difference entails computing the energy difference between the REF and IEM signal components (resulting from the decomposition of the REF and IEM speech signals into signal components spanning multiple frequency bands), for each frequency band. In another embodiment, computing the difference entails computing the amplitude difference between the REF and IEM signal components for each frequency band. In yet another embodiment, computing the difference entails computing the amplitude and phase difference between the REF and IEM signal components for each frequency band. In yet another embodiment, computing the difference entails computing the energy difference between consecutive frames of the REF and IEM signal components for each frequency band. It should be understood that computing the difference may also comprise one or more of computing the energy difference, the amplitude difference, and the amplitude and phase difference mentioned above.
[0079] The filtering model 108 then predicts corrective gains which are the closest to the ground truth corrective gains. In particular, the filtering model 108 is configured to produce time-varying corrective gains that estimate the ground truth time-varying corrective gains from a plurality of time-varying features determined from the IEM speech signal received from the lEM-equipped device (e.g., the intra-aural device 18 of Fig. 1A). Forthis purpose, and referring now to Fig. 6 in addition to Fig. 1 B, the filtering model 108 is configured to map the feature matrix from the features extraction module 104. In this manner, the filtering model 108 outputs the time-varying corrective gain CG1 , CG2, CG3, or CG4 corresponding to each frequency band determined by the cross-over filter 106. Fig. 6 depicts an example in which the filtering model fits are illustrated across four (4) frequency bands. Plots 602, 604, 606, 608 illustrate the model fits on training data for the first, second, third, and fourth frequency bands, respectively. Plots 610, 612, 614, 616 illustrate the model fits on testing data for the first, second, third, and fourth frequency bands, respectively.
[0080] Fig. 7A illustrates the variation over time of the corrective gains for the four (4) frequency bands of Fig. 6. Plot 700 illustrates the ground truth corrective gains (i.e. based on training data) and plot 710 illustrates the predicted corrective gains (i.e. based on testing data). On each plot 700 and 710, the corrective gain for the first frequency band is illustrated by curve 702, the corrective gain for the second frequency band is illustrated by curve 704, the corrective gain for the third frequency band is illustrated by curve 706, and the corrective gain for the fourth frequency band is illustrated by curve 708, and the IEM speech signal is illustrated by curve 709.
[0081] Fig. 7B shows a plot 720 that illustrates band-wise corrective gains for four (4) frequency bands. In the example of Fig. 7B, for the first frequency band (e.g., below 500 Hz), the filtering model 108 predicts a corrective gain having a first value 722. The corrective gain forthe second frequency band (e.g., between 500 Hz and 1 kHz) is predicted to have a second value 724 greater than the first corrective gain value 722. The corrective gain for the third frequency band (e.g., between 1 kHz and 2 kHz) is predicted to have a third value 726 greater than the second corrective gain value 724. The corrective gain for the fourth frequency band (e.g., above 2 kHz) is predicted to have a fourth value 728 greater than the third value 726. It should however be understood that any suitable corrective gain value may be used, depending on the application. For each frequency band, the corresponding corrective gain is then applied to the IEM speech signal (illustrated by curve 730). In particular, for each IEM speech frame, the corrective gains are applied to corresponding regions of the IEM speech signal 730 to obtain a processed signal.
[0082] Referring back Fig. 1 B, in one embodiment, the predicted corrective gains are output by the filtering model 108 and fed as input to the interpolation / smoothing module 110 where smoothing and interpolation are applied on the corrective gains to upsample the latter prior to their application to the filtered IEM speech signals. The interpolation / smoothing module 110 may be configured to apply any suitable interpolation and smoothing technique. In one embodiment,
[0083] The output of the interpolation / smoothing module 110 is then sent to the multiplier 112, which also receives the filtered IEM speech signals IEM_1 , IEM_2, IEM_3, and IEM_4 from the cross-over filter 106. The multiplier 1 12 computes the product of each corrective gain CG1 , CG2, CG3, CG4 and its respective filtered IEM speech signal (or sample) IEM_1 , IEM_2,IEM_3, and IEM_4 and outputs the result, i.e. processed IEM signals (labelled IEMp_1 , IEMp_2, IEMp_3, and IEMp_4 in Fig. 1 B). The processed IEM signals are then summed at summation block 114 in order to obtain a processed IEM speech signal (labelled “lEMp frames” in Fig. 1 B) that is output by the signal processing unit 100. The processed IEM speech signal may be output to another device (e.g., a smart phone or other suitable communication device, not shown) to which the signal processing system 100 is communicatively coupled via suitable communications means (e.g., wired and / or wireless).
[0084] It should be understood that, in some embodiments, the processed IEM speech signal may be further processed through post-processing steps including, but not limited to, dynamic compression / expansion, static filtering, denoising methods, transient boosting, harmonic excitations, or similar temporal or spectral effects.
[0085] It should also be understood that, in some embodiments, the reference dataset may be processed prior to establishing the model. The reference dataset may indeed be processed to maximize intelligibility and / or subjective quality, or to maximize the accuracy of a speech- to-text system.
[0086] Referring now to Fig. 8, the systems and methods described herein may apply any suitable deep-neural network (DNN) approach. In one embodiment, a fully connected neural network-based approach may be used. Using such an approach, the filtering model (reference 108 in Fig. 1 B) may develop inherent features directly from the FFT of the IEM frames (as computed by the FFT block 406 of Fig. 4A) to predict the time-varying corrective gains. Fig. 8 shows an example neural network architecture 800 that may be used. In the example of Fig. 8, the neural network 800 has an input layer with 128 FFT bins and an output layer providing four (4) time-varying corrective gains CG1 , CG2, CG3, and CG4. In this embodiment, the filtering model 108 may predict all four (4) corrective gains in one traversal, thus alleviating the need to train one model per corrective gain.
[0087] Referring now to Figs. 9A and 9B, a method 900 for speech signal processing will now be described, in accordance with one embodiment. As can be seen in Fig. 9A, step 902 comprises obtaining a first IEM speech signal and a reference speech signal. Step 904 comprises decomposing each of the first IEM speech signal and the reference speech signal obtained at step 902 into a plurality of IEM and reference speech signal components spanninga plurality of frequency bands. Step 904 may be performed using the cross-over filter 106 described herein above with reference to Fig. 1 B. Step 906 comprises determining, for each frequency band, a plurality of time-varying ground truth corrective gains based on a difference between corresponding features (e.g., temporal and / or spectral features) of the IEM and reference speech signal components. Step 906 may be performed in the manner described above with reference to Fig. 1 B. Step 908 comprises determining a plurality of features (e.g., time-varying features) of the first IEM speech signal. The features may comprise temporal features and / or spectral features different from the temporal and / or spectral features based on which the ground truth corrective gains are determined at step 906. Step 908 may be performed using the features extraction module 104 described above with reference to Fig. 1 B. Step 910 comprises configuring at least one model to predict, from the plurality of features, a plurality of time-varying corrective gains that estimate the ground truth corrective gains determined at step 906. Step 910 may be performed using the filtering model 108 described above with reference to Fig. 1 B. In some embodiments, steps 904 and 906 may be performed in parallel to step 906. Steps 902 to 910 correspond to the process of establishing the at least one model used to predict the time-varying corrective gains.
[0088] As can be further seen in Fig. 9B, the next step 912 comprises obtaining a second IEM speech signal from an lEM-equipped device. Step 914 comprises decomposing the second IEM speech signal into a plurality of IEM speech signal components that span the plurality of frequency bands. Step 916 comprises determining a plurality of time-varying features of the second IEM speech signal. The plurality of features determined at step 916 matches (or corresponds to) the plurality of features determined at step 908. Step 918 comprises executing the at least one model for predicting the plurality of corrective gains based on the plurality of features determined at step 916. Step 920 then comprises applying the plurality of corrective gains predicted at step 918 to the second IEM speech signal to obtain a processed IEM speech signal. Step 920 may be performed using the multiplier 1 12 and summation block 114 described above with reference to Fig. 1 B. Step 922 comprises outputting the processed IEM speech signal. Steps 912 to 922 correspond to the process of applying the at least one model (established in steps 902 to 910) to the lEM-equipped device.
[0089] Referring now to Fig. 10, the system 10 of Fig. 1A and / or the method 900 of Figs. 9A and 9B may be implemented using a computing device 1000. For simplicity only onecomputing device 1000 is shown but the system 10 and / or the method 900 may involve more computing devices 1000 which may be the same or different types of devices. The computing device 1000 comprises a processing unit 1002 and a memory 1004 which has stored therein computer-executable instructions 1006. The processing unit 1002 may comprise any suitable devices configured to implement the system 10 and / or the method 900 such that instructions 1006, when executed by the computing device 1000 or other programmable apparatus, may cause the functions / acts / steps of the system 10 and / or the method 1000 described herein to be executed. The processing unit 1002 may comprise, for example, any type of general- purpose microprocessor or microcontroller, a digital signal processing (DSP) processor, a central processing unit (CPU), an integrated circuit, a field programmable gate array (FPGA), a reconfigurable processor, other suitably programmed or programmable logic circuits, or any combination thereof.
[0090] The memory 1004 may comprise any suitable known or other machine-readable storage medium. The memory 1004 may comprise non-transitory computer readable storage medium, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. The memory 1014 may include a suitable combination of any type of computer memory that is located either internally or externally to device, for example random-access memory (RAM), read-only memory (ROM), compact disc read-only memory (CDROM), electro-optical memory, magneto-optical memory, erasable programmable read-only memory (EPROM), and electrically-erasable programmable read-only memory (EEPROM), Ferroelectric RAM (FRAM) or the like. Memory 1004 may comprise any storage means (e.g., devices) suitable for retrievably storing machine-readable instructions 1006 executable by processing unit 1002.
[0091] The above description is meant to be exemplary only, and one skilled in the art will recognize that changes may be made to the embodiments described without departing from the scope of the invention disclosed. Still other modifications which fall within the scope of the present invention will be apparent to those skilled in the art, in light of a review of this disclosure.
[0092] Various aspects of the systems and methods described herein may be used alone, in combination, or in a variety of arrangements not specifically discussed in the embodimentsdescribed in the foregoing and is therefore not limited in its application to the details and arrangement of components set forth in the foregoing description or illustrated in the drawings. For example, aspects described in one embodiment may be combined in any manner with aspects described in other embodiments. Although particular embodiments have been shown and described, it will be apparent to those skilled in the art that changes and modifications may be made without departing from this invention in its broader aspects. The scope of the following claims should not be limited by the embodiments set forth in the examples, but should be given the broadest reasonable interpretation consistent with the description as a whole.
Claims
WHAT IS CLAIMED IS:1 . A method for speech signal processing, the method comprising: at a computing device: obtaining a first in-ear microphone (IEM) speech signal and a reference speech signal; decomposing the first IEM speech signal into a plurality of first IEM speech signal components, and the reference speech signal into a plurality of reference speech signal components, the first IEM speech signal components and the reference speech signal components spanning a plurality of frequency bands; determining, for each frequency band, a plurality of time-varying ground truth corrective gains based on a difference between corresponding ones of a first plurality of features associated with the first IEM speech signal components and the reference speech signal components, the first plurality of features comprising at least one of first temporal features and first spectral features; determining a second plurality of time-varying features of the first IEM speech signal, the second plurality of features comprising at least one of second temporal features and second spectral features; configuring at least one model to predict, from the second plurality of features, a plurality of time-varying corrective gains that estimate the plurality of ground truth corrective gains; obtaining a second IEM speech signal from an lEM-equipped device; decomposing the second IEM speech signal into a plurality of second IEM speech signal components spanning the plurality of frequency bands; determining a third plurality of time-varying features of the second IEM speech signal, the third plurality of features matching the second plurality of features; executing the at least one model for predicting the plurality of corrective gains based on the third plurality of features; applying the plurality of corrective gains to the second IEM speech signal to obtain a processed IEM speech signal; and outputting the processed IEM speech signal.
2. The method of claim 1 , further comprising pre-processing each of the first IEM speech signal, the reference speech signal, and the second IEM speech signal prior to the decomposing.
3. The method of claim 1 or 2, further comprising applying, prior to decomposing the second IEM speech signal, voice activity detection (VAD) to the second IEM speech signal to detect a presence or an absence of voice activity in the second IEM speech signal, wherein the plurality of corrective gains are applied to speech frames obtained based on applying the VAD to the second IEM speech signal, the speech frames indicative of the presence of voice activity in the second IEM speech signal.
4. The method of any one of claims 1 to 3, wherein determining each of the second plurality of features and the third plurality of features comprises determining at least one of a pitch, a harmonic ratio, an energy, a transience, a zero-crossing rate, a spectral skewness, a spectral entropy, a spectral centroid, a spectral kurtosis, a spectral roll-off point, a spectral slope, a spectral crest, a spectral decrease, a spectral flux, a spectral flatness, a spectral spread, and a plurality of Fast Fourier Transform (FFT) bins obtained by applying a FFT to windowed frames of a respective one of the first IEM speech signal and the second IEM speech signal.
5. The method of any one of claims 1 to 4, wherein the second plurality of features and the third plurality of features are determined based on at least one of: a respective one of the first IEM speech signal components and the second IEM speech signal components, and a full bandwidth of a respective one of the first IEM speech signal and the second IEM speech signal.
6. The method of any one of claims 1 to 5, wherein decomposing each of the first IEM speech signal and the second IEM speech signal comprises applying a filter to a respective one of the first IEM speech signal and the second IEM speech signal, the filter spanning a frequency spectrum including the plurality of frequency bands.
7. The method of any one of claims 1 to 6, wherein the plurality of time-varying ground truth corrective gains are determined based on at least one of an energy differencebetween the first IEM speech signal components and the reference speech signal components, an amplitude difference between the first IEM speech signal components and the reference speech signal components, an amplitude and phase difference between the first IEM speech signal components and the reference speech signal components, and an energy difference between consecutive frames of the first IEM speech signal components and the reference speech signal components.
8. The method of any one of claims 1 to 7, wherein each of the first IEM speech signal and the reference speech signal is obtained from a dataset generated based on at least one multi-subject recording.
9. The method of any one of claims 1 to 7, wherein each of the first IEM speech signal and the reference speech signal is obtained from at least one field recording, the first IEM speech signal and the reference speech signal being respectively captured by an IEM and an OEM in response to the IEM and the OEM detecting disturbance-free user speech.
10. The method of any one of claims 1 to 9, further comprising fine-tuning the at least one model based on a difference between corresponding features of the second IEM speech signal and the reference speech signal, wherein each of the second IEM speech signal and the reference speech signal is obtained from at least one field recording, the second IEM speech signal and the reference speech signal being respectively captured by an IEM and an OEM in response to the IEM and the OEM detecting disturbance-free user speech.1 1 . The method of any one of claims 1 to 10, wherein configuring the at least one model comprises configuring a plurality of models, and further wherein executing the at least one model comprises selecting a given model among the plurality of models based on the third plurality of features and executing the given model.
12. The method of any one of claims 1 to 11 , wherein configuring the at least one model comprises one of training at least one machine learning model and defining at least one empirical model based on a set of empirical rules.
13. A system for speech signal processing, the system comprising: a processing unit; and a non-transitory computer-readable medium having stored thereon program instructions executable by the processing unit for: obtaining a first in-ear microphone (IEM) speech signal and a reference speech signal; decomposing the first IEM speech signal into a plurality of first IEM speech signal components, and the reference speech signal into a plurality of reference speech signal components, the first IEM speech signal components and the reference speech signal components spanning a plurality of frequency bands; determining, for each frequency band, a plurality of time-varying ground truth corrective gains based on a difference between corresponding ones of a first plurality of features associated with the first IEM speech signal components and the reference speech signal components, the first plurality of features comprising at least one of first temporal features and first spectral features; determining a second plurality of time-varying features of the first IEM speech signal, the second plurality of features comprising at least one of second temporal features and second spectral features; configuring at least one model to predict, from the second plurality of features, a plurality of time-varying corrective gains that estimate the plurality of ground truth corrective gains; obtaining a second IEM speech signal from an lEM-equipped device; decomposing the second IEM speech signal into a plurality of second IEM speech signal components spanning the plurality of frequency bands; determining a third plurality of time-varying features of the second IEM speech signal, the third plurality of features matching the second plurality of features; executing the at least one model for predicting the plurality of corrective gains based on the third plurality of features; applying the plurality of corrective gains to the second IEM speech signal to obtain a processed IEM speech signal; and outputting the processed IEM speech signal.
14. The system of claim 13, wherein the program instructions are further executable by the processing unit for applying, prior to decomposing the second IEM speech signal, voice activity detection (VAD) to the second IEM speech signal to detect a presence or an absence of voice activity in the second IEM speech signal, further wherein the program instructions are further executable by the processing unit for applying the plurality of corrective gains to speech frames obtained based on the applying the VAD to the second IEM speech signal, the speech frames indicative of the presence of voice activity in the second IEM speech signal15. The system of claim 13 or 14, wherein the program instructions are executable by the processing unit for determining each of the second plurality of features and the third plurality of features comprising determining at least one of a pitch, a harmonic ratio, an energy, a transience, a zero-crossing rate, a spectral skewness, a spectral entropy, a spectral centroid, a spectral kurtosis, a spectral roll-off point, a spectral slope, a spectral crest, a spectral decrease, a spectral flux, a spectral flatness, a spectral spread, and a plurality of Fast Fourier Transform (FFT) bins obtained by applying a FFT to windowed frames of a respective one of the first IEM speech signal and the second IEM speech signal.
16. The system of any one of claims 13 to 15, wherein the program instructions are executable by the processing unit for determining the plurality of time-varying ground truth corrective gains based on at least one of an energy difference between the first IEM speech signal components and the reference speech signal components, an amplitude difference between the first IEM speech signal components and the reference speech signal components, an amplitude and phase difference between the first IEM speech signal components and the reference speech signal components, and an energy difference between consecutive frames of the first IEM speech signal components and the reference speech signal components.
17. The system of any one of claims 13 to 16, wherein each of the first IEM speech signal and the reference speech signal is obtained from one of: a dataset generated based on at least one multi-subject recording, and at least one field recording, with the first IEM speech signal and the reference speech signal being respectively captured during the atleast one field recording by an IEM and an OEM in response to the IEM and the OEM detecting disturbance-free user speech.
18. The system of any one of claims 13 to 17, wherein the program instructions are further executable by the processing unit for fine-tuning the at least one model based on a difference between corresponding features of the second IEM speech signal and the reference speech signal, wherein each of the second IEM speech signal and the reference speech signal is obtained from at least one field recording, the second IEM speech signal and the reference speech signal being respectively captured by an IEM and an OEM in response to the IEM and the OEM detecting disturbance-free user speech.
19. The system of any one of claims 13 to 18, wherein the program instructions are executable by the processing unit for configuring the at least one model comprising configuring a plurality of models, and further wherein the program are instructions executable by the processing unit for executing the at least one model comprising selecting a given model among the plurality of models based on the third plurality of features and executing the given model.
20. A non-transitory computer readable medium having stored thereon program code executable by at least one processor for: obtaining a first in-ear microphone (IEM) speech signal and a reference speech signal; decomposing the first IEM speech signal into a plurality of first IEM speech signal components, and the reference speech signal into a plurality of reference speech signal components, the first IEM speech signal components and the reference speech signal components spanning a plurality of frequency bands; determining, for each frequency band, a plurality of time-varying ground truth corrective gains based on a difference between corresponding ones of a first plurality of features associated with the first IEM speech signal components and the reference speech signal components, the first plurality of features comprising at least one of first temporal features and first spectral features;determining a second plurality of time-varying features of the first IEM speech signal, the second plurality of features comprising at least one of second temporal features and second spectral features; configuring at least one model to predict, from the second plurality of features, a plurality of time-varying corrective gains that estimate the plurality of ground truth corrective gains; obtaining a second IEM speech signal from an lEM-equipped device; decomposing the second IEM speech signal into a plurality of second IEM speech signal components spanning the plurality of frequency bands; determining a third plurality of time-varying features of the second IEM speech signal, the third plurality of features matching the second plurality of features; executing the at least one model for predicting the plurality of corrective gains based on the third plurality of features; applying the plurality of corrective gains to the second IEM speech signal to obtain a processed IEM speech signal; and outputting the processed IEM speech signal.