Speech Detection and Enhancement in Binaural Recording
By employing a method that classifies and processes self-speech and external speech separately within binaural recordings, the method addresses the challenges of uneven volume and timbre, enhancing the quality of binaural recordings.
Patent Information
- Application Number
- JP2023541746
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-09-17
- Filing Date
- 2022-01-12
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2042-01-12
AI Technical Summary
Binaural recordings of speech face challenges in distinguishing and processing self-speech and external speech due to differences in sound propagation and microphone placement, leading to uneven volume and timbre.
A method that utilizes binaural signals to identify self-speech and external speech through time-frequency conversion, feature extraction, and classification, followed by independent enhancement processing for each segment using optimal settings.
This approach effectively balances the volume and timbre of self-speech and external speech, improving the overall quality of binaural recordings by compensating for the inherent differences in sound capture.
Smart Images

Figure 0007695365000001 
Figure 0007695365000002
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority based on U.S. Provisional Patent Application Nos. 63 / 162,289 and 63 / 245,548, filed on March 17, 2021 and September 17, 2021, respectively; and Spanish Patent Application No. P202130013, filed on January 12, 2021, each of which is incorporated by reference in its entirety.
[0002] [Technical Field] The present disclosure relates to a method for enhancing speech in binaural recording, a system for implementing this method, and a non - transitory computer - readable medium storing instructions for implementing this method.
Background Art
[0003] Earphones or earbuds are wireless in - ear headphones that pair with smart devices such as phones and tablets. They are becoming a common option for smartphone users to listen to audio and, with the addition of a built - in microphone, to capture audio for real - time communication and recording voice messages. Earphones are a convenient alternative for people who want to conduct interviews, create vlog or podcast content, or simply record voice memos to record speech without a dedicated microphone.
Summary of the Invention
[0004] In the present disclosure, the expression "self - speech" is used to refer to the speech of a person wearing earphones, and the expression "external speech" is used to refer to the speech from a person other than the person wearing the earphones.
[0005] Since the microphone is located inside the ear of a person wearing earphones, when recording self-speech, the propagation of sound from the mouth to the earphones, combined with the mouth's directivity, causes a significant change in the spectrum of the sound, i.e., a loss of high-frequency energy compared to what a conventional microphone positioned in front of the mouth would pick up. When recording external speech, the distance of each external speaker results in a loss of level compared to the volume of self-speech. Both of these factors (level loss and high-frequency loss) lead to a significant difference in volume and tonality or timbre between self-speech and external speech. Compensating for these effects benefits from the identification of self-speech and external speech, the segmentation of the recording, and the processing of each part using optimal settings.
[0006] Speaker segmentation and diarization have been an active research area for many years, using well-established statistical techniques such as the Bayes Information Criterion (BIC) and recent AI-based techniques. These techniques are effective in detecting changes in speakers or acoustic conditions but do not provide additional information such as whether the speech is self or external. In particular, they operate on monaural signals (single-channel recordings) and thus do not consider the spatial aspects of sound as embedded in binaural recordings. It can be seen that spatial aspects such as the similarity between signals from left and right binaural microphones and the direction of arrival contain important information for the task of distinguishing self-speech from external speech, but such cues are usually ignored for the purpose of segmentation.
[0007] There are automatic solutions for speech enhancement, but they do not detect or use speaker segmentation information and thus do not enable optimal adjusted processing of self-speech and external speech to achieve a balanced timbre and volume.
[0008] The present disclosure describes a method for improving binaural recordings of speech by identifying portions corresponding to self-speech and external speech, segmenting the recording accordingly, and then applying independent enhancements to each segment with optimal settings according to self-speech conditions or external speech conditions.
[0009] Taking the binaural signal as an input, time-frequency conversion is applied to divide the signal into frequency bands. In parallel, the signal is sent to a Voice Activity Detector to identify which parts of the signal contain speech and avoid processing non-speech parts.
[0010] Spectral features are extracted from the time-frequency representation of the signal and used to classify self-speech and external speech on a frame-by-frame basis. In parallel, some of these features are sent to a Dissimilarity Segmentation unit that uses statistical methods to find frames where speaker identification or acoustic conditions have changed. The segmentation unit receives information from the classification and dissimilarity segmentation units and combines them by majority voting to determine (self or external) for each segment. The segmentation is used to process the recording as multiple independent recordings, each with appropriate settings derived from the classification into self and external speech.
Brief Description of the Drawings
[0011] Embodiments of the present invention will be described in detail with reference to the accompanying drawings.
Figure 1
Figure 2
Modes for Carrying Out the Invention
[0012] [Time-Frequency Conversion and Feature Extraction] In FIGS. 1 and 2, a binaural signal s(t) having left and right signals l(t), r(t) was obtained. The binaural signal may be acquired in various ways, including recording by earphones worn by the user. Thereafter, the binaural signal is received by a device in which the speech enhancement system operates. This device may be part of a device worn by the user or may be a separate device. In the latter case, the binaural signal is transmitted to this separate device.
[0013] The system of FIG. 1 includes a frame splitter 1 connected to receive the binaural signal and split it into frames. A time-frequency conversion unit 2 is connected to receive the frames, followed by a feature extraction unit 3. A voice activity detector (VAD) 4 is connected in parallel to units 2 and 3 and is also connected to receive the frames of the binaural signal. The outputs of the feature extraction unit 3 and the VAD 4 are both connected to two blocks, namely, a self-classification block 5 and a dissimilarity segmentation block 6. The outputs from blocks 5 and 6 are both supplied to a segmentation unit 7. Further, a speech enhancement chain 8 is connected to receive the frames of the binaural signal from the frame splitter 1 and the output from the segmentation unit 7. The speech enhancement chain 8 outputs a modified binaural signal. The operation of the system and its various components will be described in more detail below with reference also to the flowchart of FIG. 2.
[0014] In step 1, the binaural signal s(t) is split into frames n. Thereafter, in step S2, the time-frequency conversion unit 2 receives the frames and generates signals L(i,f), R(i,f) having frame indices i = 1:N and frequencies f = 1:M. The time-frequency conversion can be, for example, a discrete Fourier transform, a QMF filter bank, or another conversion.
[0015] In step S3, the frame division signals L(i, f) and R(i, f) are grouped for each frequency band, and the following feature quantities are calculated by the feature extraction unit 3 in each frequency band b: - Energy per band E(i, b) = Σ f∈b (L 2 (i, f) + R 2 (i, f)); - Inter-channel coherence IC(i, b); - Mel-frequency cepstral coefficients MFCC(i, b)
[0016] Since this analysis focuses on speech, typically only the frequency band between, for example, 80 Hz and 4 kHz, which is the frequency range of speech, is retained.
[0017] Furthermore, the spectral gradient SS(i) is calculated as the gradient of the linear approximation of E(i, b) in the target frequency range.
[0018] The spectral gradient is a measure of how much the high frequencies are attenuated, and thus it is suitable for the task of distinguishing self and external speech.
[0019] The inter-channel coherence is a measure of the similarity between L and R; considering the symmetry of the propagation from the mouth to the L microphone and the R microphone, L and R can be expected to be almost the same in self-speech, while in a typical situation, dissimilarity of external speech is expected.
[0020] MFCC is a feature commonly used for speech-related analysis and classification.
[0021] In parallel with step S3 (not shown in FIG. 2), the frame of s(t) is sent to the VAD 4, and the VAD 4 outputs the probability V(i) of containing speech for each frame of the audio i, where 0 ≤ V(i) ≤ 1. When the VAD operates on a monaural signal, a downmix such as l(t) + r(t) is used instead of s(t).
[0022] [Self and External Speech Classification] In step S4, the external classification unit 5 receives the feature quantities E(i,b), SS(i,b), IC(i,b) from the feature extraction unit 3 and generates a binary classification result C(i), that is, C(i)=1 for self speech and C(i)=0 for external speech. The classification is performed by a trained classifier such as a support vector machine (SVM). The training of the classifier can be performed on a set of labeled content where the input is the aforementioned feature vector and the output class is pre-given for each frame of the audio. Since the SVM is a powerful non-linear classifier that requires less training data than a deep neural network, the SVM is selected.
[0023] For improved performance, only the frames containing audio are passed to the SVM both during training and during classification. In the illustrated example, the classification unit 5 also receives the speech probability V from the VAD 4. Thereby, the classifier 5 can pass only the frames having a probability V exceeding a given threshold to the SVM.
[0024] The accuracy of the classifier can vary depending on the presence of noise, different speaker types, etc. Since it is a decision for each frame, a method for segmenting the signal based on this classification can be provided.
[0025] Alternatively or additionally, the external classification unit 5 receives bone conduction vibration sensor data from a bone conduction sensor (not shown) and generates a binary classification result C(i) based at least in part on the bone conduction vibration sensor data. For example, the classification based on the bone conduction vibration sensor data may be performed by determining whether the bone conduction vibration sensor data exceeds a predetermined threshold, and the predetermined threshold is such that the audio is self speech, but the bone conduction vibration sensor data that does not exceed the predetermined threshold can indicate external speech. The bone conduction vibration sensor data can be used as an alternative or supplement to the features output from the feature extraction unit 3 and the speech probability V output from the VAD 4.
[0026] [Dissimilarity Segmentation] This dissimilarity segmentation unit 6 also receives the MFCC(i,b) features and the VAD information V(i), and in step S5, defines a threshold th for voice detection such that all frames where V(i) < th are discarded. The rows k of the discarded frames are removed from the matrix MFCC(i,b), and the Bayesian Information Criterion (BIC) method is applied to the remaining frames j, and a dissimilarity function D(j) is obtained according to the conventional notation. A BIC window length corresponding to the minimum length (e.g., 2s) to be segmented can be used. Then, the transitions of the speech signal are: i) The peak must be higher than a pre-defined threshold th D and ii) The peak should be separated by a minimum number of frames Δj, usually corresponding to the BIC window length, and are obtained by finding the peaks of D(j) under these conditions.
[0027] After finding the peaks in the speech-only frames, their positions are mapped back to the full set of frames, and as a result, the transitions are based on the time of the original signal.
[0028] Note that the dissimilarity segmentation detects not only transitions between the self-speaker and the external speaker, but also any other changes in the speaker or acoustic conditions; even for transitions between self-speech and external speech, it does not provide information on which is which.
[0029] [Segmentation] The segmentation unit 7 receives the self and external speech classifications for each frame C(i) from the classification unit 5 and receives a set of frames j where the speech transitions are identified by the dissimilarity segmentation unit 6. In step S6, unit 7 segments the binaural into segments based on the transition frames j. Then, in step S7, unit 7 provides the final segmentation of the audio into self and external speech segments of sufficient length and classification confidence.
[0030] For each segment k provided by the dissimilarity segmentation unit 6, the multiple frames belonging to the segment are considered to determine whether the segment is considered self-speech.
[0031] For example, applying "majority voting" to the classification for each frame, if the number of frames classified as self-speech CS(k) is greater than the number of frames classified as external speech CE(k), segment k can be regarded as self-speech, and vice versa. The confidence σ(k) of segment k is determined based on the relative difference between the number of frames CS(k) of segment k classified as self-speech and the number of frames CE(k) of segment k classified as external speech: σ(k)=|CS(k)-CE(k)| / N(k) where N(k) is the total number k of frames in segment k including non-speech frames, that is, N(k)=CS(k)+CE(k).
[0032] Threshold th σ is defined such that segments with σ<th σ are considered uncertain.
[0033] The segmentation unit 7 can further merge adjacent segments in certain situations. For example, adjacent segments that are classified into the same class (self or external) and are considered reliable according to the trust criteria can be merged into a single segment. Similarly, adjacent uncertain segments can be merged to form a single uncertain segment. Segments shorter than a predetermined duration can be merged with larger adjacent frames. An uncertain segment surrounded by two segments of the same class can be merged into one segment together with the adjacent segments. An uncertain segment surrounded by two certain segments of different classes (i.e., one self-speech and one external speech) can be merged with the longest adjacent segment.
[0034] Furthermore, an uncertain segment surrounded by two certain segments of different classes (i.e., one self-speech and one external speech) can be used as a transition region in the subsequent speech emphasis chain. For example, a short uncertain segment can be used as a cross-fade region for the transition between different processes applied to adjacent segments.
[0035] The final segmentation obtained by unit 7 is passed to the speech emphasis unit in a format that includes the transition points and the inferred class (self-speech or external speech) of each segment. Alternative representations such as the start point and duration of the segment are also possible.
[0036] [Segmentation-based Speech Emphasis] The speech enhancement chain 8 may comprise signal processing blocks that perform sibilance reduction, equalization, dynamic range compression, noise reduction, de-reverberation, and other processing. In many cases, the optimal amounts and settings of each processing block may vary depending on the characteristics of the signal: typically, self and external speech may benefit from different equalizations, independent leveling, different amounts of reverberation suppression, etc.
[0037] Accordingly, by using the segmentation of self and external speech provided by the segmentation unit 7, two classes of speech can be processed separately to achieve optimal sound quality.
[0038] Examples of segmentation-based processing are: - Equalization to compensate for high-frequency losses in recordings of self-speech; the correction curve (gain per frequency band) can be measured, estimated, or obtained by simulation and then applied only to self-speech segments. - Leveling: Aligning the levels and dynamic ranges of self and external speech can be difficult when the content is considered as a whole unit. By segmentation, each segment can be leveled independently, thus ensuring the volume and dynamic range required for each speaker. - Ambience suppression: Ambience usually enhances immersion but reduces intelligibility. Ambience and reverberation suppression can be applied heavily to external speech to increase intelligibility and lightly to self-speech to maintain immersion. - Binaural signal rotation to stabilize the perceived image by compensating for the effects of head movement during recording: Self-speech does not require stabilization (which is actually perceived as an unwanted rotation), but external speech benefits from stabilization. - Channel imbalance correction: Earphones can have inter-channel imbalance in the high-frequency range depending on how well the buds are positioned in each ear canal. This causes the voiceless part of self-speech (sibilance) to be located in a sound source direction slightly away from the center, producing a less solid sound than monaural recording. Compensating for the level difference between the left and right channels in the affected high-frequency band can improve self-speech quality, but applying the same process to external speech may affect its spatial cue.
[0039] When segmented data becomes available, the entire signal is split into segments, and each segment is processed according to the inferred class. The segments can include extra frames at the boundaries for cross-fading by overlap when recombining the processed segments. The settings used to process each frame are based on different presets for self and external speech classes (e.g., for processes where different processing is required for self and external speech, such as ambient suppression), or on the same settings (e.g., for cases where the goal is to obtain a uniform result, such as leveling).
[0040] In some implementations, - Classification of self and external speech can be achieved by a bone conduction vibration sensor. In such implementations, the classifier can perform classification based on bone conduction vibration sensor data in addition to or instead of using features. For example, the classifier can classify audio as self-speech in response to detecting bone vibrations corresponding to speech, or as external speech in response to detecting the absence of bone vibrations based on data from the bone conduction vibration sensor. Thus, bone conduction vibration sensor data can complement or replace features. - The frame sizes used for MFCC, VAD, and other features may differ; in such cases, when combining different features or different metrics derived from features, coarser features can be "upsampled" to the resolution of the finest features by interpolation or simple nearest - neighbor repetition.
[0041] Aspects of the systems described herein can be implemented in a suitable computer - based sound processing network environment for processing digital or digitized audio files. Portions of the adaptive audio system can include one or more networks with any desired number of individual machines, including one or more routers (not shown) that function to buffer and route data transmitted between computers. Such networks may be built on various different network protocols and may be the Internet, a wide - area network (WAN), a local - area network (LAN), or any combination thereof.
[0042] One or more of the components, blocks, processes, or other functional components may be implemented through a computer program that controls the execution of a processor - based computing device of the system. Also note that the various functions disclosed herein can be described as data and / or instructions embodied in various machine - readable or computer - readable media, using any number of combinations of hardware, firmware, and / or from the perspective of their behavior, register transfers, logic components, and / or other characteristics. Computer - readable media in which such formatted data and / or instructions can be embodied include various forms of physical (non - transitory) non - volatile storage media such as optical storage media, magnetic storage media, or semiconductor storage media.
[0043] Although specific embodiments have been described by way of example with respect to one or more embodiments, it should be understood that one or more embodiments are not limited to the disclosed embodiments. On the contrary, as will be apparent to those skilled in the art, various modifications and similar configurations are intended to be included. Accordingly, the appended claims should be construed as broadly as possible to encompass all such deviations and similar configurations.
Claims
1. A method comprising: - splitting a binaural audio signal into frames; - applying time-frequency transformation to each frame; - calculating features of the frame based on the time-frequency representation; - classifying each frame as self-speech or external speech by a classifier based at least in part on a subset of the features, wherein the subset of the features includes inter-channel coherence which is a measure of similarity between L and R, and performing the classification taking into account the symmetry of propagation from the mouth to the L microphone and the R microphone; - calculating a dissimilarity function between the L and the R based on the subset of the features; - segmenting the binaural audio signal at a peak of the dissimilarity function; - for each segment, determining an overall class of each of self-speech or external speech by aggregating classifier data of the frames belonging to the segment; - processing each segment with a speech enhancement chain, wherein the setting of the speech enhancement chain is based on the overall class determined for such segment, A method.
2. Calculating a respective speech probability for each frame using voice activity detection (VAD), wherein only frames for which the speech probability is greater than a predetermined value are considered for classification and segmentation, The method according to claim 1.
3. The features include at least one of energy per frequency band, spectral gradient in a predetermined frequency range, inter-channel coherence per frequency band, or Mel-frequency cepstral coefficients. The method according to claim 1.
4. wherein the classifier is a support vector machine The method according to any one of claims 1 to 3.
5. wherein the dissimilarity function is obtained by applying the Bayesian information criterion (BIC) to a subset of the features The method according to any one of claims 1 to 3.
6. including the step of retaining the peak of the dissimilarity function, on condition that the value of the dissimilarity function is greater than a predetermined value and the distance to the nearest peak is greater than another predetermined value The method according to claim 5.
7. The step of determining each overall class comprises: calculating the number of frames (CE) classified as external speech within the segment; calculating the number of frames (CS) classified as self-speech within the segment; assigning the class self-speech if CS≧CE, and assigning the class external speech if CE>CS The method according to claim 1.
8. further including the step of assigning a respective classification confidence value to each segment using the formula abs(CE - CS) / N, where N is the total number of frames within the segment The method according to claim 7.
9. including the step of designating as uncertain segments for which the confidence value is less than a predetermined value The method according to claim 8.
10. merging adjacent segments of the same class into a single segment of the same class; merging a segment designated as uncertain and surrounded by two segments of the same class (self or external) with the surrounding segments The method according to claim 9.
11. The step of processing each segment in the speech enhancement chain is: Noise estimation and noise reduction; Equalization including specific filters for self-speech and external speech; Levelling including specific target levels and dynamic ranges for self-speech and external speech; Ambience balance including different amounts of boost or attenuation for self-speech and external speech; Spatial rotation including different amounts of rotation for self-speech and external speech; and Channel imbalance correction including different amounts of correction for self-speech and external speech, including one or more of: The method according to any one of claims 1 to 10.
12. - One segment is - Specified as uncertain; - Surrounded by two segments of different classes (one self, one external); and - Shorter than a predetermined length; Processed in both settings and used as a crossfade region, - One segment is: - Specified as uncertain; - Surrounded by two segments of different classes (one self, one external); and - Longer than a predetermined length, Processed in a neutral setting or merged with the longest adjacent segment, The method according to claim 11.
13. The step of recombining the processed segments into a sequence according to their order in the original input; and The step of reducing audible discontinuities by applying a crossfade at the transition points, including: The method according to claim 11.
14. A system comprising: one or more processors; a non-transitory computer-readable storage medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 13. A system.
15. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 13. A non-transitory computer-readable storage medium.
Citation Information
Patent Citations
Device and method for processing acoustic signal
JP2005203981A
Sound signal processing device and method
JP2010038943A
Driving device of piezoelectric motor, driving method of piezoelectric motor, electronic component transportation device, electronic component inspection device, robot hand and robot
JP2013121196A
Hearing device comprising an own voice detector
US20190075406A1