Context-aware speech intelligibility enhancement

Through multi-band speech and noise correction, short clip analysis, long clip drawing and global gain analysis, the equipment limitation and signal problems in speech intelligibility processing are solved, and the effect of improving speech intelligibility in a noisy environment is achieved.

CN114402388BActive Publication Date: 2025-06-06DTS INC(US)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080063374.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-11
Filing Date
2020-09-09
Publication Date
2025-06-06
Estimated Expiration
2040-09-09

AI Technical Summary

Technical Problem

The prior art fails to effectively solve the problems of physical limitations, signal clearance and long-term voice characteristics of playback devices and noise capture devices when improving speech intelligibility.

Method used

The digital-to-acoustic level conversion of multi-band speech and noise correction, short-segment speech intelligibility analysis, long-segment speech and noise drawing, and global and per-band gain analysis, can achieve relative gain adjustment of speech signals through these technical means.

Benefits of technology

It improves the intelligibility of speech in noisy environments, overcomes practical challenges, realizes the natural conversion from unprocessed speech to post-processed speech, and improves the quality of voice playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114402388B_ABST
    Figure CN114402388B_ABST
Patent Text Reader

Abstract

A method includes: detecting noise in an environment with a microphone to generate a noise signal; receiving a speech signal to be played into the environment through a speaker; performing multi-band correction of the noise signal based on a microphone transfer function of the microphone to generate a corrected noise signal; performing multi-band correction of the speech signal based on a speaker transfer function of the speaker to generate a corrected speech signal; and calculating a multi-band speech intelligibility result based on the corrected noise signal and the corrected speech signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority declaration

[0002] This application claims priority to U.S. Provisional Application No. 62 / 898,977, filed on September 11, 2019, which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure relates to speech intelligibility processing. Background Art

[0004] Voice playback devices such as artificial intelligence (AI) speakers, mobile phones, teleconferencing, Internet of Things (IoT) devices, etc. are often used in acoustic environments including high-level background noise. The voice played by the voice playback device may be masked by the background noise, resulting in reduced voice intelligibility. There are many technologies available to improve voice intelligibility. Some of these technologies also utilize noise capture devices to enhance voice intelligibility in noisy environments. However, these technologies do not specify and solve practical challenges associated with implementation-specific limitations, such as the physical limitations of playback devices, the physical limitations of noise capture devices, the signal headroom of voice intelligibility processing, and long-term voice characteristics. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 is a high-level block diagram of an example system in which embodiments directed to speech intelligibility processing may be implemented.

[0006] Figure 2 is Figure 1 Block diagram of an example speech intelligibility processor (VIP) and associated speech and noise processing implemented in a system of FIG.

[0007] Figure 3 An example graph of a frequency band-importance function of the Speech Intelligibility Index (SII) is shown.

[0008] Figure 4 Example speaker frequency responses for two different speakers are shown.

[0009] Figure 5 is an example idealized microphone frequency response and an example idealized loudspeaker frequency response, as well as frequency graphs for various frequency analysis ranges determined based on the relationship between the two frequency responses.

[0010] Figure 6 A graph showing a short segment of a speech signal and its corresponding spectrum.

[0011] Figure 7 Another short segment of a speech signal and a graph of its corresponding spectrum are shown.

[0012] Figure 8A graph showing a long segment of a speech signal and its corresponding spectrum.

[0013] Fig. 9 A high-level block / signal flow diagram of a part of VIP's Speech Enhancer.

[0014] Fig.10 is a flow chart of an example method of multi-band speech intelligibility analysis / processing and speech intelligibility enhancement performed by VIP. DETAILED DESCRIPTION

[0015] Example Embodiments

[0016] Addressing the above challenges and concerns can achieve optimal performance for natural conversion from unprocessed speech to processed speech. Therefore, the embodiments provided herein introduce novel features and improvements for speech intelligibility analysis that improve speech intelligibility in noisy environments and overcome the practical challenges described herein. Embodiments include, but are not limited to: (1) digital to acoustic level conversion in combination with multi-band speech and noise correction, (2) short segment speech intelligibility analysis, (3) long segment speech and noise profiling, and (4) global and per-band gain analysis. Because the analysis results performed in the embodiments produce relative gain adjustment parameters for the speech signal for playback, both broadband and per-band, the processing in the embodiments is not limited to specific audio signal processing and can include any combination of known dynamic processing such as compressors, expanders, and formant enhancement.

[0017] As used herein, the terms: "voice", "speech" and "voice / voice" are synonyms and can be used interchangeably; "frame", "segment" and "time segment" are synonyms and can be used interchangeably; "voice (or speech) intelligibility" and "intelligibility" are synonyms and can be used interchangeably; "bin" and "frequency band" are synonyms and can be used interchangeably; and "bandwidth (BW)" and "passband" are synonyms and can be used interchangeably.

[0018] Figure 1 1 is an example system 100 in which the embodiments presented herein may be implemented. System 100 is an example, and many variations are possible. Such variations may omit or add audio components. System 100 may represent a voice communication device that supports voice communications (e.g., voice calls) with a remote communication device (not shown). System 100 may also represent a multimedia playback device coupled to a communication device. Non-limiting examples of system 100 include phones (e.g., mobile phones, smart phones, voice over Internet Protocol (IP) (VoIP) phones, etc.), computers (e.g., desktop computers, laptop computers, tablet computers, etc.), and home theater sound systems equipped with voice communication devices.

[0019] The system 100 is deployed in an acoustic environment, such as a room, an open space, etc. The system 100 includes a voice transmission path, a voice playback path, and a media playback path coupled to each other. The voice transmission includes a microphone 104, an acoustic echo canceller 106, and a noise preprocessor 108, which are coupled to each other and represent a voice / noise capture device (also referred to as a "noise capture device" for short). The microphone 104 converts the sound in the acoustic environment into a sound signal representing the sound. The sound signal represents the background noise (referred to as "noise") in the acoustic environment, and may also represent the voice from the speaker. For example, the acoustic echo canceller 106 and the noise preprocessor 108 (collectively referred to as "preprocessors") respectively eliminate echoes and reduce noise in the sound signal, and send the processed sound signal (e.g., processed speech) for playback at, for example, a remote station.

[0020] The voice playback path includes a voice intelligibility processor (VIP) 120, a system volume control 122, and a speaker 124 (more generally, a playback device). In the voice playback path, the VIP 120 receives a voice signal (i.e., a voice playback signal) to be played back through the speaker 124. For example, the voice signal may have been transmitted from the above-mentioned remote communication device (e.g., a remote mobile phone) to the system 100 for playback. In addition, the VIP 120 receives a noise signal representing noise in the acoustic environment from the microphone 104. The noise signal received by the VIP 120 can be an echo-cancelled noise signal generated by the acoustic echo canceller 106 to avoid self-activation of the VIP. According to the embodiments presented herein, the VIP 120 simultaneously processes the voice signal for playback with the noise signal (e.g., the noise sensed by the microphone 104) to enhance the intelligibility of the voice signal, thereby generating an intelligibility-enhanced voice signal. VIP 120 provides the intelligibility enhanced speech signal to speaker 124 (via system volume control 122) for playback by the speaker into the acoustic environment.

[0021] The media playback path includes an audio post processor 130, a system volume control 122, and a speaker 124. The audio post processor 130 processes the media signal for playback by the speaker 124 (via the system volume control 122). The system 100 may also include a switch 140 to selectively direct voice playback or media playback to the speaker 124.

[0022] The system 100 also includes a controller 150 coupled to the microphone 104 and the speaker 124. The controller 150 can be configured to implement, for example, the acoustic echo canceller 106, the noise preprocessor 108, the VIP 120, the audio postprocessor 130, the switch 140, and the system volume control 122. The controller 150 includes a processor 150a and a memory 150b. The processor 150a may include, for example, a microcontroller or microprocessor configured to execute software instructions stored in the memory 150b. The memory 150b may include a read-only memory (ROM), a random access memory (RAM), or other physical / tangible (e.g., non-volatile) memory storage device. Therefore, in general, the memory 150b may include one or more computer-readable storage media (e.g., memory devices) encoded with software, the software including computer-executable instructions, and when the software is executed (by the processor 150a), it is operable to perform the operations described herein. For example, the memory 150b stores or is encoded with instructions for control logic to implement the VIP 120 (e.g., in conjunction with the following Figure 2-9 Modules of the VIP described above) and other modules of the above-mentioned system 100, and perform overall control of the system 100.

[0023] The memory 150b also stores information / data 150c used and generated by the control logic as described herein.

[0024] Figure 2 1 is an example high-level block diagram of a VIP 120 and the processing performed by the VIP according to an embodiment. The VIP includes a speech and noise analyzer 202 coupled to a speech enhancer 204. The speech and noise analyzer 202 receives a noise signal from a microphone 104. The speech and noise analyzer 202 also receives a speech signal for playback. In the example, the noise signal and the speech signal are time domain signals and can each be in a pulse code modulation (PCM) format, but other formats are also possible. The speech and noise analyzer 202 simultaneously analyzes / processes the noise signal and the speech signal to produce a multi-band speech intelligibility result 205 and provides it to the speech enhancer 204. The speech enhancer 204 processes the speech signal based on the multi-band speech intelligibility result 205 to enhance or improve the intelligibility of the speech signal, thereby producing an intelligibility-enhanced speech signal. The intelligibility-enhanced speech signal is played back through the system volume control 122 and the speaker 124.

[0025] The speech and noise analyzer 202 includes a noise correction path 206, a speech correction path 208, a speech intelligibility calculator 210 after both correction paths, and a gain determiner 212 after the speech intelligibility calculator 210. The noise correction path 206 includes a noise digital to acoustic level converter (DALC) 222 and a multi-band noise corrector 224 after the noise DALC. The speech correction path 208 includes a speech DALC 226 and a multi-band speech corrector 228 after the speech DALC. The speech intelligibility calculator 210 includes a short segment analyzer 230, a long segment analyzer 232, and a silence / pause detector 234. The noise correction path 206 receives pre-measured and / or derived noise pickup device parameters 240 (e.g., known microphone parameters) that characterize the microphone 104 or are associated with the microphone 104. The speech correction path 208 receives pre-measured and / or derived playback device parameters 242 (e.g., known speaker parameters) that characterize the speaker 124 or are associated with the speaker 124.

[0026] In general, the noise correction path 206 applies multi-band noise correction to the noise signal based on the noise pickup device parameters 240. Specifically, based on the noise pickup device parameters 240, the noise DALC 222 performs a digital to acoustic level conversion (e.g., scaling) of the noise signal, and the noise corrector 224 performs multi-band noise correction on the converted or scaled noise signal to generate a corrected noise signal. The noise correction path 206 provides the corrected noise signal to the speech intelligibility calculator 210. Similarly, the speech correction path 208 applies multi-band speech correction to the speech signal. Specifically, based on the playback device parameters 242, the speech DALC 226 performs a digital to acoustic level conversion (e.g., scaling) of the speech signal, and the speech corrector 228 applies multi-band correction to the converted / scaled speech signal to generate a corrected speech signal. The speech correction path 208 provides the corrected speech signal to the speech intelligibility calculator 210.

[0027] The speech intelligibility calculator 210 performs a multi-band speech intelligibility analysis on the corrected noise signal and the corrected speech signal to generate a multi-band speech intelligibility result (MVIR), and provides it to the gain determiner 212. More specifically, the short segment analyzer 230 performs a multi-band speech intelligibility analysis on the short / medium length frames / segments of the corrected noise / speech to generate short / medium length segment multi-band speech intelligibility results (also referred to as "short-term speech intelligibility results" or simply "short-term results"). The short-term results include a sequence of per-band speech intelligibility values, global speech intelligibility values, per-band noise power values, and per-band speech power values ​​corresponding to a sequence of short / medium length segments of noise / speech.

[0028] On the other hand, the long segment analyzer 232 performs long-term noise and speech description (including speech intelligibility analysis) on the corrected long frames / segments of noise / speech (which are longer than the short / medium length segments) to produce long segment speech intelligibility results (also referred to as "long-term speech intelligibility results" or simply "long-term results"), such as long-term per-band speech intelligibility values ​​and long-term global gain values. For example, the long-term noise and speech description can perform a moving average (e.g., over a time segment equal in length to the long segment) on the values ​​in the short-term result sequence to produce a long-term result. In addition, the long-term noise and speech description can employ other types of long-term processing of the short-term results, such as peak hold and reset of noise / speech power values ​​across multiple short / medium length segments, for example as described below.

[0029] The silence / pause detector 234 detects silence / pause in the corrected speech signal to interrupt the intelligibility analysis during silence, for example, to prevent the intelligibility analysis from being activated during silence, etc.

[0030] The speech intelligibility results provided to the gain determiner 212 may include a combination of short-term results and long-term results. The gain determiner 212 derives the global gain and the gain per frequency band for the short / medium length segments based on the aforementioned speech intelligibility results, and provides the gains to the speech enhancer 204. The speech enhancer 204 may include a speech compressor, a speech expander, a formant enhancer, etc. The speech enhancer 204 performs speech enhancement processing on the (uncorrected) speech signal based in part on the analysis results 205. For example, the speech enhancer 204 applies the gain to the speech signal to produce an intelligibility-enhanced speech signal, which is played back through the system volume control 122 and the speaker 124.

[0031] Embodiments presented herein include, but are not limited to, multi-band noise and speech correction performed by noise and speech correction paths 206, 208, short / medium length segment speech intelligibility analysis performed by short segment analyzer 230, long term noise and speech delineation performed by long segment analyzer 232, and global and per-band gain analysis performed by gain determiner 212. Various embodiments will be described more fully below.

[0032] Multi-band noise and speech correction

[0033] Multi-band noise and speech analysis is known. One form of such analysis includes a speech intelligibility index (SII). The SII analysis receives a multi-band speech signal to be played back into an acoustic environment through a loudspeaker, and a noise signal representing the noise in the acoustic environment detected by a microphone. The SII analysis calculates (i) the difference between the level of the speech signal and the noise signal for each frequency band of the speech signal, e.g., calculating the speech-to-noise ratio (SNR) for each frequency band of the speech signal, multiplying the SNR of each frequency band by the frequency band-importance function of the corresponding frequency band, and summing the results.

[0034] Figure 3 Different graphs of the band-importance function of the speech intelligibility index are shown. The band-importance function basically applies different weights to the frequency bands of the speech signal according to their contribution / importance to the speech / speech intelligibility. In addition to the band-importance function, the research also discussed that the fundamental formant and the first formant of human speech may not affect the intelligibility of the voice / speech compared to the second formant and other factors. These are important factors to consider when calculating speech intelligibility.

[0035] Directly manipulating the frequency response of a speech signal based on an intelligibility index or an intelligibility contribution factor for each frequency band may degrade speech quality when played back through a loudspeaker. For example, such manipulation may cause speech to sound unnatural when changing the frequency balance and / or introducing time-varying fluctuations. In addition, if the transducer frequency response (e.g., the frequency response of the microphone and the loudspeaker) is not compensated prior to the intelligibility analysis, the results of the above-mentioned intelligibility analysis (e.g., SII) will be inaccurate. In addition, if the limitations of the loudspeaker (e.g., its small size or small driver) prevent the loudspeaker from reproducing the full frequency band of speech, the loudspeaker may further degrade speech / voice quality to change the frequency balance and result in inaccurate speech intelligibility results. Increasing the gain of speech frequencies that the loudspeaker cannot reproduce does not solve the problem and may result in nonlinear distortion and / or may stress the loudspeaker's driver.

[0036] Figure 4 The speaker frequency responses of two different speakers, denoted spk1 and spk2, are shown.Since the transducer characteristics of different speakers and different microphones are different, the speaker compensation and microphone compensation of a given system should be considered when calculating multi-band speech intelligibility.

[0037] Therefore, in addition to the intelligibility contribution factor for each frequency band, the multi-band noise and speech correction performed by the noise and speech correction paths 206, 208 also corrects the frequency bands of noise and speech used to calculate the multi-band speech intelligibility result based on the characteristics of the speaker and microphone, respectively. As an example, the noise correction path 206 corrects the noise signal (H) based on the noise pickup device parameters 240. ns) of the frequency band (e.g., adjusting the power level of the frequency band) to generate a corrected noise signal (H An_ns ), and the speech correction path 208 corrects the speech signal (H) based on the playback device parameters 242 spch ) of the frequency band (eg, adjusting the power level of the frequency band) to generate a corrected speech signal (H An_spch Then, the speech intelligibility calculator 210 calculates the corrected noise signal (H An_ns ) and the corrected speech signal (H An_spch ) performs multi-band speech intelligibility analysis.

[0038] Examples of noise pickup device parameters 240 include the transfer function H of the microphone mic (e.g., a known microphone transfer function), a gain G associated with the microphone mic (i.e., the output gain of the noise signal), the acoustic-to-digital conversion gain C of the noise signal mic , and the sensitivity of the microphone. Examples of playback device parameters 242 include the transfer function H of the speaker spk (i.e., the known loudspeaker transfer function), the gain G associated with the loudspeaker spk (i.e., the output gain of the speech signal), the acoustic-to-digital conversion gain C of the speech signal spk , and the sensitivity of the speaker (which may be provided separately or incorporated into the other parameters). The transfer function may include a frequency domain representation of the time domain impulse response of the corresponding transducer (e.g., a microphone or speaker), including magnitude and phase information over multiple continuous frequency bands across the transfer function.

[0039] As an example, the speech correction path 208 uses the playback device parameters 242 to correct the speech signal (H) according to the following formula: spch ) (e.g., the spectrum of the speech signal) to generate a corrected speech signal (H An_spch ):

[0040] H An_spch (z) = H spch (z)*H spk (z)*g spk *c spk Formula (1)

[0041] For example, the voice DALC 226 is based on the parameter g spk and c spk The speech signal is scaled, and the speech corrector 228 is based on the speaker transfer function H spk (z) Perform multi-band correction on the scaled speech signal.

[0042] Similarly, the noise correction path 206 uses the noise pickup device parameters 240 to correct the noise signal (H) according to the following formula: ns ) to produce a corrected noise signal (H An_ns ):

[0043] H An_ns (z) = H ns (z)*H mic (z) -1 *g mic *c mic Formula (2)

[0044] For example, the noise DALC 222 is based on the parameter g mic and c mic The noise signal is scaled, and the noise corrector 224 is based on the microphone transfer function H mic (z) Perform multi-band correction on the scaled noise signal. This produces an accurate estimate of the noise in the acoustic environment.

[0045] The above-mentioned scaling of noise and speech signals may include scaling based in part on microphone sensitivity and speaker sensitivity, respectively. In one example, the scaled noise / speech value is given by:

[0046] Scale_val=10 (A / 20) / 10 (D / 20) =10 ((A-D) / 20) .

[0047] Where A = Acoustic Level (dB) and D = Equivalent Digital Level (dB)

[0048] This scaling is performed separately for microphone 104 and speaker 124 in order to match the respective input signals (i.e., noise or speech) to their corresponding acoustic levels (in dB). Alternatively, the scaling may be performed to align the noise and speech levels to the microphone and speaker sensitivities. Since the speech intelligibility calculations subsequently performed on the scaled values ​​use the ratio of the (corrected) speech signal and the (corrected) noise signal from the same acoustic environment, the intelligibility calculations will be accurate if the delta caused by the different microphone and speaker sensitivities is adjusted.

[0049] in this case:

[0050] Scale_val_mic=10 (Aspk / 20) / 10 (Amic / 20)

[0051] Among them, A spk and A micThe acoustic level (dB) is measured / calculated based on the numerical level (dBFS) of the same level.

[0052] Because scaling adjusts the relative increments, the scaled values ​​can be applied only to noise signals.Alternatively, the inverse of Scale_val_mic can be applied only to speech signals.

[0053] The speech and noise signal correction of formulas (1) and (2) improves subsequent multi-band speech intelligibility analysis. In addition to speech and noise correction, the embodiments provided herein perform multi-band (frequency) regional analysis on the frequency response of microphone 104 and speaker 124. The multi-band regional analysis can be performed in noise correction path 206, speech correction path 208 and / or speech intelligibility calculator 210, or by a separate module of speech and noise analyzer 202. The multi-band regional analysis checks / determines the overlapping and non-overlapping interrelationships between the frequency range of the microphone and the frequency range of the speaker, and based on the interrelationships determined by them, the frequency bands used for multi-band speech intelligibility analysis are divided into different frequency analysis regions / ranges. Then, multi-band speech intelligibility analysis is performed based on (i.e., taking into account) the different frequency analysis regions established by the multi-band regional analysis. For example, the multi-band speech intelligibility analysis can apply different types of intelligibility analysis to speech analysis bands within different frequency analysis ranges, as described below.

[0054] Figure 5 Frequency graphs of an idealized (brick wall) microphone frequency response 502 and an idealized loudspeaker frequency response 504 are shown, along with various frequency analysis ranges (a)-(g) determined by multi-band regional analysis based on the interrelationship between the two frequency responses. The microphone frequency response 502 has a useful / responsive microphone frequency range or bandwidth (BW) / frequency passband (e.g., 3 dB BW, although other measures of what is considered a useful microphone passband may be used) that extends from a minimum ("min") / starting frequency fmic1 of the microphone frequency response to a maximum ("max") / stop frequency fmic2. Similarly, the loudspeaker frequency response 504 has a useful / responsive loudspeaker frequency range or BW / frequency passband (e.g., 3 dB BW, although other measures of what is considered a useful loudspeaker passband may be used) that extends from a minimum / starting frequency fspk1 of the loudspeaker frequency response to a maximum / stop frequency fspk2.

[0055] exist Figure 5In the example of , the relationship of the minimum or starting frequencies fspk1, fmic1 is fspk1>fmic1, and the relationship of the maximum or stopping frequencies fmic2, fspk2 is fmic2>fspk2. Therefore, the microphone passband is larger than and completely includes the speaker passband, that is, the speaker passband is completely within the microphone passband. In this case, the speaker passband and the microphone passband overlap only on the speaker passband. In another example, vice versa, that is, the relationship of the minimum frequencies is fmic1>fspk1, and the relationship of the maximum frequencies is fspk2>fmic2, so that the speaker passband is larger than and completely includes the microphone passband, that is, the microphone passband is completely within the speaker passband. In this case, the speaker passband and the microphone passband overlap only on the microphone passband.

[0056] exist Figure 5 In the example of FIG. 1 , for performing multi-band speech intelligibility in a region, the multi-band regional analysis may classify the frequency analysis regions (a)-(g) (referred to as “regions (a)-(g)” for short) according to the following:

[0057] a. Regions (a) and (b) may be defined as regions that remain unchanged through speech intelligibility analysis, or as attenuation regions for headroom preservation (ie, preserving headroom).

[0058] b. Regions (c) and (g) should not be included in the speech intelligibility analysis because the noise capture device (e.g., microphone) cannot provide accurate analysis results. The frequency region / region below fmic1 and above fmic2 includes unstable capture frequency regions / bands, where H mic The inverse of (i.e., H mic -1 ) is not stable enough to be applied to noisy signals for noise correction.

[0059] c. Regions (d) and (f) should be included in the speech intelligibility analysis used to calculate the (global) noise level and masking threshold, but not used in the per-band speech intelligibility analysis; for example, any per-band speech level increase in regions (d) and (f) resulting from the speech intelligibility analysis cannot be accommodated by playback devices that are not responsive in these regions.

[0060] d. For Figure 5 In the arrangement of the opposite loudspeaker and microphone frequency responses shown in , i.e., the loudspeaker passband is larger than the microphone passband, the noise signal level in region (d) (i.e., between fspk1 and fmic1) can be approximated using the noise signal level in the frequency band adjacent to this region (e.g., the frequency band above / adjacent to fmic1). In this case, the corrected noise signal can be calculated as: H An_ns (k) = alpha * HAn_ns (k+1), where alpha is an approximate coefficient in the range of 0 to 1.0, although the minimum value is preferably greater than 0.

[0061] exist Figure 5 In the example of , where the microphone passband is wider than and includes the speaker passband, regions (d) and (f) should be included in the global noise level and masking threshold calculations because the level of the noise signal is considered accurate after the correction of formula (2) is applied to the noise signal. However, in the alternative / inverted example, where the speaker passband is wider than and includes the microphone passband, the processing of regions (d) and (f) should be different because the level of the noise signal in the region is inaccurate, while the level of the speech signal is accurate. In this case, regions (d) and (f) can be excluded from both the global analysis and the per-band analysis.

[0062] Considering the frequency analysis range as described above improves the accuracy of the speech intelligibility analysis because frequency bands with inaccurate noise levels are removed from the analysis. The speech intelligibility analysis also provides the best global speech intelligibility results and per-band speech intelligibility results by addressing the differences in the frequency band ranges / passbands of the speakers and microphones.

[0063] Then, the speech correction and noise correction can be combined with the intelligibility contribution factor of each frequency band (i.e., each speech analysis band). For example, using speech / noise correction, the intelligibility value V per frequency band (speech) can be calculated according to the following formula idx (i) (for frequency bands i=1 to N):

[0064] V idx (i)=I(i)*A(i), i=from max(fmic1, fspk1) to min(fmic2, fspk2)

[0065] Formula (3), where: i = a band index identifying a given band (e.g., band i = 1 to band i = 21);

[0066] I = importance factor;

[0067] A = frequency band audibility value; and

[0068] The functions max(fmic1, fspk1) to min(fmic2, fspk2) determine / define the frequency overlap between the speaker and microphone passbands (eg, the "overlap passband" over which the speaker and microphone passbands overlap).

[0069] The speech and noise analyzer 202 uses the above relationship to determine the overlapping passbands based on the start and stop frequencies of the speaker and microphone.

[0070] The frequency band audibility value A is based on the corrected speech signal and the corrected noise signal speech obtained from formulas (1) and (2), respectively. For example, the frequency band audibility value A can be proportional to the ratio of the corrected speech signal power to the corrected noise signal power in a given frequency band. The frequency analysis range per frequency band is defined / corrected based on the noise pickup device parameters 240 and the playback device parameters 242, as described above.

[0071] From the above, it can be concluded that formula (3) produces speech intelligibility results from speech analysis band 1 to N based on different frequency analysis areas as follows:

[0072] a. From band 1 (ie, lowest band) to max(fmic1,fspk1) => intelligibility N / A.

[0073] b. From fspk1 to fspk2 => the speech intelligibility value per frequency band is given by formulas (1) and (2).

[0074] c. From min(fmic2,fspk2) to band N (ie, highest frequency band) => intelligibility N / A.

[0075] If max(fmic1,fspk1) is fspk1, then Figure 5 The area (a) shown in can be attenuated to reserve headroom for processing. If max(fmic1,fspk1) is fmic1, the area below fspk1 can be used to reserve headroom. This headroom may be critical in certain situations where the speech signal reaches the maximum (or near-maximum) output level of the system (e.g., a speaker). In this case, intelligibility cannot be improved because there is no headroom for speech intelligibility analysis. Alternatively, a compressor / limiter can be introduced to increase the root mean square (RMS) value while retaining the peaks of the speech signal; however, if the amount of compression exceeds a certain level, this can introduce compression artifacts such as unnatural sounds and "pumping". Therefore, if a speaker cannot fully reproduce a particular frequency range in a certain area, the speech signal in that area can be attenuated to reserve more headroom.

[0076] Using speech correction and its analysis area calculation, the global speech intelligibility value (also called the global speech-to-noise ratio (SNR) (Sg), equivalently called the global speech-to-noise ratio) can be calculated according to the following formula:

[0077]

[0078] Where: If fmic1 ≥ fspk1, C1 = fmic1, or if fmic1 <fspk1,C1=fspk1;

[0079] If fmic2 < fspk2, C2 = fmic2, or if fmic2 ≥ fspk2, C2 = fspk2;

[0080] ∝ is a normalization coefficient; and

[0081] H spch (j) and H noise (i) are the corrected speech and noise signals for the j-th and i-th frequency bands, respectively.

[0082] According to the above relationship, the frequency C1 = max(fmic1, fspk1), the frequency C2 = min(fmic2, fspk2), and the frequency range from C1 to C2 is the overlapping frequency range (i.e., overlapping passband) between the microphone passband and the speaker passband. The numerator of formula (4) accumulates / sums the corrected speech power only over the overlapping frequency range, while the denominator accumulates / sums the corrected noise power only over the frequency range / passband of the microphone.

[0083] The short segment analyzer 230 generates a sequence of per-band speech intelligibility values calculated according to formula (3), and a sequence of global speech-to-noise ratios (Sg) calculated according to formula (4). The long segment analyzer 232 processes (e.g., averages) the stored values (i.e., value sequences) of the noise power and speech power from the short segment analyzer 230 over a plurality of short / medium-length segments equal to the long segment to generate per-band intelligibility values for the long segment and a global intelligibility value for the long segment. The long segment analyzer 232 can perform further operations on the short-term stored values, such as peak holding and reset, as described below.

[0084] The above combination Figure 5 described embodiments determine the analysis region used as the frequency range setting or limitation for formulas (3) and (4). In another embodiment, the corresponding weight coefficients can be directly applied to H An_spch and H An_ns , to substantially calculate formulas (3) and (4) without limiting the range, since the limitation is included in the respective weight coefficients. In this embodiment, the correction is applied according to the following formula:

[0085] H An_spch = W sp * H An_spch .

[0086] H An_ns = W ns * H An_ns ,

[0087] where W sp and W ns are the weighting coefficients applied to the speech and noise for each frequency band (0 to pi).

[0088] In summary, the embodiments provide a comprehensive method for computing noise / speech intelligibility using noise / speech correction, as follows:

[0089] a. Using playback and noise capture device characteristics, correct speech and noise signals and define the frequency band or range of speech and noise for analysis.

[0090] b. Cross-check the devices’ speech intelligibility contribution weighting factors and frequency ranges.

[0091] c. Given the speech and noise input to the speech and noise analyzer 202, analysis is performed to produce speech intelligibility values ​​with processing gain parameters per frequency band and / or a global processing gain value.

[0092] Note that for the analysis described herein, the frequency band is not limited to a specific frequency band. The frequency band may be an octave band, a one-third octave band, a critical band, etc.

[0093] Intelligibility analysis of short / medium length speech segments

[0094] Many speech playback use cases require minimal latency. Therefore, it is impractical to use long segments of about one second or more (e.g., long speech / noise segments) for speech intelligibility analysis (referred to as "long segment analysis") because long segment analysis may introduce too much latency. Instead, typically, short / medium length segments used for analyzing and processing speech / noise have a duration of about 2 to 32ms. In addition, noise may not be static, but dynamic, for example, consider a dog barking, a noisy car passing by, and so on. Therefore, multi-band speech intelligibility analysis of short / medium length segments that are relatively shorter than long segments (referred to as "short / medium length segment analysis") is preferred. That is, short / medium length segment analysis is generally preferred over long analysis.

[0095] The problem with short / medium length segment analysis is that it can produce unwanted artifacts when combined with other processing (e.g., gain processing). For example, too fast an adjustment of the processing gain can result in unnatural speech fluctuations and frequent changes in the speech frequency balance. A common way to mitigate such artifacts is to add smoothing to the gain changes by setting the attack and decay times.

[0096] However, this smoothing of speech intelligibility results introduces a trade-off between accuracy and stability. In order to obtain the best accuracy while maintaining stable speech sound, long-term sound and noise rendering can improve the results. Different from traditional methods, the embodiments provided herein combine traditional short / medium length segment analysis with long-term speech and noise rendering, as described below.

[0097] Long-term speech and noise profiling

[0098] Compared to short / medium length segments of 2 to 32 ms, the long segments analyzed by long-term speech and noise profiling can be two words to several sentences in length (e.g., about 1 to 30 seconds). For long-term speech and noise profiling, it is not necessary to store the noise / speech signal for a long time. Instead, long-term speech and noise profiling uses a sliding window to accumulate short-term results (i.e., short / medium length segment characteristics) over time (i.e., over long segments). The long-term analysis generated by long-term speech and noise profiling does not increase the latency of the speech intelligibility results because the long-term analysis uses past samples of speech and noise.

[0099] Figure 6 , Figure 7 and Figure 8 Different time segments of a speech signal and their corresponding spectra are shown. Figure 6 The top graph includes a short time segment of a speech signal (ie, a "short segment") and a bottom graph showing the spectrum of the short segment. The short segment includes 1024 speech samples spanning a short segment of approximately 23ms. Similarly, Figure 7 Included is a top graph showing another short segment of a speech signal and a bottom graph showing a second spectrum of the short segment. Figure 6 and Figure 7 The short segments shown in the top graph of are each periodic, as is typical of speech. Figure 6 and Figure 7 The spectra shown in the bottom graph of are different because the different phonemes they represent have different formant frequencies.

[0100] Figure 8 Included is a top graph showing a long time segment of a speech signal (ie, a "long segment") and a bottom graph showing the spectrum of the long segment. The long segment includes 1024 speech samples spanning approximately 4.24 seconds. Figure 6 and Figure 7 Short clips and Figure 8 The long segments capture common data, including the fundamental frequency of speech, but the long segments show the spectral characteristics of speech over longer time segments. Therefore, speech intelligibility analysis including long-term speech and noise rendering can benefit from wider band analysis values ​​and capture the long-term characteristics of speech signals over long segments, rather than just trying to dynamically allocate narrowband frequency gains based on per-band analysis that may change rapidly over time. In addition, long-term speech and noise rendering also captures the temporal characteristics of speech over long segments.

[0101] Examples of persistent noise in an environment include fan noise or humming plus occasional transient / dynamic noise, such as dogs barking and cars passing by. In this case, long-term speech and noise profiling can identify the characteristics of static / persistent noise, while short / medium length segment analysis can identify dynamic noise. Long-term speech and noise profiling can capture peak noise, which can then be reset by comparing the long-term results with the short-term results to identify that the persistent background noise has changed or has been removed. For example, long-term speech and noise profiling can include peak retention of speech / noise for a long segment, but then use the short-term results to determine whether to reset the peak, for example, when the speech playback changes to another speaker or synthesized speech. Another example is to use segments of a few words in length for analysis so that a sliding window can slowly capture the transition from one speaker to another.

[0102] Global and per-band gain analysis

[0103] The gain determiner 212 calculates multi-band gain values, including per-band gains (adjustments) and global gains (adjustments) to be applied to the (uncorrected) speech signal, based on the results produced by the short segment analyzer 230. The gain determiner 212 provides the gains to the speech enhancer 204, which applies the gains to the speech signal. The gain calculation can be flexible, depending on the processing applied to improve intelligibility. If there are computational resource limitations, the analysis bands can be grouped to effectively reduce the number of analysis bands to be processed, or some analysis bands can be omitted from the processing. If the processing already contains some intelligence, such as formant position enhancement or spectral peak enhancement, then based on the above-mentioned analysis method, the processing can use this intelligence to provide intelligibility information and appropriate global gain parameters about the frequency positions whose gains the processing selectively increases / decreases.

[0104] In an example, the gain may be calculated according to the following or similar relationship:

[0105] Global gain (g_Global) = Wg*St_g / Sc

[0106] Gain per band (g_perband(i)) = Wpb*St_pb / Sc(i)

[0107] Among them: g_Global and g_perband are applied to the speech output signal;

[0108] Wg and Wpb are the global and per-band weight coefficients;

[0109] St_g and St_pb are per-band and global intelligibility values ​​(eg, speech-to-noise ratio (SNR) values) for short / medium length segments; and

[0110] Sc is the current SNR.

[0111] The weights Wg and Wpb can be determined based on a threshold of the intelligibility value so that the weights vary with the current speech intelligibility value (e.g., when the intelligibility value is relatively high, more weight (Wg) is applied to g_Global and less weight (Wpb) is applied to g_perband, and vice versa).

[0112] Fig. 9 9 is a high level block / signal flow diagram of a portion of the speech enhancer 204 according to an embodiment. In this example, the speech enhancer 204 includes a multi-band compressor 904 that applies per-band gain values ​​g_pb(i) and a global gain g_Global to the speech signal to produce an intelligibility enhanced speech signal.

[0113] Fig.10 is a flow chart of an example method 1000 for performing speech intelligibility processing performed by VIP 120. The operations of method 1000 are based on the operations described above.

[0114] At 1002 , a microphone detects noise in an acoustic environment to generate a noise signal.

[0115] At 1004, an input of the VIP 120 receives a speech signal for playback into an acoustic environment via a speaker.

[0116] At 1006, VIP 120 performs a digital to acoustic level (DAL) conversion of the noise signal and performs a multi-band correction of the noise signal based on a known or derived microphone transfer function of the microphone to produce a corrected noise signal. The multi-band correction adjusts the spectrum of the noise signal to compensate for the microphone transfer function.

[0117] At 1008, VIP 120 performs a DAL conversion of the speech signal and performs a multi-band correction of the speech signal based on a known or derived speaker transfer function of the speaker to produce a corrected speech signal. The multi-band correction adjusts the frequency spectrum of the speech signal to compensate for the speaker transfer function.

[0118] At 1010, VIP 120 determines a frequency analysis region for multi-band speech intelligibility calculations based on a relationship between the microphone transfer function and the speaker transfer function. For example, VIP 120 determines an overlapping passband where a microphone passband of the microphone transfer function overlaps a speaker passband of the speaker transfer function based on the start and stop frequencies of the passbands. For example, the start and stop frequencies of a given passband may correspond to opposite 3 dB down points (or other suitable "X" dB down points) of the transfer function corresponding to the given passband.

[0119] At 1012, VIP 120 performs multi-band speech intelligibility analysis on multiple speech analysis bands based on the noise signal (e.g., on the corrected noise signal) and based on the speech signal (e.g., on the corrected speech signal) to calculate multi-band speech intelligibility results. For example, the analysis can be limited to speech analysis bands in overlapping passbands. The results include per-band speech intelligibility values ​​and a global sound / speech noise ratio. The multi-band speech intelligibility analysis includes analysis of / analysis based on short / medium length segments / frames to produce short-term results, and analysis of / analysis based on longer segments to produce long-term results.

[0120] At 1014, the VIP 120 calculates per-band gains and a global gain based on the per-band speech intelligibility values ​​and the global sound / speech noise ratio.

[0121] At 1016, the VIP enhances the intelligibility of the speech signal based on the gain and plays the enhanced speech signal through a speaker.

[0122] In various embodiments, some operations of method 1000 may be omitted, and / or operations of method 1000 may be reordered / permuted. For example, conversion / correction operations 1006 and 1008 may be omitted, such that operation 1012 performs multi-band speech intelligibility analysis based on the noise signal (uncorrected) and the speech signal (uncorrected) over multiple speech analysis bands to calculate multi-band speech intelligibility results. In another example, operations 1006 and 1008 may be modified to omit their respective multi-band corrections, thereby leaving only their respective DAL conversions.

[0123] In an embodiment, a method includes: detecting noise in an environment with a microphone to generate a noise signal; receiving a speech signal to be played into the environment through a speaker; determining a frequency analysis region for multi-band speech intelligibility calculation based on a relationship between a microphone transfer function of the microphone and a speaker transfer function of the speaker; and calculating a multi-band speech intelligibility result on the frequency analysis region based on the noise signal and the speech signal. The method also includes: performing multi-band correction of the noise signal based on the microphone transfer function to generate a corrected noise signal; and performing multi-band correction of the speech signal based on the speaker transfer function to generate a corrected speech signal, wherein the calculation includes calculating the multi-band speech intelligibility result on the frequency analysis region based on the corrected noise signal and the corrected speech signal.

[0124] In another embodiment, a device includes: a microphone for detecting noise in an environment to generate a noise signal; a speaker for playing a speech signal into the environment; and a controller coupled to the microphone and the speaker and configured to perform: multi-band correction of the noise signal based on a microphone transfer function of the microphone to generate a corrected noise signal; multi-band correction of the speech signal based on a speaker transfer function of the speaker to generate a corrected speech signal; calculating a multi-band speech intelligibility result based on the corrected noise signal and the corrected speech signal; calculating a multi-band gain value based on the multi-band speech intelligibility result; and enhancing the speech signal based on the multi-band gain value.

[0125] In another embodiment, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium is encoded with instructions, which, when executed by a processor, cause the processor to perform: receiving a noise signal representing noise in an environment from a microphone; receiving a speech signal to be played into the environment through a speaker; performing a digital-to-acoustic level conversion on the noise signal and performing a multi-band correction on the noise signal based on a microphone transfer function to generate a corrected noise signal; performing a digital-to-acoustic level conversion on the speech signal and performing a multi-band correction on the speech signal based on a speaker transfer function to generate a corrected speech signal; and calculating a multi-band speech intelligibility result based on the corrected noise signal and the corrected speech signal, including a per-band speech intelligibility value and a global speech-to-noise ratio.

[0126] While the techniques are illustrated and described herein as implemented in one or more specific examples, it is not intended to be limited to the details shown, since various modifications and structural changes may be made within the scope and range of equivalents of the claims.

[0127] Each of the claims set forth below represents a separate embodiment, and embodiments that combine different claims and / or different embodiments are within the scope of the disclosure and will be apparent to those of ordinary skill in the art after reviewing this disclosure.

Claims

1. A method for speech intelligibility processing, include: Detecting noise in the environment with a microphone to generate a noise signal; receiving a speech signal to be played into the environment via a speaker; determining a frequency analysis region for multi-band speech intelligibility calculation based on a relationship between a microphone transfer function of the microphone and a speaker transfer function of the speaker; performing a multi-band correction of the noise signal based on the microphone transfer function to generate a corrected noise signal; as well as performing multi-band correction of the speech signal based on the speaker transfer function to generate a corrected speech signal; and A multi-band speech intelligibility result is calculated over the frequency analysis region based on the corrected noise signal and the corrected speech signal.

2. The method according to claim 1, further comprising: include: The intelligibility of the speech signal is enhanced using the multi-band speech intelligibility results.

3. The method according to claim 1, in: The determining includes determining an overlapping passband where a microphone passband of the microphone transfer function overlaps a speaker passband of the speaker transfer function as the frequency analysis region; and The calculating includes calculating per-band speech intelligibility values ​​over a speech analysis band limited to the overlapping passbands.

4. The method according to claim 3, in, The calculating also includes calculating a global speech-to-noise ratio of (i) speech power based on the speech signal over a speech analysis band limited to the overlapping passband and (ii) noise power based on the noise signal over the microphone passband.

5. The method according to claim 3, further comprising: include: determining whether a starting frequency of the speaker passband is greater than a starting frequency of the microphone passband; as well as When the start frequency of the speaker passband is greater, the voice signal is attenuated in a frequency band lower than the start frequency of the microphone passband.

6. The method according to claim 3, in, The determination includes: identifying a start frequency and a stop frequency that define the microphone passband and the speaker passband, respectively; and The overlapping passband is calculated as the passband extending from the maximum start frequency to the minimum stop frequency.

7. The method according to claim 1, in: The calculating of the multi-band speech intelligibility results includes calculating the speech intelligibility value of each frequency band and the global speech-to-noise ratio.

8. The method according to claim 1, in, The calculation of the multi-band speech intelligibility result includes: performing a multi-band speech intelligibility analysis based on short / medium length segments of the speech signal and the noise signal to produce a short-term speech intelligibility result; and A multi-band speech intelligibility analysis is performed based on the speech signal and long segments of the noise signal that are longer than the short / medium length segments to generate a long-term speech intelligibility result.

9. A device for speech intelligibility processing, include: A microphone for detecting noise in the environment to generate a noise signal; A speaker for playing a speech signal into the environment; as well as A controller is coupled to the microphone and the speaker and is configured to: performing a multi-band correction of the noise signal based on a microphone transfer function of the microphone to generate a corrected noise signal; performing multi-band correction of the speech signal based on a speaker transfer function of the speaker to generate a corrected speech signal; Calculating a multi-band speech intelligibility result based on the corrected noise signal and the corrected speech signal; Calculating a multi-band gain value based on the multi-band speech intelligibility result; as well as The speech signal is enhanced based on the multi-band gain values.

10. The device according to claim 9, in, The controller is further configured to perform: The intelligibility of the speech signal is enhanced using the multi-band speech intelligibility results.

11. The device according to claim 9, in, The controller is further configured to perform: determining an overlapping passband where a microphone passband of the microphone transfer function overlaps a loudspeaker passband of the loudspeaker transfer function, Wherein the controller is configured to perform the calculation by calculating a per-band speech intelligibility value over a speech analysis band limited to the overlapping passband.

12. The device according to claim 9, in: The controller is configured to perform the calculating of the multi-band speech intelligibility results by calculating a per-band speech intelligibility value and a global speech-to-noise ratio.

13. The device according to claim 9, in, The calculation of the multi-band speech intelligibility result includes: performing a multi-band speech intelligibility analysis on short / medium length segments of the corrected speech signal and the corrected noise signal to produce a short-term speech intelligibility result; and A multi-band speech intelligibility analysis is performed on long segments of the corrected speech signal and the corrected noise signal that are longer than the short / medium length segments to generate a long-term speech intelligibility result.

14. The device according to claim 9, further comprising: include: Prior to the multi-band correction of the noise signal, performing a digital level to acoustic level conversion of the noise signal based on the sensitivity of the microphone; as well as Prior to the multi-band correction of the speech signal, a digital level to acoustic level conversion of the speech signal is performed.

15. A non-transitory computer-readable medium encoded with instructions that, when executed by a processor, cause the processor to: receiving a noise signal representing ambient noise from a microphone; receiving a speech signal to be played into the environment via a speaker; performing a digital level to acoustic level conversion of the noise signal, and performing a multi-band correction of the noise signal based on a microphone transfer function to generate a corrected noise signal; performing a digital level to acoustic level conversion of the speech signal, and performing a multi-band correction of the speech signal based on a speaker transfer function to generate a corrected speech signal; as well as Based on the corrected noise signal and the corrected speech signal, a multi-band speech intelligibility result is calculated, wherein the multi-band speech intelligibility result includes a per-band speech intelligibility value and a global speech-to-noise ratio.

Citation Information

Patent Citations

  • Signal processing method and device adaptive to noise environment and terminal device employing same

    CN109416914A

  • System for giving intelligibility feedback to a speaker

    EP1818912A1