Real-time speech clarity online detection and real-time feedback system
Through the combination of microphones, embedded computers and displays, real-time non-intrusive detection and feedback of speech clarity in a classroom environment are achieved, solving the problem of difficulty in performing real-time online speech clarity evaluation in a classroom environment in existing technologies, and providing a real-time feedback mechanism to improve speech quality.
Patent Information
- Application Number
- CN202210642799.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2042-06-08
AI Technical Summary
Existing speech intelligibility evaluation methods are difficult to achieve non-invasive real-time online detection and feedback in classroom environments. Traditional methods are limited to telephone scenarios and cannot be widely applied to other voice communication scenarios.
Using a combination of microphones, embedded computers and displays, the system calculates the speech signal-to-noise ratio and transfer index through short-time Fourier transform, key frequency band division, multi-resolution discrete Fourier transform and auditory correction, thereby achieving real-time detection and feedback of speech clarity.
It provides real-time, non-intrusive detection and feedback of speech clarity in classroom environments. It can analyze the spectrum envelope and modulation of speech signals online to help speakers adjust the way they wear microphones or the parameters of sound amplification equipment to improve speech quality.
Smart Images

Figure CN115083444B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to a speech recognition system, and in particular relates to a real-time speech clarity online detection and real-time feedback system. Background Art
[0002] Language is a vital means of human communication, and speech quality assessment is widely used in fields such as education, communication, and justice. Speech quality can be categorized into three aspects: speech clarity, intelligibility, and naturalness. Speech clarity primarily depends on the ratio of the speech signal to the noise interference, known as the signal-to-noise ratio (SNR), and is also affected by individual sensory differences. Speech clarity assessment can be divided into subjective and objective evaluations. Subjective evaluation involves a human assessor rating a test speech under pre-defined rules and conditions, and statistical analysis is performed. For example, the Mean Opinion Score (MOS) method evaluates speech quality based on human listeners' perception of speech, with 1 being the basic unit, 2 being unacceptable, 3 being fair, 4 being good, and 5 being excellent. This method is labor-intensive and time-consuming, and is primarily used in medical research and standard setting. Objective evaluation involves quantitative analysis of the physical properties of speech signals and combines this with principles of auditory perception to derive the evaluation results. Objective evaluation can be categorized as either intrusive or non-intrusive. Intrusive evaluation involves comparing clear speech with the test speech to determine the evaluation results, while non-intrusive evaluation involves using no clear speech as a reference.
[0003] In voice communications, the intrusive speech evaluation standard includes ITU-T P.862 (PESQ, Perceptual Evaluation of Speech Quality), and the non-intrusive standard includes ITU-T P.563. Other relevant standards include the ANSIS 3.5 Speech Intelligibility Index (SII) and the IEC 60268-16 Speech Transmission Index (STI).
[0004] The ANSI S3.5 Speech Intelligibility Index (SII) test method divides the speech signal into several frequency bands, known as critical bands. Speech intelligibility is determined by measuring the signal-to-noise ratio between the speech spectrum level and the noise spectrum level in each frequency band. However, the speech spectrum level primarily contains information about speech tones and does not adequately reflect the variations in phonemes, such as vowels and consonants. These variations are primarily contained in the spectrum envelope. According to Weber's law, the human perception threshold for signal variation depends on the signal's modulation. In speech intelligibility, this is measured as the ratio of low-frequency fluctuations in the spectrum envelope to the spectrum level. The Speech Transmission Index (STI) test method effectively measures the modulation of the speech spectrum envelope. However, because the STI test method requires the injection of a test signal into the test environment, it is intrusive and unsuitable for online speech quality measurement in classroom environments.
[0005] ITU-T P.563 uses a single-point objective measurement method to evaluate narrowband speech signals in communication networks. This test method is primarily limited to telephone scenarios. Although its measurement principle has certain reference value, further research and improvement are still needed to extend it to applications beyond telephones. Summary of the Invention
[0006] In view of the above, the object of the present invention is to provide a real-time online detection and feedback system for speech clarity, so as to detect the speech clarity of speech signals at any point in a classroom environment or similar indoor environment in real time and provide timely feedback.
[0007] To achieve the above-mentioned object of the invention, the embodiment provides a real-time speech clarity online detection and real-time feedback system, including a microphone, an embedded computer and a display;
[0008] The microphone is used to collect voice signals in real time;
[0009] The embedded computer is used to perform a short-time Fourier transform on an active speech signal extracted from a speech signal in real time, divide the obtained active speech spectrum into key frequency bands and calculate the spectrum envelope and average peak spectrum of each key frequency band, perform a multi-resolution discrete Fourier transform on the average peak spectrum after low-pass filtering, divide the transformation result into fluctuation frequency bands and calculate the fluctuation spectrum mean of each fluctuation frequency band; calculate the speech modulation degree of the fluctuation frequency band based on the fluctuation spectrum mean, perform auditory correction on the speech modulation degree and then calculate the signal-to-noise ratio; and calculate the transfer index, modulation transfer index and speech transfer index in sequence based on the signal-to-noise ratios of all fluctuation frequency bands;
[0010] The display is used to display the speech transfer index obtained by the embedded computer analysis as a speech clarity analysis result in a graphical manner.
[0011] In one embodiment, the active speech signal extracted from the speech signal includes: determining whether the current segment is an active speech signal based on the short-time zero-crossing rate, energy level or tone period detection of the speech signal in the time domain, truncating and removing the inactive speech signal, and the remaining speech signal is the active speech signal.
[0012] In one embodiment, for each key frequency band, the spectrum envelope is calculated by:
[0013] If the spectrum amplitude value of frequency point k is greater than the spectrum amplitude values of the adjacent frequency points k-1 and k+1, then frequency point k is considered to be a local peak point, and the spectrum amplitude value of frequency point k is used as the local peak value. The outline formed by all local peaks in each key frequency band is the spectrum envelope;
[0014] For each key frequency band, the average peak spectrum is calculated by constructing the average peak spectrum using the average peak value of each spectrum envelope.
[0015] In one embodiment, performing a multi-resolution discrete Fourier transform after low-pass filtering the average peak spectrum includes:
[0016] When performing Fourier transform, the spectral resolution used in the low-frequency band is higher than the spectral resolution used in the high-frequency band to obtain the transformation result. Among them, the setting of the spectral resolution must ensure that the divided fluctuation frequency band contains the required fluctuation spectrum amplitude.
[0017] In one embodiment, dividing the transformation result into fluctuation frequency bands and calculating the fluctuation spectrum mean of each fluctuation frequency band includes:
[0018] After dividing the transformation result into frequency bands according to the required frequency spectrum amplitude, the frequency spectrum mean of each frequency band is calculated according to the following formula:
[0019]
[0020] Among them, b1 is the index of the key frequency band, b2 is the index of the fluctuation frequency band, S b1,b2 represents the mean value of the fluctuation spectrum of the b2th fluctuation band contained in the b1th key frequency band, E′ b1(m) represents the result of low-pass filtering of the frequency of the average peak spectrum of the key frequency band, where m is the time sequence index of the average peak spectrum, NumSE is the number of average peak spectra, M and N are the spectrum sequence numbers of the multi-resolution discrete Fourier transform results in the b2th fluctuating frequency band, where M is the first spectrum sequence number in the fluctuating frequency band and N is the last spectrum sequence number, and M, N ∈ {1, 2,... NumSE}; the relationship between M and N is: fl < M * fbin ≤ N * fbin < fu, where fl is the lower limit of the fluctuating frequency band and fu is the upper limit of the fluctuating frequency band. When there is only one spectrum data in the fluctuating frequency band, M = N.
[0021] In one embodiment, the following formula is used to calculate the speech modulation degree of the fluctuating frequency band based on the mean value of the fluctuating spectrum:
[0022] M b1,b2 = S b1,b2 / A b1
[0023] Where, M b1,b2 represents the speech modulation degree, S b1,b2 represents the mean value of the fluctuating spectrum of the b2th fluctuating frequency band included in the b1th key frequency band, and A b1 represents the sum of the average peak spectra of each group of key frequency bands, and the calculation formula is:
[0024]
[0025] Where, E b1 (m) represents the average peak spectrum of the key frequency band, m is the index of the average peak spectrum, and NumSE is the number of average peak spectra in each group.
[0026] In one embodiment, the calculation formula for the signal-to-noise ratio of the fluctuating frequency band is:
[0027] SNR b1,b2 = 10lg(M b1,b2 / (1 - M b1,b2 )
[0028] Where, SNR b1,b2 represents the signal-to-noise ratio of the b2th fluctuating frequency band included in the b1th key frequency band, and M b1,b2 represents the speech modulation degree of the b2th fluctuating frequency band included in the b1th key frequency band;
[0029] When the calculated SNR b1,b2 is not within the range [-15, 15] dB, the closest boundary value of -15 dB or 15 dB is taken.
[0030] [[ID=
[0031] TI b1,b2 =(SNR b1,b2 +15) / 30
[0032]
[0033]
[0034] Among them, SNR b1,b2 It represents the signal-to-noise ratio of the b2th fluctuation band contained in the b1th key frequency band, TI b1,b2 N represents the transmission index of the b2th fluctuation band contained in the b1th key band, b2 Indicates the total number of fluctuation bands in a single key frequency band, MTI b1 represents the modulation transfer index of the b1th key frequency band, W b1 represents the weight coefficient of the b1th key frequency band, and the sum of the weight coefficients of all key frequency bands is equal to 1.
[0035] In one embodiment, when the microphone collects speech signals, the sampling frequency is at least 8 kHz and the quantization accuracy is at least 1 byte; when performing short-time Fourier transform and multi-resolution discrete Fourier transform, the window is a Hanning window or a Hamming window.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] The real-time online detection and feedback system for speech clarity provided uses a non-invasive method to detect speech activity in the collected speech signals at any location in a classroom environment or similar indoor environment, intercept the audio signal time period with speech activity, record and analyze the spectrum envelope modulation of key speech frequency bands, and use this to calculate the speech signal-to-noise ratio and the corresponding speech transmission index, thereby obtaining an objective evaluation of the clarity. The clarity of the speech signal is fed back to the speaker in the form of numbers, colors or graphical interfaces, which helps the speaker take necessary response measures, such as adjusting the microphone wearing method or the parameters of the amplification equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0039] Figure 11 is a schematic structural diagram of a real-time speech clarity online detection and real-time feedback system provided by an embodiment;
[0040] Figure 2 This is a specific module diagram of the real-time speech clarity online detection and real-time feedback system provided by the embodiment;
[0041] Figure 3 This is a detection process implemented by an embedded computer provided in an embodiment;
[0042] Figure 4 is a physical diagram of the system provided in the above embodiment;
[0043] Figure 5 This is a visualization of the analysis results obtained by using the above system to perform speech intelligibility analysis. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.
[0045] With the development of internet technology, live online broadcasting has become widely used in teaching, placing higher demands on classroom voice quality. This requires a non-invasive detection device that can provide real-time feedback on voice quality in a classroom environment. To meet this demand, an embodiment of the present invention provides a real-time online voice clarity detection and feedback system to meet such application scenarios.
[0046] Figure 1 Schematic diagram of the structure of the real-time speech clarity online detection and real-time feedback system provided by the embodiment. Figure 1 As shown, the system provided in the embodiment includes a microphone, an embedded computer and a display, wherein the microphone is in array form and consists of audio acquisition, AD conversion, and digital signal processing modules. It has functions such as far-field audio acquisition, noise reduction, and directionality, and is used to collect voice signal data. The embedded computer performs voice clarity analysis and calculation on the collected voice signal. In this embodiment, the embedded computer uses a multi-core processor, storage media such as memory and flash memory, and required peripheral interfaces. The display presents the voice clarity analysis results to the user in a graphical interface. This embodiment uses a 7-inch IPS LCD display with a resolution of 1024x600 pixels. Of course, a digital display tube or LED indicator light can also be used to replace the display to achieve a visual presentation of the voice clarity analysis results.
[0047] Specifically, if Figure 2As shown, the entire system can adopt the Linux Ubuntu operating system, the graphical display interface Grafana module, the real-time database Prometheus module, the HTTP service Django module, the speech clarity STI calculation module, and the speech interface file module.
[0048] like Figure 3 As shown, based on the above system, the speech intelligibility analysis process implemented by the embedded computer includes:
[0049] Step 1: Obtain a speech signal collected for a fixed time period.
[0050] In this embodiment, the speech signal is acquired by a microphone array to obtain speech intensity data, with a sampling frequency of at least 8 kHz and a quantization accuracy of at least 8 bits (1 byte). Specifically, the sampling frequency fs is set to fs = 16 kHz, and the quantization accuracy is 16 bits. The fixed duration can be 10 seconds, that is, the signal acquisition data of a speech segment is obtained every 10 seconds.
[0051] Step 2: Perform voice activity detection on the voice signal.
[0052] In an embodiment, whether the current segment is an active speech signal can be determined based on the short-term zero-crossing rate, energy level or tone period detection of the speech signal in the time domain. If speech activity is detected, the collected sound intensity data is considered valid; otherwise, the collected sound intensity data is invalid and is an inactive speech signal. The inactive speech signal is truncated and discarded. To avoid introducing truncation noise, the truncation point should be selected with an amplitude value close to 0 as much as possible, and the remaining speech signal is an active speech signal.
[0053] In the embodiment, intonation period detection may adopt methods such as autocorrelation function (ACF), short-time mean amplitude difference function (ST-AMDF), and probabilistic Yin algorithm (PYIN).
[0054] Step 3: Perform short-time Fourier transform on the active speech signal to obtain the active speech spectrum.
[0055] In an embodiment, the active speech signal obtained through voice activity detection is framed and windowed, wherein the frame length is preferably 512 bytes, the window adopts a Hanning window or a Hamming window, and the sliding step size is preferably 256 bytes. Then, a short-time Fourier transform (STFT) is performed on each windowed data according to the following formula to obtain the active speech spectrum.
[0056]
[0057] Among them, n is the time scale, x(n) is the time domain speech signal, k is the frequency scale, X(k) is the spectrum of the speech signal, N stftis the number of speech signal points participating in STFT calculation. When the sampling frequency is 16000Hz, N stft =512 when the period of each frame data is T f =N / fs=32ms.
[0058] Step 4: Divide the active speech spectrum into key frequency bands and calculate the spectrum envelope and average peak spectrum of each key frequency band.
[0059] In the embodiment, the active speech spectrum collected at each fixed time length is divided into several key frequency bands according to the characteristics of human speech and hearing. In the embodiment, 7 octave bands are used as key frequency bands b1, and the center frequency of each octave band is b1∈{125, 250, 500, 1000, 2000, 4000, 8000} (Hz). The key frequency bands can also be divided according to other scales, such as the Bark scale or the Mel scale, or according to 1 / 3 octave bands. The spectrum range of the key frequency band is generally 50 to 8000 Hz according to the language spectrum, but is not limited to this range and can be wider or narrower.
[0060] For the correlation between the frequency point k in step 3 and each key frequency band, the value range is: b1 = 125 Hz, k∈{3, 4, 5}; b1 = 250 Hz, k∈{6, 7...11}; b1 = 500 Hz, k∈{12, 13...22}; b1 = 1000 Hz, k∈{23, 24...45}; b1 = 2000 Hz, k∈{46, 47...90}; b1 = 4000 Hz, k∈{91, 92...180}; b1 = 8000 Hz, k∈{181, 182...256}.
[0061] In the embodiment, for each key frequency band, the spectrum envelope is calculated as follows: if the spectrum amplitude value of frequency point k is greater than the spectrum amplitude values of adjacent frequency points k-1 and k+1, the amplitude value of frequency point k is considered to be a local peak value, which is recorded as P b1 (k)=|X(k)|. If frequency point k is not a local peak point, then P b1 (k) = 0, k is not within the critical frequency band, and P b1 (k) = 0. The formula is:
[0062]
[0063] After all local peak points are obtained, the outline formed by all local peaks is the spectrum envelope.
[0064] Of course, the spectrum envelope can also be calculated using methods such as cepstral.
[0065] For each key frequency band, the average peak spectrum is calculated by taking the average of all local peaks within the key frequency band as a data point for constructing the average peak spectrum. The formula is expressed as follows:
[0066]
[0067] Among them, Num is the number of non-zero local peak points in the key frequency band. To prevent data overflow, when Num=0, E b1 The value of is set to zero.
[0068] Step 5: Repeat steps 1-4 multiple times to accumulate multiple time series average peak spectrum data of each key frequency band. The time period of each data is T f The average peak spectrum data of the time series is low-pass filtered and then subjected to a multi-resolution discrete Fourier transform. The transformation result is divided into fluctuation bands and the fluctuation spectrum mean of each fluctuation band b2 is calculated.
[0069] In the embodiment, a multi-resolution discrete Fourier transform (MR-DFT) is performed on the average peak spectrum of each key frequency band to determine the fluctuation spectrum of the speech spectrum envelope in each key frequency band as the vocal tract (the vocal organs above the vocal cords, including lips, tongue, teeth, jaw, larynx, etc.) moves. The frequency of change of the vocal tract movement is low frequency, ranging from 0 to 20 Hz, so the fluctuation spectrum of the speech spectrum envelope is also mainly low frequency. Calculating a lower fluctuation spectrum requires a higher spectral resolution, and thus more sampling data is required to participate in the DFT calculation, which prolongs the calculation amount and update time of the DFT data. Therefore, a multi-resolution (MR: multi-resolution) method is adopted to solve this problem, and a higher resolution DFT calculation is adopted when the fluctuation band is low, and a lower resolution DFT is adopted when the fluctuation band is high. Spectral resolution f bin The calculation formula is as follows:
[0070] f bin =1 / (T f *NumSE)
[0071] Among them, T f is the data period of each frame in step 3, NumSE is the required number of average peak spectrum samples, and the value of NumSE should ensure that the required fluctuation spectrum data is obtained through DFT calculation for each divided fluctuation frequency band.
[0072] In the embodiment, the fluctuation frequency band adopts the 1 / 3 octave band with a center frequency of 0.63Hz to 12.5Hz adopted in the STI measurement convention, but other embodiments are not limited to this frequency band division. For example, an equally divided frequency band is adopted, and the center frequency can also be adjusted in the range of 0 to 20Hz according to different frequency band division methods. In order to calculate the speech spectrum fluctuation of the 1 / 3 octave band of the high segment b2∈{12.5Hz, 10Hz, 8Hz, 6.3Hz} of the fluctuation frequency band, NumSE is set to 64, that is, repeat the above steps 1-4 64 times to obtain 64 average peak spectrum data of each group of 7 key frequency bands, and the spectrum resolution of each group of data subjected to DFT is f bin =1 / (0.032*64)=0.488Hz, the frame length for DFT is 64*32=2048ms≈2s, and the sum of the average peak spectrum of each key frequency band is calculated, that is, the sum of the amplitudes is:
[0073]
[0074] In this embodiment, the voice spectrum fluctuations of the 1 / 3 octave band of the low-frequency band b2∈{4.5Hz, 4Hz, 3.15Hz, 2.5Hz, 2Hz, 1.6Hz, 1.25Hz, 1Hz, 0.8Hz, 0.63Hz} are calculated, which requires a higher spectrum resolution. When NumSE is 256, at least one DFT data can be obtained for each 1 / 3 octave band of the above low-frequency band. Therefore, the above steps 1-4 are repeated 256 times to obtain 7 groups of 256 key frequency band spectrum envelope data. The spectrum resolution f of the DFT of each group of data is 100. bin =1 / (0.032*256)=0.122Hz, the total frame length of each data set is about 8s, and the sum of the amplitudes of the average peak spectrum of each key frequency band is calculated according to NumSE=256. b1 .
[0075] In this embodiment, the average peak spectrum E of each group of key frequency bands is b1 (m) Perform low-pass filtering and windowing to obtain E' b1 (m), the cutoff frequency of the low-pass filter is set to 15Hz, and the window function can be selected from either Hanning or Hamming window. NumSE is set to 64 or 256 according to the 16kHz speech signal sampling rate in step 1 and 512 data per frame in step 3. When the embodiment uses different sampling rates fs and frame lengths T f NumSE should be adjusted according to the specific embodiment to obtain the appropriate spectrum resolution f bin And ensure that each fluctuation frequency band obtains at least one fluctuation spectrum amplitude.
[0076] In the embodiment, after obtaining the fluctuation spectrum of the required fluctuation band, the mean value of the fluctuation spectrum of each fluctuation band is calculated according to the following formula:
[0077]
[0078] where b1 is the index of the key band, b2 is the index of the fluctuation band, and S b1,b2 represents the mean value of the fluctuation spectrum of the b2th fluctuation band included in the b1th key band, and E′ b1 (m) represents the result of low-pass filtering of the frequency of the average peak spectrum of the key band, m is the index of the average peak spectrum, NumSE is the number of average peak spectra participating in the DFT calculation, M and N are the spectral sequence numbers of the multi-resolution discrete Fourier transform results within the b2th fluctuation band, where M is the first spectrum within this fluctuation band and N is the last spectrum. The relationship between M and N is: fl < M * fbin < N * fbin < fu, where fl is the lower limit of this 1 / 3 octave band and fu is the upper limit of this 1 / 3 octave band. When there is only one spectral data within this 1 / 3 octave band, M = N.
[0079] In the embodiment, the center frequencies of the fluctuation band b2 are 12.5 Hz, 10 Hz, 8 Hz, 6.3 Hz, 4.5 Hz, 4 Hz, 3.15 Hz, 2.5 Hz, 2 Hz, 1.6 Hz, 1.25 Hz, 1 Hz, 0.8 Hz, 0.63 Hz, a total of 14 1 / 3 octave bands. The key band b1 is 7 octave bands with center frequencies of 125 Hz, 250 Hz, 500 Hz, 1000 Hz, 2000 Hz, 4000 Hz, and 8000 Hz. When b1 is any one of the above 7 octave bands and b2 is 12.5 Hz, substitute M = 24, N = 27, and NumSE = 64 into the above formula to calculate S b1,12.5 value; when b2 is 10 Hz, substitute M = 19, N = 22, and NumSE = 64 into the above formula to calculate S b1,10 value; when b2 is 8 Hz, substitute M = 15, N = 18, and NumSE = 64 into the above formula to calculate S b1,8 value; when b2 is 6.3 Hz, substitute M = 12, N = 14, and NumSE = 64 into the above formula to calculate S b1,6.3 value; when b2 is 4.5 Hz, substitute M = 35, N = 38, and NumSE = 256 into the above formula to calculate S b1,4.5 value; when b2 is 4 Hz, substitute M = 31, N = 34, and NumSE = 256 into the above formula to calculate S b1,4 value; when b2 is 3.15 Hz, substitute M = 24, N = 27, and NumSE = 256 into the above formula to calculate S b1,3.15When b2 is 2.5Hz, substitute M=19, N=22, NumSE=256 into the above formula to calculate S b1,2.5 When b2 is 2Hz, substitute M=15, N=18, NumSE=256 into the above formula to calculate S b1,2 When b2 is 1.6Hz, substitute M=12, N=14, NumSE=256 into the above formula to calculate S b1,1.6 When b2 is 1.25Hz, substitute M=10, N=11, NumSE=256 into the above formula to calculate S b1,1.25 When b2 is 1Hz, substitute M=8, N=9, NumSE=256 into the above formula to calculate S b1,1 When b2 is 0.8Hz, substitute M=6, N=7, NumSE=256 into the above formula to calculate S b1,0.8 When b2 is 0.63Hz, substitute M=N=5, NumSE=256 into the above formula to calculate S b1,0.63 value.
[0080] Step 6: Calculate the speech modulation degree of the fluctuation frequency band based on the fluctuation spectrum mean.
[0081] In the embodiment, the voice modulation degree of the fluctuation frequency band is calculated using the following formula:
[0082] M b1,b2 =S b1,b2 / A b1
[0083] Among them, M b1,b2 Indicates the voice modulation degree, S b1,b2 A represents the mean value of the fluctuation spectrum of the b2th fluctuation band contained in the b1th key frequency band, b1 Represents the sum of the average peak spectra for each group of key frequency bands.
[0084] In the embodiment, in step 3 and step 4, the Fourier transform may also use a wavelet transform method to calculate the spectrum envelope modulation of each frequency band.
[0085] Step 7: Perform auditory correction on the speech modulation before calculating the signal-to-noise ratio.
[0086] In the embodiment, considering the influence of masking effect and absolute hearing threshold on speech clarity perception, it is equivalent to b1 Add noise, that is, according to A b1 The value of M b1,b2 Perform auditory correction. The specific correction process is: M′ b1,b2 =A cf,b1 *M b1,b2
[0087] Among them A cf,b1 is the correction factor, A cf,b1 The value is determined by the sum of the average peak spectrum amplitude A of the current octave band b1 and the sum of the average peak spectrum amplitude of the previous octave band A' b1 Sure.
[0088] In the embodiment, after the speech modulation degree is auditorily corrected, the signal-to-noise ratio calculation formula of the fluctuation frequency band is:
[0089] SNR b1,b2 =101g(M′ b1,b2 / (1-M′ b1,b2 )
[0090] Among them, SNR b1,b2 M′ represents the signal-to-noise ratio of the b2th fluctuation band contained in the b1th key band. b1,b2 Indicates the speech modulation degree of the b2th fluctuation frequency band contained in the b1th key frequency band;
[0091] When calculating the SNR b1,b2 When the value exceeds the range [-15, 15]dB, the nearest boundary value -15dB or 15dB is taken.
[0092] Step 8: Calculate the transmission index based on the signal-to-noise ratio of all fluctuation frequency bands.
[0093] In the embodiment, the signal-to-noise ratio SNR b1,b2 Normalized to the range of 0 to 1 to obtain the transfer index TI b1,b2 The calculation formula is:
[0094] TI b1,b2 =(SNR b1,b2 +15) / 30
[0095] Step 9: Calculate the modulation transfer index of a single key frequency band based on the transfer indexes of all fluctuation frequency bands.
[0096] In the embodiment, the modulation transfer index MTI b1 The transfer index TI b1,b2 The calculation formula for averaging the fluctuation frequency band is as follows:
[0097]
[0098] Among them, N b2 Indicates the total number of fluctuation bands in a single key frequency band. In this embodiment, N b2 =14.
[0099] Step 10: Calculate the speech transmission index according to the modulation transmission index of all key frequency bands.
[0100] In the embodiment, the calculation formula of the speech transmission index STI is:
[0101]
[0102] Among them, W b1 represents the weight coefficient of the b1th key frequency band, and the sum of the weight coefficients of all key frequency bands is equal to 1.
[0103] Figure 4 is a physical diagram of the system provided in the above embodiment, Figure 5 This is a visualization of the MOS value converted from the speech intelligibility analysis results obtained using the aforementioned system. Compared to PESQ and STI technologies, the system provided by the present invention does not require reference signal acquisition or test signal injection, enabling online speech intelligibility analysis. Compared to the P563, the present invention does not rely on speech enhancement technology to generate reference signals, enabling more accurate speech intelligibility testing in a variety of scenarios, not limited to voice and telephone communication scenarios.
[0104] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A real-time speech clarity online detection and real-time feedback system, characterized in that: Includes microphone, embedded computer and display; The microphone is used to collect voice signals in real time; The embedded computer is used to perform a short-time Fourier transform on the active speech signal extracted from the speech signal in real time, divide the obtained active speech spectrum into key frequency bands and calculate the spectrum envelope and average peak spectrum of each key frequency band, perform a multi-resolution discrete Fourier transform on the average peak spectrum after low-pass filtering, divide the transformation result into fluctuation frequency bands and calculate the fluctuation spectrum mean of each fluctuation frequency band; The speech modulation degree of the fluctuation frequency band is calculated based on the mean value of the fluctuation spectrum, and the speech modulation degree is auditorily corrected before the signal-to-noise ratio is calculated; the transfer index, modulation transfer index and speech transfer index are calculated in turn based on the signal-to-noise ratio of all fluctuation frequency bands; The method of performing a multi-resolution discrete Fourier transform after low-pass filtering the average peak spectrum includes: when performing the Fourier transform, the spectrum resolution used in the low-frequency band is higher than the spectrum resolution used in the high-frequency band, so as to obtain a transformation result, wherein the spectrum resolution is set to ensure that the divided fluctuation frequency band contains the required fluctuation spectrum amplitude; The step of dividing the transformation result into frequency bands and calculating the mean value of the fluctuation spectrum of each frequency band includes dividing the transformation result into frequency bands according to the required fluctuation spectrum amplitudes and then calculating the mean value of the fluctuation spectrum of each frequency band according to the following formula: Among them, b1 is the index of the key frequency band, b2 is the index of the fluctuation frequency band, S b1,b2 represents the mean value of the fluctuation spectrum of the b2th fluctuation band contained in the b1th key frequency band, E′ b1 (m) represents the low-pass filtering result of the average peak spectrum of the key frequency band, m is the time series index of the average peak spectrum, NumSE is the number of average peak spectra, M and N are the spectrum serial numbers of the multi-resolution discrete Fourier transform results in the b2th fluctuation frequency band, where M is the first spectrum serial number in the fluctuation frequency band and N is the last spectrum serial number, M, N∈{1, 2, ...NumSE}; the relationship between M and N is: fl <M*f bin ≤N*f bin <fu,f bin is the spectrum resolution, fl is the lower limit of the fluctuation band, fu is the upper limit of the fluctuation band, when there is only one spectrum data in the fluctuation band, M=N; The display is used to display the speech transfer index obtained by the embedded computer analysis as a speech clarity analysis result in a graphical manner.
2. The real-time speech clarity online detection and real-time feedback system according to claim 1 is characterized in that: The active voice signal extracted from the voice signal includes: determining whether the current segment is an active voice signal based on the short-time zero-crossing rate, energy level or tone period detection of the voice signal in the time domain, truncating and removing the inactive voice signal, and the remaining voice signal is the active voice signal.
3. The real-time speech clarity online detection and real-time feedback system according to claim 1 is characterized in that: For each key frequency band, the spectrum envelope is calculated by: If the spectrum amplitude value of frequency point k is greater than the spectrum amplitude values of the adjacent frequency points k-1 and k+1, then frequency point k is considered to be a local peak point, and the spectrum amplitude value of frequency point k is used as the local peak value. The outline formed by all local peaks in each key frequency band is the spectrum envelope; For each key frequency band, the average peak spectrum is calculated by constructing the average peak spectrum using the average peak value of each spectrum envelope.
4. The real-time speech clarity online detection and real-time feedback system according to claim 1 is characterized in that: The following formula is used to calculate the speech modulation degree of the fluctuation frequency band based on the mean value of the fluctuation spectrum: M b1,b2 =S b1,b2 / A b1 Among them, M b1,b2 Indicates the voice modulation degree, S b1,b2 A represents the mean value of the fluctuation spectrum of the b2th fluctuation band contained in the b1th key frequency band, b1 It represents the sum of the average peak spectra of each group of key frequency bands and is calculated as: Among them, E b1 (m) represents the average peak spectrum of the key frequency band, m is the index of the average peak spectrum, and NumSE is the number of average peak spectra in each group.
5. The real-time speech clarity online detection and real-time feedback system according to claim 1 is characterized in that: The signal-to-noise ratio calculation formula of the fluctuation frequency band is: SNR b1,b2 =10lg(M b1,b2 / (1-M b1,b2 ) Among them, SNR b1,b2 M represents the signal-to-noise ratio of the b2th fluctuation band contained in the b1th key frequency band. b1,b2 Indicates the speech modulation degree of the b2th fluctuation frequency band contained in the b1th key frequency band; When calculating the SNR b1,b2 If the value is not within the range [-15, 15]dB, the nearest boundary value -15dB or 15dB is used.
6. The real-time speech clarity online detection and real-time feedback system according to claim 1, characterized in that: The transfer index, modulation transfer index, and speech transfer index are calculated in turn based on the signal-to-noise ratio of all fluctuation frequency bands according to the following formula: TEN b1,b2 =(SNR b1,b2 +15) / 30 Among them, SNR b1,b2 It represents the signal-to-noise ratio of the b2th fluctuation band contained in the b1th key frequency band, TI b1,b2 N represents the transmission index of the b2th fluctuation band contained in the b1th key band, b2 Indicates the total number of fluctuation bands in a single key frequency band, MTI b1 represents the modulation transfer index of the b1th key frequency band, W b1 represents the weight coefficient of the b1th key frequency band, and the sum of the weight coefficients of all key frequency bands is equal to 1.
7. The real-time speech clarity online detection and real-time feedback system according to claim 1, characterized in that: When the microphone collects speech signals, the sampling frequency is at least 8 kHz and the quantization accuracy is at least 1 byte; when performing short-time Fourier transform and multi-resolution discrete Fourier transform, the window uses a Hanning window or a Hamming window.
Citation Information
Patent Citations
Sound amplifying system with classroom speech intelligibility measuring function
CN111757235A