Method of determining a perceptual impact of reverberation on the perceptual quality of a signal, and computer program product

By windowing the audio signal and calculating the energy-time curve, peak values ​​are identified, solving the problem of inaccurate reverberation estimation in existing technologies. This achieves more accurate audio signal quality assessment and is applicable to various sound processing and assessment methods.

CN116134801BActive Publication Date: 2026-06-02NEDERLANDSE ORG VOOR TOEGEPAST NATUURWETENSCHAPPELIJK ONDERZOEK TNO

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NEDERLANDSE ORG VOOR TOEGEPAST NATUURWETENSCHAPPELIJK ONDERZOEK TNO
Filing Date
2021-07-19
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing POLQA methods are inaccurate in estimating the impact of reverberation on perceived quality in audio signals due to factors such as signal length, noise, impulse distortion, and time-shift distortion.

Method used

By windowing the degraded audio signal and the reference audio signal, the energy-time curves of the local and global impulse response signals are calculated, the peak values ​​are identified, and the reverberation or echo amount is estimated based on the energy of the peak values. Combined with Fourier transform and autocorrelation operations, the effects of local and global interference are compensated.

Benefits of technology

It provides a more accurate estimate of reverberation or echo, improving the accuracy of audio signal quality assessment. It is applicable to a variety of sound processing and assessment methods, and is particularly relevant in the POLQA method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004113654300000151
    Figure BDA0004113654300000151
  • Figure BDA0004113654300000161
    Figure BDA0004113654300000161
  • Figure BDA0004113654300000221
    Figure BDA0004113654300000221
Patent Text Reader

Abstract

The invention relates to a method of determining a perceptual impact of an amount of echo or an amount of reverberation in a degraded audio signal on a perceptual quality of the degraded audio signal, wherein the degraded audio signal is received from an audio transmission system, the degraded audio signal being obtained by a transmission of a reference audio signal by the audio transmission system to provide the degraded audio signal. The method comprises performing a windowing operation on the degraded audio signal and the reference audio signal by multiplying the degraded audio signal and the reference audio signal with a window function to produce degraded digital audio samples and reference digital audio samples. A local estimate of the amount of echo or the amount of reverberation is determined based on the degraded digital audio samples and the reference digital audio samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for determining the perceptual effect of echo or reverberation in a degraded audio signal on the perceived quality of the degraded audio signal, wherein the degraded audio signal is received from an audio transmission system, obtained by providing the degraded audio signal by transmitting a reference audio signal through the audio transmission system, and obtained at a computer program product. Background Technology

[0002] Over the past few decades, objective speech quality measurement methods have been developed and deployed using perceptual measurement approaches. In these methods, perceptual algorithms simulate the behavior of a test subject who rates the quality of audio segments in an auditory test. For speech quality, the so-called absolute category rating auditory test is primarily used, where the test subject judges the quality of degraded speech segments without access to a clear reference segment. The auditory tests conducted by the International Telecommunication Union (ITU) primarily use the Absolute Category Rating (ACR) 5-point opinion scale, and are therefore also used in objective speech quality measurement methods. Objective speech quality measurement methods are standardized by the ITU using the following algorithms: Perceptual Speech Quality Measure (PSQM (ITU-T Recommendation P.861, 1996)) and its subsequent Perceptual Evaluation of Speech Quality (PESQ (ITU-T Recommendation P.862, 2000)). These measurement standards focus on narrowband speech quality (audio bandwidth 100-3500Hz), although a wideband extension (50-7000Hz) was designed in 2005. For narrowband speech data, PESQ and subjective listening tests show good correlation, and for wideband data, PESQ and subjective listening tests show acceptable correlation.

[0003] As the telecommunications industry rolls out new broadband voice services, there is a need for an advanced measurement standard with proven performance that can accommodate larger audio bandwidths. Therefore, ITU-T (ITU-Telecom) Study Group 12 proposed a standardization of a new voice quality assessment algorithm as a technological update to PESQ. The new third-generation measurement standard, POLQA (Perceptual Objective Listening Quality Assessment), overcomes the shortcomings of the PESQ P.862 standard, such as incorrect assessment of the impact of linear frequency response distortion, time stretching / compression found in Voice-over-IP services, certain types of codec distortion, and reverberation.

[0004] Compared to the previous quality assessment algorithms PSQM (P.861) and PESQ (P.862), POLQA (P.863) offers numerous improvements, and the current version of POLQA also proposes many improvements, such as correctly assessing the impact of linear frequency response distortion, time stretching / compression found in Voice-over-IP services, certain types of codec distortion, reverberation, and playback levels.

[0005] One of the factors affecting perceived speech and sound quality is the presence of echo and reverberation in the audio signal, the latter being the superposition of echoes. The amount of reverberation, or echo, can be determined, for example, by performing autocorrelation on the digitized audio signal to estimate the energy-time curve. When both a reference signal and a degraded signal are available, as in the case of POLQA, the energy-time curve can be determined based on the estimated transfer function of the system under test. POLQA uses the latter method; however, the accuracy of the obtained estimate is affected by the signal length and the presence of certain types of noise, impulses, or time-shift distortion, leading to inaccurate determination of the perceptual impact of reverberation on perceived audio quality. Summary of the Invention

[0006] The purpose of this invention is to eliminate the above-mentioned defects and to provide a method for accurately estimating the perceptual effect of reverberation in an audio signal on the perceived quality of the audio signal.

[0007] To this end, this paper provides a method for determining the perceptual impact of echo or reverberation in a degraded audio signal on the perceived quality of the degraded audio signal. The method involves receiving a degraded audio signal from an audio transmission system and obtaining the degraded audio signal by transmitting a reference audio signal through the audio transmission system to provide the degraded audio signal. The method includes: obtaining at least one degraded digital audio sample from the degraded audio signal and at least one reference digital audio sample from the reference audio signal by a controller; determining a local impulse response signal by the controller based on the at least one degraded audio sample and the at least one reference audio sample; determining an energy-time curve by the controller based on the impulse response signal, wherein the energy-time curve is proportional to the square root of the absolute value of the impulse response signal; identifying one or more peaks in the energy-time curve based on the impulse response, wherein the one or more peaks occur temporally at a delay in the energy-time curve after the start of the energy-time curve; and determining an estimate of the echo or reverberation based on the amount of energy in the one or more peaks. The method of obtaining at least one degraded digital audio sample includes: sampling a degraded audio signal in a time-domain segment, wherein sampling includes performing a windowing operation on the degraded audio signal by multiplying the degraded audio signal with a window function to generate a degraded digital audio sample; and wherein obtaining at least one reference digital audio sample includes: sampling a reference audio signal in a time-domain segment, wherein sampling includes performing a windowing operation on the reference audio signal by multiplying the reference audio signal with a window function to generate a reference digital audio sample; wherein the window function used to obtain at least one reference digital audio sample and at least one degraded digital audio sample has a non-zero value in the time-domain segment to be sampled and a zero value outside the time-domain segment.

[0008] This invention is based on the understanding that many interferences in a signal affect the accurate determination or estimation of the perceived effect of reverberation. These interferences include different types of noise, different types of impulse distortion, and different types of time-shift distortion, some of which impair the determination of reverberation at a global or overall level, while others are primarily harmful or present at a local level. This invention overcomes this problem by performing windowing on the degraded and reference signals before determining the reverberation. For example, a set of perceived reverberation effect parameters can be calculated based on a single frame or a series of consecutive frames, which can (through windowing) constitute audio samples of the degraded and reference audio signals. First, windowing allows for the calculation of reverberation location estimates, which are then considered in the final reverberation estimate. Second, windowing enables local compensation and optimization of processing parameters. The latter can even be done based on the duration of a temporal segment of the sample or the relative position of the sample within the complete signal (or relevant portion). Therefore, due to the windowing operation, the method of this invention provides a more accurate estimate of the reverberation or echo level. This can be applied to many different kinds of sound processing and evaluation methods. However, the method of the present invention is significantly relevant in assessing the quality or intelligibility of degraded speech signals, such as the POLQA method described above, and therefore this application provides a preferred embodiment of the method.

[0009] The step of obtaining at least one digital audio sample preferably includes: sampling the audio signal in a time-domain segment by performing the windowing operation described above, thereby obtaining multiple digital audio samples from the audio signal. In this case, the time-domain segments of at least two consecutive digital audio samples among the multiple digital audio samples may overlap. For example, the overlap between at least two consecutive digital audio samples is in the range of 10% to 90% overlap between time-domain segments, preferably in the range of 25% to 75% overlap, more preferably in the range of 40% to 60% overlap, such as 50% overlap. This may depend on the type of window function applied, for example, as part of optimization.

[0010] In some embodiments, the window function is at least one of the following: Hamming window, von Hann window, Tukey window, cosine window, rectangular window, B-spline window, triangular window, Bartlett window, Parzen window, Welch window, and nth power cosine window. thThe window function can be a power-of-cosine window, a Kaiser window, a Nuttall window, a Blackman window, a Blackman Harris window, a Blackman Nuttall window, or a flattop window, where n>1. This invention is not limited to a specific type of window function, and window functions different from those mentioned herein can be used. Furthermore, new optimized window functions that can be used in the methods of this invention can be developed without departing from the inventive concept of this invention.

[0011] To determine an estimate of the amount of reverberation, in some embodiments, the invention may include weighting the amount of energy in each peak of the energy-time curve based on the amplitude of each peak and / or the (relative) delay position of each peak along the time axis. This is based on the understanding that peaks with the largest amplitudes typically have a significant impact on the perceived level of reverberation, and how peaks with the largest amplitudes may impede the intelligibility or quality of speech or sound.

[0012] In some preferred embodiments, the method further includes: obtaining a digital signal by a controller, the digital signal representing at least a portion of an audio signal and having a duration longer than a time-domain segment of at least one digital audio sample; performing an autocorrelation operation on the digital signal by the controller to generate a total impulse response signal; determining a total energy-time curve by the controller based on the impulse response signal, wherein the energy-time curve is proportional to the square root of the total impulse response signal; and determining one or more peaks in the energy-time curve based on the total impulse response signal, the one or more further peaks occurring temporally at a delay in the energy-time curve after the start of the energy-time curve, and determining a further estimate of the amount of echo or reverberation based on the amount of energy in the one or more further peaks.

[0013] The preferred embodiments described above provide a way to correctly include and compensate for local and global interferences (i.e., interferences that have a local effect on the reverberation level and interferences that impair the estimation of the more global overall level of the sound signal (or a portion of the signal). Furthermore, similar to the reverberation estimation methods described above for local applications, further estimations of the reverberation amount at the global or overall level can also include weighting the amount of energy in each peak based on the amplitude of each peak.

[0014] In other or further embodiments, the method may further include at least one of the following steps: calculating a partial reverberation indicator value by the controller based on the estimated echo or reverberation amount; calculating a global reverberation indicator value by the controller based on the further estimated echo or reverberation amount; or calculating a final reverberation indicator value by the controller based on the estimated echo or reverberation amount and the further estimated echo or reverberation amount.

[0015] Furthermore, in the above method, determining the (local or global) impulse response signal based on the audio sample or the digital signal as described includes the following steps: the controller converts the audio sample or digital signal from the time domain to the frequency domain by applying a Fourier transform to the audio sample or digital signal; the controller determines the transfer function based on the power spectrum signal from the audio sample or the digital signal in the frequency domain; and the controller converts the power spectrum signal from the frequency domain to the time domain to generate the local impulse response signal or the global impulse response signal.

[0016] In a preferred embodiment, the present invention provides a method for evaluating the quality or intelligibility of a degraded speech signal received from an audio transmission system, wherein a reference speech signal is transmitted through the audio transmission system to provide the degraded speech signal, wherein the method includes: - sampling the reference speech signal into a plurality of reference signal frames, sampling the degraded speech signal into a plurality of degraded signal frames, and forming frame pairs by associating the reference signal frames and the degraded signal frames with each other; - providing a difference function for each frame pair, the difference function representing the difference between the degraded signal frame and the associated reference signal frame; - compensating the difference function for one or more types of interference, thereby providing an interference density function for each frame pair, the interference density function being suitable for a human auditory perception model; - obtaining an overall quality parameter based on the interference density functions of the plurality of frame pairs, the quality parameter indicating at least the quality or intelligibility of the degraded speech signal; wherein the method further includes: - determining the reverberation amount in at least one of the degraded speech signal and the reference speech signal, wherein the reverberation amount is determined by applying the method according to any of the above embodiments.

[0017] In embodiments of the above categories, the method according to the invention has been applied to methods for determining the quality or intelligibility of degraded speech signals. Since the presence of reverberation significantly affects perceived quality or intelligibility, the method according to the invention for determining the estimation of reverberation amount is particularly useful in such methods of assessing quality or intelligibility.

[0018] In some embodiments of the above embodiments, the step of obtaining at least one digital audio sample by the controller can be performed by forming audio samples from multiple consecutive signal frames, including one or more degraded signal frames or one or more reference signal frames. For example, the number of signal frames to be included in the multiple signal frames can depend on the duration of the temporal segment of the at least one digital audio sample, wherein the duration is greater than 0.3 seconds, preferably between 0.4 seconds and 5.0 seconds, for example, the duration is at least one of the following: 0.5 seconds, 1.0 seconds, 1.5 seconds, 2.0 seconds, 2.5 seconds, 3.0 seconds, 3.5 seconds, 4.0 seconds, 4.5 seconds, or 5.0 seconds. In some applications, such as POLQA, a single frame is often too short to determine the amount of reverberation, but an audio signal segment shorter than one second may be long enough to be analyzed to provide a local estimate of the amount of reverberation.

[0019] Therefore, in some embodiments, a first estimate of the reverberation amount is obtained by performing local estimation using, for example, a 0.5-second digital audio sample. This involves obtaining one or more second estimates for each of a plurality of digital audio samples formed from a plurality of consecutive signal frames providing a longer duration of audio signal, and calculating a reverberation indicator value based on at least one of the second estimates and the first estimate.

[0020] In some embodiments, for each frame pair, the compensation step is performed by setting a determined reverberation amount in at least one of the degraded speech signal and the reference speech signal to one of one or more interference types, and compensating the reverberation amount associated with the respective frame pair for each frame pair based on the formation of the digital audio samples. Here, reverberation estimation can be considered at a local level associated with the frame pair. These are the frame pairs that constitute the degraded signal samples.

[0021] In some embodiments, prior to the step of determining the impulse response signal, the method further includes a noise suppression step, the noise suppression comprising: performing a first scaling on at least one of the degraded speech signal or the reference speech signal to obtain a similar average volume; processing the degraded speech signal to remove local signal peaks from the degraded speech signal; and performing a second scaling on at least one of the degraded speech signal or the reference speech signal to obtain a similar average volume.

[0022] Furthermore, as described above, to evaluate the quality or intelligibility of speech or sound signals, the method can be well limited to a lower frequency range, i.e., the range of interest associated with the speech or sound signal. For example, the method can be performed on audio signals within a predetermined frequency range, such as a frequency range below a threshold frequency or a frequency range corresponding to the speech signal, for example, a frequency range below 5 kHz. Preferably, for speech signals, the frequency range is between 200 Hz and 4 kHz, or for other sound signals, the frequency is up to 20 kHz. Attached Figure Description

[0023] Referring to the accompanying drawings, the invention will be further illustrated by describing some specific embodiments. These specific embodiments provide examples of possible implementations of the invention, but should not be considered as the only embodiments falling within the scope of the invention. The scope of the invention is defined in the claims, and the embodiments should be considered illustrative rather than limiting. In the drawings:

[0024] Figure 1 An overview of the first part of the POLQA perception model according to an embodiment of the present invention is provided.

[0025] Figure 2 A schematic overview of the frequency alignment used in the POLQA sensing model according to an embodiment of the present invention is provided.

[0026] Figure 3 A POLQA perception model based on an embodiment of the present invention is provided, located in Figure 1 An overview of the second part following the first part shown.

[0027] Figure 4 This is an overview of the third part of the POLQA perception model according to an embodiment of the present invention.

[0028] Figure 5 is a schematic overview of the masking methods used in the POLQA model.

[0029] Figure 6 This is a schematic diagram illustrating the method of compensating for overall quality parameters.

[0030] Figures 7A-7C The illustration schematically depicts the windowing operation performed on a speech signal as applied in an embodiment of the invention.

[0031] Figure 8 The calculation of the reverberation indicator according to an embodiment is illustrated schematically. Detailed Implementation

[0032] POLQA Perception Model

[0033] The basic approach of POLQA (ITU-T Recommendation P.863) is the same as that used in PESQ (ITU-T Recommendation P.862), which employs a human perception model to map both the reference input and degraded output speech signals to internal representations. The perception model uses the difference between these two internal representations to predict the perceived speech quality of the degraded signal. A key new approach implemented in POLQA is an idealization method that removes low-level noise from the reference input signal and optimizes timbre. Other major improvements in the POLQA perception model include modeling the impact of the playback level on the perceived quality and separating the handling of low-level distortion from the handling of high-level distortion.

[0034] Figures 1 to 4 An overview of the perception models used in POLQA is provided. Figure 1 The first part of the perception model is provided, which is used to compute the internal representations of the reference input signal X(t)3 and the degraded output signal Y(t)5. Both the reference input signal X(t)3 and the degraded output signal Y(t)5 are scaled by 17 and 46, respectively, and the internal representations of pitch-loudness-time are computed according to the following steps 13 and 14, followed by the computation of the difference function 12. Figure 1 The difference function is represented by operator 7. Two different types of perceptual difference functions were calculated: one for the overall interference introduced by the tested system using operators 7 and 8, and the other for the additional interference using operators 9 and 10. This modeled the asymmetry between the degradation effect caused by omitting time-frequency components from the reference signal and the degradation caused by introducing new time-frequency components. In POLQA, the two types of perceptual difference functions are calculated in two different ways: one focusing on degradation in the normal range, and the other focusing on loudness degradations, which cause... Figure 1 The four difference functions marked in the figure are used to calculate 7, 8, 9 and 10.

[0035] For the degraded output signal 49 with frequency domain distortion, an alignment algorithm 52 was used, which in Figure 2 The information is provided in the text. Figure 3 and Figure 4 The final processing method for obtaining the MOS-LQO (Mean Opinion Score - Listening Quality Objective) score is given.

[0036] POLQA begins by calculating some basic constant settings, then derives the pitch power density of the reference signal and the pitch power density (power as a function of time and frequency) of the degraded signal from the time- and frequency-aligned signal. Based on the pitch power density, internal characterizations of the reference signal and the degraded signal are obtained through multiple steps. Furthermore, these densities are used to derive the first three POLQA quality indicators: the quality indicator 41 for frequency response distortion (FREQ), the quality indicator 42 for additive noise (NOISE), and the quality indicator 43 for room reverberation (REVERB). These three quality indicators 41, 42, and 43 are calculated separately based on the primary interference indicator to enable balanced impact analysis of various distortion types. These indicators can also be used to perform a more detailed analysis of the types of degradation present in the speech signal using a degradation decomposition approach.

[0037] As described above, four different variations of the internal characterization of the reference signal and the degraded signal were calculated in steps 7, 8, 9, and 10. Two variations focus on interference for normal and large distortions, while the other two variations focus on interference increased by normal and large distortions. These four variations, 7, 8, 9, and 10, are the inputs for calculating the final interference density.

[0038] The internal representation of the reference input signal 3 is called the ideal representation because the low-level noise in the reference input signal is removed (step 33) and the timbre distortion that may be caused by the suboptimal timbre of the original reference recordings is partially compensated for in the degraded signal (step 35).

[0039] Four different variants of the ideal internal characterization and the degraded internal characterization calculated using operators 7, 8, 9 and 10 were used to calculate two final disturbance densities 142 and 143, one representing the final disturbance 142 as a function of time and frequency, focusing on the overall degradation, and the other representing the final disturbance 143 as a function of time and frequency, but focusing on the treatment of the increased degradation.

[0040] Figure 4 An overview of how the MOS-LQO score (i.e., the objective MOS score) is calculated based on two final interference densities 142 and 143, as well as FREQ indicator 41, noise indicator 42, and reverberation indicator 43.

[0041] Pre-calculation of constant settings

[0042] The FFT (Fast Fourier Transform) window size depends on the sampling frequency.

[0043] POLQA operates at different sampling rates of 8, 16, and 48 kHz, with window sizes W set to 256, 512, and 2048 samples respectively for each rate to match the temporal analysis window of the human auditory system. When using the Hann window, the overlap between consecutive frames is 50%. For both the reference and degraded signals, the power spectrum (the sum of the squares of the real and imaginary parts of the complex FFT components) is stored in separate real-valued arrays. The POLQA algorithm discards phase information within a single frame, and all calculations are based solely on the power representation.

[0044] Start and end point calculation

[0045] In subjective testing, noise typically occurs before the start of speech activity in the reference signal. However, it can be anticipated that leading steady-state noise in subjective testing reduces the effect of steady-state noise, while in objective measurements that take leading noise into account, it will increase that effect; therefore, it can be anticipated that omitting leading and trailing noise is the correct perceptual approach. Thus, after validating the expectation using available training data, the start and end points used in POLQA processing are calculated according to the beginning and end of the reference file. The sum of five consecutive absolute samples (using the normal 16-bit PCM range of ±32,000) from the beginning to the end of the original speech file must exceed 500 to designate that position as the start or end. The interval between the start and end is defined as the activity processing interval. Distortion outside this interval is ignored in POLQA processing.

[0046] Power scaling factor SP and loudness scaling factor SL

[0047] To calibrate the FFT time-frequency transform, a sine wave with a frequency of 1000 Hz and an amplitude of 40 dB SPL is generated using a reference signal X(t) calibrated towards 73 dB SPL. In steps 18 and 49, a windowed FFT is used to transform this sine wave to the frequency domain with lengths determined by the sampling frequencies for X(t) and Y(t), respectively. In steps 21 and 54, the frequency axis is converted to the Bark scale, and the peak amplitude of the resulting pitch power density is normalized to a power value of 10 by multiplying it by power scaling factors SP 20 and 50 for X(t) and Y(t), respectively. 4 .

[0048] A reference tone of the same 40 dB SPL was used to calibrate the psychoacoustic (sone) loudness scale. After distorting the intensity axis to the loudness scale using Zwicker's law, the integral of the loudness density on the Buck frequency scale was normalized to 1 sone in 30 and 58 using loudness scaling factors SL 31 and 59 for X(t) and Y(t), respectively.

[0049] Scaling and calculation of pitch power density

[0050] The degraded signal Y(t)5 is multiplied by a calibration factor C47 of 46, and then transformed to the time-frequency domain using a 50% overlap FFT frame. The calibration factor maps the dB overload in the digital domain to the dB SPL in the acoustic domain. Before being transformed to the time-frequency domain, the reference signal X(t)3 is scaled towards a predetermined fixed optimal level, approximately equivalent to 73 dB SPL. This calibration step is entirely different from the calibration step used in PESQ, where both the degraded and reference signals are scaled towards a predetermined fixed optimal level. PESQ presupposes that all playbacks are performed at the same optimal playback level, while in POLQA subjective testing, a level between 20 dB and +6 relative to the optimal level is used. Therefore, scaling towards a predetermined fixed optimal level is not used in the POLQA perceptual model.

[0051] After horizontal scaling, the reference and degraded signals are transformed to the time-frequency domain using a windowed FFT method (18, 49). For files where the frequency axis of the degraded signal is distorted compared to the reference signal, frequency domain dedistortion is performed on the FFT frame. In the first step of this dedistortion, both the reference and degraded FFT power spectra are preprocessed to reduce the impact of their very narrow frequency response distortion, along with the overall spectral shape difference, on the following calculations. Preprocessing 77 may include smoothing, compressing, and flattening the power spectrum. In 78, the smoothing operation is performed using a sliding window average of the power over the FFT band, while compression is simply done by taking the logarithm of the smoothed power in each band (79). In 80, the overall shape of the power spectrum is further flattened by applying a sliding window normalization to the smoothed logarithmic power over the FFT band. Next, the stochastic sub-harmonic pitch algorithm is used to calculate the pitch of the current reference frame and the degraded frame. The ratio 74 of the reference pitch to the degraded pitch ration is then used (in step 84) to determine the range of possible distortion factors. If possible, this search range is expanded using the pitch ratios of the previous and next frame pairs.

[0052] The frequency alignment algorithm then iterates within the search range, using the distortion factor of the current iteration to distort the degraded power spectrum by 85, and processes the distorted power spectrum using the aforementioned preprocessing 77. Then, for frequency bands below 1500 Hz, the correlation between the processed reference spectrum and the processed and distorted degraded spectra is calculated (in step 89). After completing the iterations within the search range, the "optimal" (i.e., the one producing the highest correlation) distortion factor is obtained in step 90. The correlation between the processed reference spectrum and the optimal distorted degraded spectrum is then compared to the correlation between the original processed reference spectrum and the degraded spectrum. If the correlation increases by a set threshold, the "optimal" distortion factor is retained (97). If necessary, the distortion factor is limited in 98 to have the largest relative change relative to the distortion factor determined for the previous frame.

[0053] After performing the necessary de-distortion to align the frequency axes of the reference signal and the degraded signal, in steps 21 and 54, the frequency scale in Hz is distorted to a pitch scale in barks, reflecting the fact that the human hearing system has a finer frequency resolution for low frequencies compared to high frequencies. This is achieved by binning the FFT bands and summing the corresponding frequencies of the FFT bands while normalizing the summation portion. For this purpose, the values ​​given in the literature are approximated by a distortion function that maps the frequency scale in Hertz to the pitch scale in barks, an approximation well known to those skilled in the art. The resulting reference signal and degraded signal are referred to as the pitch power density PPX(f). n (not in) Figure 1 (shown in the image) and PPY(f) n 56, where f is the frequency in bars and n represents the frame index.

[0054] Calculation of voice activity frames, silence frames, and super-silent frames (step 25)

[0055] In step 25, POLQA operates on three types of frames, the differences of which are as follows:

[0056] Voice activity frames, where the frame level of the reference signal is higher than the average level by approximately 20 dB;

[0057] A silent frame, wherein the frame level of the reference signal is approximately 20 dB below the average; and

[0058] Super-silent frames, in which the level of the reference signal is approximately 35 dB lower than the average.

[0059] Calculation of frequency, noise, and reverberation indicators

[0060] In step 40, the global effects of frequency response distortion, noise, and room reverberation are quantified separately. For the effect of overall global frequency response distortion, an indicator 41 is calculated based on the average spectra of the reference and degraded signals. To make the estimation of the effect of frequency response distortion independent of additive noise, the average noise spectral density of the degraded signal on the silent frames of the reference signal is subtracted from the pitch loudness density of the degraded signal. Then, the resulting pitch loudness density of the degraded signal is averaged with the pitch loudness density of the reference signal in each Barker band across all speech activity frames for both the reference and degraded files. The difference between these two densityes is then integrated over the pitch to obtain the indicator 41 (frequency) used to quantify the effect of frequency response distortion.

[0061] For the effect of additive noise, indicator 42 is calculated based on the average spectrum of the degraded signal on the silence frame of the reference signal. The difference between the average pitch loudness density of the degraded signal on the silence frame and the pitch loudness density of the zero reference signal determines the noise loudness density function that quantifies the effect of additive noise. The noise loudness density function is then integrated over pitch to obtain the average noise effect indicator 42 (noise). Therefore, indicator 42 is calculated based on ideal silence so that the transparent chain measured using a noisy reference signal does not provide the maximum MOS score in the final POLQA end-to-end speech quality measurement.

[0062] Regarding the impact of room reverberation, the energy over time (ETC) function is calculated based on the reference and degradation time series. ETC represents the envelope of the impulse response h(t) of the system H(f), and is defined as Y. a (f) = H(f)·X(f), where Y a (f) is the spectrum of the horizontally aligned degraded signal, and X(f) is the spectrum of the reference signal. Horizontal alignment (noise suppression) is used to suppress the global and local gain differences between the reference and degraded signals. In the first step, horizontal alignment is performed, for example, by scaling the degraded speech signal (or the reference signal, or both); smoothing is then performed by removing or suppressing peaks or spikes in the degraded signal. A second scaling step is then performed to level the volume in both signals to complete the horizontal alignment. The impulse response h(t) is calculated using the inverse discrete Fourier transform based on H(f). The ETC is calculated using normalization and clipping based on the absolute value of h(t).

[0063] exist Figures 7A to 7C The diagram illustrates an example of windowing a speech signal using a Hamming window. Figure 7A This is a schematic diagram of the Hamming window function 300. The Hamming window function is a bell-shaped function with a maximum value of 1.0 and values ​​of 0.0 at both ends. Figure 7B The image shows an arbitrary speech signal 301. For example... Figure 7C As shown, a windowing operation 320 can be performed on the speech signal 301 by performing a local convolution between the Hamming window 300 and the speech signal 301. Figure 8 The Hamming window 300 has a width corresponding to the temporal segment 305 of the audio sample to be created through the convolution step. The subsequent Hamming window 300 is applied to the speech signal 301 to produce multiple overlapping digital audio samples 308. Figure 7CIn this diagram, 50% overlap is illustrated using digital audio samples 308 in the interlacing diagram. This 50% overlap allows each portion of the signal to be fully considered on two consecutive samples 308. In this invention, depending on embodiments with or without portions of a fully degraded speech signal, the reverberation indicator 43 can be calculated using audio samples such as sample 308 obtained through a windowing operation performed on the degraded signal 5 and the reference signal 3. This windowing operation is performed on equivalent portions of the reference signal and the degraded signal. In POLQA, the duration of the temporal segment 305 used for windowing is significantly longer than the duration of a single frame. Figure 8 The method used is illustrated schematically.

[0064] According to some embodiments of the invention, the reverberation indicator 43 to be calculated may be based on a global or global reference speech signal 3 and a degraded speech signal 5, as well as multiple local samples 309 and 310 of the reference speech signal 3 and the degraded speech signal 5. For calculating the global estimate, the global or global reference speech signal 3 and the degraded speech signal 5 may be considered as a whole, or may be divided into long-duration signal segments (e.g., any suitable duration, such as >5 seconds or >10 seconds). Short local samples 309 and 310 may be obtained by performing windowing operations 320a and 320b on the reference speech signal 3 and the degraded speech signal 5 or their long-duration signal segments, or by integrating or combining multiple signal frames from the reference signal X(t) 3 and the degraded signal Y(t) 5. For example, short local samples 309 and 310 may include sound segments with a duration of, for example, 0.5 or 1.0 seconds (occasionally referred to herein as time-domain segment 305). Smaller segments may provide too little reverberation information. Short-duration local segments 309 and 310, obtained using windowing operations 320 (i.e., 320a and 320b), have been obtained, for example, by applying a Hamming window 300 with 50% overlap between them. Short-duration local samples 309 and 310 are formed by multiplying the degraded speech signal 5 with the applied windowing function 300 (e.g., a Hamming window function). For optimal determination of the local reverberation indicator, if the speech of the corresponding reference sample is below a threshold (which indicates a perceived silence interval), a weighting factor can be used to give lower weights to earlier degraded samples in the window. This weighting is performed in 321a and 321b. Subsequently, a Fast Fourier Transform (FFT) is performed on samples 309 and 310 and the overall degraded speech signal 5 in steps 322a and 322b. The global reference speech signal and the degraded speech signal 5 are processed by performing a Fast Fourier Transform (FFT) on the reference digital signal 3 and the degraded digital signal 5 in steps 340a and 340b. The FFT in steps 322a / b and 340a / b can be performed on a portion of the frequency range that includes the contribution of the speech signal (e.g., below 5 kHz or between 200 Hz and 4 kHz).

[0065] In steps 324 and 342, the transfer function H(f) is calculated from the transformed signal in the frequency domain. In steps 326 and 344, the impulse response signal is obtained by inverse FFT, thereby allowing the ETC to be calculated in steps 328 and 346. In steps 328 and 346, the ETC is determined for these long-duration signal portions (or the overall reference signal and the degraded signal) 3 and 5, and for the short-duration local samples 309 and 310, in the manner described above. In each ETC, one or more peaks are identified in steps 330 and 348, which appear delayed after the start of the energy-time curve based on the impulse response. For example, three maximum peaks appearing at least 60 milliseconds after the start of the energy-time curve can be identified. In steps 332 and 350, the energy in these peaks is determined and used in conjunction with the delayed position of these peaks on the time axis to calculate local and global reverberation indicators. For local samples and global portions, partial and global reverberation indicators can be calculated in steps 332 and 350, and the partial and global reverberation indicators can be combined in step 360 to produce a good estimate of the reverberation indicator 43 to be used.

[0066] Based on global and local samples of the ETC, multiple reflections are searched in each ETC in steps 330 and 348. In the first step, the loudest reflection is calculated by simply determining the maximum value of the ETC curve after the direct sound. In the POLQA model, the direct sound is defined as all sounds arriving within 60 milliseconds. Next, a second loudest reflection is determined based on the loudest reflection within an interval without direct sound and without considering reflections arriving within 100 milliseconds. Then, a third loudest reflection is determined based on the loudest reflection and the second loudest reflection within an interval without direct sound and without considering reflections arriving within 100 milliseconds. The energy and time delay of the three reflections are then combined into a single reverberation indicator 43 (reverberation).

[0067] Optionally, in the calculation of the reverberation indicator 43, a reverberation estimate within one standard deviation of the average of the partial reverberation estimates can be used. This estimate can then be weighted in a specific manner. In a computer program product developed to implement the method described herein, this can be achieved, for example, by a program such as:

[0068]

[0069] As an alternative to the above, the reverberation indicator can also be estimated based solely on short-duration local samples, which has provided an improvement over the traditional method of estimating the amount of reverberation in a signal.

[0070] Global and local scaling of the reference signal to the degraded signal (step 26)

[0071] At this point, according to step 17, the reference signal is at an internal ideal level, i.e., equivalent to approximately 73 dB SPL, while the degraded signal is represented as being at a level consistent with the playback level due to step 46. Global level difference compensation is performed in step 26 before comparing the reference and degraded signals. Furthermore, small changes in local levels are also partially compensated to take into account the fact that sufficiently small level variations are imperceptible to the listener in a listening-only situation. Global level equalization 26 is performed using frequency components between 400 and 3500 Hz based on the average power of the reference and degraded signals. The reference signal is globally scaled towards the degraded signal, and thus the effect of the global playback level difference is preserved at this stage of processing. Similarly, for slowly changing gain distortion, local scaling is performed using the full bandwidth of both the reference and degraded audio files for level variations up to approximately 3 dB.

[0072] Partial compensation for the original pitch power density to address linear frequency response distortion (step 27)

[0073] To accurately model the effects of linear frequency response distortion caused by filtering in the system under test, partial compensation is used in step 27. To model the imperceptibility of moderate linear frequency response distortion in subjective testing, the reference signal is partially filtered using the transfer characteristics of the system under test. This is achieved by calculating the average power spectrum of the original pitch power density and the degraded pitch power density for all speech activity frames. A partial compensation factor for each Bark bin is calculated based on the ratio of the degraded spectrum to the original spectrum.

[0074] Modeling of the masking effect and calculation of pitch loudness density activation

[0075] In steps 30 and 58, the masking is modeled by calculating a fuzzy representation of the pitch power density. Temporal fuzzing and frequency fuzzing are performed as follows: Figures 5a to 5c The principles illustrated are taken into account. Time-frequency domain fuzzification uses a convolutional approach. Based on this fuzzified representation, the representations of the reference pitch power density and the degraded pitch power density are recalculated, thereby suppressing low-amplitude time-frequency components that are partially masked by neighboring high-loudness components in the time-frequency plane. Suppression is achieved in two ways: subtracting the fuzzified representation from the unfuzzified representation; and dividing the unfuzzified representation by the fuzzified representation. Then, the resulting sharpened representation of the pitch power density is transformed into a pitch loudness density representation using a modified version of the Zwicker power law described below:

[0076]

[0077] Where SL is the loudness scaling factor, P0(f) is the absolute hearing threshold, and fB and Pfn are frequency- and level-related corrections defined as follows:

[0078] f B = -0.03*f + 1.06 when f < 2.0 barcks

[0079] f B =1.0 when 2.0≤f≤22bark

[0080] f B = -0.2*(f-22.0)+1.0 (When f>22.0) (Barks)

[0081] P fn =(PPX(f)) n +600) 0.008

[0082] Where f represents the frequency in bars, PPX(f) n Let f be the tone power density in the frequency time cell f,n. The two-dimensional array LX(f) is obtained at the output of step 30 for the reference signal X(t) and at the output of step 58 for the degraded signal Y(t), respectively. n and LY(f) n It is called pitch loudness density.

[0083] Global low-level noise suppression in reference and degraded signals

[0084] Due to the absolute classification rating test procedure, the test subject attributes low-level noise in the reference signal that is not affected by the tested system (e.g., a transparent system) to the tested system. Therefore, this low-level noise must be suppressed during the calculation of the internal characterization of the reference signal. In step 33, the average steady-state noise loudness density LX(f) of the reference signal as a pitch function is calculated for the super-silent frame. n Then, an "idealization process" is performed. Next, the average noise loudness density is partially subtracted from all the tone loudness density frames of the reference signal. At the output of step 33, the result is the idealized internal representation of the reference signal.

[0085] The audible steady-state noise in a degraded signal has a lower impact compared to non-steady-state noise. This applies to all noise levels, and the effect can be modeled by partially removing the steady-state noise from the degraded signal. This is done in step 60 by calculating the average steady-state noise loudness density LY(f) of the degraded signal as a function of pitch for a subset of frames. nTo achieve this, for these frames, the frames corresponding to these frames of the reference signal are classified as super-silent. Then, the average noise loudness density is partially subtracted from all the tonal loudness density frames of the degraded signal. Different strategies are used for partial compensation for low-level noise and high-level noise. For low-level noise, the compensation is negligible, while the suppression used becomes stronger for high-loudness additive noise. The result is an internal characterization 61 of the degraded signal with additive noise suitable for representing the subjective effects observed in listening tests using an idealized noise-free representation of the reference signal.

[0086] In step 33 above, in addition to performing global low-level noise suppression, a loudness indicator 32 is determined for each reference signal frame. The loudness indicator, or loudness value, can be used to determine a loudness-based weighting factor for weighting a specific type of distortion. Once the final interference densities 142 and 143 are provided, weighting can be implemented in steps 125 and 125' for the four representations of distortion provided by operators 7, 8, 9, and 10.

[0087] Here, the loudness level indicator has been determined in step 33; however, it should be understood that the loudness level indicator can be determined for each reference signal frame in other parts of the method. In step 33, determining the loudness level indicator is possible because the average steady-state noise loudness density LX(f) of the reference signal has already been determined for the super-silent frame. n The super-silent frame is then used to construct a noise-free reference signal for all reference frames. However, while this is possible in step 33, it is not the optimal implementation.

[0088] Optionally, a loudness level indicator (loudness) can be obtained from the reference signal in an additional step following step 35. This additional step... Figure 1 The loudness level indicator is represented by a dashed box 35' with a dashed output (loudness) of 32'. As will be understood by those skilled in the art, if step 35' is performed, it is no longer necessary to obtain the loudness level indicator from step 33.

[0089] Local scaling of distortion pitch loudness density for time-varying gain between the degraded signal and the reference signal (steps 34 and 63)

[0090] Slow changes in gain are inaudible, and small changes are already compensated for during the calculation of the reference signal representation. Before the internal characterization is correctly calculated, the required residual compensation is performed in two steps: first, in step 34, the reference signal is compensated for a signal level where the loudness of the degraded signal is lower than that of the reference signal; second, in step 63, the degraded signal is compensated for a signal level where the loudness of the reference signal is lower than that of the degraded signal.

[0091] For the portion of the degraded signal exhibiting severe signal loss (e.g., in the case of time clipping), the first compensation 34 scales the reference signal towards a lower level. This scaling ensures that the residual difference between the reference signal and the degraded signal represents the effect of time clipping on local perceived speech quality. The portion of the reference signal with a loudness lower than that of the degraded signal is not compensated for; therefore, additive noise and loud clicks are not compensated for in this first step.

[0092] For the portion of the degraded signal that exhibits clicking sounds and for the portion of the signal containing noise within silence intervals, the second compensation 63 scales the degraded signal towards a lower level. This scaling ensures that the residual difference between the reference signal and the degraded signal represents the impact of clicking sounds and slowly varying additive noise on the local perceived speech quality. Although clicking sounds are compensated in both silence and speech activity portions, noise is compensated only in silence portions.

[0093] Partial compensation for the original pitch loudness density to address frequency response distortion (step 35)

[0094] In step 27, imperceptible linear frequency response distortion has been compensated for by partially filtering the reference signal in the pitch power density domain. To further correct the fact that linear distortion is less objectionable than nonlinear distortion, in step 35, the reference signal is partially filtered in the pitch loudness domain. This is achieved by calculating the average power spectrum of the original pitch loudness density and the degraded pitch loudness density for all speech activity frames. A partial compensation factor for each Barker band is calculated based on the ratio of the degraded loudness spectrum to the original limit spectrum. This partial compensation factor is used to filter the reference signal, which has a smoothed, lower-amplitude frequency response of the system under test. After this filtering, the difference between the reference pitch loudness density and the degraded pitch loudness density caused by linear frequency response distortion is reduced to a level that represents the impact of linear frequency response distortion on perceived speech quality.

[0095] Final scaling of pitch loudness density and noise suppression

[0096] Up to this point, as used in subjective experiments, all calculations for the signals are performed at the playback level. At low playback levels, this will produce a small difference between the reference pitch loudness density and the degraded pitch loudness density, and generally yield a much more optimistic estimate of the listening speech quality. To compensate for this effect, the degraded signal is now scaled toward a “virtual” fixed internal level in step 64. After this scaling, the reference signal is scaled toward the degraded signal level in step 36, and at this point both the reference and degraded signals are ready for the final noise suppression operations in steps 37 and 65, respectively. Noise suppression processes the final portion of the steady-state noise level in the loudness domain that still has a significant impact on speech quality calculations. The resulting signals 13 and 14 are in the perceptual-relevant internal representation domain and are based on the ideal pitch-loudness-time function LX. 理想 (f) n 13 and degraded pitch-loudness-signal function LY 劣化 (f) n 14. Interference densities 142 and 143 can be calculated. Four different variations of the ideal pitch-loudness-time function and the degraded pitch-loudness-signal function are calculated in 7, 8, 9, and 10. Two variations (7 and 8) focus on interference for normal and large distortion, and two variations (9 and 10) focus on interference for increased normal and large distortion.

[0097] Calculation of final interference density

[0098] Calculate two different interference densities, 142 and 143. The first, the normal interference density, is calculated in 7 and 8 according to the ideal pitch-loudness-time function LX. 理想 (f) n With degraded pitch-loudness-signal function LY 劣化 (f) n The difference is obtained. The second method, used in 9 and 10, is an optimized version relative to the introduced degradation, obtained according to the ideal pitch-loudness-time function and the degraded pitch-loudness-signal function, and is referred to as the added interference. In the calculation of this added interference, the signal portion with a degradation power density greater than the reference power density is weighted by a factor (asymmetric factor) that depends on the power ratio in each pitch-time unit.

[0099] To handle a wide range of distortion, two different processing methods were implemented: one based on 7 and 9, focusing on small to medium distortion, and the other based on 8 and 10, focusing on medium to large distortion. Switching between the two was based on a first estimate derived from interferences focusing on small to medium levels of distortion. This processing method necessitated the calculation of four different ideal pitch-loudness-time functions and four different degraded pitch-loudness-time functions to be able to calculate individual interferences and individual added interference functions (see...). Figure 3 While a single disturbance and a single added disturbance function compensate for a large number of specific distortions of various types.

[0100] Significant deviations from the optimal listening level are quantified in 127 and 127' using an indicator derived directly from the signal level of the degraded signal. A global indicator (level) is also used in the MOS-LQO calculation.

[0101] The severe distortion introduced by frame repetition is quantified in 128 and 128' by an indicator obtained by comparing the correlation of consecutive frames of the reference signal with the correlation of consecutive frames of the degraded signal.

[0102] The significant deviation from the optimal "ideal" timbre of the degraded signal is quantified in 129' and 129' using an indicator derived from the loudness difference between the higher and lower frequency bands. The timbre indicator is calculated based on the loudness difference between 2-12 bark in the low-frequency portion of the degraded signal's bark band and 7-17 bark in the higher range (i.e., using 5-bark overlap), which "penalizes" any significant imbalance regardless of the fact that this might be a result of an incorrect timbre in the reference signal. Compensation is performed for each frame and at a global level. This compensation calculates the power in both the lower and higher bark bands of the degraded signal (less than 12 bark and greater than 7 bark, i.e., using 5-bark overlap), and the loudness difference "penalizes" any significant imbalance regardless of the fact that this might be a result of an incorrect timbre in the reference signal. It should be noted that in POLQA end-to-end speech quality measurements, transparent chains using poorly recorded reference signals, containing excessive noise and / or incorrect timbre, will not provide the maximum MOS score. This compensation also has an impact when measuring the quality of transparent devices. When the reference signal used shows a significant deviation from the optimal "ideal" timbre, the tested system will be judged as opaque, even if the system has not introduced any degradation into the reference signal.

[0103] The impact of severe peaks in the interference is quantified in 130 and 130' by the flatness indicator, which is also used in the calculation of MOS-LQO.

[0104] The subject's attention is focused on the changes in the severity of the noise level, which are quantified in 131 and 131' by a noise contrast indicator obtained from a degraded signal frame with a corresponding silent reference frame.

[0105] In steps 133 and 133', a weighting operation is performed to weight the interference based on whether it is consistent with the actual spoken voice. To assess the quality or intelligibility of the degraded signal, interference perceived during the silence phase is not considered as detrimental as interference perceived during the actual spoken voice. Therefore, a weighting value is determined to weight the interference based on the loudness indicator determined according to the reference signal in step 33 (or alternatively step 35'). The weighting value is used to weight the difference function (i.e., the interference) to incorporate the impact of the interference on the quality or intelligibility of the degraded speech signal into the evaluation. In particular, since the weighting value is determined based on the loudness indicator, it can be expressed as a loudness-based function. The loudness-based weighting value can be determined by comparing the loudness to a threshold. If the loudness indicator exceeds the threshold, the perceived interference is fully considered during the evaluation. On the other hand, if the loudness value is less than the threshold, the weighting is based on the loudness level indicator; that is, in this example, the weighting is equal to the loudness level indicator (in a system where the loudness is less than the threshold). The advantage is that for weak portions of a speech signal, such as at the end of a spoken word before a pause or silence, interference is partially considered detrimental to quality or intelligibility. As an example, it should be understood that a certain amount of noise perceived when saying the letter "f" at the end of a word might cause the listener to perceive it as the letter "s." This would be detrimental to quality or intelligibility. On the other hand, those skilled in the art should understand that when the loudness value is below the aforementioned threshold, any noise during silence or pause can also be simply ignored by setting the weighting to zero.

[0106] Back to Figure 3 During the alignment process, severe jumps that occur during the alignment process are detected, and the impact is quantified by a compensation factor in steps 136 and 136'.

[0107] Finally, the interference density and the increased interference density are clipped to the maximum level at 137 and 137', and the specific temporal structure of the interference is compensated using the variance of the interference at 138 and 138' and the effect of the jumps at 140 and 140' on the loudness of the reference signal.

[0108] This yields the final interference density D(f) for regular interference. n 142 and the final interference density DA(f) for increased interference. n 143.

[0109] The mapping of interference with pitch, burst, and time to intermediate MOS scores.

[0110] The final interference density D(f) for each frame on the pitch axis. n 142 and the final increased interference density DA(f) n 143. Integrating this yields two distinct per-frame interferences: one obtained using L1 integration 153 and derived from the interference itself, and the other obtained using L1 integration 159 and derived from the added interference (see [link to relevant documentation]). Figure 4 ):

[0111] D n =∑ f=1,.. Number of Buck bands |D(f)n|w f

[0112] DA n =∑ f=1,.. Number of Buck bands | DA(f) n |W f

[0113] Among them, W f These are a series of constants proportional to the Barker band.

[0114] Next, the average of the two interferences per frame is calculated using L4 155 weighting for the interference and L1 160 weighting for the added interference over 6 consecutive speech frames, and is defined as a speech burst.

[0115]

[0116]

[0117] Finally, for each file, the interference and added interference are calculated based on the averaging over time of L2 156 and 161:

[0118]

[0119]

[0120] In step 161, reverberation indicator 42 and noise indicator 43 are used to compensate for the increased interference for loud reverberation and loud additive noise. However, the two interferences are combined with frequency indicator 41 (frequency) 170 to obtain an internal indicator, which is linearized by a third-order regression polynomial to obtain a MOS-like intermediate indicator 171.

[0121] Final POLQA MOS-LQO calculation

[0122] In step 175, the following four different compensations are used to obtain the unprocessed POLQA score based on the MOS intermediate indicator:

[0123] Two compensation methods are used to address the specific time-frequency characteristics of interference, one of which is applied at frequency 148, burst 149, and time 150. 511 The calculation is performed by aggregation, and an L is used at frequency 145, burst 146, and time 147. 313 Aggregation is used for computation;

[0124] A form of compensation for using level indicators for very low presentation levels;

[0125] A compensation method using a flatness indicator in the frequency domain to address large tonal distortions.

[0126] The training of this mapping is performed on a large set of degradations, including degradations that are not part of the POLQA benchmark. These unprocessed MOS scores 176 target the principal part, which has been linearized by the third-order polynomial mapping used in the calculation of the MOS-like intermediate indicator 171.

[0127] Finally, in step 180, a third-order polynomial is used to map the unprocessed POLQA MOS score 176 to the MOS-LQO score 181. This polynomial is optimized for the 62 databases available in the final stage of POLQA standardization. The maximum POLQA MOS-LQO score is 4.5 in narrowband mode and 4.75 in the ultra-wideband model. An important consequence of this idealization is that in some cases, when the reference signal contains noise or when the sound timbre is severely distorted, the transparent chain will not provide the maximum MOS score of 4.5 in narrowband mode or 4.75 in ultra-wideband mode.

[0128] The consonant-vowel-consonant compensation according to the present invention can be implemented as follows. Figure 1 In this process, reference signal frame 220 and degraded signal frame 240 can be obtained as described above. For example, reference signal frame 220 can be obtained from step 21, which twists the reference signal to Buck, while degraded signal frame can be obtained from the corresponding step 54 performed for the degraded signal. Figure 1 The precise locations of the reference signal frame and / or degraded signal frame shown, obtained according to the method of the present invention, are merely examples. The reference signal frame 220 and the degraded signal frame 240 can be obtained from... Figure 1The degradation signal frame can be obtained from any other step in the process, particularly from somewhere between the input of the reference signal X(t)3 and the global and local scaling to the degraded level in step 26.

[0129] Consonant-vowel-consonant compensation, such as Figure 6 As shown. First, in step 222, the signal power of the reference signal frame 220 is calculated in the desired frequency domain. Ideally, this frequency domain for the reference frame includes only the speech signal (e.g., a frequency range between 300 Hz and 3500 Hz). However, in step 224, a selection is made as to whether to include the reference frame as the active speech reference frame by comparing the calculated signal power with a first threshold 228 and a second threshold 229. As described in POLQA (ITU-T Recommendation P.863), when scaling the reference signal is used, the first threshold may, for example, be equal to 7.0 × 10⁻⁶. 4 And similarly, the second threshold can be equal to 2.0 × 2 × 10. 8 In step 225, a reference signal frame corresponding to the soft speech reference signal (the key part of the consonant) is selected for processing by comparing the calculated signal power with a third threshold 230 and a fourth threshold 231. The third threshold 230 may, for example, be equal to 2.0 × 10⁻⁶. 7 And the fourth threshold can be equal to 7.0 × 10 7 .

[0130] Steps 224 and 225 yield reference signal frames corresponding to the active frame portion and the soft speech frame portion, respectively, namely, active speech reference signal portion frame 234 and soft speech reference signal portion frame 235. These frames are provided to step 260, which will be discussed below.

[0131] Similar to the calculation of the correlated signal portion of the reference signal, firstly, in step 242, the degraded signal frame 240 is analyzed to calculate the desired signal power in the frequency domain. For the degraded signal frame, it is advantageous to calculate the signal power over a frequency range that includes the range of spoken sound frequencies and over a frequency range where most audible noise exists, for example, a frequency range between 300 Hz and 8000 Hz.

[0132] Based on the signal power calculated in step 242, relevant frames (i.e., frames associated with relevant reference frames) are selected. Selection occurs in steps 244 and 245. In step 245, for each degraded signal frame, it is determined whether it is time-aligned with the reference signal frame selected as the soft speech reference signal frame in step 225. If the degraded frame is time-aligned with the soft speech reference signal frame, it is identified as a soft speech degraded signal frame, and the calculated signal power is used in the calculation in step 260. Otherwise, the frame is discarded as a soft speech degraded signal frame for calculating the compensation factor in step 247. In step 244, for each degraded signal frame, it is determined whether it is time-aligned with the reference signal frame selected as the active speech reference signal frame in step 224. If the degraded frame is time-aligned with the active speech reference signal frame, it is identified as an active speech degraded signal frame, and the calculated signal power is used in the calculation in step 260. Otherwise, the frame is discarded as an active speech degraded signal frame for calculating the compensation factor in step 247. This results in the soft speech degradation signal portion frame 254 and the active speech degradation signal portion frame 255 being provided to step 260.

[0133] Step 260 receives the following as inputs: active speech reference signal portion frame 234, soft speech reference signal portion frame 235, soft speech degradation signal portion frame 254, and active speech degradation signal portion frame 255. In step 260, the signal power of these frames is processed to determine the average power for the active speech reference signal portion and the soft speech reference signal portion, as well as for the active speech degradation signal portion and the soft speech degradation signal portion, and based on this (also in step 260), the consonant-vowel-consonant signal-to-noise ratio compensation parameter (CVC) is calculated. SNR_因子 )as follows:

[0134]

[0135] Parameters Δ1 and Δ2 are constant values ​​used to adapt the model's behavior to the subject's behavior. Other parameters in the formula are as follows: P 活动,参考,平均 The average active speech reference signal power. Parameter P 软,参考,平均 The average soft speech reference signal power. Parameter P 活动,劣化,平均 The average active speech degradation signal power, and parameter P 软,劣化,平均 This is used to average the signal power of the degraded soft speech signal. At the output of step 260, the consonant-vowel-consonant signal-to-noise ratio compensation parameter CVC is provided. SNR_因子 .

[0136] In step 262, CVC SNR_因子 Compared to the threshold of 0.75 in this example. If CVC SNR_因子If the value exceeds this threshold, the compensation factor is set to 1.0 in step 265 (no compensation occurs). In CVC SNR_因子 If the value is less than the threshold (0.75 in this case), the compensation factor is calculated in step 267 as follows: Compensation factor = (CVC) SNR_因子 +0.25) 1 / 2 (Note that the value 0.25 is obtained by taking the range 1.0-0.75, where 0.75 is used for comparing CVC.) SNR_因子 The resulting compensation factor of 270 is at the threshold. Figure 4 In step 182, it is used as a multiplier for the MOS-LQO score (i.e., the overall quality parameter). As will be understood, compensation (e.g., multiplication) does not necessarily have to occur in step 182, but can be incorporated into one of steps 175 or 180 (in which case step 182 will be removed from step 182). Figure 4 (The figure disappears). However, in this example, compensation is achieved by multiplying the MOS-LQO score by the compensation factor calculated as described above. It should be understood that compensation can also take another form. For example, it may also depend on CVC. SNR_因子 A variable is subtracted from or added to the obtained MOS-LQO. Those skilled in the art will understand and recognize other meanings of the compensation in accordance with the teachings of this invention.

[0137] The invention has been described according to certain specific embodiments thereof. It should be understood that the embodiments shown in the drawings and described herein are for illustrative purposes only and are not intended to limit the invention in any way or by any means. It should be understood that the operation and structure of the invention will be apparent from the foregoing description and drawings. It will be apparent to those skilled in the art that the invention is not limited to any of the embodiments described herein, and that modifications should be considered within the scope of the appended claims. Furthermore, kinematic inversion is considered inherently disclosed and is within the scope of the invention. Moreover, any components and elements of the various disclosed embodiments may be combined or incorporated into other embodiments where deemed necessary, desirable, or preferred, without departing from the scope of the invention as defined in the claims.

[0138] In the claims, any reference numerals shall not be construed as limiting the claims. When the term "comprising / including" is used in this specification or the appended claims, it should not be interpreted in an exclusive or exhaustive sense, but rather in an inclusive sense. Therefore, the expression "comprising" as used herein does not exclude the presence of other elements or steps besides those listed in any of the claims. Furthermore, the term "a / an" should not be construed as limited to "only one," but is used to mean "at least one," and does not exclude a plurality. Features not specifically or explicitly described or claimed may be additionally included in the structure within the scope of the invention. Expressions such as "meaning..." should be understood as "component configured for..." or "component constructed as..." and should be construed as including equivalents of the disclosed structure. Expressions such as "critical," "preferred," and "particularly preferred" are used. This is not intended to limit the invention. Additions, deletions, and modifications can generally be made within the scope of a person skilled in the art without departing from the spirit and scope of the invention, as determined by the claims. The invention may be practiced in other ways as specifically described herein and is limited only by the appended claims.

[0139] Figure Labels

[0140] 3. Reference signal X(t)

[0141] 5. Degraded signal Y(t), amplitude-time

[0142] 6. Delay identifiers to form frame pairs

[0143] 7. Difference Calculation

[0144] The first variation of 8 difference calculation

[0145] 9. The second variation of difference calculation

[0146] The third variation of the 10-difference calculation

[0147] 12 bad signal

[0148] 13 Internal Ideal Pitch-Loudness-Time LX 理想 (f)n

[0149] 14 Internal Degradation Pitch-Loudness-Time LY 劣化 (f)n

[0150] 17 Global zoom towards a fixed level

[0151] 18. Windowed FFT

[0152] 20 scaling factor SP

[0153] 21 Twist to Bark

[0154] 25 (Super) Silent Frame Detection

[0155] 26 Global and Local Scaling to Degradation Level

[0156] 27 Partial Frequency Compensation

[0157] 30 Activation and Twist to Song

[0158] 31 Absolute Threshold Scaling Factor SL

[0159] 32 Loudness

[0160] 32’ (Determined according to alternative step 35’) Loudness

[0161] 33 Global Low-Level Noise Suppression

[0162] 34 Local Compensation in the Case of Y < X

[0163] 35 Partial Frequency Compensation

[0164] 35’ (Alternative) Determine Loudness

[0165] 36 Scaling towards Degradation Level

[0166] 37 Global Low-Level Noise Suppression

[0167] 40 Frequency, Noise, Reverberation Indicator

[0168] 41 Frequency Indicator

[0169] 42 Noise Indicator

[0170] 43 Reverberation Indicator

[0171] 44 PW_R 总体 Indicator (Overall Audio Power Ratio between Degraded Signal and Reference Signal)

[0172] 45 PW_R 总体 Indicator (Audio Power Ratio per Frame between Degraded Signal and Reference Signal)

[0173] 46 Scaling towards Playback Level

[0174] 47 Calibration Factor C

[0175] 49 Windowed FFT

[0176] 52 Frequency Alignment

[0177] 54 Twist to Bark

[0178] 55 scaling factor SP

[0179] 56 Degraded Signal Tone-Power-Time (PPY) (f)n

[0180] 58 Activation and Twisting to Song

[0181] 59 Absolute threshold scaling factor SL

[0182] 60 Global high-level noise suppression

[0183] 61 Degraded signal pitch-loudness-time

[0184] 63 Local compensation when Y>X

[0185] 64. Scaling towards a fixed internal level

[0186] 65. Global high-level noise suppression

[0187] 70 Reference Spectrum

[0188] 72 Degraded Spectrum

[0189] 74. Ratio of reference pitch to degraded pitch in the current frame and surrounding frames (+ / -1).

[0190] 77 Preprocessing

[0191] 78 Narrow peaks and drops in smoothed FFT spectra

[0192] 79. Take the logarithm of the spectrum and apply a threshold for the minimum intensity.

[0193] 80. Use a sliding window to flatten the shape of the total logarithmic spectrum.

[0194] 83 Optimization Cycle

[0195] 84. Range of distortion factor: [Minimum pitch ratio <= 1 <= Maximum pitch ratio]

[0196] 85 Distorted and Degraded Spectrum

[0197] 88 Application Preprocessing

[0198] 89. Calculate the correlation of the spectrum for frequency bands below 1500Hz.

[0199] 90 Tracking the best distortion factor

[0200] 93 Distorted and degraded spectrum

[0201] 94 Application Preprocessing

[0202] 95. Correlation of the spectrum calculated for frequency bands below 3000Hz

[0203] 97. If the correlation is sufficiently high, retain the distorted, degraded spectrum; otherwise, restore the original spectrum.

[0204] 98. Limit the variation of the distortion factor from one frame to the next.

[0205] 100 Ideal Convention

[0206] 101 Deterioration of Standard

[0207] 104. Ideals Greatly Distorted

[0208] 105. Significant distortion due to degradation.

[0209] 108 Ideal increase

[0210] 109 Deterioration increased

[0211] 112 The significant distortion added to the ideal

[0212] 113. Significant distortion due to degradation

[0213] 116. Conventional selection of interference density

[0214] 117 High Interference Density Distortion Selection

[0215] 119. Increased selection of interference density

[0216] 120 Increased interference density, large distortion selection

[0217] 121 Switching function 123's PW_R 总体 enter

[0218] 122 Switching function 123's PW_R 帧 enter

[0219] 123 Large Distortion Detection (Switch)

[0220] 125 Correction factors for a large number of specific distortions

[0221] 125' Correction factor for a large number of specific distortions

[0222] 127 level

[0223] 127' Level

[0224] 128 frames repeat

[0225] 128' frame repeat

[0226] 129 timbres

[0227] 129' Tone

[0228] 130 Spectral flatness

[0229] 130' Spectral flatness

[0230] 131 Noise contrast during quiet periods

[0231] 131' Noise contrast during silent periods

[0232] 133 Loudness-based interference weighting

[0233] 133' Loudness-based interference weighting

[0234] 134 Loudness of the reference signal

[0235] Loudness of the 134' reference signal

[0236] 136 Alignment Jump

[0237] 136' Alignment Jump

[0238] 137 Peak shaving to maximum degradation

[0239] 137' Peak clipping to maximum degradation

[0240] 138. Disturbance variance

[0241] 138' Interference Variance

[0242] 140 loudness jump

[0243] 140' loudness jump

[0244] 142 Final Interference Density D (f)n

[0245] 143 The final increase in interference density DA (f)n

[0246] 145L3 Frequency Integral

[0247] 146L1 Burst Integrals

[0248] 147L3 Time Integration

[0249] 148L5 Frequency Integral

[0250] 149L1 Burst Integrals

[0251] 150L1 Time Integration

[0252] 153L1 Frequency Integral

[0253] 155L4 Burst Integrals

[0254] 156L2 Time Integral

[0255] 159L1 Frequency Integral

[0256] 160L1 Burst Integrals

[0257] 161L2 Time Integration

[0258] 170 mapped to intermediate MOS fraction

[0259] 171 Type MOS Intermediate Indicators

[0260] 175 MOS Scale Compensation

[0261] 176 Unprocessed MOS score

[0262] 180 mapped to MOS-LQO

[0263] 181 MOS LQO

[0264] 182 CVC Intelligibility Compensation (Intelligibility Model Only)

[0265] 185. Changes in intensity of a short sine tone over time

[0266] 187 short sine tone

[0267] 188 Masking threshold of the second short sine tone

[0268] 195. Changes in intensity of a short sine tone with frequency

[0269] 198 short sine tones

[0270] 199 Masking threshold of the second short sine tone

[0271] 205. Intensity variation with frequency and time in a 3D plot.

[0272] 211 The masking threshold used as the suppression intensity leads to sharpening of the internal characterization.

[0273] 220 Reference signal frame (see also) Figure 1 )

[0274] 222 Determine the signal power in the speech domain (e.g., 300Hz-3500Hz).

[0275] 224. Compare the signal power with the first and second thresholds; if it is within the range, then select...

[0276] 225. Compare the signal power with the third and fourth thresholds; if it is within the range, then select...

[0277] 228 First Threshold

[0278] 229 Second Threshold

[0279] 230 Third Threshold

[0280] 231 Fourth Threshold

[0281] 234 Power average of active voice reference signal frames

[0282] 235 Power Average of Soft Speech Reference Signal Frames

[0283] 240 Degraded signal frames (see also) Figure 1 )

[0284] 242 Determine the signal power in the domain for speech and audible interference (e.g., 300Hz–8000Hz).

[0285] 244 degraded frames are time-aligned with the selected active speech reference signal frames.

[0286] 245 Degraded frames are time-aligned with the selected soft speech reference signal frames.

[0287] 247 frames were discarded as active / soft speech degradation signal frames.

[0288] Average power of 254 soft speech degradation signal frames

[0289] Average power of 255 active voice degradation signal frames

[0290] 260. Calculate the consonant-vowel-consonant signal-to-noise ratio compensation factor (CVC). SNR_因子 )

[0291] 262 CVC SNR_因子 Is it less than the compensation threshold (e.g., 0.75)?

[0292] 265 No → Compensation factor = 1.0 (No compensation)

[0293] 267 is → Compensation factor is (CVC) SNR_因子 +0.25) 1 / 2

[0294] 270 Provides a compensation value to step 182 for compensating MOS-LQO

Claims

1. A method for determining the perceptual impact of echo or reverberation in a degraded audio signal on the perceived quality of the degraded audio signal, wherein, The method includes the following steps: receiving the degraded audio signal from an audio transmission system and obtaining the degraded audio signal by transmitting a reference audio signal through the audio transmission system to provide the degraded audio signal; The controller obtains at least one degraded digital audio sample from the degraded audio signal and at least one reference digital audio sample from the reference audio signal; The controller determines the local impulse response signal based on the at least one degraded digital audio sample and the at least one reference digital audio sample; The controller determines an energy-time curve based on the local impulse response signal, wherein the energy-time curve is proportional to the square root of the absolute value of the local impulse response signal; and Based on the local impulse response signal, one or more peaks in the energy-time curve are identified, the one or more peaks occurring in time at a delay in the energy-time curve after the start of the energy-time curve, and an estimate of the echo or reverberation is determined based on the amount of energy in the one or more peaks. The degraded audio signal includes multiple degraded signal frames, and the reference audio signal includes multiple reference signal frames. The step of obtaining the at least one degraded digital audio sample includes: sampling the degraded audio signal from multiple consecutive degraded signal frames in a time-domain segment, the time-domain segment having a duration greater than 0.3 seconds, the sampling including performing a windowing operation on the degraded audio signal by multiplying the degraded audio signal by a window function to generate the degraded digital audio sample; and The step of obtaining the at least one reference digital audio sample includes: sampling the reference audio signal from multiple consecutive reference signal frames in the time domain segment, wherein the sampling includes performing a windowing operation on the reference audio signal by multiplying the reference audio signal with the window function to generate the reference digital audio sample; The window function used to obtain the at least one reference digital audio sample and the at least one degraded digital audio sample has a non-zero value in the time-domain segment to be sampled and a zero value outside the time-domain segment.

2. The method according to claim 1, wherein, The step of obtaining the at least one digital audio sample includes: obtaining a plurality of digital audio samples from the audio signal, wherein each of the plurality of digital audio samples is obtained by performing the windowing operation, and the time-domain segments of at least two consecutive digital audio samples of the plurality of digital audio samples overlap.

3. The method according to claim 2, wherein, The overlap between the at least two consecutive digital audio samples is within the range of 10% to 90% overlap between the time-domain segments.

4. The method according to any one of claims 1 to 3, wherein, The window function is at least one of the following: Hamming window, von Hahn window, Tugi window, cosine window, rectangular window, B-spline window, triangular window, Bartlett window, Patzen window, Welch window, cosine to the power of n window, Kaiser window, Nuttor window, Blackman window, Blackman-Harris window, Blackman-Nuttor window, or flat-top window, where n>1.

5. The method according to any one of claims 1 to 3, wherein, The step of determining the estimate of the reverberation amount includes: weighting the amount of energy in each peak based on the amplitude of each peak or the delayed position of each peak along the time axis.

6. The method according to any one of claims 1 to 3, wherein, The method further includes the following steps: The controller obtains a degraded digital signal, the degraded digital signal representing at least a portion of the degraded audio signal, and having a duration longer than the time-domain segment of the at least one degraded digital audio sample; The controller obtains a reference digital signal, which represents at least a portion of the reference audio signal and has a duration longer than the time-domain segment of the at least one reference digital audio sample; The controller determines the global impulse response signal based on the at least one degraded digital signal and the at least one reference digital signal; The controller determines a global energy-time curve based on the global impulse response signal, wherein the global energy-time curve is proportional to the square root of the absolute value of the global impulse response signal; and Based on the global impulse response signal, one or more further peaks in the global energy time curve are identified, the one or more further peaks occurring at a delay in the global energy time curve after the start of the global energy time curve, and a further estimate of the echo or reverberation is determined based on the amount of energy in the one or more further peaks.

7. The method according to claim 6, wherein, The step of determining a further estimate of the reverberation amount includes: weighting the amount of energy in each peak based on the amplitude of each further peak or the delayed position of each further peak along the time axis.

8. The method of claim 6, further comprising at least one of the following steps: The controller calculates a partial reverberation indicator value based on the estimated echo or reverberation amount obtained from the at least one degraded digital audio sample and the at least one reference digital audio sample; The controller calculates the global reverberation indicator value based on the further estimated echo or reverberation amount; The controller calculates the final reverberation indicator value based on the estimated echo or reverberation amount and the further estimated echo or reverberation amount.

9. The method according to claim 6, wherein, The steps of determining the local impulse response signal based on the digital audio samples, or determining the global impulse response signal based on the digital signal, include: The controller applies a Fourier transform to the digital audio sample or the digital signal, converting the digital audio sample or the digital signal from the time domain to the frequency domain. The controller determines the transfer function based on the power spectrum signal from the digital audio sample or the digital signal in the frequency domain; and The controller converts the transfer function from the frequency domain to the time domain to generate the local impulse response signal or the global impulse response signal.

10. The method according to any one of claims 1-3, wherein, The step of determining the local impulse response signal includes: using a weighting factor to indicate a perceived silence interval if the speech of the corresponding reference digital audio sample is below a threshold, thereby giving lower weights to the earlier degraded digital audio samples in the window.

11. A method for evaluating the quality or intelligibility of a degraded speech signal received from an audio transmission system, wherein a reference speech signal is transmitted through the audio transmission system to provide the degraded speech signal, wherein, The method includes: - The reference speech signal is sampled into multiple reference signal frames, the degraded speech signal is sampled into multiple degraded signal frames, and frame pairs are formed by associating the reference signal frames and the degraded signal frames with each other; - Provide a difference function for each frame pair, the difference function representing the difference between the degraded signal frame and the associated reference signal frame; - Compensate the difference function for one or more types of interference to provide an interference density function for each frame pair, the interference density function being suitable for a human auditory perception model; - An overall quality parameter is obtained based on the interference density function of multiple frame pairs, the quality parameter indicating at least the quality or intelligibility of the degraded speech signal; The method further includes the following steps: - Determine the reverberation amount in at least one of the degraded speech signal and the reference speech signal, wherein the reverberation amount is determined by applying the method according to any one of claims 1 to 10.

12. The method according to claim 11, wherein, The step of obtaining the at least one degraded digital audio sample and the at least one reference digital audio sample by the controller is performed by forming the degraded audio sample and the reference audio sample from a plurality of consecutive signal frames, the signal frames including a plurality of the degraded signal frames and a plurality of the reference signal frames.

13. The method according to claim 12, wherein, The number of signal frames to be included in the plurality of consecutive signal frames depends on the duration of the time-domain segment of the at least one digital audio sample, wherein the duration is between 0.4 seconds and 5.0 seconds.

14. The method according to claim 12, wherein, For each frame pair, the compensation step is performed by setting the determined reverberation amount of at least one of the degraded speech signal and the reference speech signal to one of the one or more interference types, and compensating the reverberation amount associated with the corresponding frame pair for each frame pair according to the formation of the digital audio samples.

15. The method according to any one of claims 11-14, further comprising a noise suppression step before the step of determining the impulse response signal, the noise suppression comprising the following steps: At least one of the degraded speech signal or the reference speech signal is first scaled to obtain a similar average volume. The degraded speech signal is processed to remove one or more of local signal peaks, peak clipping, and signal loss from the degraded speech signal; At least one of the degraded speech signal or the reference speech signal is subjected to a second scaling to obtain a similar average volume.

16. The method according to any one of claims 11-14, wherein, The method is performed on the audio signal within a predetermined frequency range, which is below a threshold frequency or a frequency range corresponding to the speech signal.

17. The method according to claim 16, wherein, The frequency range is below 5 kHz.

18. A computer program product suitable for loading into a memory of a computer system, the computer program product comprising instructions that, when loaded into the memory and processed by a controller of the computer system, cause the computer system to perform the method according to any one of the preceding claims.