Method for determining the perceptual effect of reverberation on the perceived quality of a signal and computer program product

JP2023535694A5Pending Publication Date: 2025-12-15NEDERLANDSE ORG VOOR TOEGEPAST NATUURWETENSCHAPPELIJK ONDERZOEK TNO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023503439
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-07-20
Filing Date
2021-07-19
Publication Date
2025-12-15

AI Technical Summary

Technical Problem

Existing methods for determining the perceptual impact of reverberation on audio quality are inaccurate due to noise, pulse distortions, and time-shift distortions, especially in broadband audio signals.

Method used

A method involving windowing operations on both degraded and reference audio signals to estimate the amount of reverberation, using various window functions like Hamming, Von Hann, and Tukey windows, and calculating energy-time curves to accurately determine the perceptual impact of reverberation.

Benefits of technology

The method provides a more accurate estimation of reverberation impact, improving the assessment of audio quality and intelligibility by compensating for local and global disturbances, enhancing the precision of perceptual quality measurements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000028_0000
    Figure 00000028_0000
  • Figure 00000028_0001
    Figure 00000028_0001
  • Figure 00000028_0002
    Figure 00000028_0002
Patent Text Reader

Abstract

The present disclosure relates to a method for determining the perceptual effect of the amount of echo or reverberation in a degraded audio signal on the perceived quality of the degraded audio signal, wherein the degraded audio signal is obtained by transmitting a reference audio signal through the audio transmission system to provide the degraded audio signal. The method includes performing a windowing operation on the degraded digital audio samples and the reference digital audio samples by multiplying the degraded digital audio samples and the reference digital audio samples by a window function to obtain the degraded digital audio samples and the reference digital audio samples. A local estimate of the amount of echo or reverberation is determined based on these samples.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for determining the perceptual effect of the amount of echo or reverberation in a degraded audio signal on the perceived quality of the degraded audio signal, the degraded audio signal being received from an audio transmission system, and the degraded audio signal being obtained by transmitting a reference audio signal through the audio transmission system to provide the degraded audio signal; therefore, the present invention also relates to computer program products. [Background technology]

[0002] Over the past several decades, objective methods for measuring speech quality have been developed and implemented using perceptual measurement techniques. These techniques use perception-based algorithms to mimic the behavior of subjects in listening tests, where they rate the quality of audio fragments. For speech quality, so-called absolute category scale listening tests are frequently used, in which subjects judge the quality of degraded speech fragments without access to a clear reference speech fragment. Listening tests conducted by the International Telecommunication Union (ITU) primarily use a five-point absolute category rating scale (ACR), which is therefore also used in the ITU-standardized objective speech quality measurement method, the Perceptual Speech Quality Measure (PSQM (ITU-T Recommendation P.861, 1996)), and its successor, the Perceptual Speech Quality Measure (PESQ (ITU-T Recommendation P.862, 2000)). These measurement standards target narrowband speech quality (audio bandwidth 100–3500Hz), but a broadband extension (50–7000Hz) was devised in 2005. PESQ enables very good correlation with subjective listening tests for narrowband speech data, and satisfactory correlation for broadband data.

[0003] As the telecommunications industry deploys new broadband voice services, there is a need for advanced measurement standards that can verify performance and accommodate higher audio bandwidths. Therefore, ITU-T (ITU-Telecom Division) Research Group 12 initiated the standardization of a new voice quality assessment algorithm as a technical update to PESQ. The new, third-generation measurement standard, POLQA (Perceptual Objective Listening Quality Assessment), overcomes the shortcomings of the PESQ P.862 standard, such as the effects of linear frequency response distortions, time expansion / compression as seen in Voice-over-IP, certain codec distortions, and inaccurate assessment of reverberation.

[0004] POLQA (P.863) offers several improvements over the previous quality assessment algorithms PSQM (P.861) and PESQ (P.862), and the current version of POLQA addresses several improvements, such as the effects of linear frequency response distortion, time stretching / compression as seen in Voice-over-IP, certain codec distortions, reverberation, and accurate assessment of the effects of playback levels.

[0005] One of the factors that influences perceived speech and sound quality is the presence of echo and reverberation in the audio signal, the latter being superimposed echo. Determining the amount of reverberation or echo can be achieved, for example, by performing autocorrelation of the digitized audio signal to estimate the energy-time curve. When both a reference signal and a degraded signal are available, as in the case of POLQA, the energy-time curve can be determined from the estimated transfer function of the system under test. Although this latter method is used in POLQA, the accuracy of the estimation obtained is affected by the length of the signal and the presence of some kind of noise, pulse, or time-shift distortion, resulting in an inaccurate determination of the perceived impact of the amount of reverberation on perceived audio quality.

Summary of the Invention

Problems to be Solved by the Invention

[0006] An object of the present invention is to solve the above-mentioned disadvantages and to provide a method for accurately estimating the perceptual effect of reverberation in an audio signal on the perceived quality of the audio signal.

Means for Solving the Problems

[0007] For this purpose, a method is provided herein for determining the perceptual effect of the amount of echo or reverberation in a degraded audio signal on the perceived quality of the degraded audio signal, wherein the degraded audio signal is received from an audio transmission system, and the degraded audio signal is obtained by transmitting a reference audio signal through the audio transmission system to provide the degraded audio signal, the method comprising: a controller obtaining at least one degraded digital audio sample from the degraded audio signal and at least one reference digital audio sample from the reference audio signal; the controller determining a local impulse response signal based on the at least one degraded audio sample and the at least one reference audio sample; the controller determining an energy time curve based on the impulse response signal, wherein the energy time curve is proportional to the square root of the absolute value of the impulse response signal, and within the energy time curve there are 1 or more peaks The steps include identifying a peak, wherein the one or more peaks occur delayed in the energy time curve after the beginning of the energy time curve based on the impulse response, and determining an estimate of the amount of echo or reverberation based on the amount of energy in the one or more peaks; obtaining the at least one degraded digital audio sample, which includes sampling the degraded audio signal within a time-domain fraction, wherein the sampling includes performing a windowing operation on the degraded audio signal by multiplying the degraded audio signal by a window function to yield the degraded digital audio sample; obtaining the at least one reference digital audio sample, which includes sampling the reference audio signal within a time-domain fraction, wherein the sampling includes performing a windowing operation on the reference audio signal by multiplying the reference audio signal by a window function to yield the reference digital audio sample;The window function used to obtain the at least one reference digital audio sample and the at least one degraded digital audio sample has non-zero values within the time domain fraction to be sampled and zero values outside the time domain fraction.

[0008] The present invention is based on the insight that many disturbances in a signal affect the accurate determination or estimation of the perceptual effect of the amount of reverberation. These disturbances include various types of noise, various types of pulse distortion, and various types of time-shift distortion, some of which impair the determination of the amount of reverberation at an overall or global level, and some of which are mainly harmful at a local level or exist at a local level. The present invention makes it possible to overcome this problem by performing window processing on the degraded signal and the reference signal before determining the amount of reverberation. For example, a set of perceptual reverberation effect parameters can be calculated from one frame, or from a set of consecutive frames, that constitute audio samples of the degraded audio signal and the reference audio signal (by its window processing). First, the use of window processing makes it possible to calculate an estimated value of reverberation and take it into account in the final reverberation estimation. Second, the use of window processing enables local compensation and local optimization of the processing parameters. The latter can even be performed according to the duration of the time domain fraction of the sample, or its relative position within the complete signal (or relevant part). Thus, by the window processing operation, the method of the present invention provides a more accurate estimated value of the amount of reverberation or echo. This can be applied in many different types of sound processing and evaluation methods. However, it has a significant relevance in the assessment of the quality or intelligibility of a degraded audio signal in combination with, for example, the POLQA method described herein, and therefore its application provides a preferred embodiment of the method.

[0009] The step of obtaining the at least one digital audio sample preferably includes the step of obtaining a plurality of digital audio samples from the audio signal by sampling the audio signal within the time domain fraction using the windowing operation described above. In this case, the time domain fractions of at least two successive digital audio samples may overlap. For example, the overlap between the at least two successive digital audio samples is in the range of 10% to 90% overlap between the time domain fractions, preferably in the range of 25% to 75% overlap, more preferably in the range of 40% to 60% overlap, for example, 50% overlap. This may depend on the type of window function applied, for example, as part of optimization.

[0010] In some embodiments, the window function is the Hamming window, Von Hann window, Tukey window, cosine window, rectangular window, B-spline window, triangular window, Bartlett window, Parzen window, Welch window, and cosine to the power of n window (n thThe window function may be at least one of the following: a power-of-cosine window (where n>1), a Kaiser window, a Nuttall window, a Blackman window, a Blackman-Harris window, a Blackman-Nuttall window, and a Flattop window. The present invention is not limited to any particular type of window function and may be applied using window functions other than those mentioned herein. New optimized window functions may even be developed that may be useful in the method of the present invention without departing from the inventive concept of the present invention.

[0011] In some embodiments, to determine an estimate of the amount of reverberation, the present invention may include weighting the amount of energy at each peak of the energy-time curve based on the magnitude of each peak and / or its (relative) delay position on the time axis. This is based on the insight that the largest magnitude peaks typically have a significant impact on the perceived level of reverberation and how it may impair the clarity or quality of speech or sound.

[0012] In some preferred embodiments, the method further includes: the controller obtaining a digital signal that represents at least a portion of the audio signal and has a duration longer than the time-domain fraction of the at least one digital audio sample; the controller performing autocorrelation operations on the digital signal to yield an overall impulse response signal; the controller determining an overall energy-time curve based on the impulse response signal, wherein the energy-time curve identifies one or more peaks within the energy-time curve that are proportional to the square root of the overall impulse response signal, wherein the one or more further peaks occur delayed in the energy-time curve after the beginning of the energy-time curve based on the overall impulse response; and determining a further estimate of the amount of echo or reverberation based on the amount of energy in the one or more further peaks.

[0013] The preferred embodiments described above provide means for compensating for both local and global disturbances, i.e., disturbances that have a local effect on the reverberation level and disturbances that impair the estimation of the sound signal (or signal portion) at a more global overall level. Furthermore, similar to the locally applied reverberation estimation method described above, determining such further estimates of the amount of reverberation at a global or overall level may similarly involve weighting the amount of energy at each peak based on the magnitude of each peak.

[0014] In other or further embodiments, the method may further include at least one of the following steps: the controller calculates a partial reverberation index value based on the estimated amount of echo or reverberation; the controller calculates a global reverberation index value based on the further estimated amount of echo or reverberation; or the controller calculates a final reverberation index value based on the estimated and further estimated amounts of echo or reverberation.

[0015] Furthermore, in the method described above, determining the (local or global) impulse response signal based on the audio sample, or, if so, the digital signal, includes, by the controller, converting the audio sample or the digital signal from the time domain to the frequency domain by applying a Fourier transform to the audio sample or the digital signal; by the controller, determining a transfer function from the power spectral signal to the audio sample or the digital signal in the frequency domain; and by the controller, converting the power spectral signal from the frequency domain to the time domain to yield the local or global impulse response signal.

[0016] In a preferred embodiment, the present invention provides a method for evaluating the quality or clarity of a degraded audio signal received from an audio transmission system by transmitting a reference audio signal through the audio transmission system to provide the degraded audio signal, the method comprising: sampling the reference audio signal into a plurality of reference signal frames; sampling the degraded audio signal into a plurality of degraded signal frames; forming frame pairs by relating the reference signal frames and the degraded signal frames to each other; providing for each frame pair a difference function representing the difference between the degraded signal frame and the associated reference signal frame; compensating the difference function for one or more disturbance types to provide a disturbance density function for each frame pair that is fitted to a human auditory perception model; deriving an overall quality parameter from the disturbance density functions of the plurality of frame pairs, wherein the quality parameter indicates at least the quality or clarity of the degraded audio signal, and the method further comprises determining the amount of reverberation in at least one of the degraded audio signal and the reference audio signal, the amount of reverberation being determined by applying a method such as that described according to any of the above embodiments.

[0017] In the embodiments described above, the method according to the present invention has been applied in a method for determining the quality or clarity of a degraded audio signal. The method for determining an estimate of the amount of reverberation according to the present invention is particularly useful in this method for evaluating quality or clarity, since the presence of reverberation significantly affects the perceived quality or clarity.

[0018] In some embodiments, the controller may obtain the at least one digital audio sample by forming the audio sample from a plurality of consecutive signal frames, the signal frames comprising one or more of the degraded signal frames or one or more of the reference signal frames. For example, the number of signal frames to be included in the plurality of signal frames may depend on the duration of the time-domain fraction of the at least one digital audio sample, the duration being longer than 0.3 seconds, preferably between 0.4 and 5.0 seconds, e.g., at least one of 0.5 seconds, 1.0 seconds, 1.5 seconds, 2.0 seconds, 2.5 seconds, 3.0 seconds, 3.5 seconds, 4.0 seconds, 4.5 seconds, or 5.0 seconds. In some applications, e.g., POLQA, a single frame is typically too short to be significant in determining the amount of reverberation, and an audio signal fragment shorter than 1 second may be long enough to be analyzed to provide a local estimate of the amount of reverberation.

[0019] Therefore, in some embodiments, a first estimate of the amount of reverberation is obtained by performing a local estimation using, for example, a 0.5-second digital audio sample, and one or more second estimates are obtained for each of a plurality of digital audio samples formed from a plurality of consecutive signal frames providing longer-duration audio signals, and a reverberation index value is calculated based on the first estimate and at least one of the second estimates.

[0020] In some embodiments, for each frame pair, the compensation step is performed by setting the determined reverberation amount in at least one of the degraded audio signal and the reference audio signal as one or more disturbance types, and then compensating each frame pair for the amount of reverberation associated with each frame pair based on the formation of the digital audio sample. Here, the reverberation estimate may be taken into consideration at a local level associated with the frame pair. These are frame pairs consisting of frames that constitute the degraded signal sample.

[0021] In some embodiments, the method further includes a noise suppression step prior to the step of determining the impulse response signal, the noise suppression including performing a first scaling of at least one of the degraded audio signal or the reference audio signal to obtain a similar average volume; processing the degraded audio signal to remove local signal peaks from the degraded audio signal; and performing a second scaling of at least one of the degraded audio signal or the reference audio signal to obtain a similar average volume.

[0022] Furthermore, in order to assess the quality or clarity of a voice or sound signal, the method may be limited to a low frequency range, i.e., a range of the object of interest related to the voice or sound signal. For example, the method may be performed on an audio signal within a predetermined frequency range, e.g., a frequency range below a threshold frequency, or a frequency range corresponding to a voice signal, where the frequency range is, for example, below 5 kHz, preferably 200 Hz to 4 kHz for a voice signal, and up to 20 kHz for other sound signals.

[0023] The present invention will be further described below by describing several specific embodiments thereof with reference to the accompanying drawings. While the detailed description provides examples of possible implementations of the invention, it should not be considered to describe only embodiments that fall within the scope. The scope of the invention is defined in the claims, and the description should be considered illustrative and not limiting to the invention. [Brief explanation of the drawing]

[0024] [Figure 1] Figure 1 provides an overview of the first part of the POLQA perceptual model in an embodiment according to the present invention. [Figure 2] Figure 2 provides an illustrative overview of the frequency alignment used in the POLQA perception model in an embodiment according to the present invention. [Figure 3] Figure 3 provides an overview of a second portion of the POLQA perceptual model in an embodiment according to the present invention, following the first portion shown in Figure 1. [Figure 4] Figure 4 is an overview of the third portion of the POLQA perceptual model in an embodiment according to the present invention. [Figure 5] Figure 5 is a schematic diagram of the masking technique used in the POLQA model. [Figure 6] Figure 6 is a schematic diagram of a method for compensating for overall quality parameters. [Figure 7A] Figure 7A schematically illustrates the windowing operation performed on an audio signal as applied in an embodiment according to the present invention. [Figure 7B] Figure 7B schematically illustrates the windowing operation performed on an audio signal as applied in an embodiment according to the present invention. [Figure 7C] Figure 7C schematically illustrates a windowing operation performed on an audio signal as applied in an embodiment according to the present invention. [Figure 8] Figure 8 schematically illustrates the calculation of the reverberation index according to one embodiment. [Modes for carrying out the invention]

[0025] POLQA Perceptual Model The basic method of POLQA (ITU-T Recommendation P.863) is the same as that used in PESQ (ITU-T Recommendation P.862), namely, that the reference input signal and the degraded output audio signal are mapped to internal representations using a model of human perception. The difference between the two internal representations is used by the cognitive model to predict the perceived audio quality of the degraded signal. A key new concept implemented in POLQA is an idealization technique that removes low levels of noise in the reference input signal and optimizes the timbre. Further major changes in the perception model include modeling the effect of playback levels on perceived quality, as well as a significant division in the handling of low and high levels of distortion.

[0026] An overview of the perceptual model used in POLQA is given in Figures 1 to 4. Figure 1 provides the first part of the perceptual model used to calculate the internal representations of the reference input signal X(t)(3) and the degraded output signal Y(t)(5). Both are scaled (17, 46), and the internal representations as pitch-loudness-time (13 and 14) are calculated in several steps described below, after which a difference function (12) is calculated, which is shown in Figure 1 together with the difference calculation operator (7). Two different flavors of the perceptual difference function are calculated, one for the overall disturbance introduced by the system using operators 7 and 8 under test, and the other for the corresponding added portion of the disturbance using operators (9 and 10). This models the asymmetry of effects between the degradation caused by removing the time-frequency component from the reference signal and the degradation caused by introducing a new time-frequency component. In POLQA, both flavors are calculated using two different methods: one for the range of normal degradation and the other for large degradation, resulting in four difference function calculations (7, 8, 9, and 10) shown in Figure 1.

[0027] In the case of a degraded output signal with frequency domain distortion (49), the alignment algorithm (52) shown in Figure 2 is used. The final processing for obtaining the MOS-LQO score is shown in Figures 3 and 4.

[0028] POLQA begins with the calculation of several basic constant settings, after which the pitch power density (power as a function of time and frequency) of the reference signal and the degraded signal is derived from the time signal, which is time and frequency aligned. From this pitch power density, the internal representation of the reference signal and the degraded signal is derived in several steps. Furthermore, these densities are also used to derive the first three POLQA quality indices for frequency response distortions (41) (FREQ), additive noise (42) (NOISE), and room reverb (43) (REVERB) (40). These three quality indices (41, 42, and 43) are calculated separately from the main disturbance indices to enable a balanced impact analysis across a wide range of distortion types. These indices can also be used to analyze in more detail the type of degradation observed in the audio signal using degradation decomposition techniques.

[0029] As described, four different variations of the internal representation of the reference signal and the degraded signal are calculated in (7, 8, 9, and 10), two of which deal with disturbances related to normal and large distortions, and two deal with added disturbances related to normal and large distortions. These four variations (7, 8, 9, and 10) serve as input to the calculation of the final disturbance density.

[0030] The internal representation of the reference (3) is referred to as the ideal representation. This is because low levels of noise in the reference are removed (step 33), and any tonal distortion present in the degraded signal that may have resulted from the suboptimal timbre of the original reference recording is partially compensated for (step 35).

[0031] The four different variations of the ideal and degraded internal representations, calculated using operators (7, 8, 9, and 10), are used to calculate two final disturbance densities (142 and 143), one representing the final disturbance (142) as a function of time and frequency relating to the overall degradation, and the other representing the final disturbance (143) as a function of time and frequency relating to the processing of the added degradation.

[0032] Figure 4 provides an overview of the calculation of the objective MOS score, MOS-LQO, from the two final disturbance densities (142 and 143) and the FREQ (41), NOISE (42), and REVERB (43) indices.

[0033] Pre-calculation of constant settings FFT window size according to sample frequency POLQA operates at three different sample rates: 8, 16, and 48 kHz sampling, for which the window size W is set to 256, 512, and 2048 samples, respectively, to match the time analysis window of the human auditory system. The overlap between consecutive frames is 50% when using the Hann window. The power spectrum, which is the sum of the squared real and squared imaginary parts of the complex FFT components, is stored in separate real-valued arrays for both the reference signal and the degraded signal. Phase information within a single frame is discarded in POLQA, and all calculations are based solely on the power representation.

[0034] Calculation of start / stop points In subjective testing, noise typically begins before the start of speech activity in the reference signal. However, while leading steady-state noise in subjective testing reduces the effect of other steady-state noise, objective measurements that consider leading noise can be expected to increase that effect; therefore, it is expected that excluding leading and trailing noise is a proper perceptual technique. Accordingly, after confirming that this expectation is correct in the available training data, the start and stop points used in POLQA processing are calculated from the beginning and end of the reference file. The sum of five consecutive absolute sample values ​​(using the usual 16-bit PCM range - +32,000) must exceed 500 from the beginning and end of the original audio file for that position to be designated as the start or end. The interval between this start and end is defined as the effective processing interval. Distortion outside this interval is ignored in the POLQA processing.

[0035] Power scaling coefficient SP and loudness scaling coefficient SL For the calibration of the FFT time-to-frequency conversion, a sine wave with a frequency of 1000 Hz and an amplitude of 40 dB SPL is generated using the calibration of a reference signal X(t) directed to 73 dB SPL. This sine wave is converted to the frequency domain in steps 18 and 49, respectively, using a windowed FFT with a length determined by the sampling frequency for X(t) and Y(t). After converting the frequency axis to the Bark scale in (21 and 54), the resulting peak amplitude of the pitch power density is then multiplied by the power scaling coefficients SP(20 and 55) for X(t) and Y(t), respectively, to 10 4 It is normalized to the power value.

[0036] The same 40 dB SPL reference tone is used to calibrate the psychoacoustic (Sone) loudness scale. After distorting the intensity axis to the loudness scale using Zwicker's law, the integral of the loudness density according to the Bark frequency scale is normalized to 1 Sone in (30) and (58) using the loudness scaling coefficients SL(31) and (59) for X(t) and Y(t), respectively.

[0037] Scaling and calculation of pitch power density The degraded signal Y(t)(5) is multiplied by the calibration coefficient C(47)(46), which deals with the mapping from dB overload in the digital domain to dB SPL in the acoustic domain, and is then converted to the time-frequency domain using 50% overlapping FFT frames(49). The reference signal X(t)(3) is scaled toward a predetermined, fixed optimal level corresponding to approximately 73 dB SPL before being converted to the time-frequency domain(18)17. This calibration procedure is fundamentally different from the procedure used in PESQ, where both the degraded signal and the reference signal are scaled toward a predetermined, fixed optimal level. PESQ assumes that all playback is performed at the same optimal playback level, whereas the subjective test of POLQA uses levels of 20 dB to +6 relative to the optimal level. Therefore, the perceptual model of POLQA cannot use scaling toward a predetermined, fixed optimal level.

[0038] After the level scaling, the reference signal and the degraded signal are transformed into the time-frequency domain using a windowed FFT technique (18, 49). If the frequency axis of the degraded signal is distorted when compared to the reference signal, distortion correction in the frequency domain is performed on the FFT frame. In the first step of this distortion correction, both the reference FFT power spectrum and the degraded FFT power spectrum are preprocessed to reduce the impact of both very narrow frequency response distortion and overall spectral shape differences on subsequent calculations. Preprocessing (77) may consist of smoothing, compressing, and flattening the power spectrum. The smoothing operation is performed using a sliding window average of the power across the FFT bands in (78), while the compression is performed simply by taking the logarithm of the smoothed power in each band (79). The overall shape of the power spectrum is further flattened by sliding window normalization of the smoothed logarithmic power across the FFT bands in (80). Next, the pitches of the current reference frame and the degraded frame are calculated using a stochastic low-harmonic pitch algorithm. Then, the ratio of the reference pitch to the degraded pitch (74) is used to determine the range of possible distortion coefficients (step 84). Where possible, this search range is expanded by using the pitch ratio of the previous frame pair and the subsequent frame pair.

[0039] The frequency alignment algorithm then iterates through the search range, distorts the degraded power spectrum by the distortion coefficient of the current iteration (85), and processes the distorted power spectrum using the preprocessing 77 described above (88). The correlation between the processed reference spectrum and the processed distorted degraded spectrum is then calculated for bins below 1500 Hz (step 89). After iterating through the entire search range, the "best" (i.e., the highest resulting correlation) distortion coefficient is extracted in step 90. The correlation between the processed reference spectrum and the best distorted degraded spectrum is then compared to the correlation between the original processed reference spectrum and the degraded spectrum. The "best" distortion coefficient is retained if the correlation increases by a set threshold (97). If necessary, the distortion coefficient is limited in (98) by the maximum relative change to the distortion coefficient determined for the previous frame pair.

[0040] After distortion correction, which may be necessary to align the frequency axes of the reference signal and the degraded signal, the frequency scale in Hz is distorted in steps 21 and 54 toward Bark's pitch scale, reflecting that the human auditory system has finer frequency resolution at lower frequencies than at higher frequencies. This is achieved by binning the FFT bands and summing the corresponding powers of the FFT bands together, along with normalizing the summed portions. The distortion function that maps the frequency scale in Hertz to the pitch scale in Bark approximates a value known to the author, given in the literature for this purpose. The resulting reference signal and degraded signal have a pitch power density PPX(f) n (Not shown in Figure 1) and PPY(f) n This is known as (56), where f is the frequency in Bark units and the subscript n represents the frame index.

[0041] Calculation of audio active frames, silent frames, and super silent frames (Step 25) POLQA operates on three types of frames, which are distinguished in process 25. • The frame level of the reference signal is higher than approximately 20 dB below the average level of the audio active frame. • Silent frames in which the frame level of the reference signal falls below approximately 20 dB below the average level, and • An ultra-silent frame in which the frame level of the reference signal is approximately 35 dB lower than the average.

[0042] Calculation of frequency index, noise index, and reverberation index The global effects of frequency response distortion, noise, and room reverberation are quantified separately in step 40. For the effect of the total global frequency response distortion, an index (41) is calculated from the average spectra of the reference signal and the degraded signal. To estimate the effect of frequency response distortion regardless of added noise, the average noise spectral density of the degraded signal over the silent frame of the reference signal is subtracted from the pitch loudness density of the degraded signal. The resulting pitch loudness density of the degraded signal and the pitch loudness density of the reference signal are then averaged over each Bark band over all audio active frames for the reference file and the degraded file. The difference in pitch loudness density between these two densities is then integrated over pitch to derive an index (41) for quantifying the effect of frequency response distortion (FREQ).

[0043] Regarding the effect of added noise, an index (42) can be calculated from the average spectrum of the degraded signal over the silent frame of the reference signal. The difference between the average pitch loudness density of the degraded signal over the silent frame and the zero reference pitch loudness density determines a noise loudness density function that quantifies the effect of added noise. This noise loudness density function is then integrated over pitch to derive the average noise effect index (42) (NOISE). Thus, this index (42) is calculated from ideal silence so that a transparent chain measured using a noisy reference signal does not give the maximum MOS score in the final POLQA inter-terminal speech quality measurement.

[0044] In the case of the effects of room reverberation, the energy function with respect to time (ETC) is calculated from the reference time series and the degraded time series. This ETC represents the envelope of the impulse response h(t) of system H(f), which is Y a (f) is defined as H(f)·X(f), where Y a (f) is the spectrum of the level-aligned representation of the degraded signal, and X(f) is the spectrum of the reference signal. Level alignment (noise suppression) is performed to suppress global and local gain differences between the reference signal and the degraded signal. This is done, for example, by a first step of scaling the degraded audio signal (or the reference signal or both), followed by smoothing by removing or suppressing peaks or spikes in the degraded signal. A second scaling step is then performed to flatten the volume in both signals in order to finalize the level alignment. The impulse response h(t) is calculated from H(f) using the inverse discrete Fourier transform. The ETC is calculated from the absolute value of h(t) through normalization and clipping.

[0045] An example of a windowing operation on an audio signal using a Hamming window is schematically shown in Figures 7A and 7C. Figure 7A is a schematic diagram of the Hamming window function (300). The Hamming window function is a bell-shaped function with a maximum value of 1.0 and values ​​of 0.0 at both ends. An arbitrary audio signal (301) is shown in Figure 7B. The windowing operation (320) (Figure 8) on the audio signal (301) can be performed by taking a local convolution between the Hamming window (300) and the audio signal (301), as shown in Figure 7C. The Hamming window (300) has a width that matches the time-domain fraction (305) of the audio sample that will be created in the convolution process. A subsequent Hamming window (300) is applied to the audio signal (301) to obtain a plurality of overlapping digital audio samples (308). In Figure 7C, a 50% overlap is shown by alternating the digital audio samples (308) in the figure. This 50% overlap allows every portion of the signal to be fully considered across two subsequent samples (308). In the present invention, the audio samples obtained by the windowing operation performed on the degraded signal 5 and the reference signal 3, e.g., sample (308), may be used to calculate the reverberation index (43) with or without using a complete section of the degraded audio signal, depending on the embodiment. This windowing is performed on a substantial portion of both the reference signal and the degraded signal. The duration of the time-domain fraction (305) used for windowing is significantly longer than the duration of a single frame in POLQA. The method applied is schematically shown in Figure 8.

[0046] According to several embodiments of the present invention, the calculated reverberation index (43) may be based on both a global or overall reference signal (3) and a degraded audio signal (5) and a plurality of local samples thereof (309 and 310). To calculate a global estimate, the global or overall reference signal and the degraded audio signal (3 and 5) may be considered as a whole or divided into long-duration signal portions (e.g., longer than 5 seconds or longer than 10 seconds, or any other preferred duration). Short local samples (309 and 310) may be obtained by performing windowing operations 320a and 320b on the reference signal and the degraded audio signal (3 and 5) or their long-duration signal portions, or by integrating or combining a plurality of signal frames from the reference signal X(t)(3) and the degraded signal Y(t)(5). For example, the short local samples 309 and 310 may include sound fragments having a duration of, for example, 0.5 or 1.0 seconds (sometimes referred to herein as time-domain fractions (305)). Smaller fragments may provide insufficient information regarding reverberation. The short-duration local fractions (309 and 310) obtained using the windowing operation (320) (i.e., 320a and 320b) are obtained, for example, by applying a Hamming window 300 with a 50% overlap with each other. The short-duration local samples (309 and 310) are formed by multiplying the degraded audio signal (5) by the applied window function (300) (e.g., a Hamming window function). For optimal determination of the local reverberation index, a weighting coefficient may be used which gives lower weight to samples that degrade earlier in the window when the audio of the corresponding reference audio sample falls below a threshold and perceptually indicates an interval of silence. This weighting is performed in (321a and 321b). Subsequently, in steps 322a and 322b, a Fast Fourier Transform (FFT) is performed on the samples (309 and 310) and the overall degraded audio signal (5).The global reference signal and the degraded audio signals (3 and 5) are processed in steps 340a and 340b by performing a Fast Fourier Transform (FFT) on the reference signal (3) and the degraded digital signal (5). The FFT in steps 322a / b and 340a / b may be performed over a portion of the frequency range that includes the contribution of the audio signal (e.g., below 5 kHz, or 200 Hz to 4 kHz).

[0047] In steps 324 and 342, the transfer function H(f) is calculated from the transformed signal in the frequency domain. The impulse response signal is obtained by inverse FFT in steps 326 and 344, from which the ETC can be calculated in steps 328 and 346. The ETC is determined in steps 328 and 346 for both the long-duration signal portions (or the entire reference signal and degraded signal) (3 and 5) and the short-duration local samples (309 and 310) in the manner described above. In each of the ETC, one or more peaks are identified in steps 330 and 348, which occur with a time delay after the beginning of the energy-time curve based on the impulse response. For example, three largest peaks may be determined that occur at least 60 milliseconds after the beginning of the curve. The energies at these peaks are determined and, in steps 332 and 350, are used in combination with their respective delay positions on the time axis to calculate local and global reverberation indices. For both the local sample and the global portion, local and global reverberation indices may be calculated in steps 332 and 350, which are combined in step 360 to obtain a good estimate of the reverberation index (43) to be used thereafter.

[0048] Based on the ETC of the global portion and local sample, in steps 330 and 348, multiple reflections may be explored within each ETC. In the first step, the loudest reflection is calculated by simply determining the maximum value of the ETC curve after the direct sound. In the POLQA model, the direct sound is defined as all sounds arriving within 60 milliseconds. Next, the second loudest reflection is determined over the direct sound-free interval and without considering reflections arriving within 100 milliseconds of the loudest reflection. Then, the third loudest reflection is determined over the direct sound-free interval and without considering reflections arriving within 100 milliseconds of the loudest and second loudest reflections. The energies and delays of these three loudest reflections are then combined to form the partial and global reverberation index values, which may then be combined to form a single reverberation index (43)(REVERB).

[0049] Optionally, in calculating the reverberation index (43), only reverberation estimates that are within one standard deviation of the mean of the partial reverberation estimates may be taken. These may then be weighted in a particular manner. In a computer program product developed to carry out the method described herein, this may be carried out, for example, as follows:

[0050] [Table 1]

[0051] As an alternative to the above, the reverberation index may be estimated based solely on the short-duration local sample, which already offers an improvement over the conventional method of estimating the amount of reverberation in a signal.

[0052] Global and local scaling of the reference signal toward the degraded signal (step 26) The reference signal is at this point at the internal ideal level, i.e., approximately 73 dB SPL, according to step 17, while the degraded signal is represented at a level matching the playback level as a result of (46). The global level difference is compensated in step 26 before the comparison between the reference signal and the degraded signal is made. Furthermore, small local level changes are partially compensated to take into account that sufficiently small level fluctuations are not perceptible to the subject in listening-only situations. The global level equalization (26) is performed using frequency components from 400 to 3500 Hz based on the average power of the reference signal and the degraded signal. The reference signal is globally scaled toward the degraded signal, so that the effect of the global playback level difference is maintained at this stage of processing. Similarly, for slowly fluctuating gain distortion, local scaling is performed over level changes of up to approximately 3 dB using the full bandwidth of both the reference audio file and the degraded audio file.

[0053] Partial compensation of the original pitch power density with respect to linear frequency response distortion (Step 27) To accurately model the effects of linear frequency response distortion induced by filtering within the system under test, a partial compensation technique is used in step 27. To model the imperceptibility of moderate linear frequency response distortion in subjective testing, the reference signal is partially filtered using the transfer characteristics of the system under test. This is done by calculating the average power spectra of the original and degraded pitch power densities over all speech active frames. For each Bark bin, a partial compensation coefficient is calculated from the ratio of the degraded spectrum to the original spectrum (27).

[0054] Modeling of mask effects, calculation of pitch loudness density excitation In steps 30 and 58, masking is modeled by calculating a smoothed representation of the pitch power density. Smoothing in both the time domain and the frequency domain is considered according to the principles shown in FIGS. 5a to 5c. Time-frequency domain smoothing uses a convolution technique. From this smoothed representation, representations of the reference pitch power density and the degraded pitch power density are recalculated to suppress low-amplitude time-frequency components that are partially masked by adjacent high-volume components in the time-frequency plane. This suppression is performed in two different ways: subtraction of the smoothed representation from the non-smoothed representation and division of the non-smoothed representation by the smoothed representation. The resulting sharpened pitch power density representation is then converted to a pitch loudness density representation using a modified version of Zwicker's power law as follows,

[0055] [Number] SL is the loudness scaling factor, P0(f) is the absolute listening threshold, and f B and P fn are

[0056] [Number] is a correction that depends on the frequency and level defined by, where f represents the frequency in Bark units and PPX(f) n represents the pitch power density in the frequency-time cell f,n. The resulting two-dimensional arrays LX(f) n and LY(f) n are called pitch loudness density and are at the output of step 30 for the reference signal X(t) and the output of step 58 for the degraded signal Y(t), respectively.

[0057] Suppression of global low-level noise in the reference signal and the degraded signal Low levels of noise in the reference signal that are not affected by the system under test (e.g., a transmission chain) are attributed to the system under test by the subject, given that this is an absolute category scale test procedure. Therefore, these low levels of noise must be suppressed when calculating the internal representation of the reference signal. This "idealization process" in step 33 is performed by representing the reference signal LX(f) as a function of pitch over the ultra-silent frame. n This is done by calculating the average steady-state noise loudness density. This average noise loudness density is then partially subtracted from all pitch loudness density frames of the reference signal. The result is an idealized internal representation of the reference signal at the output of step 33.

[0058] In the degraded signal, audible steady-state noise has a less significant impact than transient steady-state noise. This holds true for all levels of noise, and the effect of this phenomenon can be modeled by partially removing steady-state noise from the degraded signal. This is done in step 60, where the corresponding frame of the reference signal is classified as ultra-silent in the degraded signal LY(f) n This is done by calculating the average steady-state noise loudness density of the frame as a function of pitch. This average noise loudness density is then partially subtracted from all pitch loudness density frames of the degraded signal. This partial compensation uses different strategies for low-level and high-level noise. For low-level noise, the compensation is minimal, while the suppression used is more aggressive for high-volume added noise. The result is an internal representation (61) of the degraded signal with added noise adapted to the subjective effects observed in listening tests, using an idealized, noise-free representation of the reference signal.

[0059] In step 33 described above, in addition to suppressing the global low-level noise, a loudness index (32) is also determined for each of the reference signal frames. The loudness index or loudness value may be used to determine loudness-dependent weighting coefficients for weighting specific types of distortion. The weighting itself may be performed in steps 125 and 125' on the four distortion representations given by operators (7, 8, 9, and 10) to give the final disturbance densities (142 and 143).

[0060] Therefore, although the loudness level index is determined in step 33, it can be recognized that the loudness level index may be determined in a separate part of the method corresponding to each reference signal frame. In step 33, determining the loudness level index means that the average steady-state noise loudness density is over the ultra-silent frame of the reference signal LX(f) n This is possible because it has already been determined and is then used to construct a noise-free reference signal for all reference frames. However, although it is possible to do this in step 33, it is not the most preferred implementation method.

[0061] Alternatively, the loudness level index (LOUDNESS) may be taken from the reference signal in an additional step following step 35. This additional step is also shown in Figure 1 as a dotted box (35') with a dotted output (LOUDNESS) (32'). If implemented in step 35', it is no longer necessary to receive the loudness level index from step 33, as the reader of this art may recognize.

[0062] Local scaling of distorted pitch loudness density with respect to time-varying gain between the degraded signal and the reference signal (steps 34 and 63) Slow fluctuations in gain are inaudible, and small changes are already compensated for during the calculation of the reference signal representation. Any remaining compensation required before the correct internal representation can be calculated is performed in two steps: firstly, in step 34, the reference signal is compensated for signal levels where the loudness of the degraded signal is less than the loudness of the reference signal; and secondly, in step 63, the degraded signal is compensated for signal levels where the loudness of the reference signal is less than the loudness of the degraded signal.

[0063] The first compensation (34) scales the reference signal to a lower level for portions of the degraded signal that exhibit significant signal loss, for example, in the context of time clipping. The scaling is such that the remaining difference between the reference signal and the degraded signal represents the effect of time clipping on locally perceived audio quality. Portions where the loudness of the reference signal is less than the loudness of the degraded signal are not compensated, and therefore additional noise and loud clicks are not compensated in this first step.

[0064] The second compensation (63) scales the degraded signal to a lower level for portions of the signal in which the degraded signal exhibits a click and portions of the signal in which there is noise during silent intervals. The scaling is such that the remaining difference between the reference signal and the degraded signal represents the effect of the click and slowly changing additional noise on the locally perceived speech quality. The click is compensated for in both silent and speech-active portions, while the noise is compensated for in silent portions only.

[0065] Partial compensation of the original pitch loudness density with respect to linear frequency response distortion (step 35) Imperceptible linear frequency response distortion has already been compensated for in step 27 by partially filtering the reference signal in the pitch power density region. To further compensate for the fact that linear distortion is less unpleasant than nonlinear distortion, the reference signal is then partially filtered in the pitch loudness region in step 35. This is done by calculating the average loudness spectrum of the original and degraded pitch loudness densities over all speech active frames. For each Bark bin, a partial compensation coefficient is calculated from the ratio of the degraded loudness spectrum to the original loudness spectrum. This partial compensation coefficient is used to filter the reference signal with a smoothed, lower amplitude version of the frequency response of the system under test. After this filtering, the difference between the reference pitch loudness density and the degraded pitch loudness density resulting from linear frequency response distortion is reduced to a level that represents the effect of linear frequency response distortion on the perceived speech quality.

[0066] Final scaling and noise suppression of the pitch loudness density. Up to this point, all calculations for the signal are performed at the playback level used in subjective experiments. At low playback levels, this results in a small difference between the reference pitch loudness density and the degraded pitch loudness density, which is generally an overly optimistic estimate of the perceived speech quality. To compensate for this effect, the degraded signal is then scaled in step 64 toward a "virtual" fixed internal level. After this scaling, the reference signal is scaled toward the degraded signal level in step 36, so that both the reference signal and the degraded signal are now ready for the final noise suppression operation in (37) and (65), respectively. This noise suppression deals with the last portion of the steady-state noise level in the loudness region, which still has too large an impact on the speech quality calculation. The resulting signals (13 and 14) are now in the perceptually relevant internal representation region, and represent the ideal pitch-loudness-time LX 理想 (f)n Function (13) and degraded pitch-loudness-time LY 劣化 (f) n From function (14), disturbance densities (142 and 143) can be calculated. Four different variations of the ideal and degraded pitch-loudness-time function are calculated in (7, 8, 9 and 10), two variations (7 and 8) for disturbances with normal and large distortions, and two (9 and 10) for added disturbances with normal and large distortions.

[0067] Final disturbance density calculation Two different flavors of the disturbance density (142 and 143) are calculated. The first flavor, the normal disturbance density, is calculated in (7 and 8) as the ideal pitch-loudness-time LX 理想 (f) n and the degraded pitch-loudness-time function LY 劣化 (f) n It is derived from the difference between the two. The second flavor is derived from the ideal pitch-loudness-time and the degraded pitch-loudness-time functions using an optimized version with respect to the introduced degradation in (9 and 10), and is called the added disturbance. In the calculation of this added disturbance, the signal portions in which the degraded power density is greater than the reference power density are weighted by a coefficient that depends on the power ratio in each pitch-time cell, i.e., an asymmetry coefficient.

[0068] To address a wide range of distortions, two different versions of the processing are performed: one for small to moderate distortions based on (7 and 9), and the other for moderate to large distortions based on (8 and 10). Switching between these two is performed based on a first estimate from the disturbance for small to moderate levels of distortion. This processing technique leads to the need to compute four different ideal pitch-loudness-time functions and four different degraded pitch-loudness-time functions so that a single disturbance function and a single additional disturbance function (see Figure 3) can be computed, and these disturbance functions are then compensated for specific amounts of distortion of several different kinds of severity.

[0069] A significant deviation from the optimal listening level is quantified at (127 and 127') by an index directly derived from the signal level of the degraded signal. This global index (level) is also used in the calculation of the MOS-LQO.

[0070] The severe distortion introduced by frame repetition is quantified in (128 and 128') by an index derived from a comparison of the correlation of consecutive frames of the reference signal with the correlation of consecutive frames of the degraded signal.

[0071] Severe deviations of the degraded signal from its optimal "ideal" timbre are quantified at (129 and 129') by an index derived from the difference in loudness between the higher and lower frequency bands. The timbre index is calculated from the difference in loudness in the Bark bands of the degraded signal, which are 2–12 Barks in the lower frequency range and 7–17 Barks in the higher range (i.e., using a 5-Bark overlap), and this "punishes" any serious imbalance, regardless of whether it may be a result of inaccurate vocal timbre in the reference audio file. Compensation is performed frame by frame at a global level. This compensation calculates the power in the lower and upper Bark bands of the degraded signal (below 12 Barks and above 7 Barks, i.e., using a 5-Bark overlap), and "punishes" any serious imbalance, regardless of whether it may be a result of inaccurate vocal timbre in the reference audio file. Therefore, it should be noted that a transparent chain using a poorly recorded reference signal containing excessive noise and / or inaccurate timbre will not yield the highest MOS score in POLQA's inter-terminal voice quality measurement. This compensation also affects the measurement of the quality of a transparent device. If a reference signal exhibiting a significant deviation from the optimal "ideal" timbre is used, the system under test will be judged as non-transparent, even if the system does not introduce any degradation into the reference signal.

[0072] The impact of severe peaks in disturbances is quantified in (130 and 130') as a flatness index, which is also used in the calculation of the MOS-LQO.

[0073] Severe noise level fluctuations, which target the subject's attention to noise, are quantified in (131 and 131') by a noise contrast index derived from a degraded signal frame in which the corresponding reference signal frame is silent.

[0074] In steps 133 and 133', a weighting operation is performed to weight the disturbance depending on whether it matches the actual spoken voice. Disturbances perceived during silence are not considered as detrimental as disturbances perceived during the actual spoken voice in order to assess the quality or intelligibility of the degraded signal. Therefore, a weighting value is determined to weight the disturbance based on the loudness index determined from the reference signal in step 33 (or alternatively, step 35'). The weighting value is used to weight the difference function (i.e., the disturbance) in order to incorporate the effect of the disturbance on the quality or intelligibility of the degraded speech signal into the assessment. In particular, since the weighting value is determined based on the loudness index, the weighting value may be expressed by a loudness-dependent function. The loudness-dependent weighting value may be determined by comparing the loudness value to a threshold. If the loudness index exceeds the threshold, the perceived disturbance is fully taken into account when making the assessment. In contrast, if the loudness index is less than the threshold, the weighting value is made to depend on the loudness level index; that is, in this example, the weighting value becomes equal to the loudness level index (in situations where the loudness is below the threshold). The advantage is that, for example, weak portions of the speech signal at the end of a word spoken immediately before a pause or silence are partially considered as disturbances that are detrimental to the quality or clarity. As an example, it can be recognized that a certain amount of noise perceived when saying the letter "f" at the end of a word may cause the listener to perceive it as the letter "s". This can be detrimental to the quality or clarity. In contrast, a person skilled in the art can recognize that it is also possible to simply ignore all noise during silence or pauses by setting the weighting value to zero when the loudness value is below the threshold mentioned above.

[0075] Continuing with Figure 3, a serious jump in the alignment is detected during the alignment process, and this effect is quantified by a compensation factor in steps 136 and 136'.

[0076] Finally, the disturbance and the added disturbance density are clipped to the maximum level at (137 and 137'), and the variance of the disturbance (138 and 138') and the effects of the jumps (140 and 140') in the loudness of the reference signal are used to compensate for the specific temporal structure of the disturbance.

[0077] This is the final disturbance density D(f) for normal disturbances. n (142) and the final disturbance density DA(f) for the added disturbances n (143) brings about.

[0078] Aggregation of disturbances against pitch, spurt, and time, and mapping to intermediate MOS scores. The final disturbance D(f) n (142) and the added disturbance DA(f) n The density (143) is integrated frame by frame along the pitch axis using the L1 integrals (153 and 159) (see Figure 4), resulting in two distinct disturbances per frame, one derived from the disturbance and the other from the added disturbance:

[0079]

number

[0080] Next, these two disturbances per frame are averaged over a concatenation of six consecutive audio frames, defined as an audio spurt, and L4(155) and L1(160) weights are applied to the disturbance and the added disturbance, respectively.

[0081]

number

[0082] Finally, the disturbances and added disturbances are calculated for each file from the time average of L2 (156 and 161):

[0083]

number

[0084] The added disturbances are compensated in step 161 for loud reverberation and loud added noise using the REVERB index (42) and noise index (43). The two disturbances are then combined with the frequency index (41) (FREQ) (170) to derive an internal index, which is linearized with a cubic regression polynomial to obtain an intermediate index (171) similar to MOS.

[0085] Final calculation of POLQA MOS-LQO The raw POLQA score is derived from an intermediate index similar to the MOS, using all four different compensations in process 175: Two compensations for the specific time-frequency characteristics of the disturbance: one calculated by L511 aggregation over frequency (148), spurt (149), and time (150), and the other calculated by L313 aggregation over frequency (145), spurt (146), and time (147). • One compensation for very low presentation levels using level indicators • One compensation for large tonal distortion using a flatness index in the frequency domain.

[0086] This mapping is trained on a large set of degradations, including degradations that were not part of the POLQA benchmark. These raw MOS scores (176) are for the major portion that have already been linearized by the cubic polynomial mapping used in the calculation of an intermediate index (171) similar to the MOS.

[0087] Finally, the unprocessed POLQA MOS score (176) is mapped to the MOS-LQO score (181) in (180) using a cubic polynomial optimized for 62 databases, as were available in the final stages of POLQA standardization. In narrowband mode, the maximum POLQA MOS-LQO score is 4.5, while in ultra-wideband mode, this point is 4.75. One important consequence of the idealization process is that, under certain circumstances, when the reference signal contains noise or the timbre is severely distorted, the transmission chain will not yield either the maximum MOS score of 4.5 in narrowband mode or 4.75 in ultra-wideband mode.

[0088] Consonant-vowel-consonant compensation according to the present invention may be carried out as follows. In Figure 1, the reference signal frame (220) and the degraded signal frame (240) may be obtained as shown. For example, the reference signal frame (220) may be obtained from step 21 of distorting the reference signal to bark, while the degraded signal frame may be obtained from corresponding (54) performed on the degraded signal. The exact locations where the reference signal frame and / or the degraded signal frame are obtained from the method of the present invention, as shown in Figure 1, are merely examples. The reference signal frame (220) and the degraded signal frame (240) may be obtained from any of the other steps in Figure 1, in particular somewhere between the input of the reference signal X(t)(3) and the global and local scaling to the degraded level in step 26. The degraded signal frame may be obtained at any point between the input of the degraded signal Y(t)(5) and step 54.

[0089] Consonant-vowel-consonant compensation proceeds as shown in Figure 6. First, in step 222, the signal power of the reference signal frame (220) is calculated within a desired frequency domain. For the reference frame, this frequency domain, in the most optimal circumstances, includes only speech signals (e.g., the frequency range of 300 Hz to 3500 Hz). Next, in step 224, the selection of whether or not to include this reference signal frame as an active speech reference signal frame is made by comparing the calculated signal power with a first threshold (228) and a second threshold (229). The first threshold is, for example, 7.0 × 10⁻¹⁰ when using the scaling of the reference signal as described in POLQA (ITU-T Recommendation P.863). 4 It may be equal to, and the second threshold is 2.0 × 2 × 10 8 It may be equal to. Similarly, in step 225, the calculated signal power is compared with a third threshold (230) and a fourth threshold (231) to select the reference signal frame corresponding to the soft speech reference signal (consonant important portion) for processing. The third threshold (230) is, for example, 2.0 × 10 7 It may be equal to 7.0 × 10⁻⁶, and the fourth threshold is 7.0 × 10⁻⁶. 7 It can be equivalent to...

[0090] Steps 224 and 225 yield reference signal frames corresponding to the active voice portion and the soft voice portion, respectively, the active voice reference signal portion frame (234) and the soft voice reference signal portion frame (235). These frames are provided to step 260, which is described below.

[0091] In exactly the same manner as the calculation of the relevant signal portion of the reference signal, the degraded signal frame (240) is also first analyzed in step 242 to calculate the signal power in a desired frequency domain. In the case of the degraded signal frame, it is advantageous to calculate the signal power within a frequency range that includes the frequency range of the spoken voice and the frequency range in which most of the audible noise exists, for example, the frequency range of 300 Hz to 8000 Hz.

[0092] From the calculated signal power in step 242, the relevant frame, i.e., the frame associated with the relevant reference frame, is selected. Selection is performed in steps 244 and 245. In step 245, for each degraded signal frame, it is determined whether it is temporally aligned with the reference signal frame selected as the soft voice reference signal frame in step 225. If the degraded frame is temporally aligned with the soft voice reference signal frame, the degraded frame is identified as a soft voice degraded signal frame, and its calculated signal power is used in the calculation of step 260. Otherwise, the frame is discarded as a soft voice degraded signal frame for the calculation of the compensation coefficient in step 247. In step 244, for each degraded signal frame, it is determined whether it is temporally aligned with the reference signal frame selected as the active voice reference signal frame in step 224. If the degraded frame is temporally aligned with the active voice reference signal frame, the degraded frame is identified as an active voice degraded signal frame, and its calculated signal power is used in the calculation of step 260. Otherwise, the frame is discarded as an active voice-degraded signal frame for the calculation of the compensation coefficient in step 247. This results in the soft voice-degraded signal subframe (254) and the active voice-degraded signal subframe (255), which are provided to step 260.

[0093] Step 260 receives the active voice reference signal portion frame (234), the soft voice reference signal portion frame (235), the soft voice degraded signal portion frame (254), and the active voice degraded signal portion frame 255 as input. In step 260, the signal power of these frames is processed to determine the average signal power corresponding to the active voice and soft voice reference signal portions and the active voice and soft voice degraded signal portions, from which (also in step 260) the consonant-vowel-consonant signal-to-noise ratio compensation parameter (CVC) is determined. SNR_係数 ) is calculated as follows:

[0094]

number

[0095] Parameters Δ1 and Δ2 are constant values ​​used to adapt the model's behavior to the subject's behavior. The other parameters in this equation are as follows: P アクティブ,参照,平均 This is the average active audio reference signal partial signal power. Parameter P ソフト,参照,平均 This is the average soft voice reference signal partial signal power. Parameter P アクティブ,劣化した,平均 This is the average active audio degradation signal partial signal power, and parameter P ソフト,劣化した,平均 This is the average soft voice degradation signal portion signal power. At the output of process 260, the consonant-vowel-consonant signal-to-noise ratio compensation parameter CVC is used. SNR_係数 It is given

[0096] The CVC SNR_係数 In step 262, the CVC is compared to a threshold value, which in this example is 0.75. SNR_係数 If this threshold is greater, the compensation coefficient for process 265 is determined to be equal to 1.0 (no compensation is performed). SNR_係数 If the value is less than the threshold (in this case, 0.75), the compensation coefficient is calculated in step 267 as follows: Compensation coefficient = (CVC SNR_係数 (+0.25) 1 / 2(The value 0.25 is interpreted as being equal to 1.0 to 0.75, and here 0.75 is the CVC) SNR_係数 Note that this is a threshold used for comparison. The compensation coefficient (270) thus obtained is used as a multiplier for the MOS-LQO score (i.e., the overall quality parameter) in step 182 of Figure 4. As will be recognized, the compensation (by multiplication) does not necessarily have to be performed in step 182, but can be incorporated into either step 175 or step 180 (in which case step 182 will be absent from the schematic diagram in Figure 4). Furthermore, in this example, compensation is achieved by multiplying the MOS-LQO score by the compensation coefficient calculated as shown. It will be recognized that compensation can take other forms. For example, the CVC SNR_係数 Depending on the circumstances, it may also be possible to subtract or add variables to the obtained MOS-LQO. Those skilled in the art will understand and recognize the other meanings of compensation in accordance with this teaching.

[0097] The present invention has been described in terms of several specific embodiments. It will be recognized that the embodiments shown in the drawings and described herein are for illustrative purposes only and are not intended in any way to limit the present invention. The operation and configuration of the present invention are expected to be apparent from the foregoing description and the accompanying drawings. Those skilled in the art will see that the present invention is not limited to any embodiment described herein and that modifications are possible that are considered to be within the scope of the appended claims. Furthermore, kinematic inversion is essentially disclosed and is considered to be within the scope of the present invention. Moreover, any of the components and elements of the various embodiments disclosed may be combined or incorporated into other embodiments if deemed necessary, desirable or preferred, without departing from the scope of the present invention as defined in the claims.

[0098] In the claims, reference numerals should not be construed as limiting the claims. The words “equipped with” and “including,” when used in this description or in the appended claims, should be construed as inclusive, not exclusive or exclusionary. Thus, the expression “equipped with” as used herein does not exclude the presence of other elements or processes in addition to those enumerated in any claim. Furthermore, the words “a” and “an” should not be construed as limiting to “only one,” but rather as meaning “at least one,” and do not exclude the plural. Features not specifically or explicitly described or claimed may be additionally included in the structures of the invention within the scope of the invention. Expressions, for example, “means for,” should be read as “components configured for” or “members configured to,” and should be construed to include corresponding structures of the disclosed invention. The use of expressions such as “important,” “preferred,” or “particularly preferred” is not intended to limit the invention. Additions, deletions, and modifications within the understanding of those skilled in the art may be made in general without departing from the spirit and scope of the invention as defined by the claims. The invention may be carried out in a manner different from that specifically described herein and is limited by the appended claims.

[0099] Reference number 3: Reference signal X(t) 5: Degraded signal Y(t), amplitude-time 6: Delay identification, frame pair formation 7: Calculating the difference 8: The first variation of difference calculation 9: The second variation of difference calculation 10: The third variation of difference calculation 12: Difference signal 13: Ideal pitch-loudness-time LX for internal representation 理想 (f) n 14: Degraded pitch-loudness-time LY of internal representation 劣化 (f) n 17: Global scaling towards a fixed optimal level 18: FFT with window 20: Scaling coefficient SP 21: Distort with respect to Bark 25: (Ultra) silent frame detection 26: Global and local for degraded levels 27: Partial frequency compensation 30: Excitation and distortion for psychoacoustics (Sone) 31: Absolute threshold scaling coefficient SL 32: LOUDNESS 32’: LOUDNESS (determined according to alternative process 35’) 33: Global low-level noise suppression 34: Local scaling, when Y < X 35: Partial frequency compensation 35’: (Alternative) loudness determination 36: Scaling for degraded levels 37: Global low-level noise suppression 40: FREQ NOISE REVERB index 41: FREQ index 42: NOISE index 43: REVERB index 44: PW_R 全体 Index (overall audio ratio between degraded signal and reference signal) 45: W_R フレーム Index (audio power ratio per frame between degraded signal and reference signal) 46: Scaling for playback level 47: Calibration coefficient C 49: FFT with window 52: Frequency alignment 54: Distort with respect to Bark 55: Scaling coefficient SP 56: Pitch-power-time PPY(f)n of the degraded signal 58: Excitation and distortion for psychoacoustics (Sone) 59: Absolute threshold scaling coefficient SL 60: Global high-level noise suppression 61: Pitch-Loudness-Time of a Degraded Signal 63: Local scaling, when Y > X 64: Scaling against fixed internal levels 65: Global high-level noise suppression 70: Reference Spectrum 72: Degraded Spectrum 74: Ratio of the current and + / -1 enclosed frame reference pitch to the degraded pitch 77: Preprocessing 78: Smoothing out narrow spikes and drops in the FFT spectrum. 79: Take the logarithm of the spectrum and apply the minimum intensity threshold. 80: Use a sliding window to flatten the overall logarithmic spectral shape. 83: Optimization Loop 84: Range of distortion coefficient: [Minimum pitch ratio <= 1 <= Maximum pitch ratio] 85: Distorted and degraded spectrum 88: Apply preprocessing Calculate spectral correlation for bins below 89:1500Hz 90: Track the best distortion coefficient 93: Degraded Spectrum 94: Apply preprocessing Calculate spectral correlation for bins below 95:3000Hz 97: If the correlation is sufficient, the distorted degraded spectrum is retained; otherwise, the original is restored. 98: Limit the change in the distortion coefficient from one frame to the next. 100: Ideal, normal 101: Deteriorated, normal 104: Ideal, large distortion 105: Deteriorated, severely distorted 108: Ideal, added 109: Degraded, added 112: Ideal, with significant distortion added. 113: Degraded, heavily distorted 116: Disturbance density is normally selected. 117: Disturbance density, large strain selection 119: Selection of added disturbance density 120: Added disturbance density, large strain selection 121:PW_R 全体 , Input to switching function 123 122:PW_R フレーム , input to switching function 123 123: Determining large distortion (switching) 125: Coefficient for a specific amount of distortion 125': Correction factor for a severe amount of specific distortion 127: Level 127': Level 128: Frame Repeat 128': Frame repeat 129: Tone 129': Tone 130: Spectral flatness 130': Spectral flatness 131: Noise contrast during silent periods 131': Noise contrast during periods of noise silence 133: Loudness dependent on distortion weighting 133': Loudness dependent on distortion weighting 134: Loudness of the reference signal 134': Loudness of the reference signal 136: Alignment Jump 136': Alignment Jump 137: Clip to maximum degradation 137': Clip to maximum degradation 138: Disturbance dispersion 138': Disturbance dispersion 140: Loudness Jump 140': Loudness Jump 142: Final disturbance density D (f) n 143: Final added disturbance density DA (f) n 145: L3 frequency integration 146: L1 Spurt Integration 147: L3 Time Integration 148: L5 frequency integration 149: L1 Spurt Integration 150: L1 Time Integration 153: L1 frequency integration 155: L4 Spurt Integration 156: L2 Time Integration 159: L1 frequency integration 160: L1 Spurt Integration 161: L2 Time Integration 170: Mapping to intermediate MOS score 171: Intermediate metrics like MOS 175: MOS scale compensation 176: Raw MOS score 180: Mapping to MOS-LQO 181:MOS LQO 182: CVC Clarity Compensation (Clarity Model Only) 185: Intensity of a short sinusoidal tone over time 187: Short sine wave tone 188: Mask the threshold for the second short sine wave tone. 195: Intensity of frequency for short sine wave tones 198: Short sine wave tone 199: Set the threshold for the second short sine wave tone. 205: Intensity against frequency and time in 3D plot 211: Masking the threshold used as the suppression strength results in a sharper internal representation. 220: Reference signal frame (see Figure 1 again) 222: Measure signal power in the audio domain (e.g., 300Hz to 3500Hz). 224: Compare the signal power to the first threshold and the second threshold, and select if it is within the range. 225: Compare the signal power to the third and fourth thresholds, and select if it is within the range. 228: First threshold 229: Second threshold 230: Third threshold 231: The fourth threshold 234: Power Average of Active Audio Reference Signal Frames 235: Power average of soft voice reference signal frames 240: Degraded signal frame (see Figure 1 again) 242: Determine the signal power in the domain for speech and audible disturbances (e.g., 300Hz to 8000Hz). 244: Does the degraded frame time match the selected active audio reference signal frame? 245: Does the degraded frame time match the selected soft audio reference signal frame? 247: Frames discarded as degraded signal frames for active / soft voice. 254: Power averaging of soft-voice degraded signal frames 255: Power averaging of active audio degraded signal frames 260: Consonant-Vowel-Consonant Signal-to-Noise Ratio Compensation Parameter (CVC) SNR_係数 ) calculate 262: CVC SNR_係数 Is it below the compensation threshold (e.g., 0.75)? 265: No -> Compensation coefficient = 1.0 (No compensation) 267: Yes -> The compensation coefficient is (CVC SNR_係数 (+0,25) ½ 270: Provide a compensation value to process 182 to compensate for MOS-LQO.

Claims

1. 1. A method for determining the perceptual impact of an amount of echo or reverberation in a degraded audio signal on the perceived quality of the degraded audio signal, wherein the degraded audio signal is received by a computer system from an audio transmission system, the degraded audio signal being obtained by transmitting, by the computer system, a reference audio signal through the audio transmission system to provide the degraded audio signal, the method comprising: obtaining, by a controller of the computer system, at least one degraded digital audio sample from the degraded audio signal and at least one reference digital audio sample from the reference audio signal; determining, by the controller, a local impulse response signal based on the at least one degraded digital audio sample and the at least one reference digital audio sample; determining, by the controller, an energy-time curve based on the impulse response signal, wherein the energy-time curve is proportional to the square root of the absolute value of the impulse response signal; and identifying, by the controller, one or more peaks in the energy-time curve, where the one or more peaks in time occur delayed in the energy-time curve after a beginning of the energy-time curve based on the impulse response; and determining, by the controller, an estimate of the amount of echo or reverberation based on the amount of energy in the one or more peaks. The process includes the steps of: obtaining the at least one degraded digital audio sample comprises sampling the degraded audio signal in a time domain fraction, the sampling comprising performing a windowing operation on the degraded audio signal by multiplying the degraded audio signal by a window function to result in the degraded digital audio sample; and obtaining the at least one reference digital audio sample comprises sampling the reference audio signal within a time-domain fraction, the sampling comprising performing a windowing operation on the reference audio signal by multiplying the reference audio signal with the window function to result in the reference digital audio sample; the window function used to obtain the at least one reference digital audio sample and the at least one degraded digital audio sample has a non-zero value within the time domain fraction to be sampled and a zero value outside the time domain fraction. The method.

2. 2. The method of claim 1, wherein the obtaining step comprises obtaining a plurality of digital audio samples from the audio signal, wherein each sample of the plurality of digital audio samples is obtained by performing a windowing operation, and wherein the time-domain fractions of at least two successive digital audio samples of the plurality of digital audio samples overlap.

3. 3. The method of claim 2, wherein the overlap between the at least two successive digital audio samples is in the range of 10% to 90% overlap between the time domain fractions, preferably in the range of 25% to 75% overlap, more preferably in the range of 40% to 60% overlap, such as 50% overlap.

4. 4. The method of claim 1, wherein the window function is at least one of the group consisting of a Hamming window, a von Hann window, a Tukey window, a cosine window, a rectangular window, a B-spline window, a triangular window, a Bartlett window, a Parzen window, a Welch window, a cosine power n window (where n>1), a Kaiser window, a Nuttall window, a Blackman window, a Blackman-Harris window, a Blackman-Nuttal window, and a flattop window.

5. 5. A method according to claim 1, wherein determining the estimate of the amount of reverberation comprises weighting the amount of energy in each peak based on the magnitude of each peak or the delayed position of each peak along a time axis.

6. obtaining, by the controller, a degraded digital signal representing at least a portion of the degraded audio signal and having a duration longer than the time-domain fraction of the at least one degraded digital audio sample; obtaining, by the controller, a reference digital signal representing at least a portion of the reference audio signal and having a duration longer than the time-domain fraction of the at least one reference digital audio sample; determining, by the controller, a global impulse response signal based on the at least one degraded digital signal and the at least one reference digital signal; determining, by the controller, a global energy-time curve based on the global impulse response signal, wherein the global energy-time curve is proportional to the square root of the absolute value of the global impulse response signal; and identifying, by the controller, one or more peaks in the energy-time curve, where the one or more additional peaks in time occur delayed in the energy-time curve after a beginning of the energy-time curve based on the global impulse response signal; and determining, by the controller, a further estimate of the amount of echo or reverberation based on an amount of energy in the one or more additional peaks. The method according to any one of claims 1 to 5, further comprising the step of:

7. 7. The method of claim 6, wherein determining the further estimate of the amount of reverberation comprises weighting the amount of energy in each further peak based on the magnitude of each further peak or on the delay position of each further peak on the time axis.

8. calculating, by the controller, a partial reverberation index value based on the estimate of the amount of echo or reverberation obtained from the at least one degraded audio sample and the at least one reference audio sample; calculating by said controller and insofar as dependent on claim 6 a global reverberation index value based on said further estimate of said amount of echo or reverberation. calculating by said controller and insofar as dependent on claim 6 a final reverberation index value based on said estimate and said further estimate of said amount of echo or reverberation. The method according to any one of claims 1 to 7, comprising at least one of the steps:

9. The step of determining the local impulse response signal according to claim 1 or the step of determining the global impulse response signal according to claim 6, converting, by the controller, the at least one degraded digital audio sample and the at least one reference digital audio sample or the at least one degraded digital signal and the at least one reference digital signal from the time domain to the frequency domain by applying a Fourier transform to the at least one degraded digital audio sample and the at least one reference digital audio sample according to claim 1 or the at least one degraded digital signal and the at least one reference digital signal according to claim 6; determining, by the controller, a transfer function from a power spectrum signal from the at least one degraded digital audio sample and the at least one reference digital audio sample or the at least one degraded digital signal and the at least one reference digital signal in the frequency domain; and Transforming, by the controller, the power spectrum signal from the frequency domain to the time domain to yield the local impulse response signal of claim 1 or the global impulse response signal of claim 6. The method according to any one of claims 1 to 8, comprising the steps of:

10. 10. The method of claim 1, wherein the step of determining the local impulse response signal comprises using weighting factors that give lower weight to earlier degraded samples within a window when the voice of the corresponding reference audio sample is below a threshold, indicating a perceptually silent interval.

11. 1. A method for assessing the quality or intelligibility of a degraded speech signal received from an audio transmission system by transmitting a reference speech signal through the audio transmission system to provide the degraded speech signal, the method comprising: sampling, by the controller, the reference speech signal into a plurality of reference signal frames; sampling, by the controller, the degraded speech signal into a plurality of degraded signal frames; and forming frame pairs by correlating, by the controller, the reference signal frames and the degraded signal frames; providing, by the controller, for each frame pair, a difference function that represents a difference between the degraded signal frame and the associated reference signal frame; compensating, by the controller, the difference function for one or more disturbance types to provide, for each frame pair, a disturbance density function that is adapted to a human auditory perception model; deriving, by the controller, an overall quality parameter from the disturbance density function of a plurality of frame pairs, wherein the quality parameter is at least indicative of the quality or intelligibility of the degraded speech signal; Including, The method comprises: determining, by the controller, an amount of reverberation in at least one of the degraded speech signal and the reference speech signal, wherein the amount of reverberation is determined by applying a method according to any one of claims 1 to 10. The method.

12. 12. The method of claim 11, wherein obtaining, by the controller, the at least one degraded digital audio sample and the at least one reference digital audio sample is performed by forming the degraded audio sample and the reference audio sample from a plurality of successive signal frames, wherein the signal frames include one or more of the degraded signal frames and one or more of the reference signal frames.

13. 13. The method of claim 12, wherein the number of signal frames to be included in the plurality of signal frames depends on a duration of the time-domain fraction of the at least one degraded digital audio sample and / or the at least one reference digital audio sample, wherein the duration is greater than 0.3 seconds, preferably between 0.4 seconds and 5.0 seconds, such as at least one of 0.5 seconds, 1.0 seconds, 1.5 seconds, 2.0 seconds, 2.5 seconds, 3.0 seconds, 3.5 seconds, 4.0 seconds, 4.5 seconds, or 5.0 seconds.

14. 14. The method of claim 12 or 13, wherein for each frame pair, the compensating step is performed by setting the determined amount of reverberation in at least one of the degraded speech signal and the reference speech signal as one of the one or more disturbance types, and compensating each frame pair for an amount of reverberation associated with the respective frame pair based on the forming of the digital audio samples.

15. The method further comprises a step of noise suppression by the controller before the step of determining the local impulse response signals according to claim 1 and / or the step of determining the global impulse response signal according to claim 6, wherein the noise suppression comprises: performing a first scaling of at least one of the degraded audio signal or the reference audio signal to obtain a similar average loudness; processing the degraded speech signal to remove one or more of local signal peaks, clipping, and signal loss from the degraded speech signal; performing a second scaling of at least one of the degraded audio signal or the reference audio signal to obtain a similar average loudness; The method according to any one of claims 11 to 14, comprising:

16. 16. The method according to any one of claims 1 to 15, wherein the method is performed on an audio signal within a predetermined frequency range, such as a frequency range below a threshold frequency or a frequency range corresponding to a speech signal, for example the frequency range is below 5 kilohertz, preferably the frequency range is between 200 hertz and 4 kilohertz.

17. A computer program for causing a computer system to carry out the method according to any one of claims 1 to 16.