Speech quality evaluation device, speech quality evaluation method, and program

The speech quality evaluation device addresses the challenge of acoustic echo detection in loudspeaker systems by using DCR or BS.1116 listening tests, enhancing the E-model's ability to accurately assess sound quality.

JP7732588B2Active Publication Date: 2025-09-02NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024520177
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-12
Publication Date
2025-09-02
Estimated Expiration
2042-05-12

AI Technical Summary

Technical Problem

Existing speech quality evaluation methods, such as the E-model, struggle to accurately assess the quality of loudspeaker communication systems due to the difficulty in detecting and evaluating acoustic echo, which is not easily perceived as noise and thus not recognized as a quality issue.

Method used

A speech quality evaluation device and method that incorporates acoustic echo into the E-model evaluation by using DCR or BS.1116 listening tests to detect acoustic echo, allowing for a more accurate estimation of sound quality by comparing test sounds with reference sounds.

Benefits of technology

The proposed method effectively evaluates loudspeaker communication systems by treating acoustic echo as distortion, providing a more accurate estimation of sound quality through methods like DCR and BS.1116, ensuring consistency with PESQ evaluations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007732588000002
    Figure 0007732588000002
  • Figure 0007732588000003
    Figure 0007732588000003
  • Figure 0007732588000004
    Figure 0007732588000004
Patent Text Reader

Abstract

According to the present invention, this call quality evaluation device comprises: an Ie value calculation unit that calculates an Ie value in an R value calculation expression of an E-model, on the basis of a subjective evaluation scale including both wording that indicates a difference between a test sound and a reference sound and wording that indicates the audibility of the test sound, on the basis of DCR or ITU-R BS.1116, on the basis of the evaluation result of the test sound evaluated by an evaluator; and an R value calculation unit that calculates an R value of the E-model on the basis of the calculated Ie value.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a speech quality evaluation device, a speech quality evaluation method, and a program for estimating the sound quality of a loudspeaker communication system using an E-model. [Background technology]

[0002] As a quality evaluation index for IP telephone services, there is an E-model evaluation method that estimates a subjective evaluation value (speech MOS) from a subjective evaluation value (listening MOS) for the output signal of a voice signal processing device or a subjective evaluation value (listening MOS) estimated from the results of physical measurements (Non-Patent Document 1). [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] "JJ-201.01 IP Telephone Speech Quality Evaluation Method," 9th Edition, Information and Communications Technology Committee, August 29, 2018 Summary of the Invention [Problem to be solved by the invention]

[0004] The method in Non-Patent Document 1 can be applied to the quality evaluation of IP telephone services by evaluating the sound quality of the voice without considering delays and line echoes, which have little effect on quality. However, when evaluating voice that is affected by acoustic echo, such as hands-free calls used in cars or remote conferences, the subjective evaluation of the acoustic echo, which is the returned voice, is difficult for anyone other than the person making the call to perceive, and if the acoustic echo, which is supposed to be an unwanted sound, is not distorted, it will not be recognized as noise and will not be evaluated properly, making it difficult to apply this method.

[0005] Therefore, an object of the present invention is to provide a speech quality evaluation device that can estimate the sound quality of a loudspeaker communication system using an E-model. [Means for solving the problem]

[0006] The speech quality evaluation device of the present invention comprises an Ie value calculation unit and an R value calculation unit. The Ie value calculation unit calculates the Ie value in the E-model R value calculation formula based on the evaluation results of the test sounds evaluated by an evaluator in accordance with DCR or ITU-R BS.1116, based on a subjective evaluation scale including both wording indicating the difference between the test sounds and a reference sound and wording indicating the audibility of the test sounds. The R value calculation unit calculates the E-model R value based on the calculated Ie value. [Effects of the Invention]

[0007] The speech quality evaluation device of the present invention can estimate the sound quality of a loudspeaker communication system using an E-model. [Brief explanation of the drawings]

[0008] [Figure 1] Diagram showing E-model and R-value. [Figure 2] FIG. 1 is a diagram showing categories used in evaluation. [Figure 3] FIG. 1 is a diagram showing factors that cause sound quality degradation related to AEC processing. [Figure 4] FIG. 10 is a diagram showing test conditions for AEC processing by computer simulation. [Figure 5] FIG. 10 is a diagram showing the results of an evaluation test, which are the results of a listening test. [Figure 6] FIG. 10 is a diagram showing the results of an evaluation test using PESQ. [Figure 7] FIG. 1 is a graph showing the results of the evaluation test, illustrating the relationship between the listening test and the evaluation by PESQ. [Figure 8] FIG. 10 is a diagram showing the relationship between the degree of AEC processing and received sound. [Figure 9] A diagram showing the test conditions for AEC processed sound using an actual device. [Figure 10] FIG. 10 is a diagram showing the results of a DCR test, which is the result of an evaluation test. [Figure 11] FIG. 1 is a graph showing the relationship between PESQ and DCR as a result of an evaluation test. [Figure 12]FIG. 1 is a block diagram showing the configuration of an E-model evaluation system according to a first embodiment. [Figure 13] 3 is a flowchart showing the operation of the E-model evaluation system according to the first embodiment. [Figure 14] A diagram illustrating Ie, eff, and other parameters in the R value of the E-model. [Figure 15] FIG. 2 is a diagram showing an example of the functional configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described in detail. Components having the same functions are given the same numbers, and duplicated explanations will be omitted.

[0010] <Quality assessment method for public address systems> Below, we will clarify the differences and similarities between the call environments for handset calls and loudspeaker calls, based on the differences in parameters based on the E-model, in light of the general method of evaluating IP telephone call quality.We will also describe the results of an examination of an effective listening test method for loudspeaker calls, based on the E-model, which estimates the conversational MOS from the listening MOS in IP telephone quality evaluation.

[0011] <Quality evaluation method for IP telephones> For IP telephony, there are conversation tests that evaluate overall quality and listening tests that focus only on sound quality. By using the E-model calculation formula (1) shown in Figure 1, the conversation MOS can be estimated from the listening MOS (Reference Non-Patent Document 1, Reference Non-Patent Document 2).

[0012] (Reference Non-Patent Document 1: ITU-T Recommendation G.107, “The E-model: a computational model for use in transmission planning,” June 2015.) (Reference Non-Patent Document 2: Takahashi Rei, Yoshino Hideaki, Kitawaki Nobuhiko, "Speech Quality Assessment Technology for IP Telephony Services," IEICE Transactions on Electronics, Information and Communication Engineers, Vol. J88-B, No. 5, pp. 863-874, 2005.) The E-model is a model developed based on circuit switching technology for the purpose of checking actual service quality, and is composed of five psychological factor parameters shown in Figure 1 (Noisiness, Loudness, Delay and echo, Distortion, and Advantage factor).

[0013] The R value is used to estimate the conversational MOS, and it is common to use measured values ​​for distortion, delay, and echo, and specified values ​​for the other parameters (see Non-Patent Document 1 and Non-Patent Document 2).

[0014] Initially, it was thought that the E-model would be difficult to apply to IP telephony, but in Japan the Telecommunications Technology Committee (TTC) recommended the use of the E-model using listening tests, stating that "When providing IP telephony services, the quality factors that must be considered are sound quality, delay, and line echo, and it is desirable to evaluate the sound quality of the actual system from the perspective of checking the quality of the actual service" (Non-Patent Document 1).

[0015] Here, distortion is a factor caused by the encoding process, while delay and echo are network-related conditions. Furthermore, the TTC stipulates that listening tests should be conducted under network conditions that "guarantee the absence of fluctuations other than intentionally introduced delays and packet loss" (Non-Patent Document 1). In response to this, it is common to conduct listening tests under the simpler condition of "no network effects." Among network-related conditions, for example, delay has no effect on quality unless the test is a conversation test, and echo (line echo) has little effect on quality, being perceived as being similar to telephone sidetone if the delay is sufficiently short (Non-Patent Document 1). Furthermore, packet loss has a significant impact on sound quality degradation, but it is included in the evaluation criteria for the encoding process, so it does not need to be considered here. Therefore, for IP telephones, we assume no network effects and evaluate only distortion (especially encoded voice).

[0016] <Proposal of a quality assessment method for public address calls> To estimate the overall quality of a public address call using the E-model, it is desirable that the only parameter involved in the listening test is distortion (AEC processed audio, including encoding), just as with IP telephones. The biggest difference between public address calls and IP telephones in terms of call conditions is the presence of acoustic echo. While network-related conditions and line echo can be considered to have no effect on quality, just like with IP telephones, acoustic echo is a condition that has a significant impact on quality when evaluating public address calls and cannot be ignored.

[0017] In this specification, acoustic echo in listening tests is assumed to be "interfering audio caused by a third party." Therefore, we proposed incorporating acoustic echo into sound quality evaluations together with distortion, and confirmed the theoretical validity of this evaluation. As with IP phone evaluations, we assumed no network effects and evaluated the sound quality of AEC-processed audio, including encoding. Treating acoustic echo as distortion in listening tests requires careful selection and ingenuity of the evaluation method. For IP phones, the Absolute Category Rating (ACR), also known as the MOS test (Reference Non-Patent Document 3), is commonly used, and the listening MOS estimated by PESQ is also the result of the ACR.

[0018] (Reference Non-Patent Document 3: ITU-T Recommendation P.800, “Methods for subjective determination of transmission quality,” Aug. 1996.) ACR is a suitable method for evaluating speech quality because it evaluates the impression received by the evaluator directly. However, when using ACR to evaluate "received sound with superimposed acoustic echo," there is a risk that the acoustic echo, which should be unwanted sound, will not be recognized as noise and will receive a high rating if it is not distorted. This is why it has been said that detecting acoustic echo is difficult unless it is a conversation test. In order to evaluate public address calls in a listening test, some ingenuity is required to enable the evaluator to detect the "acoustic echo."

[0019] In this specification, we have examined acoustic echo detection methods with reference to official tests conducted in test laboratories when the coding method, which is the main quality factor of IP telephony, was recommended as the ITU-T international standards G.729 (Reference Non-Patent Document 4) and G.711.1 (Reference Non-Patent Document 5).

[0020] (Reference Non-Patent Document 4: ITU-T Recommendation G.729, “Coding of speech at 8 kbit / s using conjugate-structure algebraic-code-excited linear prediction (CS-ACELP),” June 2016.) (Reference Non-Patent Document 5: ITU-T Recommendation G.711.1, “Wideband embedded extension for ITU-T G.711 pulse code modulation,” Sep. 2012.) In official tests, ACR is used to evaluate basic performance and clean speech, but for conditions where ambient noise is superimposed or coding methods that require high-quality performance, where the scores tend to be biased downward or upward, DCR (Degradation Category Rating, Reference Non-Patent Document 3) or ITU-R BS.1116 (Double-blind triple-stimulus with hidden reference, Reference Non-Patent Document 6) are often used (Reference Non-Patent Document 7).

[0021] (Reference Non-Patent Document 6: ITU-R Recommendation BS.1116-3, “Methods for the subjective assessment of small impairments in audio systems,” February 2015.) (Reference Non-Patent Document 7: Shoko Kurihara, Akitoshi Kataoka, Shinji Hayashi, Takao Kaneko, "Quality Evaluation for ITU-T G.729 Speech Coding Scheme Extensions," IEICE Transactions on Information and Communication Engineers, D-II 87(2), pp.416-426, 2004.) DCR is a method added to ITU-T P.800 to improve the evaluation accuracy of conditions where differences are difficult to detect with ACR, and BS.1116 is a method for evaluating DCR with even greater precision. Like PESQ, DCR and BS.1116 evaluate by comparing a reference sound (ideal sound) with the sound to be evaluated (received sound), so they can detect sound quality differences that are difficult to detect with ACR. The result is the Degradation Mean Opinion Score (DMOS, see Non-Patent Document 3), which indicates the quality difference from the reference sound. Figure 2 shows the categories used for ACR and DCR BS.1116.

[0022] The ACR, DCR, and BS.1116 disclosed here use five-level categories with the same meaning. Although the evaluation range changes when the evaluation method is changed, all categories are ordinal scales, and unless different evaluation methods are compared, the superiority or inferiority of sound quality will not change regardless of which evaluation method is used, and the distribution and quality differences will be reflected as they are. In the official tests when the coding method, which is the main quality of IP telephones, was recommended as an ITU-T international standard, ACR, DCR, and BS.1116 were adopted to match the evaluation conditions (Reference Non-Patent Document 7, Reference Non-Patent Document 8).

[0023] (Reference Non-Patent Document 8: ITU-T SG 12 Q.7 / 12 Rapporteurs, “Superwideband extension to G.711.1 and G.722 Qualification Quality Assessment Test Plan,” October 2008.) When evaluating normal sound quality, DCR, which is designed for general evaluators, is used, but for evaluation targets where the difference in sound quality is slight, BS.1116, which is designed for expert listeners and can detect even subtle differences with high accuracy, is often used (Reference Non-Patent Document 7, Reference Non-Patent Document 8).

[0024] Following this example, we propose adopting DCR or BS.1116 for listening tests of public address calls. By comparing the speech uttered by the far-end talker as a reference sound, even if the speech is clear and not distorted by acoustic echo, if acoustic echo is superimposed, it can be determined to be unwanted sound.

[0025] Considering consistency with PESQ, evaluation results for "satisfaction with calls (ease of listening)" such as those obtained from ACR are necessary. Therefore, in order to obtain results that focus on "ease of listening" while evaluating with DCR or BS.1116, which can present reference sounds, we propose the unique categories shown in Figure 2. In this specification, we will confirm the relationship with PESQ as a method for evaluating public address calls in listening tests, including the change from ACR to DCR and BS.1116.

[0026] <Quality evaluation of AEC processed audio generated by computer simulation> <AEC processed audio by computer simulation> Here, we conduct a sound quality evaluation test on computer-generated AEC-processed speech to confirm the effectiveness of the listening test and PESQ evaluation for loudspeaker calls proposed in <Quality evaluation method for loudspeaker calls>, as well as the relationship between the proposed listening test and PESQ evaluation.

[0027] The audio used in the test assumed the sound quality degradation factors related to AEC processing, "without AEC processing" and "with AEC processing (insufficient / appropriate / excessive)," as shown in Figure 3. Using two commonly used degradation measures, SER (Signal to Distortion Ratio) and SDR (Signal to Echo Ratio; see Non-Patent Document 9), as objective evaluation measures for AEC processing, the test created double-talk conditions by computer-generated "audio distortion" and "residual echo."

[0028] (Reference non-patent document 9: M. Fukui, S. Shimauchi, and A. Nakagawa, "Convolutive residual echo power estimation for acoustic echo reduction," Journal of Signal Processing, vol.24, no.6, pp.237-245, Nov. 2020.) As shown in Figure 4, 90 conditions are targeted, combining 9 conditions of SER: -6 to 42 dB (in 6 dB increments) and 10 conditions of SDR: 3 to 30 dB (in 3 dB increments). The equations used for processing are shown in (2) and (3).

number

[0029] For evaluation testing, it is necessary to prepare test sounds that combine sound quality that anyone would consider good (rating 5) with sound quality that anyone would consider bad (rating 1). Therefore, in addition to AEC-processed sound from a computer simulation, we decided to evaluate "speech emitted by the far-end talker (ideal sound)" as a condition for sound quality that anyone would consider good, and "speech with superimposed acoustic echo without AEC processing" as a condition for sound quality that anyone would consider bad.

[0030] <Quality evaluation of AEC processed audio created by computer simulation> In our <Quality Assessment Method for Public Address Communication>, we proposed the use of DCR / BS.1116, which is a listening test for public address communication, to enable the evaluator to detect acoustic echoes by comparing the sound with a reference sound.

[0031] The test sounds to be evaluated here were created by computer simulation, processing the SER and SDR values ​​in 6 dB and 3 dB increments, and the quality difference between each sound was minimal. Therefore, we decided to use BS.1116, a method for expert listeners that can detect subtler differences than DCR (see Non-Patent Document 8). BS.1116 compares three sounds—a reference sound (ideal sound), a hidden reference sound (sound identical to the reference sound), and the sound to be evaluated (received sound)—as many times as necessary until the listener is satisfied. The hidden reference sound is identified and evaluated using the same five-point degradation category scale as DCR, in 0.1 increments. While BS.1116 typically uses the difference between the reference sound (5 points) and the score assigned by the evaluator as the evaluation value, in this specification, we consider the evaluation value to be the score assigned by the evaluator itself, as with DCR, because the purpose of this study is to investigate the relationship with PESQ evaluation.

[0032] The audio used in the test consisted of 368 conditions: four speakers (two females, two males) with 9 SER conditions, 10 SDR conditions, ideal sound (score of 5), and audio with acoustic echo superimposed without AEC processing (score of 1). The evaluators were 64 members of the public. This number was 2.5 times larger than the usual number to achieve evaluation accuracy comparable to that of expert listeners, based on the official G.711.1 testing conducted at a test lab when the coding method, which is a key quality factor for IP telephony, was recommended as the ITU-T international standards G.729 and G.711.1 (see Non-Patent Documents 5 and 8). Furthermore, the order in which the test sounds were presented is an important factor in conducting the evaluation test. Because the timing of the test sounds can affect the evaluation scores, the test sounds were presented randomly, with each group of four evaluators presenting the sounds in a different order.

[0033] The results of the proposed listening test and PESQ for AEC-processed speech are shown in Figures 5 to 7. The data plotted here are for 90 conditions of AEC-processed speech (9 SER conditions × 10 SDR conditions) created through computer simulation, and each point is the average of 256 data points (4 speakers × 64 evaluators).

[0034] Figure 5 shows the results of the BS.1116 listening test, and Figure 6 shows the results of the PESQ evaluation. Here, the vertical axis shows the evaluation value, the horizontal axis shows SER, and each line shows SDR. To read the graph, the vertical axis indicates quality, with higher values ​​indicating higher quality. Figure 5 shows the BS.1116 evaluation value (the actual score given by the evaluator), and Figure 6 shows the PESQ evaluation value (raw score). The horizontal axis SER and line SDR are conditions related to sound quality, with higher values ​​indicating better sound quality. In theory, the combination of SER: 42 dB, SDR: 30 dB provides the highest sound quality, while the combination of SER: -6 dB, SDR: 3 dB provides the lowest sound quality.

[0035] The results of the BS.1116 listening test and the PESQ evaluation showed similar trends, and as expected, the higher the SER and SDR values, the higher the evaluation. The listening evaluation values ​​do not follow a smoother line than the PESQ evaluation values, but this is presumably because the test sounds used here had only slight differences in quality, making evaluation difficult.

[0036] In the results for BS.1116 in Figure 5, the ranking of the subjective quality assessment values ​​corresponding to SDR and SER did not change, and so the results obtained are consistent with the assumption that the proposed method can treat "acoustic echo" in a loudspeaker call as the same condition as "interference sound from a third party superimposed on the far-end terminal" in an IP phone call.

[0037] Similarly, the PESQ evaluation results in Figure 6 show that the results are consistent with the assumption that "acoustic echo" in a loudspeaker call can be considered to be the same condition as "ambient noise on the transmitting side" in an IP phone call.

[0038] Figure 7 shows the relationship between BS.1116 listening scores and PESQ scores. Here, the vertical axis represents the BS.1116 scores (the scores given by the evaluators) and the horizontal axis represents the PESQ scores (raw scores). The plotted data represents the actual measured values ​​of BS.1116 listening scores and PESQ scores, and the solid line represents the estimated values ​​calculated by regression analysis (regression line).

[0039] From the results of the evaluation test, the coefficient of determination R adjusted for the degrees of freedom of the estimation model (hereinafter referred to as the BS.1116 estimation model) for the listening evaluation value (measured value of BS.1116) and the PESQ evaluation value was 2 was 0.97, which indicates that the BS.1116 estimation model is 97% consistent with the actual measurements, and it was confirmed that PESQ can explain 97% of the proposed listening test BS.1116.

[0040] As shown in the figure, the relationship between the listening evaluation value and the PESQ evaluation value can be approximated by a linear function Fy = a·x + b. x is the PESQ value, y is the listening evaluation value, a is 1.3 or close to 1.3, and b is -0.3 or close to -0.3. Near α means a value in the range α-δ1 or greater and α-δ2 or less. Note that δ1 and δ2 are positive values, and δ1 may be equal to δ2 or may not be equal to δ2. Examples of δ1 and δ2 are values ​​that are 10% or 20% of |α|. For example, a = 1.33 and b = -0.27.

[0041] The results of this test, which involved conducting the proposed listening test (BS.1116) and PESQ evaluation on AEC-processed speech in which "speech distortion," "acoustic echo," and "residual echo (residual acoustic echo)" were gradually superimposed by a computer, showed that the results were consistent with the assumptions underlying the proposed listening test. It was also confirmed that the BS.1116 estimation model estimates the measured BS.1116 values ​​with extremely high accuracy.

[0042] <AEC processed audio quality evaluation using actual equipment> <AEC processed audio from actual equipment> Here, in order to confirm the validity and relationship of the proposed listening tests and PESQ evaluations examined in <Quality evaluation methods for public address communications> and carried out in <Quality evaluation of AEC-processed speech created by computer simulation>, the evaluation target will be expanded to "public address communications sound from actual equipment (AEC-processed speech)" and the proposed listening tests and PESQ evaluation tests will be carried out. The test sounds used here are "received sounds from two-way communications" recorded in advance using seven types of communication equipment.

[0043] The speakers were four far-end speakers who uttered the sounds to be evaluated, and four near-end speakers who uttered the original sounds of the acoustic echo.To make it easier to distinguish the acoustic echo, the far-end and near-end speakers were of the opposite sex.There were three speaking conditions: single talk and two types of double talk (mixed sounds from the far-end and near-end speakers).

[0044] In listening tests, it is necessary to prepare test sounds that combine sound quality that anyone would consider good (rating 5) with sound quality that anyone would consider bad (rating 1). However, because the target here is sound recorded from an actual device, it is not possible to prepare a wide variety of test sounds like those used in the <Quality evaluation of AEC-processed sound created by computer simulation>. Therefore, in addition to AEC-processed sound from an actual device, we decided to evaluate "single talk received without AEC processing" as a condition for sound quality that anyone would consider good, and "double talk received without AEC processing with superimposed acoustic echo" as a condition for sound quality that anyone would consider bad.

[0045] Sound quality differs depending on the communication device, but theoretically, the highest sound quality is achieved when the received sound is single talk without AEC processing, and the lowest sound quality is achieved when the received sound is double talk with superimposed acoustic echo without AEC processing.

[0046] The relationship between the level of AEC processing and the quality of the received sound is shown in Figure 8. Note that the relationship between AEC processing and sound quality shown in the figure is an image of the sound quality that was aimed for when planning this test, and does not represent the actual sound quality of the test sound in this test.

[0047] Figure 9 shows the test conditions for AEC processing using an actual device. The test sounds used in this test consisted of 168 conditions: four speakers (two women, two men) x seven communication devices x three speech conditions x two AEC processing conditions (AEC ON / OFF). The test sounds were presented randomly, with a different presentation order for each group of evaluators (four people per group). To simplify the conditions, the volume and network delay were constant, and there was no packet loss. The test sounds used here included all the effects of the mobile device, transmission path, encoding process, etc., and do not indicate the performance of the AEC itself.

[0048] <AEC processed audio quality evaluation using actual equipment> The sound to be evaluated here has a rougher quality difference than the sound quality evaluation of AEC processed sound created by computer simulation, and does not require high evaluation accuracy. For this reason, DCR was used in this test, and 24 ordinary people were used as evaluators.

[0049] Figure 10 shows the results of listening tests on sounds recorded with an actual device. Here, the vertical axis represents the DCR listening score (DMOS), and the horizontal axis represents the communication device number. Each point on the graph represents the score for each AEC processing condition and speech condition, with higher quality indicated at the top of the graph. Here, black and white triangles represent single talk, black and white circles represent double talk 1 (far-end talker first), and black and white squares represent double talk 2 (near-end talker first). Here, solid black represents the condition with AEC processing (hereafter referred to as AEC ON), and open white represents the condition without AEC processing (hereafter referred to as AEC OFF). Each point is the average of 96 data points (4 speakers x 24 evaluators).

[0050] The results of the proposed listening test by DCR were generally in line with theory, with the condition of "single talk reception sound with AEC OFF" receiving the highest rating, and the condition of "double talk reception sound with acoustic echo superimposed with AEC OFF" receiving the lowest rating.

[0051] The graph shown here shows the distribution of sound quality variations evaluated, and does not indicate differences in AEC processing performance. By including AEC-off conditions (received sound of single talk and received sound of double talk with superimposed acoustic echo) in the evaluation, we were able to increase the variation.

[0052] Figure 11 shows the relationship between DMOS and PESQ evaluation scores. Here, the vertical axis is DMOS, and the horizontal axis is PESQ evaluation score (raw score). The crosses indicate the actual measured values ​​of DMOS and PESQ evaluation scores, and the solid lines indicate the estimated values ​​(regression lines) calculated by regression analysis. The plotted data here is 79 data points, obtained by excluding five PESQ error conditions from the average values ​​of seven communication devices, three speech conditions, two AEC processing conditions, and two speaker genders (male and female), and each point is the average value of 48 data points (two speakers per gender × 24 evaluators).

[0053] As a result of the evaluation test, the coefficient of determination R adjusted for the degrees of freedom of the estimation model for the DMOS measured values ​​and PESQ evaluation values ​​(hereinafter referred to as the DMOS estimation model) was 2 was 0.71. This indicates that the DMOS estimation model is 71% consistent with the measured values. This means that 71% of the variance in DMOS can be explained, and it can be said that the proposed listening test DCR can be estimated from PESQ with high accuracy.

[0054] Furthermore, detailed analysis of the five PESQ error conditions excluded from the calculation revealed that residual acoustic echo may have caused a malfunction in the PESQ's internal function for adjusting the time difference between the reference signal and the degraded signal, resulting in an error in the calculation of the PESQ evaluation score. Furthermore, analysis of other conditions performed simultaneously confirmed that, if there was no time difference, no malfunction occurred and the PESQ function operated normally. The malfunction of the time difference adjustment function within the PESQ revealed here is a problem with PESQ evaluation (subjective value estimation) for loudspeaker speech. This problem can be addressed by synchronizing the reference signal and the degraded signal in advance. For example, if an adjustment error occurs despite prior synchronization, discarding the PESQ evaluation score in question may improve the accuracy of the subjective value estimation.

[0055] From the results of this test, we conducted a proposed listening test (DCR) on actual AEC-processed audio, and obtained results consistent with the theory, confirming that the hypothesis was correct. We also confirmed the relationship between the evaluation values ​​obtained from the proposed listening test and the values ​​estimated by PESQ. [Example]

[0056] Based on the above-mentioned research results, an E-model evaluation system according to a first embodiment, capable of estimating the sound quality of a loudspeaker communication system using an E-model, will now be described. The configuration of the E-model evaluation system of this embodiment will be described with reference to FIG. 12. As shown in the figure, the E-model evaluation system 1 of this embodiment includes a data storage device 11, a subjective evaluation device 12, an objective evaluation device 13, and a speech quality evaluation device 14. The subjective evaluation device 12 includes a test sound presentation unit 121, an evaluation result acquisition unit 122, a compilation unit 123, and a compilation result storage unit 120A. The objective evaluation device 13 includes a PESQ evaluation value calculation unit 131, a linear transformation unit 132, a PESQ evaluation value storage unit 130A, and an estimated value storage unit 130B. The speech quality evaluation device 14 includes an Ie value calculation unit 141, an R value calculation unit 142, and an R value storage unit 140A. The operation of each device and component will now be described with reference to FIG. 13.

[0057] <Data storage device 11> The data storage device 11 stores test sounds in advance. An example of the test sound is preferably one that includes N×M×P speech sounds from P speakers, with N levels of speech distortion and M levels of residual echo superimposed in stages, where N, M, and P are integers greater than or equal to 1. N, M, and P can be set arbitrarily. An example where N=9, M=10, and P=4 has already been shown in FIG. 4 and the corresponding explanation. Note that the test sound may include not only computer-processed sounds (speech distortion and acoustic echo) but also sounds processed by actual equipment, such as "output speech from a communication device."

[0058] As shown in Figure 4, if the test sound is superimposed in stages with SER: -6 to 42 dB in 9 steps of 6 dB, and residual echo: SDR: 3 to 30 dB in 10 steps of 3 dB, this is suitable because it is expected that highly accurate evaluations can be obtained, as shown in Figures 5 to 7.

[0059] <Subjective assessment device 12> The speech quality evaluation unit 14, which will be described later, is a unit that calculates an R value based on a listening evaluation value or a PESQ value, and so the processing flow differs depending on whether the R value is calculated based on a listening evaluation value or a PESQ value. First, the flow of the subjective assessment unit 12 when calculating an R value based on a listening evaluation value will be explained.

[0060] <Test sound presentation unit 121> The test sound presentation unit 121 presents the test sound stored in the data storage device 11 to the evaluator (S121).

[0061] <Evaluation result acquisition unit 122> The evaluation result acquisition unit 122 acquires the evaluation result in which the evaluator evaluates the test sound based on the DCR or ITU-R BS.1116, based on a subjective evaluation scale that includes both words indicating the difference between the test sound and the reference sound and words indicating the audibility of the test sound (S122).

[0062] The phrase indicating the difference from the reference sound of the test sound is, for example, phrases such as "no difference (unknown) from the reference sound" and "there is a difference (there is a discrepancy) from the reference sound", and the phrase indicating the ease of hearing the test sound is, for example, phrases such as "easy to hear", "no problem in hearing", "a little difficult to hear", "difficult to hear", and "very difficult to hear".

[0063] An example of a subjective evaluation scale including both the phrase indicating the difference from the reference sound of the test sound and the phrase indicating the ease of hearing the test sound has already been shown in FIG. 2.

[0064] As shown in FIG. 2, if the subjective evaluation scale includes "5: the difference from the reference sound is unknown", "4: there is a difference but no problem in hearing", "3: there is a difference and a little difficult to hear", "2: there is a difference and difficult to hear", and "1: there is a difference and very difficult to hear", as shown in FIGS. 5 to 11, it is suitable because high-precision evaluation can be expected to be obtained.

[0065] <Aggregation unit 123>[[ID=X]] The aggregation unit 123 aggregates the evaluation results and stores them in the aggregation result storage unit 120A (S123).

[0066] ID=17]] <Aggregation result storage unit 120A> The aggregation result storage unit 120A stores the evaluation results obtained in step S122 and aggregated in step S123.

[0067] [[ID=2X]]<Objective evaluation device 13> Next, the flow of the objective evaluation device 13 when calculating the R value based on the PESQ evaluation value will be described.

[0068] <PESQ evaluation value calculation unit 131> The PESQ evaluation value calculation unit 131 calculates the PESQ evaluation value of the test sound and transmits the calculated PESQ evaluation value to the PESQ evaluation value storage unit 130A (S131). Regarding the calculation example of the PESQ evaluation value, the algorithm is strictly defined in ITU-T Recommendation P.862, and reference software is attached to this recommendation. It has already been shown in FIG. 6 and the corresponding description.

[0069] <PESQ evaluation value storage unit 130A> The PESQ evaluation value storage unit 130A stores the PESQ evaluation value calculated in step S131.

[0070] <Linear transformation unit 132> The linear transformation unit 132 linearly transforms the PESQ evaluation value calculated in step S131 based on the regression equation obtained by performing regression analysis on the evaluation result and the PESQ evaluation value, acquires an estimated value of the subjective evaluation value, and stores the acquired estimated value in the estimated value storage unit 130B (S132). Regarding the calculation example of the regression equation, it has already been shown in FIG. 7, FIG. 11 and the corresponding description.

[0071] <Estimated value storage unit 130B> The estimated value storage unit 130B stores the estimated value of the subjective evaluation value acquired in step S132.

[0072] <Call quality evaluation device 14> Hereinafter, the operation of the call quality evaluation device 14 in each flow of the listening evaluation and the PESQ evaluation will be described.

[0073] <Ie value calculation unit 141 (in the case of listening evaluation)> Based on a subjective evaluation scale that includes both the statement indicating the difference from the reference sound of the test sound and the statement indicating the ease of hearing the test sound, the Ie value calculation unit 141 calculates the Ie value in the R value calculation formula of the E-model based on DCR or ITU-R BS.1116 and based on the evaluation result (details have already been shown in step S122 etc.) of the test sound evaluated by the evaluator (S141). The Ie value calculation unit 141 can obtain the Ie value by using the conversion formula of ITU-T P.833 for the evaluation result.

[0074] <(In the case of PESQ evaluation, Ie value calculation unit 141)> Based on the estimated value of subjective evaluation obtained by linearly converting the PESQ evaluation value of the test sound using the regression formula previously obtained for the evaluation result of the listening test and the PESQ evaluation value of the test sound, the Ie value calculation unit 141 calculates the Ie value in the R value calculation formula of the E-model (S141). The Ie value calculation unit 141 can obtain the Ie value by using the conversion formula of ITU-T P.833 for the estimated value.

[0075] As shown in FIG. 14, the Ie value is the listening evaluation value of the codec, and eff is the information of the transmission error.

[0076] Also, as described above, it is preferable that the subjective evaluation scale is configured to include "5: The difference from the reference sound is not distinguishable", "4: There is a difference but there is no problem in hearing", "3: There is a difference and it is a little difficult to hear", "2: There is a difference and it is difficult to hear", "1: There is a difference and it is very difficult to hear".

[0077] Also, as described above, it is preferable that the test sound is configured to include voices of P speakers with N×M×P conditions in which the voice distortion is superimposed in N steps, the residual echo is superimposed in M steps, and N, M, and P are integers greater than or equal to 1.

[0078] <R value calculation unit 142> Based on the calculated Ie value, the R value calculation unit 142 calculates the R value of the E-model and stores it in the R value storage unit 140A (S142).

[0079] <R value storage unit 140A> The R value storage unit 140A stores the R value calculated in step S142. <Supplementary note> The device of the present invention, for example, as a single hardware entity, has an input unit to which a keyboard or the like can be connected, an output unit to which a liquid crystal display or the like can be connected, a communication unit to which a communication device (for example, a communication cable) that can communicate outside the hardware entity can be connected, a CPU (Central Processing Unit, which may include a cache memory and registers), a memory such as RAM and ROM, an external storage device such as a hard disk, and a bus that connects these input unit, output unit, communication unit, CPU, RAM, ROM, and external storage device so that data can be exchanged between them. Also, if necessary, a device (drive) that can read and write a recording medium such as a CD-ROM may be provided in the hardware entity. Examples of such a physical entity equipped with such hardware resources include general-purpose computers.

[0080] In the external storage device of the hardware entity, programs necessary to realize the above functions and data necessary for the processing of these programs are stored (not limited to the external storage device, for example, the program may be stored in a read-only storage device such as ROM). Also, data obtained by the processing of these programs is appropriately stored in RAM, an external storage device, or the like.

[0081] In the hardware entity, each program stored in the external storage device (or ROM, etc.) and data necessary for the processing of each program are read into the memory as needed and appropriately interpreted and executed / processed by the CPU. As a result, the CPU realizes a predetermined function (each component represented as the above-mentioned... unit,... means, etc.).

[0082] The present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the spirit of the present invention. Furthermore, the processes described in the above embodiments may not only be executed in chronological order according to the order described, but may also be executed in parallel or individually depending on the processing capacity of the device that executes the processes or as needed.

[0083] As described above, when the processing functions of the hardware entities (apparatuses of the present invention) described in the above embodiments are realized by a computer, the processing contents of the functions that the hardware entities should have are described by a program. Then, by executing this program on a computer, the processing functions of the hardware entities are realized on the computer.

[0084] The various processes described above can be implemented by loading a program that executes each step of the above method into the recording unit 10020 of the computer 10000 shown in Figure 15 and operating the control unit 10010, input unit 10030, output unit 10040, etc.

[0085] The program describing the processing contents can be recorded on a computer-readable recording medium. Examples of computer-readable recording media include magnetic recording devices, optical disks, magneto-optical recording media, and semiconductor memories. Specifically, examples of magnetic recording devices include hard disk drives, flexible disks, and magnetic tapes; optical disks include DVDs (Digital Versatile Discs), DVD-RAMs (Random Access Memory), CD-ROMs (Compact Disc Read Only Memory), and CD-Rs (Recordable) / RWs (Rewritable); magneto-optical recording media include MOs (Magneto-Optical discs), and semiconductor memories include EEP-ROMs (Electrically Erasable and Programmable-Read Only Memory).

[0086] The program may be distributed, for example, by selling, transferring, lending, etc. a portable recording medium such as a DVD or CD-ROM on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to another computer via a network, thereby distributing the program.

[0087] A computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes the process in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program. Furthermore, the computer may execute the process in accordance with the received program each time a program is transferred from a server computer to the computer. Alternatively, the server computer may not transfer the program to the computer, but may execute the process through a so-called ASP (Application Service Provider) service, which realizes the processing function by issuing an execution instruction and obtaining the results. In this embodiment, the program includes information used for processing by a computer that is equivalent to a program (such as data that is not a direct instruction to the computer but has properties that define computer processing).

[0088] In addition, in this embodiment, a hardware entity is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.

Claims

1. an Ie value calculation unit that calculates the Ie value in the R value calculation formula of the E-model based on an evaluation result obtained by an evaluator evaluating the test sound based on DCR or ITU-R BS.1116, based on a subjective evaluation scale including both a statement indicating the difference between the test sound and a reference sound and a statement indicating the audibility of the test sound; Includes an R-value calculation section that calculates the R-value of the E-model based on the calculated Ie value. Call quality assessment device.

2. 2. A speech quality evaluation device according to claim 1, The Ie value calculation unit The Ie value in the E-model R value calculation formula is calculated based on the subjective evaluation estimate obtained by linearly converting the PESQ evaluation value of the test sound using a regression equation previously obtained for the evaluation results and the PESQ evaluation value of the test sound. Call quality assessment device.

3. 3. A speech quality evaluation device according to claim 1 or 2, The subjective evaluation scale includes the following: "5: no difference from the reference sound", "4: there is a difference but it does not affect hearing", "3: there is a difference and it is slightly difficult to hear", "2: there is a difference and it is difficult to hear", and "1: there is a difference and it is very difficult to hear". Call quality assessment device.

4. 3. A speech quality evaluation device according to claim 1 or 2, The test sound is N, M, and P are integers greater than or equal to 1, and include speech under NxMxP conditions by P speakers with N levels of speech distortion and M levels of residual echo superimposed in stages, or speech output from a communication device. Call quality assessment device.

5. A speech quality evaluation method in which a speech quality evaluation device executes each step, an Ie value calculation step of calculating the Ie value in the R value calculation formula of the E-model based on an evaluation result obtained by an evaluator evaluating the test sound based on a subjective evaluation scale including both a statement indicating the difference between the test sound and a reference sound and a statement indicating the audibility of the test sound, and based on DCR or ITU-R BS.1116; Includes an R-value calculation step that calculates the R-value of the E-model based on the calculated Ie value. Call quality evaluation method.

6. 6. A speech quality evaluation method according to claim 5, The Ie value calculation step The Ie value in the E-model R value calculation formula is calculated based on the subjective evaluation estimate obtained by linearly converting the PESQ evaluation value of the test sound using a regression equation previously obtained for the evaluation results and the PESQ evaluation value of the test sound. Call quality evaluation method.

7. A program that causes a computer to function as the speech quality evaluation device according to claim 1 or 2.

Citation Information

Patent Citations

  • Quality of experience determination for multi-party VOIP conference calls that account for focus degradation effects

    US20150156314A1