Acoustic quality evaluation device, acoustic quality evaluation method, and program
The sound quality evaluation device and method provide an objective evaluation of acoustic quality in public address systems by analyzing reference and noisy signals, addressing the limitations of costly and non-reproducible subjective evaluations.
Patent Information
- Application Number
- PCT/JP2024/024361
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-05
- Publication Date
- 2026-01-08
AI Technical Summary
Existing methods for evaluating acoustic quality in public address communication systems, particularly in noisy environments, are costly and lack reproducibility, as they rely on time-consuming conversation tests or subjective evaluations that do not accurately reflect user experience.
A sound quality evaluation device and method that uses a first and second sound quality evaluation unit to analyze reference and noisy signals, calculating PESQ values and determining subjective ratings, allowing for objective evaluation with limited human intervention.
Enables accurate acoustic quality evaluation in noisy environments through objective testing, reducing costs and improving reproducibility compared to traditional methods.
Smart Images

Figure JP2024024361_08012026_PF_FP_ABST
Abstract
Description
Acoustic quality evaluation device, acoustic quality evaluation method, and program
[0001] The disclosed technology relates to a technology for evaluating speech quality, and in particular to a quality evaluation test technology for a loudspeaker communication system.
[0002] With the advancement of communication technology, opportunities to use public address communication systems such as conference systems and hands-free calls using smartphones are increasing due to the convenience of being able to make calls without carrying any equipment. Acoustic echo cancellers (AEC) are used to remove acoustic echo and ambient noise that can be a problem in public address communication systems, providing a comfortable calling environment.
[0003] Figure 1 shows a schematic diagram of acoustic echo and AEC. A near-end talker 101 and a far-end talker 102 communicate using a public address system. Reference numerals 103 and 104 denote a microphone and speaker on the near-end talker's side, and 105 and 106 denote a microphone and speaker on the far-end talker's side.
[0004] The "Hello" uttered by near-end speaker 101 is output from far-end speaker 105 (107) and reaches the ears of far-end speaker 102. In a public address system, speaker output 107 is also picked up by far-end microphone 106 (loopback 108). If the near-end speaker's voice "Hello" (acoustic echo) picked up by far-end microphone 106 is transmitted directly to the near-end side, it can make communication difficult or cause feedback. For this reason, public address system systems are equipped with AEC 109, which transmits a voice signal to the near-end side from which the voice originating from the near-end speaker has been removed or reduced. If the AEC also has a noise cancellation function, noise around the far-end speaker is also removed or suppressed.
[0005] If the AEC effect is weak, the acoustic echo remains, but if it is too strong, even the speech transmitted from the far end is removed, resulting in distortion or disappearance and difficulty in hearing. Because AEC performance depends on how accurately the acoustic echo is removed, the mainstream method of evaluating AEC performance has traditionally been objective evaluation (evaluation using computers, etc.) that focuses on the amount of acoustic echo removed. Objective evaluation is easy because it can be performed using computer processing, but there is a problem in that it does not necessarily match the quality experienced by the user in an actual call (also known as "user-perceived quality").
[0006] To evaluate acoustic echo or AEC-processed sound through subjective evaluation (human listening), the acoustic echo must be perceived, and the evaluation can only be made by the evaluator actually speaking. Therefore, for public address communication systems such as hands-free public address systems, quality evaluation through two-way conversation tests has been recommended (see Non-Patent Document 1). However, conducting conversation tests requires know-how, is time-consuming and costly, and has the problem of low reproducibility.
[0007] On the other hand, when using a handset or headset for a call, the speech transmitted from the far-end is not affected by acoustic echoes and other factors from the near-end speaker, allowing only the far-end speech to be evaluated. In this case, speech quality can be evaluated using a simplified conversation test, a listening test for one-way conversations, and this test method is commonly used for evaluating IP phone call quality. Listening tests are more convenient than conversation tests because they have higher reproducibility and shorter implementation time. Objective evaluation methods, such as the Perceptual Evaluation of Speech Quality (PESQ), which estimates subjective evaluation values (also known as "Mean Opinion Score (MOS)") based on listening tests, have also been established (see Non-Patent Document 2). Recently, methods have been proposed for applying subjective evaluations based on listening tests and objective evaluations such as PESQ to public address communication systems (Non-Patent Document 3, Patent Document 1).
[0008] JP 2016-46694 A International Publication No. WO2024 / 121962
[0009] ITU-T, "ITU-T Recommendation P.800: Methods for subjective determination of transmission quality", ITU, 1996; ITU-T, "ITU-T Recommendation P.862: Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs", ITU, 2002; Kurihara, S., Shimauchi, S., Fukui, K., and Harada, N., "Evaluation of user-perceived quality in hands-free calling: A study of a subjective evaluation method consistent with PESQ," IEICE Technical Report, vol. 117, no. 386, CQ2017-96, pp. 63-68, January 2018.
[0010] With the widespread use of smartphones and PCs, there are increasing opportunities to use hands-free calls in a noisy environment. Environmental noise includes, for example, office air conditioning noise, interior noise from a moving car, traffic noise at an intersection, insects, keyboard clicks, factory machinery noise, and multiple people's voices (chat noise), regardless of whether the noise is loud or quiet, indoors or outdoors. However, no method has yet been established for evaluating the sound quality of public address communication systems in such an environment.
[0011] When the near-end talker is in a "quiet environment," the speaker output sound from the far-end terminal (the far-end talker's voice, the far-end talker's ambient noise, acoustic echo, distortion of the far-end talker's voice due to AEC processing, residual echo, etc.) is easy to perceive in detail, so even if there is even a small amount of distortion or noise superimposition, the rating will be low. On the other hand, when the near-end talker is in an "environment with ambient noise," the speaker output sound from the far-end terminal is masked by the "near-end talker's ambient noise," making it difficult to hear the noise from the far-end terminal (the far-end talker's ambient noise, acoustic echo, distortion of the far-end talker's voice due to AEC processing, residual echo, etc.). For this reason, even if there is some distortion or noise superimposition, the impact on the rating is small, and the rating tends to be higher compared to when the near-end talker is in a "quiet environment."
[0012] Thus, even if the output sound from the speaker (i.e., the sound to be evaluated) is the same when the near-end talker is in a quiet environment without ambient noise, the evaluation of the output sound will differ between when the near-end talker is in an environment with ambient noise and when the near-end talker is in a quiet environment without ambient noise. Until now, evaluation of such environments could only be performed through conversation tests. To the best of our knowledge, the present inventors have disclosed, for the first time, an acoustic quality evaluation technology that can obtain appropriate evaluation values through listening tests without conducting conversation tests (Patent Document 2). However, listening tests require human intervention, and although the cost of evaluation is lower than that of conversation tests, it is still high.
[0013] To achieve the above object, a sound quality evaluation device according to the disclosed technology includes a first sound quality evaluation unit, an analysis unit, and a second sound quality evaluation unit. The first sound quality evaluation device presents an analysis reference signal, an analysis noisy signal, and analysis ambient noise to an evaluator to obtain an analysis subjective rating value. The analysis unit calculates an analysis PESQ value from the analysis reference signal and the analysis noisy signal, and determines the relationship between the analysis subjective rating value and the analysis PESQ value. The second sound quality evaluation device calculates a test PESQ value from the test reference signal, the test noisy signal, and the test ambient noise, and calculates, from the test PESQ value according to the relationship, an estimated subjective rating value of the test noisy signal compared with the test reference signal under the test ambient noise.
[0014] According to the disclosed technology, even in a call environment with environmental noise, appropriate acoustic quality evaluation of a public address system can be achieved through objective testing with limited human intervention, thereby reducing the cost of acoustic quality evaluation testing of public address system.
[0015] 1 is a diagram schematically showing acoustic echo and AEC. FIG. 1 is a diagram explaining a sound quality evaluation test by a listening test in a public address communication system. FIG. 2 is a functional block diagram of a sound quality evaluation system relating to the listening test of Patent Document 2. FIG. 3 is a functional block diagram of a data generating device relating to the listening test of Patent Document 2. FIG. 4 is a flowchart showing the operation of a near-end system of the data generating device. FIG. 5 is a flowchart showing the operation of a far-end system of the data generating device. FIG. 6 is a flowchart showing the operation of a data recording system of the data generating device. FIG. 7 is a functional block diagram of a sound quality evaluation device relating to the listening test of Patent Document 2. FIG. 8 is a flowchart showing the operation of a sound quality evaluation device. FIG. 9 is a diagram showing an example of a display by a display unit. FIG. 10 is a diagram explaining the relationship between a reference signal pair, a noisy signal pair, and DMOS, and between a reference signal, a noisy signal, and PESQ. FIG. 11 is a functional block diagram of a sound quality evaluation system relating to the first embodiment. FIG. 12 is a flowchart showing the operation of a correlation analysis device. FIG. 13 is a diagram explaining linear regression analysis of data consisting of an analysis PESQ and an analysis DMOS. FIG. 14 is a functional block diagram of a data generating device relating to the first embodiment. FIG. 15 is a functional block diagram of a sound quality objective evaluation device, and a flowchart showing its operation. FIG. 16 is a diagram showing an example of the functional configuration of a computer.
[0016] Hereinafter, an embodiment of the disclosed technology will be described in detail. Note that components having the same function are assigned the same numbers, and duplicate explanations will be omitted. The objective evaluation test related to the disclosed technology uses the subjective evaluation results from the listening test of Patent Document 2 and the objective evaluation results from PESQ. Therefore, the listening test of Patent Document 2 will be described first, and then the objective evaluation test related to the disclosed technology will be described.
[0017] [Outline of Listening Test] First, an acoustic quality evaluation test using a listening test in a loudspeaker communication system will be conceptually explained using Fig. 2. In this acoustic quality evaluation test, a near-end speaker 201 and a far-end speaker 202 converse through a loudspeaker communication system 2, and an evaluator 203 located on the near-end speaker 201 side evaluates the quality of the loudspeaker communication system 2. A loudspeaker communication system is a communication system that transmits and receives acoustic signals between terminal devices equipped with a microphone and a speaker, and in which at least a portion of a sound output from the speaker of the terminal device (e.g., "Hello" 210) is received by the microphone of the terminal device (e.g., a sound loop 211 occurs). Examples of loudspeaker communication systems are audio conference systems and video conference systems.
[0018] In the loudspeaker communication system 2, the speech 209 of the near-end talker is received by the microphone 204 on the near-end talker side, an acoustic signal obtained based on the speech 209 is transmitted to the far-end talker side via the network 208, and the sound represented by the acoustic signal is output from the speaker 206 on the far-end talker side. Also, the speech of the far-end talker is received by the microphone 207 on the far-end talker side, an acoustic signal obtained based on the speech 209 is transmitted to the near-end talker side via the network 208, and the sound represented by the acoustic signal is output from the speaker 205 on the near-end talker side. However, at least a part of the sound output from the speaker 206 on the far-end talker side is also received by the microphone 207 on the far-end talker side. In other words, the sound from the far-end talker received by the microphone 207 on the far-end talker side is the speech 212 "Hi" of the far-end talker superimposed with the wraparound speech 211 (acoustic echo) originating from the near-end talker. That is, the sound from the far-end talker side received by the far-end talker microphone 207 is a sound in which the sound 212 from the far-end talker is superimposed with the sound 210 from the near-end talker, which is degraded in the space on the far-end talker side, resulting in sound 211. When the near-end talker 201 is not speaking, the sound from the far-end talker is not degraded because the sound 211 from the near-end talker is not superimposed. Note that the degradation of the sound from the far-end talker side is also caused by the superimposition of ambient noise 213 on the far-end talker side.
[0019] The acoustic signal transmitted to the near-end talker may be derived from a processed signal obtained by performing predetermined processing on a signal based on sound received by the far-end talker's microphone, or may be derived without such signal processing. Examples of signal processing include at least one of echo cancellation and noise cancellation. Note that echo cancellation refers to processing by an echo canceller in a broad sense for reducing echoes. The broad definition of echo cancellation refers to all processing for reducing echoes. The broad definition of echo cancellation may be achieved, for example, solely by a narrowly defined echo canceller using an adaptive filter, by a voice switch, by echo reduction, by a combination of at least some of these technologies, or by a combination with other technologies (see Reference 1 below). The noise cancellation refers to processing that suppresses or removes noise components generated around the microphone of the far-end terminal due to all environmental noise other than the voice of the far-end talker (see Reference 2 below).
[0020] [Reference 1] Knowledge Base Knowledge Forest, Group 2, Part 6, Chapter 5, "Acoustic Echo Canceller", Institute of Electronics, Information and Communication Engineers [Reference 2] Sumio Sakauchi, Yoichi Haneda, Masashi Tanaka, Junko Sasaki, Akitoshi Kataoka, "Acoustic Echo Canceller with Noise Suppression and Echo Suppression Functions", Transactions of the Institute of Electronics, Information and Communication Engineers, Vol. J87-A, No. 4, pp. 448-457, April 2004
[0021] The disclosed technology provides an apparatus and method for a hearing test, particularly in a situation where there is environmental noise on the near-end talker side.The technology disclosed in Patent Document 1 is different in that it is a hearing test that targets a quiet environment with no noise around the near-end talker.
[0022] 3 shows a functional block diagram of an example of an acoustic evaluation system related to the listening test of Patent Document 2. The acoustic evaluation system 3 includes a data generation device 31 for the test and an acoustic quality evaluation device 32.
[0023] [Data Generating Device] Fig. 4 shows a functional block diagram of a data generating device 4 related to the listening test of Patent Document 2. The data generating device 4 comprises a near-end system 41 that simulates the near-end talker environment, a far-end system 42 that simulates the far-end talker environment, and a data recording system 43 that records simulated communications between the simulated environments. The near-end system 41 and the far-end system 42 communicate via a network 44. The simulated communications recorded by the data recording system 43 will be used in a later acoustic quality evaluation.
[0024] <Near-end system> The near-end system includes a near-end ambient noise signal storage unit 410, a near-end talker voice signal storage unit 411, reproduction units 412 and 413, a near-end terminal unit 414, and a signal processing unit 415. <Far-end system> The far-end system includes a far-end ambient noise signal storage unit 420, a far-end talker voice signal storage unit 421, reproduction units 422 and 423, speakers 424, 425, and 426, a microphone 427, a far-end terminal unit 428, and a signal processing unit 429. <Data recorder> The data recording system includes a recording processing unit 430, a time adjustment processing unit 431, a data storage unit 432, data output units 433, 434, 435, 436, 437, and 438, and a switch 439.
[0025] 5, 6, and 7 are flowcharts illustrating an example of the operation of the data generating device 4. <Operation of the Near-End System> The operation of the near-end system will be described mainly with reference to FIGS. 4 and 5. The data generating device 4 retrieves a voice signal from the near-end talker voice signal storage unit 411, reproduces it using the reproduction unit 413 (step S501), and inputs it to the near-end terminal unit 414. This input corresponds to the voice uttered by the near-end talker. At the same time, the reproduced signal is output to the output units 433, 435, and 437 of the data recording system 43 (step S504). This output (the voice uttered by the near-end talker) becomes the reference sound (described below) in the stereo listening test (described below). The data generating device 4 also retrieves a noise signal from the near-end ambient noise signal storage unit 410, reproduces it using the reproduction unit 412 (step S502), and inputs it to the near-end terminal unit 414. This input corresponds to the ambient noise of the near-end talker.
[0026] The signal processing unit 415 performs signal processing (echo cancellation and noise cancellation) on the voice and noise input to the near-end terminal unit 414 (step S503), and transmits the processed voice to the far-end terminal unit 428 via the network 44 (step S505). In parallel, the data generating device 4 outputs the far-end voice received by the near-end terminal unit to the recording processing unit 430 of the data recording system (step S507).
[0027] <Operation of the Far-End System> The operation of the far-end system will be described mainly with reference to Figures 4 and 6. The far-end terminal unit 428 outputs audio based on a signal received from the near-end terminal from the speaker 426 (step S602). This output corresponds to the audio from the near-end speaker emitted from the loudspeaker communication system and heard by the far-end speaker. The data generator 4 retrieves an audio signal from the far-end speaker audio signal storage unit 421, plays it back using the playback unit 423 (step S603), and outputs it from the speaker 425 (step S604). This output corresponds to the audio emitted by the far-end speaker. In parallel, the playback unit 423 outputs the reproduced audio to the time adjustment processing unit 431 of the data recording system (step S611). This output serves as a reference signal for later quality evaluation. The data generator 4 also retrieves a noise signal from the far-end ambient noise signal storage unit 420, plays it back using the playback unit 422 (step S605), and outputs it from the speaker 424 (step S606). This output corresponds to the ambient noise of the far-end talker.
[0028] The far-end terminal unit 428 acquires the outputs of the speakers 426, 425, and 424 with the microphone 427 (step S607) and inputs the signals to the signal processing unit 429. The signal processing unit 429 performs signal processing (echo cancellation and noise cancellation) on the input voice signal (a superposition of the near-end talker voice signal, the far-end talker voice signal, and environmental noise) as necessary (step S608), and transmits the signal to the near-end terminal unit 414 via the network 44 (step S610). At the same time, the signal processing unit transmits a signal processing presence / absence signal indicating whether signal processing has been performed to the recording processing unit 430 of the data recording system (step S609). The signal processing presence / absence signal is used when recording the sound to be evaluated (described below).
[0029] <Operation of Data Recording System> The operation of the data recording system will be described mainly with reference to Figures 4 and 7. Data recording system 43 records the audio output of near-end system 41 and the audio output of far-end system 42 in data storage unit 432 for use in later sound quality evaluation. In sound quality evaluation, the voice of the far-end talker before echo or noise is superimposed (reference signal), the voice of the far-end talker with echo or noise superimposed but not signal processed (degraded signal 1), and the voice of the far-end talker with echo or noise superimposed but signal processed (degraded signal 2) are compared.
[0030] <<Stereo Listening Test>> In addition, to allow the evaluator to perceive acoustic echo during the listening test, the test sounds are stereo. The far-end sound (sound to be evaluated), which contains acoustic echo, is presented to one ear, and the near-end sound (reference sound), which is the source of the acoustic echo, is presented to the other ear simultaneously. This simulates the situation where the evaluator is sitting next to the near-end speaker, who is the source of the acoustic echo, and listening to their conversation, which is equivalent to a pseudo-speech test. The reference sound and the sound to be evaluated can be supplied to either ear, but it is desirable to supply the reference sound to the non-dominant ear (e.g., the right ear) and the sound to be evaluated to the dominant ear (e.g., the left ear).
[0031] <<Reference Sound>> In order to record the above evaluation sounds, the data generating device 4 outputs the output of the reproduction unit 413 of the near-end terminal unit as the reference sound of the standard signal, the reference sound of the noisy signal 1, and the reference sound of the noisy signal 2 to the output units 433, 435, and 437 (step S701), and records them in the data storage unit 432 (step S707).
[0032] <<Evaluation Target Sound-Reference Signal>> Furthermore, the data generating device 4 applies a delay equivalent to the delay caused by the network to the output of the reproduction unit 423 using the time adjustment processing unit 431 (step S702), outputs the output to the output unit 438 (step S703), and records it in the data storage unit 432 (step S707). A pair of the signal obtained from the output unit 437 and the signal obtained from the output unit 438 is hereinafter referred to as a reference signal pair.
[0033] <<Sound to be Evaluated - Noise Signal 1>> The data generator 4 records the signal received by the near-end terminal unit 414 in conjunction with the processing of the far-end system. That is, if the signal processing unit 429 of the far-end terminal unit does not perform signal processing such as echo cancellation or noise cancellation, the signal processing unit 429 outputs a "signal processing OFF signal" to the recording processing unit 430 (step S609). The recording processing unit 430 controls the switch 439 in accordance with the signal processing OFF signal (step S704) and outputs the sound received from the far end to the output unit 434 (step S705). The sound signal output from the output unit 434 is recorded in the data storage unit 432 (step S707). The pair of the sound signal obtained from the output unit 433 and the sound signal obtained from the output unit 434 is hereinafter referred to as "noise signal pair 1."
[0034] <<Sound to be Evaluated--Degraded Signal 2>> When the signal processing unit 429 of the far-end terminal unit performs signal processing such as echo cancellation or noise cancellation, the signal processing unit 429 outputs a "signal processing ON signal" to the recording processing unit 430 (step S609). The recording processing unit 430 controls the switch 439 in accordance with the signal processing ON signal (step S704) and outputs the sound received from the far end to the output unit 436 (step S706). The sound signal output from the output unit 436 is recorded in the data storage unit 432 (step S707). The pair of the sound signal obtained from the output unit 435 and the sound signal obtained from the output unit 436 will be referred to as degraded signal pair 2 hereinafter.
[0035] [Sound quality evaluation device] The evaluator uses a binaural sound reproduction device such as headphones or earphones to alternately listen to the sound that should be output from the speaker on the near-end speaker when there is no sound leakage on the far-end speaker side (i.e., the reference signal) and the sound that should be output from the speaker on the near-end speaker when there is sound leakage on the far-end speaker side (i.e., the signal to be evaluated), and makes a subjective evaluation (opinion evaluation) of the call quality.
[0036] Furthermore, the test sounds are presented to the evaluator in the above-described stereo configuration, but in this embodiment, the channel of the reference sound is denoted as "Rch" and the channel of the sound to be evaluated is denoted as "Lch."
[0037] FIG. 8 shows a functional block diagram of an acoustic quality evaluation device 8 for the listening test of Patent Document 2. The acoustic quality evaluation device 8 can simultaneously conduct tests for multiple (N) evaluators. Therefore, while N acoustic processing output units, display units, input units, and acoustic playback devices are illustrated as "XXX-1...XXX-N," hereafter, the notation "XXX" without a hyphen will refer collectively to the N units. For example, "evaluator 850" will refer collectively to evaluators 850-1 to 850-N.
[0038] The sound quality evaluation device 8 acquires test sounds from the data storage unit 432 and environmental noise from the near-end environmental noise signal storage unit 410 and supplies them to the evaluator 850. The evaluator 850 wears open-type binaural sound reproduction devices such as headphones or earphones (hereinafter referred to as open-type headphones) 840. Note that "open-type" refers to a structure that provides low sound insulation to prevent the reproduced sound from leaking to the outside, and therefore allows surrounding sounds to easily reach the ears of the user of the reproduction device. A speaker 860 is provided to supply environmental noise to the evaluator 850. Note that, as described below, a configuration may be adopted in which environmental noise is presented from a speaker shared by multiple evaluators, and the number of speakers does not necessarily have to match the number of evaluators. The evaluator 850 wears the open-type headphones 840 and listens to the reference sound and the sound to be evaluated in stereo. As described above, the open headphones 840 do not block out environmental noise, so the evaluator 850 listens to the reference sound and the sound to be evaluated in a noisy environment.
[0039] 9 is a flowchart illustrating an example of the operation of the sound quality evaluation device 8. The sound quality evaluation device 8 acquires a signal from the near-end environmental noise signal storage unit 410, plays it in the playback unit 806, and outputs it from the speaker 860 (step S901). The playback control unit 801 determines an evaluation execution signal from among the signals recorded in the data storage unit (step S902). The display control unit 802 displays an evaluation input screen on the display unit 820 for evaluating the signal determined by the playback control unit (step S903). The evaluation input screen may be, for example, as shown in FIG. 10.
[0040] The playback control unit 801 acquires the reference signal pair from the data storage unit 432 and outputs it from the acoustic output processing unit 810 to the open-type headphones 840 (step S904). The playback control unit 801 then acquires the noisy signal pair 1 or 2 from the data storage unit 432 and outputs it from the acoustic output processing unit 810 to the open-type headphones 840 (step S905). The evaluator 850 inputs an evaluation using the display unit 820 and input unit 830 (step S906). The sound quality evaluation device 8 determines whether all evaluations have been completed, and if not, executes the next evaluation (No in step S907). If all evaluations have been completed, the evaluation procedure ends (Yes in step S907). The counting unit 803 counts the evaluation results and records them in the counting result storage unit 805.
[0041] The above is the explanation of the listening test in Patent Document 2.
[0042] [First embodiment] <Correlation between PESQ and listening test> The evaluation result from the listening test of Patent Document 2 is called DMOS (Degradation MOS: subjective evaluation of the degree of degradation). In the listening test of Patent Document 2, DMOS was obtained by listening to and comparing a reference signal pair (a pair of a near-end talker's voice signal and a reference signal) with degraded signal pair 1 (a pair of a near-end talker's voice signal and degraded signal 1) or degraded signal pair 2 (a pair of a near-end talker's voice signal and degraded signal 2).
[0043] The degree of degradation of degraded signal 1 or degraded signal 2 relative to the reference signal can be objectively evaluated using PESQ. Figure 11 shows the relationship between the reference signal pair, the degraded signal pair, and DMOS, as well as the relationship between the reference signal, the degraded signal, and PESQ. Figure 11 suggests a certain correlation between DMOS and PESQ. Understanding this correlation allows us to estimate DMOS from PESQ, enabling objective evaluation of the acoustic quality of hands-free public address systems. Hereinafter, "degraded signal" refers to "degraded signal 1" or "degraded signal 2," and "degraded signal pair" refers to "degraded signal pair 1" or "degraded signal pair 2."
[0044] 12 shows a functional block diagram of an example of a sound quality evaluation system 12 according to the first embodiment. The sound quality evaluation system 12 includes a correlation analysis device 121, a test data generation device 122, and a sound quality objective evaluation device 123. The correlation analysis device 121 includes the sound quality evaluation system 3 described in [listening test of Patent Document 2] and a correlation analysis unit 124.
[0045] <Correlation Analysis> Fig. 13 is a flowchart showing an example of the operation of the correlation analysis device 121. This will be explained using Fig. 12 and Fig. 13. The correlation analysis device 121 acquires a plurality of analysis DMOSs using a plurality of pairs of "analysis reference signal pairs and analysis noisy signal pairs" by the acoustic quality evaluation system 3 described in [listening test of Patent Document 2] (step S1301).
[0046] The correlation analysis unit 124 calculates the analysis PESQ for each "analysis reference signal pair and analysis impairment signal pair" used in the DMOS evaluation, using the analysis reference signal of the analysis reference signal pair and the analysis impairment signal of the analysis impairment signal pair. To this end, the following steps S1302 to S1304 are performed for all "analysis reference signal pair and analysis impairment signal pair."
[0047] The correlation analysis unit 124 calculates the SNR (signal-to-noise ratio) of the analytical degradation signal S1302 and the analytical near-end environmental noise signal S2303 used in the DMOS evaluation (step S1302). DEG is the SNR and the near-end environmental noise signal N ENV is used to make the following correction (step S1303). α is a function of SNR, and is determined as follows: if SNR>0 dB, α(SNR)=0; if SNR≦0 dB, α(SNR)=constant. DEG The analytical PESQ is calculated using (step S1304).
[0048] For each "reference signal pair for analysis and degraded signal pair for analysis," if a scatter plot is drawn with the PESQ for analysis on the x-coordinate and the DMOS for analysis on the y-coordinate, it will look like Figure 14. This data is subjected to linear regression analysis to find a and b in y = ax + b (step S1306).
[0049] <Data Generation> Test data for objective evaluation is generated. Fig. 15 is a functional block diagram of a data generating device 122 according to the first embodiment. The data generating device 4 described in [Listening Test of Patent Document 2] stores a reference signal pair and a noisy signal pair in the data storage unit 432, but the data generating device 15 according to the first embodiment only needs to store a reference signal and a noisy signal.
[0050] That is, the data generator 122 records the far-end talker voice signal unaffected by the loudspeaker communication (the signal obtained by applying a delay equivalent to the network delay to the output of the reproducing unit 423 by the time adjustment processing unit 431 and outputting it to the output unit 438) as a reference signal in the data storage unit 432. The data generator 122 also records the signal obtained by superimposing the near-end talker voice signal and near-end environmental noise on the far-end talker voice signal and far-end environmental noise in the far-end system 42 and not undergoing signal processing in the far-end system terminal unit 428 (the signal output to the output unit 434) as a noisy signal 1 in the data storage unit 432. The data generator 122 also records the signal obtained by superimposing the near-end talker voice signal and near-end environmental noise on the far-end talker voice signal and far-end environmental noise in the far-end system 42 and undergoing signal processing in the far-end system terminal unit 428 (the signal output to the output unit 436) as a noisy signal 2 in the data storage unit 432. Hereinafter, the reference signal generated for the objective evaluation test will be referred to as a test reference signal, and the degradation signal generated for the objective evaluation test will be referred to as a test degradation signal.
[0051] <Sound quality objective assessment device> Fig. 16(a) is a functional block diagram of the sound quality objective assessment device 123. The sound quality objective assessment device 123 includes a PESQ calculation unit 161 and a DMOS estimation unit 162. Fig. 16(b) is a flowchart showing the operation of the sound quality objective assessment device 123. This will be explained using Fig. 16.
[0052] The PESQ calculation unit 161 calculates the SNR of the test noisy signal and the test near-end ambient noise signal (step S1601). Next, the PESQ calculation unit 161 calculates the SNR of the test noisy signal S' used for calculating the PESQ. DEG is the SNR and the near-end environmental noise signal N' ENV is used to make correction using the following formula (step S1602). The PESQ calculation unit 161 then calculates S' DEG The PESQ is calculated using the test reference signal (step S1603). The DMOS estimation unit estimates the DMOS value y using the PESQ value x as y = ax + b obtained by the above correlation analysis (step S1605).
[0053] The above is the description of the first embodiment.
[0054] [First Modification of the First Embodiment] In the first embodiment described above, the value of α was changed according to the SNR value (magnitude of the environmental noise signal). However, α may also be changed taking into account the location and characteristics (frequency characteristics, diffuseness, periodicity, etc.) of the environmental noise. For example, human hearing is most sensitive in the horizontal plane when directly in front, and accuracy decreases as the distance from the front increases, reaching its lowest level directly to the side. In the median plane, sensitivity is highest in front and lowest at the top of the head. Therefore, α may be defined as a function of the angle between the sound to be evaluated and the source of the environmental noise. Furthermore, for example, human hearing is highly sensitive to sounds around 3000 to 4000 Hz, and sensitivity decreases as the frequency decreases or increases. Therefore, α may be defined as a function of the frequency of the environmental noise.
[0055] [Second Modification of First Embodiment] According to the findings of the inventors, when the SNR is greater than 0 dB, a falls within the range of approximately 1.3 and b falls within the range of approximately −0.3. Therefore, in an environment where the SNR is greater than 0 dB, the DMOS can be estimated by y = 1.3x −0.3.
[0056] [Supplementary Information] In the above embodiment, an example was described in which PESQ was used for objective evaluation. However, to estimate subjective evaluation values through listening tests, POLQA (Perceptual Objective Listening Quality: Reference 3) may be used instead of PESQ, or another method similar to PESQ or POLQA may be used. PESQ, POLQA, and methods similar to these are referred to as "objective perceptual evaluation," and the result of the objective perceptual evaluation is referred to as "objective perceptual evaluation value." [Reference 3] ITU-T Recommendation P.863, "Perceptual Objective Listening Quality," ITU, 2018
[0057] [Program, Recording Medium] The various processes described above can be implemented by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in FIG. 17 and operating the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc.
[0058] The program describing the processing contents can be recorded on a computer-readable recording medium, which may be, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, or any other suitable recording medium.
[0059] The program may be distributed by, for example, selling, transferring, lending, etc. portable recording media such as DVDs and CD-ROMs on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to other computers via a network, thereby distributing the program.
[0060] A computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes the process in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program. Furthermore, the computer may execute the process in accordance with each program transferred from the server computer. Alternatively, the server computer may not transfer the program to the computer, but may instead execute the process through a so-called ASP (Application Service Provider) service, which realizes the processing function by issuing an execution instruction and obtaining the results. In this embodiment, the program includes information used for processing by a computer that is equivalent to a program (e.g., data that is not a direct instruction to the computer but has properties that define computer processing).
[0061] Furthermore, in this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
Claims
1. A device for estimating a subjective assessment value of the acoustic quality of a public address communication system including a first terminal device and a second terminal device, comprising: a first acoustic quality assessment unit that presents an analytical reference signal, an analytical degraded signal, and analytical environmental noise to an assessor to obtain an analytical subjective assessment value; an analysis unit that determines an analytical objective perceptual assessment value from the analytical reference signal and the analytical degraded signal, and determines the relationship between the analytical subjective assessment value and the analytical objective perceptual assessment value; and a second acoustic quality assessment unit that determines an objective perceptual assessment test value from the test reference signal, the test degraded signal, and the test environmental noise, and calculates, from the objective perceptual assessment test value according to the relationship, an estimated subjective assessment value of the test degraded signal for the test reference signal compared under the test environmental noise.
2. An acoustic quality evaluation device according to claim 1, wherein the degradation signal for analysis is received by the first terminal and from the first terminal by the second terminal, the degradation signal for analysis is presented to the evaluator using an open-type acoustic device, and the environmental noise for analysis is presented to the evaluator so as to envelop the open-type acoustic device.
3. An acoustic quality evaluation device according to claim 1, wherein the analytical objective perceptual evaluation value is determined by correcting the analytical noisy signal using the SNR between the analytical noisy signal and the analytical environmental noise, and the objective perceptual evaluation test value is determined by correcting the test noisy signal using the SNR between the test noisy signal and the test environmental noise.
4. A sound quality evaluation method according to claim 1, wherein the objective perceptual evaluation value for analysis is determined by correcting the degraded signal for analysis using characteristics of the environmental noise for analysis (location of occurrence, frequency characteristics, diffuseness, or periodicity), and the objective perceptual evaluation test value is determined by correcting the degraded signal for analysis using the characteristics of the environmental noise for test that are used to correct the degraded signal for analysis.
5. A method for estimating a subjective assessment value of acoustic quality of a public address communication system including a first terminal device and a second terminal device, comprising: a step in which a first acoustic quality evaluation device presents an analytical reference signal, an analytical degraded signal, and analytical ambient noise to an evaluator to obtain an analytical subjective assessment value; a step in which an analysis device obtains an analytical objective perceptual assessment value from the analytical reference signal and the analytical degraded signal, and determines a relationship between the analytical subjective assessment value and the analytical objective perceptual assessment value; and a step in which a second acoustic quality evaluation device obtains an objective perceptual assessment test value from the test reference signal, the test degraded signal, and the test ambient noise, and calculates an estimated subjective assessment value of the test degraded signal for the test reference signal compared under the test ambient noise from the objective perceptual assessment test value in accordance with the relationship, wherein the analytical degraded signal is received by the first terminal and received by the second terminal from the first terminal, The acoustic quality evaluation method, wherein the degraded signal for analysis is presented to the evaluator using an open-type acoustic device, and the environmental noise for analysis is presented to the evaluator so as to surround the open-type acoustic device.
6. A sound quality evaluation method according to claim 5, wherein the analytical objective perceptual evaluation value is determined by correcting the analytical noisy signal using the SNR between the analytical noisy signal and the analytical environmental noise, and the objective perceptual evaluation test value is determined by correcting the test noisy signal using the SNR between the test noisy signal and the test environmental noise.
7. A sound quality evaluation method according to claim 1, wherein the objective perceptual evaluation value for analysis is obtained by correcting the degraded signal for analysis using the characteristics of the environmental noise for analysis (location of occurrence, frequency characteristics, diffuseness, or periodicity), and the objective perceptual evaluation test value is obtained by correcting the degraded signal for analysis using the characteristics of the environmental noise for test that are used to correct the degraded signal for analysis.
8. A program for causing a computer to function as the sound quality evaluation device according to any one of claims 1 to 4.
Citation Information
Patent Citations
Portable telephone testing device
JP2000151785A
Objective evaluation server, method and program of speech quality
JP2006345149A
Acoustic quality evaluation device, acoustic quality evaluation method, and program
WO2024121962A1