Voice quality evaluation apparatus, voice quality evaluation method, and voice quality evaluation program

The audio quality evaluation apparatus addresses the limitations of existing methods by converting and comparing text-based reference and evaluation audio to assess voice quality accurately and efficiently, focusing on recipient comprehension and network changes.

JP7706086B2Active Publication Date: 2025-07-11GREEN CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021088897
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-27
Publication Date
2025-07-11
Estimated Expiration
2041-05-27

AI Technical Summary

Technical Problem

Existing voice quality evaluation methods like POLQA are inadequate for assessing long-term voice transmission quality in video conferences and distance learning, as they are data-intensive and focus on voice differences rather than recipient comprehension, failing to accurately evaluate content accuracy and handling network-induced changes.

Method used

An audio quality evaluation apparatus and method that converts reference text to audio, transmits it over a network, acquires and recognizes evaluation target audio, and compares the recognized text with the reference text to assess quality, allowing for efficient evaluation of long-term voice data.

Benefits of technology

Enables accurate evaluation of voice quality based on recipient comprehension, reduces data processing costs, and effectively assesses temporal changes in network conditions, providing a comprehensive quality score.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007706086000001
    Figure 0007706086000001
  • Figure 0007706086000002
    Figure 0007706086000002
  • Figure 0007706086000003
    Figure 0007706086000003
Patent Text Reader

Abstract

To provide a voice quality evaluation device, a voice quality evaluation method, and a voice quality evaluation program, which simplify evaluation of transmission quality in voice data transmitted via a network.SOLUTION: A voice quality evaluation device 1 includes: a voice generation unit 12 that converts reference text into reference voice data; a reference voice transmission unit 13 that transmits the reference voice data generated by the voice generation unit to a playback device 50 via a network NW; an evaluation target voice acquisition unit 14 that acquires evaluation target voice data that is played back from the playback device; a voice recognition unit 15 that generates evaluation target text by voice-recognizing words included in the evaluation target voice data; and an evaluation unit 16 that evaluates quality of the voice received via the network based on the evaluation target text.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for evaluating voice quality.

Background Art

[0002] In recent years, opportunities to conduct video conferences, distance education, etc. via communication networks have been increasing. Therefore, a technique that can easily evaluate the quality of this voice is required.

[0003] For example, in Patent Document 1, a voice evaluation method is proposed in which the initial voice recognition result of a sample after error correction and filter processing of the input voice is compared with the original text after filter processing to calculate a voice evaluation score. Also, in Patent Document 2, a technique is proposed in which input voice recognition data is recorded in a storage unit, and the voice recognition data and dictionary data are matched by a search unit to create a voice recognition result.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Disclosure of the Invention

Problems to be Solved by the Invention

[0005] For the voice quality evaluation of telephone systems, a subjective quality evaluation method called MOS (Mean Opinion Score) has been used for a long time. Also, an objective quality method POLQA (Perceptual Objective Listening Quality Assessment) for estimating the result of subjective evaluation by MOS using a computer is known. POLQA is a method of comparing the original voice called the reference voice with the voice recorded on the receiving side and calculating the MOS value.

[0006] Here, in the evaluation of voice quality in video conferences, distance learning, etc., it is crucial whether the recipient can accurately hear the speaker's utterance content over a period of several tens of minutes to several hours. However, POLQA is not suitable for evaluating the quality of voice transmitted over a long period. POLQA has standards for reference voice and reference video, and evaluations of normal conversations and voices that do not meet the standards are not targeted. Also, since POLQA is a method for comparing voice data with each other, when trying to evaluate long voice data, the data volume becomes large and the processing cost is enormous. Therefore, for example, it is difficult to evaluate temporal changes such as "how packetized voice and video are affected by changes in the network state."

[0007] Also, since POLQA is a method for evaluating the difference between the original voice and the recipient's voice, it cannot be said that it appropriately evaluates the accuracy of the utterance content that can be heard by the recipient. For example, even when POLQA is applied to a case where voice quality improvement processing such as equalizing processing is performed, a low evaluation is given as a result of determining that there are differences in the voice. Similarly, in situations where the recorded data contains ambient noise or reverberation in video conferences, distance learning, etc., even when the noise is removed and the data is played back, it is determined that there are differences in the voice and the evaluation becomes low.

[0008] Therefore, an object of the present invention is to simply perform quality evaluation of transmission in voice data transmitted via a network.

Means for Solving the Problems

[0009] To achieve the above object, an audio quality evaluation apparatus according to one aspect of the present invention includes: an audio generation unit that converts a reference text into reference audio data; a reference audio transmission unit that transmits the reference audio data generated by the audio generation unit to a playback device via a network; an evaluation target audio acquisition unit that acquires evaluation target audio data reproduced from the playback device; an audio recognition unit that generates an evaluation target text by performing audio recognition on words included in the evaluation target audio data; and an evaluation unit that performs quality evaluation of the audio received via the network based on the evaluation target text.

[0010] The evaluation unit may perform the quality evaluation by comparing the evaluation target text with the reference text.

[0011] The audio recognition unit may generate a second reference text by performing audio recognition on words included in the reference audio data, and the evaluation unit may perform the quality evaluation by comparing the evaluation target text with the second reference text.

[0012] The audio recognition unit may generate a second reference text by performing audio recognition on words included in the reference audio data, the evaluation unit may perform a first evaluation by comparing the reference text with the second reference text, perform a second evaluation by comparing the reference text with the evaluation target text, and perform the quality evaluation based on the results of the first evaluation and the second evaluation.

[0013] To achieve the above object, an audio quality evaluation method according to another aspect of the present invention includes: an audio generation process of converting a reference text into reference audio data; a reference audio transmission process of transmitting the reference audio data generated by the audio generation process to a playback device via a network; an evaluation target audio acquisition process of acquiring evaluation target audio data reproduced from the playback device; an audio recognition process of generating an evaluation target text by performing audio recognition on words included in the evaluation target audio data; and an evaluation process of performing quality evaluation of the audio received via the network based on the evaluation target text.

[0014] To achieve the above object, a voice quality evaluation program according to still another aspect of the present invention causes a computer to execute a voice generation instruction for converting a reference text into reference voice data, a reference voice transmission instruction for transmitting the reference voice data generated by the voice generation instruction to a playback device via a network, an evaluation target voice acquisition instruction for acquiring evaluation target voice data reproduced from the playback device, a voice recognition instruction for performing voice recognition on words included in the evaluation target voice data to generate an evaluation target text, and an evaluation instruction for performing a quality evaluation of the voice received via the network based on the evaluation target text. Note that the computer program can be provided by being stored in various data-readable recording media or can be provided so as to be downloadable via a network such as the Internet.

Effect of the Invention

[0015] According to the voice quality evaluation apparatus according to the present invention, in voice data transmitted via a network, it is possible to easily perform a quality evaluation of the transmission.

Brief Description of the Drawings

[0016]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiment for Carrying Out the Invention

[0017] Hereinafter, a voice quality evaluation device, a voice quality evaluation method, and a voice quality evaluation program according to embodiments of the present invention will be described with reference to the drawings.

[0018] <First Embodiment> ● Configuration of the voice quality evaluation device As shown in FIG. 1, the voice quality evaluation device 1 is a device that evaluates the quality of voice transmitted and received via the network NW. The network NW may be, for example, the Internet, or an appropriate communication line connected by wire or wirelessly, and the format is arbitrary.

[0019] For example, the voice quality evaluation device 1 is connected to the playback device 50 via the network NW. The playback device 50 is, for example, a terminal such as a personal computer, a smartphone, or a tablet, and is a terminal that a viewer of a video call is viewing.

[0020] The voice quality evaluation device 1 transmits the voice data of the reference voice (hereinafter, also referred to as "reference voice data") to the playback device 50 via the network NW. The playback device 50 converts this reference voice data into voice and plays it back. Note that the reference voice data may be received by the playback device 50 via an appropriate device on the network NW, such as a server. The voice quality evaluation device 1 acquires the voice data transmitted via the network NW and played back by the playback device 50, and evaluates the quality of this voice.

[0021] The voice quality evaluation device 1 is composed of a storage medium such as a memory, a processor, a communication module, an input / output interface, etc. By the processor executing a computer program recorded in the storage medium, the functional blocks shown in FIG. 1 are realized. The storage medium is a computer-readable recording medium and may include storage devices such as RAM (random access memory), ROM (read only memory), disk drive, SSD (solid state drive), flash memory. Here, non-temporary storage devices such as ROM, disk drive, SSD, and flash memory may be included in the voice quality evaluation device 1 as another storage device distinct from the memory.

[0022] The voice quality evaluation device 1 mainly includes, for example, a reference text acquisition unit 11, a voice generation unit 12, a reference voice transmission unit 13, an evaluation target voice acquisition unit 14, a voice recognition unit 15, and an evaluation unit 16 with the above-described hardware configuration. Note that part or all of the configuration of the voice quality evaluation device 1 may be realized by another hardware configuration, or part or all of it may be realized by a cloud computer. Also, part of the functions of the voice quality evaluation device 1 may be configured inside the playback device 50. In this case, for example, the evaluation target voice acquisition unit 14, the voice recognition unit 15, and the evaluation unit 16 may be configured in the playback device 50.

[0023] The reference text acquisition unit 11 is a functional unit that acquires a reference text used for voice quality evaluation. The reference text is, for example, text data as shown in FIG. 2(a) and may be an appropriate language not limited to Japanese. The reference text acquisition unit 11 may acquire the reference text via an appropriate network, or may receive an input via an appropriate input means of the voice quality evaluation device 1.

[0024] As shown in the conceptual diagram of FIG. 2(b), the voice generation unit 12 is a functional unit that converts the reference text acquired by the reference text acquisition unit 11 into reference voice data. The voice generation unit 12 may read out the reference text with artificial voice and convert it into reference voice data. Also, the voice generation unit 12 may be configured to display the reference text on a display unit such as a display and prompt a speaker who makes accurate utterances such as an announcer to read it out. In this case, the voice generation unit 12 has a configuration to collect the voice read out by the speaker.

[0025] The reference voice transmission unit 13 is a functional unit that transmits the reference voice data generated by the voice generation unit 12 to the playback device 50 via the network NW. The playback device 50 plays back the received reference voice data as voice. At this time, the playback device 50 may perform appropriate signal processing on the voice data and then play it back. This signal processing may be, for example, acoustic signal processing such as equalizing processing, noise canceling processing, frequency filtering processing, or amplification processing to make the words contained in the voice more clearly audible, or it may be processing to complement information lost due to transmission by the network NW.

[0026] Note that the reference voice data may be received by the playback device 50 via an appropriate device on the network NW, such as a server. Also, the above-described signal processing may be performed on the reference voice data by the device.

[0027] The evaluation target voice acquisition unit 14 is a functional unit that acquires voice data (hereinafter also referred to as "evaluation target voice data") played back from the playback device 50. The evaluation target voice acquisition unit 14 is connected to the playback device 50 and acquires voice data. As shown in the conceptual diagram of FIG. 2(c), the voice data to be evaluated is partially different from the reference voice data. In the example of this figure, it shows that the amplitude of some data shown in region L is smaller. The voice data to be evaluated may be data that has deteriorated more than the reference voice data during the transmission process, or may be data whose utterance content has been processed to be easier to understand by the above-described appropriate signal processing performed in the playback device 50 or a device on the network NW.

[0028] The voice recognition unit 15 is a functional unit that performs voice recognition on the words included in the voice data to be evaluated and generates the text to be evaluated. In the example of FIG. 2(d), it shows that some of the text shown in region W is different from the reference text.

[0029] The evaluation unit 16 is a functional unit that performs a quality evaluation of the voice received via the network NW based on the text to be evaluated. In the present embodiment, the evaluation unit 16 compares the text to be evaluated with the reference text and performs a quality evaluation of the voice. The voice quality evaluation score is calculated, for example, by the following formula (1). Quality evaluation score = Number of sounds accurately recognized in the text to be evaluated / Number of sounds in the reference text × 100 ···(1)

[0030] The higher the value of the quality evaluation score, the better the voice quality. When all the sounds in the text to be evaluated are accurately recognized, the quality evaluation score is 100. In the example of FIG. 2, the number of sounds in the reference text is 93 characters, and the text to be evaluated can accurately recognize 89 of these characters. Therefore, the quality evaluation score is 96. The quality evaluation score is displayed on an appropriate display unit that the voice quality evaluation device 1 has or is connected to.

[0031] The evaluation unit 16 may calculate one quality evaluation score for the entire text of the text to be evaluated, or may divide the text to be processed into a plurality of parts, calculate the quality evaluation score for each part, and calculate a plurality of quality evaluation scores along the time axis for one reference text. According to the configuration of calculating a plurality of quality evaluation scores for one reference text, the time change of the transmission state of the network can be evaluated. In this case, the evaluation unit 16 may divide the text to be evaluated so that there is no overlap between them, or may divide them while partially overlapping on the time axis. Note that the speech recognition unit 15 may divide the speech data and then convert each into text data. In this case, the evaluation unit 16 evaluates each text data.

[0032] ● Processing flow Using FIG. 3, the processing flow for the voice quality evaluation apparatus 1 to evaluate the transmitted voice will be described. First, the reference text acquisition unit 11 acquires a reference text (step S11). Next, the voice generation unit 12 acquires reference voice data based on the reference text (step S12). Next, the reference voice transmission unit 13 transmits the reference voice to the playback device 50 (step S13).

[0033] Next, the voice to be evaluated acquisition unit 14 acquires the target voice data played back by the playback device 50 (step S14). Next, the speech recognition unit 15 performs speech recognition on the target voice data and generates text data (step S15). The evaluation unit 16 evaluates the coincidence rate between the reference text and the text to be evaluated, and calculates a quality evaluation score (step S16). Next, the quality evaluation score is displayed on an appropriate display unit (step S17).

[0034] According to such a voice quality evaluation apparatus according to the present invention, focusing on the viewpoint of whether the speech content of the speaker can be accurately heard on the receiving side, it is possible to evaluate the quality of transmission by the network. Further, according to the voice quality evaluation apparatus according to the present invention, in order to compare and evaluate text data with each other, the amount of data to be analyzed can be compressed compared to a configuration that compares voice data with each other. Therefore, it is possible to evaluate the quality of transmission in a long speech. Further, according to the configuration in which the voice to be evaluated is converted into text data for evaluation, even when a process for improving the voice quality is performed, it is possible to appropriately evaluate whether the speech content is accurately transmitted.

[0035] <Second Embodiment> The voice quality evaluation apparatus according to the second embodiment of the present invention will be described centering on the parts different from the first embodiment. In this embodiment, the voice recognition unit 15 recognizes the words included in the reference voice data to generate a second reference text, and the evaluation unit 16 compares the evaluation target text with the second reference text to perform voice quality evaluation. Note that the description of the same configuration as that of the first embodiment will be omitted as appropriate, and the same reference numerals are used.

[0036] As shown in FIG. 4, in the voice quality evaluation apparatus according to the second embodiment, after acquiring a reference text (step S11) and generating reference voice data (step S12), the voice recognition unit 15 performs voice recognition on the reference voice data to generate a second reference text (step S21). Further, the reference voice data is transmitted (step S13), and the evaluation target voice data is acquired via the playback device 50 (step S14), and voice recognition is performed (step S15). The order of step S21 and steps S13 to S15 is arbitrary, and they may be performed simultaneously.

[0037] Next, the evaluation unit 16 evaluates the coincidence rate between the second reference text and the evaluation target text, calculates a quality evaluation score (step S22), and displays this quality evaluation score on the display unit (step S23). The quality evaluation score in this case is represented by, for example, the following formula (2). Quality evaluation score = Number of sounds accurately recognized in the text to be evaluated / Number of sounds in the second reference text × 100 ···(2)

[0038] According to this configuration, when there is a misrecognition by the speech recognition unit 15, the same misrecognition appears in both the second reference text and the text to be evaluated, so the influence of the misrecognition on the quality evaluation score can be removed. That is, in this embodiment, the speech quality can be evaluated excluding the influence of the speech recognition unit 15.

[0039] <Third Embodiment> Regarding the speech quality evaluation apparatus according to the third embodiment of the present invention, the parts different from the second embodiment will be mainly described. In this embodiment, the speech recognition unit 15 speech - recognizes the words included in the reference speech data to generate a second reference text, and the evaluation unit 16 compares the reference text with the second reference text to perform a first evaluation, and compares the reference text with the text to be evaluated to perform a second evaluation. Then, based on the results of the first evaluation and the second evaluation, the speech quality evaluation is performed. Note that the description of the same configuration as in the first embodiment or the second embodiment will be omitted as appropriate, and the same reference numerals are used.

[0040] As shown in FIG. 5, in the speech quality evaluation apparatus according to the third embodiment, the reference text is acquired (step S11), and after generating the reference speech data (step S12), the speech recognition unit 15 performs speech recognition on the reference speech data to generate a second reference text (step S21). Next, a first evaluation for calculating the coincidence rate (hereinafter, also referred to as the "first coincidence rate") between the reference text and the second reference text is performed (step S31).

[0041] Also, the reference audio data is transmitted (step S13), the audio data to be evaluated is acquired via the playback device 50 (step S14), and speech recognition is performed (step S15). Subsequently, a second evaluation is performed to calculate the coincidence rate between the reference text and the text to be evaluated (hereinafter also referred to as the "second coincidence rate") (step S32). The order of step S21 and step S31, and step S13 to S15 and step S32 is arbitrary, and they may be performed simultaneously.

[0042] Subsequently, the evaluation unit 16 compares the first coincidence rate and the second coincidence rate to calculate a quality evaluation score (step S33), and this quality evaluation score is displayed on the display unit (step S34). For example, the quality evaluation score is represented by the following formula (3). Quality evaluation score = second coincidence rate / first coincidence rate × 100 ···(3) In addition to the quality evaluation score, the first coincidence rate and the second coincidence rate may be respectively displayed on the display unit.

[0043] According to this configuration, the accuracy of the speech recognition by the speech recognition unit 15 can be confirmed by the first coincidence rate and the second coincidence rate, and the evaluation of the audio quality before and after transmission can be confirmed by the quality evaluation score.

[0044] As described above, according to the audio quality evaluation device according to the present invention, in the audio data transmitted via the network, the quality evaluation of the transmission can be easily performed.

[0045] In this description, the evaluation of the audio quality before and after transmission by the network is described as an example. However, the audio quality evaluation device according to the present invention is not limited to the network and can be used for the evaluation of the entire mechanism for transmitting the speech content. Further, the audio quality evaluation device is not only used for the evaluation of the network, but also allows the speaker who is conducting a video conference or a distance learning to visually recognize it approximately in real time or afterwards, confirm the part that has not been accurately transmitted, or prompt the speaker to speak the part again, so as to contribute to accurate information sharing in a video conference or the like.

Explanation of Symbols

[0046] 1 Voice Quality Evaluation Device 11 Reference Text Acquisition Unit 12 Voice Generation Unit 13 Reference Voice Transmission Unit 14 Evaluation Target Voice Acquisition Unit 15 Voice Recognition Unit 16 Evaluation Unit NW Network

Claims

1. An audio generation unit that converts a reference text into reference audio data, A reference audio transmission unit that transmits the reference audio data generated by the audio generation unit to a playback device via a network, An evaluation target audio acquisition unit that acquires evaluation target audio data reproduced from the playback device, An audio recognition unit that performs audio recognition on words included in the evaluation target audio data to generate an evaluation target text, An evaluation unit that performs a quality evaluation of the audio received via the network based on the evaluation target text, comprising, The audio recognition unit performs audio recognition on words included in the reference audio data to generate a second reference text, The evaluation unit compares the evaluation target text with the second reference text to perform the quality evaluation, An audio quality evaluation device.

2. An audio generation unit that converts a reference text into reference audio data, A reference audio transmission unit that transmits the reference audio data generated by the audio generation unit to a playback device via a network, An evaluation target audio acquisition unit that acquires evaluation target audio data reproduced from the playback device, An audio recognition unit that performs audio recognition on words included in the evaluation target audio data to generate an evaluation target text, An evaluation unit that performs a quality evaluation of the audio received via the network based on the evaluation target text, comprising, The audio recognition unit performs audio recognition on words included in the reference audio data to generate a second reference text, The evaluation unit performs a first evaluation by comparing the reference text with the second reference text, performs a second evaluation by comparing the reference text with the evaluation target text, and performs the quality evaluation based on the results of the first evaluation and the second evaluation, An audio quality evaluation device.

3. An audio generation process that converts a reference text into reference audio data, A reference audio transmission process that transmits the reference audio data generated by the audio generation process to a playback device via a network, An evaluation target audio acquisition process that acquires evaluation target audio data reproduced from the playback device, An audio recognition process that performs audio recognition on words included in the evaluation target audio data to generate an evaluation target text, An evaluation process that performs a quality evaluation of the audio received via the network based on the evaluation target text, including, The audio recognition process performs audio recognition on words included in the reference audio data to generate a second reference text, The evaluation process performs the quality evaluation by comparing the text to be evaluated and the second reference text. Voice quality evaluation method. **Claim 4**: A voice generation process for converting a reference text into reference voice data, A reference voice transmission process for transmitting the reference voice data generated by the voice generation process to a playback device via a network, An evaluation target voice acquisition process for acquiring evaluation target voice data reproduced from the playback device, A voice recognition process for performing voice recognition on the words included in the evaluation target voice data to generate an evaluation target text, An evaluation process for performing a quality evaluation of the voice received via the network based on the evaluation target text, comprising: The voice recognition process performs voice recognition on the words included in the reference voice data to generate a second reference text. The evaluation process performs a first evaluation by comparing the reference text and the second reference text, performs a second evaluation by comparing the reference text and the evaluation target text, and performs the quality evaluation based on the results of the first evaluation and the second evaluation. Voice quality evaluation method. **Claim 5**: A voice generation instruction for converting a reference text into reference voice data, A reference voice transmission instruction for transmitting the reference voice data generated by the voice generation instruction to a playback device via a network, An evaluation target voice acquisition instruction for acquiring evaluation target voice data reproduced from the playback device, A voice recognition instruction for performing voice recognition on the words included in the evaluation target voice data to generate an evaluation target text, An evaluation instruction for performing a quality evaluation of the voice received via the network based on the evaluation target text, causing a computer to execute: The voice recognition instruction performs voice recognition on the words included in the reference voice data to generate a second reference text. The evaluation instruction performs the quality evaluation by comparing the evaluation target text and the second reference text. Voice quality evaluation program. **Claim 6**: A voice generation instruction for converting a reference text into reference voice data, A reference voice transmission instruction for transmitting the reference voice data generated by the voice generation instruction to a playback device via a network, An evaluation target voice acquisition instruction for acquiring evaluation target voice data reproduced from the playback device, A voice recognition instruction for performing voice recognition on the words included in the evaluation target voice data to generate an evaluation target text, An evaluation instruction for performing a quality evaluation of voice received via the network based on the text to be evaluated; causing a computer to execute; The voice recognition instruction performs voice recognition on words included in the reference voice data to generate a second reference text; The evaluation instruction performs a first evaluation by comparing the reference text and the second reference text, performs a second evaluation by comparing the reference text and the text to be evaluated, and performs the quality evaluation based on the results of the first evaluation and the second evaluation. Voice quality evaluation program.

Citation Information

Patent Citations

  • Apparatus, program, and method for evaluating speech quality

    JP2007049462A

  • Synthesized speech evaluation system and synthesized speech evaluation method

    JP2010060846A

  • Speech recognition method, speech evaluation method, speech recognition system, and speech evaluation system

    JP2016051179A

  • Voice recognition result creation device, method and program

    JP2018054717A

  • Prediction device, prediction method and prediction program

    JP2021032909A