Voice processing system, voice processing method, and voice processing program

The voice processing system improves conversation quality and recognition accuracy by segregating voice processing for conversation and recognition, addressing the trade-off in conventional systems.

JP2025114977APending Publication Date: 2025-08-06SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024009244
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-08-06

AI Technical Summary

Technical Problem

Conventional voice recognition systems face a trade-off between voice recognition accuracy and conversation quality when processing audio from multiple users, as voice processing for echo cancellation and noise reduction degrades recognition accuracy, while unprocessed voice is difficult to hear.

Method used

A voice processing system that separates voice processing for conversation quality and voice recognition, using echo cancellation and noise reduction for conversation audio and processing voice recognition on unprocessed audio to improve both quality and accuracy.

Benefits of technology

Enhances conversation quality by using processed voice for conversation and unprocessed voice for recognition, thereby improving voice recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025114977000001_ABST
    Figure 2025114977000001_ABST
Patent Text Reader

Abstract

To provide a voice processing system, a voice processing method, and a voice processing program capable of improving the accuracy of voice recognition for converting a voice which has been input into a voice device into a text while improving the quality of conversation by using the voice device.SOLUTION: A voice processing device 1 includes: an acquisition processing section 111 for acquiring a voice input into microphones of voice devices 2; a voice processing section 112 for performing voice processing for reproducing the voice from a speaker on the voice acquired by the acquisition processing section 111; a voice recognition processing section 113 for performing voice recognition processing for converting the voice into a text on the basis of the voice acquired by the acquisition processing section 111; a voice output processing section 114 for outputting the voice after the voice processing by the voice processing section 112; and a text output processing section 115 for outputting a recognition result of the voice recognition processing by the voice recognition processing section 113.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technology for controlling audio when multiple users individually use audio devices to have a conversation. [Background technology]

[0002] Conventionally, systems have been known in which multiple users can converse using audio devices each equipped with a microphone and a speaker. For example, a system is known that includes multiple audio devices (personal call devices) and a hub device installed in a conference space to which the multiple audio devices can be simultaneously connected via a local network, and that constructs a group call network in which the hub device enables simultaneous mutual calls to the connected audio devices (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2023-037813 Summary of the Invention [Problem to be solved by the invention]

[0004] In a conversation between multiple users using audio devices, the speech of each user may be converted into text and displayed. In such a case, conventional technology performs audio processing such as echo cancellation (EC processing), noise cancellation (NC processing), and gain adjustment (AGC processing) on the audio input to the microphone of the audio device, and then outputs the processed audio to a conversation server and a speech recognition server. When the conversation server acquires the audio, it outputs the audio to another audio device and outputs (plays) it from a speaker. When the speech recognition server acquires the audio, it performs speech recognition processing to convert the audio into text and displays it on a display device or the like.

[0005] Here, the voice recognition server generally performs voice recognition processing by learning voice data that includes noise, etc. Therefore, when voice data that has undergone the voice processing (EC processing, NC processing, AGC processing, etc.) for reproducing the voice from a speaker is input to the voice recognition server, a problem occurs in that the voice recognition accuracy decreases.

[0006] On the other hand, if a configuration is adopted in which a voice suitable for the voice recognition processing, for example, a voice before the voice processing, is input to the conversation server and the voice recognition server, the voice recognition accuracy in the voice recognition server improves, but the voice becomes difficult to hear as conversational voice, resulting in a problem of reduced conversation quality.

[0007] The object of the present disclosure is to provide a voice processing system, a voice processing method, and a voice processing program that can improve the quality of conversations using a voice device and improve the accuracy of voice recognition that converts voice input to a voice device into text. [Means for solving the problem]

[0008] A speech processing system according to one aspect of the present disclosure includes an acquisition processing unit, a speech processing unit, a speech recognition processing unit, a speech output processing unit, and a text output processing unit. The acquisition processing unit acquires speech input to a microphone of a speech device. The speech processing unit performs speech processing on the speech acquired by the acquisition processing unit to play the speech from a speaker. The speech recognition processing unit performs speech recognition processing to convert speech to text based on the speech acquired by the acquisition processing unit. The speech output processing unit outputs the speech after the speech processing by the speech processing unit. The text output processing unit outputs the recognition result of the speech recognition processing by the speech recognition processing unit.

[0009] According to another aspect of the present disclosure, a speech processing system includes an acquisition processor, a speech processor, and an output processor. The acquisition processor acquires speech input to a microphone of an audio device. The speech processor performs speech processing on the speech acquired by the acquisition processor to play the speech from a speaker. The output processor outputs the speech after the speech processing by the speech processor to a speech server that outputs the speech from a speaker of another audio device, and outputs the speech acquired by the acquisition processor to a speech recognition server that converts speech to text.

[0010] Another aspect of the present disclosure is a voice processing method in which one or more processors perform the following steps: acquiring voice input into a microphone of an audio device; performing voice processing on the acquired voice to play the voice from a speaker; performing voice recognition processing to convert the voice into text based on the acquired voice; outputting the voice after the voice processing; and outputting the recognition result of the voice recognition processing.

[0011] Another aspect of the present disclosure is a voice processing program for causing one or more processors to perform the following operations: acquiring voice input into a microphone of an audio device; performing voice processing on the acquired voice to play the voice from a speaker; performing voice recognition processing to convert the voice into text based on the acquired voice; outputting the voice after the voice processing; and outputting the recognition result of the voice recognition processing. [Effects of the Invention]

[0012] According to the present disclosure, it is possible to provide a voice processing system, a voice processing method, and a voice processing program that can improve the quality of conversations using a voice device and improve the accuracy of voice recognition that converts voice input to a voice device into text. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a diagram illustrating an application example of a voice processing system according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram showing a configuration of a voice processing system according to an embodiment of the present disclosure. [Figure 3] FIG. 3 is a diagram schematically illustrating an application example of the voice processing device according to the first embodiment of the present disclosure. [Figure 4] FIG. 4 is a diagram illustrating an application example of a voice processing device according to a second embodiment of the present disclosure. [Figure 5] FIG. 5 is a diagram illustrating an application example of a voice processing device according to a third embodiment of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating an application example of a voice processing device according to a fourth embodiment of the present disclosure. [Figure 7] FIG. 7 is a diagram illustrating an application example of a voice processing device according to a fifth embodiment of the present disclosure. [Figure 8] FIG. 8 is a diagram illustrating an application example of a voice processing device according to a sixth embodiment of the present disclosure. [Figure 9] FIG. 9 is a flowchart illustrating an example of a procedure of a voice control process executed in a voice processing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments are examples that embody the present disclosure and do not limit the technical scope of the present disclosure.

[0015] The voice processing system according to the present disclosure can be applied to a case where, for example, multiple users in the same space (e.g., a conference room) use audio devices each equipped with a microphone and a speaker to converse (conference) with users in other spaces. The voice processing system can also be applied to a case where multiple users in one space use audio devices to converse. Furthermore, the voice processing system can also be applied to a case where one user in one space uses an audio device to converse with a user in another space.

[0016] Fig. 1 shows an application example of the voice processing system 100 according to this embodiment. As shown in Fig. 1, users A to D participate in a conference in conference room R1, and another user (not shown) participates in the conference in conference room R2. Users A to D converse using neckband-type voice devices 2A to 2D that can be worn around the neck, respectively. The user in conference room R2 may use voice device 2A, or may use a single microphone speaker device installed in conference room R2.

[0017] Each audio device 2 in conference room R1 is wirelessly connected (connected via Bluetooth (registered trademark)) to the audio processing device 1, and audio input to the microphone of each audio device 2 is output (played) from the speaker of the audio device 2 (or microphone speaker device) of the user in conference room R2 via the audio processing device 1, the conference terminal 3, and the conference server 4. Also, audio input to the microphone of the audio device 2 (or microphone speaker device) in conference room R2 is played back from the speaker of the audio device 2 of each user in conference room R1 via the conference server 4, the conference terminal 3, and the audio processing device 1. The conference server 4 is an example of an audio server of the present disclosure.

[0018] In this way, the voice processing system 100 is a system that enables multiple users to converse in the same space (conference room R1 in FIG. 1 ) using their own voice devices 2. The voice processing system 100 may also include a display device 5 that can be used in a conference. The display device 5 displays conference information such as camera images of conference participants and conference materials using a conference application, and also displays the recognition results (text information) obtained by converting voice into text using voice recognition processing.

[0019] As shown in FIG. 1, the voice processing system 100 includes a voice processing device 1, a voice device 2, a conference terminal 3, and a conference server 4. The voice device 2 is a wirelessly connected audio device equipped with a microphone and a speaker. The voice device 2 may also have functions such as an AI speaker or a smart speaker. The voice processing system 100 includes multiple voice devices 2, and is a system that transmits and receives voice data of a user's speech between the multiple voice devices 2. The voice processing system 100 is an example of a voice processing system of the present disclosure.

[0020] The voice processing device 1 controls the voice (input voice, output voice, etc.) of the voice device 2, and when a conference starts in a conference room, for example, it executes a process of transmitting and receiving voice to and from multiple voice devices 2. For example, the voice processing device 1 controls multiple voice devices 2 arranged in the same space. The voice processing device 1 also stores the voice acquired from the voice device 2 as voice to be recorded, and executes a process of converting the acquired voice into text (voice recognition process). Note that the voice processing device 1 alone may constitute the voice processing system of the present disclosure.

[0021] The voice processing system of the present disclosure may also include various servers that provide various services such as a conference service, a subtitling (transcription) service using voice recognition, a translation service, and a meeting minutes service. In this embodiment, the system includes a conference server 4 that provides conference services. The conference server 4 provides an online meeting service using a conference application, which is a type of general-purpose software. For example, the conference application is installed in a conference terminal 3. By starting up the conference terminal 3 and logging in, it becomes possible to hold an online meeting using the conference application (for example, an online meeting in conference rooms R1 and R2).

[0022] [Speech processing device 1] 2, the audio processing device 1 is a device including a control unit 11, a storage unit 12, a communication unit 13, etc. For example, the audio processing device 1 is configured as a device (for example, a mixer box) that is connected to multiple audio devices 2 and has a function of mixing or splitting audio input from the multiple audio devices 2 or conference terminals 3.

[0023] The communication unit 13 is a communication unit that connects the audio processing device 1 to a communication network by wire or wirelessly and executes data communication in accordance with a predetermined communication protocol via the communication network with external devices such as the audio device 2 and the conference terminal 3. For example, the communication unit 13 executes pairing processing using the Bluetooth system to establish a wireless connection with each audio device 2.

[0024] The storage unit 12 is a non-volatile storage unit such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory that stores various types of information. Specifically, the storage unit 12 may store data such as information that can identify the audio device 2 (such as a device number or device ID).

[0025] The storage unit 12 also stores control programs such as a voice control program (an example of a voice processing program of the present disclosure) for causing the control unit 11 to execute a voice control process (see FIG. 9 ) described below. For example, the voice control program may be non-temporarily recorded on a computer-readable recording medium such as a CD or a DVD, read by a reading device (not shown) such as a CD drive or a DVD drive provided in the voice processing device 1, and stored in the storage unit 12.

[0026] The control unit 11 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various arithmetic processes. The ROM is a non-volatile storage unit that pre-stores control programs such as a BIOS and an OS that cause the CPU to execute various arithmetic processes. The RAM is a volatile or non-volatile storage unit that stores various information and is used as a temporary storage memory (work area) for various processes executed by the CPU. The control unit 11 controls the audio processing device 1 by having the CPU execute various control programs pre-stored in the ROM or the storage unit 12.

[0027] Specifically, as shown in Fig. 2, the control unit 11 includes various processing units such as an acquisition processing unit 111, a voice processing unit 112, a voice recognition processing unit 113, a voice output processing unit 114, and a text output processing unit 115. The control unit 11 functions as the various processing units by executing various processes in accordance with the control program using the CPU. Some or all of the processing units may be configured with electronic circuits. The control program may be a program for causing multiple processors to function as the processing units.

[0028] Here, the voice processing device 1 may have the configurations shown in the following Examples 1 to 6. The specific configurations of each of Examples 1 to 6 will be described below.

[0029] [Example 1] 3 schematically illustrates an application example of the voice processing device 1 according to the first embodiment. For example, when a conference starts and a user speaks, the acquisition processing unit 111 acquires the spoken voice (voice Va) input to the microphone of the voice device 2 of the user. The acquisition processing unit 111 performs a routing process to output the voice Va to a predetermined output destination. Here, the acquisition processing unit 111 outputs the acquired voice Va (voice data) to a voice processing unit 112 for generating voice for the conference and a voice recognition processing unit 113 for converting the voice into text.

[0030] The audio processing unit 112 performs audio processing on the audio Va acquired by the acquisition processing unit 111 to reproduce the audio from a speaker. Specifically, the audio processing unit 112 performs at least one of echo cancellation (EC processing), noise cancellation (NC processing), and gain adjustment (AGC processing) on the audio Va. The audio processing unit 112 outputs the audio (audio Va1) that has been subjected to the audio processing to the audio output processing unit 114.

[0031] The voice recognition processing unit 113 executes voice recognition processing for converting voice into text based on the voice Va acquired by the acquisition processing unit 111. The voice recognition processing unit 113 converts voice into text using a predetermined voice recognition engine (trained model). Note that the voice recognition engine is generated by learning various voice data (teacher data), and the voice data is data that has not been subjected to voice processing such as echo cancellation, noise cancellation, gain adjustment, etc. The voice processing device 1 is equipped with the voice recognition engine.

[0032] In this way, the speech recognition processing unit 113 executes the speech recognition processing based on the speech Va that has not been subjected to the speech processing. The speech recognition processing unit 113 outputs the recognition result (text information Ta1) of the speech recognition processing for the speech Va to the text output processing unit 115.

[0033] The voice output processing unit 114 outputs the voice Va1 after the voice processing by the voice processing unit 112 to the conference terminal 3 in the conference room R1. The text output processing unit 115 outputs the recognition result (text information Ta1) of the voice recognition processing by the voice recognition processing unit 113 to the conference terminal 3 in the conference room R1.

[0034] When the conference terminal 3 in conference room R1 receives the voice Va1 from the voice processing device 1, it outputs the voice Va1 to the conference server 4 (see FIG. 1). When the conference server 4 receives the voice Va1 from the conference terminal 3 in conference room R1, it outputs the voice Va1 to the conference terminal 3 in conference room R2. The conference terminal 3 in conference room R2 outputs the voice Va1 to the users in conference room R2.

[0035] Furthermore, when the conference terminal 3 in the conference room R1 receives the text information Ta1 from the voice processing device 1, it displays the text information Ta1 on the display device 5 (see FIG. 1). In another embodiment, the conference terminal 3 may accumulate the text information Ta1 and create minutes of the conference. Furthermore, the conference terminal 3 may display the text information Ta1 on the user terminal (not shown) of each user.

[0036] In another embodiment, when the audio device 2 has an echo cancellation function, the audio processing unit 112 may omit the echo cancellation process. Also, when the audio device 2 does not have an echo cancellation function, the audio recognition processing unit 113 may acquire audio Va that has been subjected to echo cancellation and perform the audio recognition process. That is, the audio recognition processing unit 113 may acquire audio that has been subjected to echo cancellation on the audio acquired by the acquisition processing unit 111 and perform the audio recognition process.

[0037] Note that the user may be able to select whether to input the voice Va that has been subjected to echo cancellation to the voice recognition processing unit 113, or whether to input the voice Va that has not been subjected to echo cancellation to the voice recognition processing unit 113. In other words, the user may be able to set whether to perform echo cancellation on the voice Va acquired by the acquisition processing unit 111. Furthermore, the user may be able to set, for each audio device 2, whether to perform echo cancellation on the voice Va acquired by the acquisition processing unit 111.

[0038] According to the configuration of the first embodiment, the speech Va1 that has undergone speech processing (EC processing, NC processing, AGC processing) to reproduce the speech from the speaker is used as the speech for the conference, and the speech Va that has not undergone the speech processing is used as the speech for speech recognition. As a result, the quality of the conversation can be improved because the speech that has undergone speech processing can be used for conversation. In addition, the accuracy of speech recognition can be improved because the speech that has not undergone speech processing is converted into text.

[0039] [Example 2] 4 schematically shows an application example of the voice processing device 1 according to Example 2. Example 2 is an example in which the voice of a user in conference room R2 is reproduced by the voice device 2 in conference room R1.

[0040] For example, when a user in conference room R2 speaks, the conference terminal 3 in conference room R2 acquires the spoken voice (voice Vb) via the audio device 2 and outputs it to the conference server 4. The conference server 4 outputs the voice Vb to the conference terminal 3 in conference room R1, and the conference terminal 3 outputs the voice Vb to the voice processing device 1 upon acquiring it.

[0041] In the voice processing device 1, when the acquisition processing unit 111 acquires the voice Vb, it outputs the acquired voice Vb to the voice processing unit 112 for generating voice for the conference and to the voice recognition processing unit 113 for converting the voice into text.

[0042] The audio processing unit 112 performs the audio processing (for example, NC processing, AGC processing) on the audio Vb acquired by the acquisition processing unit 111. The audio processing unit 112 outputs the audio (audio Vb1) that has been subjected to the audio processing to the audio output processing unit 114.

[0043] The speech recognition processing unit 113 executes the speech recognition processing based on the speech Vb acquired by the acquisition processing unit 111. In this way, the speech recognition processing unit 113 executes the speech recognition processing based on the speech Vb that has not been subjected to the speech processing. The speech recognition processing unit 113 outputs the recognition result (text information Tb1) of the speech recognition processing for the speech Vb to the text output processing unit 115.

[0044] The voice output processing unit 114 outputs the voice Vb1 after the voice processing by the voice processing unit 112 to each voice device 2 in the conference room R1. The text output processing unit 115 outputs the recognition result (text information Tb1) of the voice recognition processing by the voice recognition processing unit 113 to the conference terminal 3 in the conference room R1.

[0045] When each audio device 2 acquires the audio Vb1, it outputs (plays) it from a speaker. The conference terminal 3 in the conference room R1 displays the text information Tb1 on the display device 5 (see FIG. 1).

[0046] According to the configuration of the second embodiment, the same effects as those of the first embodiment can be obtained. That is, the voice Vb1 after the voice processing (EC processing, NC processing, AGC processing) is used as the voice for the conference, and the voice Vb without the voice processing is used as the voice for voice recognition. Therefore, the quality of the conversation can be improved because the voice that has been voice-processed can be used for conversation. In addition, the accuracy of voice recognition can be improved because the voice that has not been voice-processed is used for text conversion.

[0047] By adopting the configurations of the first and second embodiments, a high-quality online conference can be realized.

[0048] [Example 3] 5 schematically shows an application example of a voice processing device 1 according to a third embodiment. In the third embodiment, the voice processing device 1 does not have the voice recognition processing function (voice recognition engine), and a voice recognition server 6 having the voice recognition processing function is provided outside the voice processing device 1. The voice recognition server 6 is an example of the voice recognition server of the present disclosure.

[0049] The speech recognition server 6 is arranged so as to be able to communicate data with the speech processing device 1 and the conference terminal 3. The speech recognition server 6 may be arranged in the conference room R1 or outside the conference room R1. For example, the speech recognition server 6 may be configured as a cloud server. Alternatively, the conference server 4 and the speech recognition server 6 may be configured as a single device, with the conference server 4 having the speech recognition processing function.

[0050] In the example shown in FIG. 5, the acquisition processing unit 111 outputs the voice Va acquired from the voice device 2 to the voice processing unit 112 and also to the voice recognition server 6.

[0051] The audio processing unit 112 performs the audio processing on the audio Va, and outputs the processed audio Va1 to the audio output processing unit 114. The audio output processing unit 114 outputs the audio Va1 after the audio processing by the audio processing unit 112 to the conference terminal 3. That is, the audio output processing unit 114 outputs the audio Va1 after the audio processing by the audio processing unit 112 to the conference terminal 3 (conference server 4), which outputs the audio from the speaker of another audio device 2.

[0052] The speech recognition server 6 executes the speech recognition process based on the speech Va. The speech recognition server 6 executes the speech recognition process using the speech recognition engine (trained model). The speech recognition server 6 is equipped with the speech recognition engine.

[0053] In this way, the voice recognition server 6 executes the voice recognition process based on the voice Va that has not been subjected to the voice processing. The voice recognition server 6 outputs the recognition result (text information Ta1) of the voice recognition process on the voice Va to the conference terminal 3.

[0054] When the conference terminal 3 in the conference room R1 receives the speech Va1 from the speech processing device 1, it outputs the speech Va1 to the conference server 4. Also, when the conference terminal 3 in the conference room R1 receives the text information Ta1 from the speech recognition server 6, it displays the text information Ta1 on the display device 5 (see FIG. 1).

[0055] According to the configuration of the third embodiment, in addition to the effects of the first embodiment, it is possible to simplify the configuration of the speech processing device 1. Furthermore, in the speech processing system 100, an existing speech recognition server 6 can be used.

[0056] The speech recognition server 6 of the third embodiment is an example of the speech recognition processing unit and the text output processing unit of the present disclosure. That is, the speech recognition processing unit and the text output processing unit of the present disclosure may be realized by the speech recognition server 6. That is, the speech processing device 1 according to the third embodiment includes an acquisition processing unit 111 that acquires speech input to a microphone of a speech device 2, a speech processing unit 112 that performs speech processing on the speech acquired by the acquisition processing unit 111 to play the speech from a speaker, and an output processing unit (acquisition processing unit 111) that outputs the speech after the speech processing by the speech processing unit 112 to a conference terminal 3 (conference server 4) that outputs the speech from the speaker of another speech device 2, and outputs the speech acquired by the acquisition processing unit 111 to the speech recognition server 6 that converts the speech into text.

[0057] [Example 4] 6 is a schematic diagram illustrating an application example of the speech processing device 1 according to the fourth embodiment. The fourth embodiment is an example in which the speech of a user in a conference room R2 is reproduced by the speech device 2 in a conference room R1. Similarly to the third embodiment, the fourth embodiment illustrates an example of a configuration in which a speech recognition server 6 is provided outside the speech processing device 1.

[0058] In the example shown in FIG. 6, the acquisition processing unit 111 outputs the voice Vb acquired from the conference terminal 3 to the voice processing unit 112 and also to the voice recognition server 6.

[0059] The audio processing unit 112 performs the audio processing on the audio Vb and outputs the processed audio Vb1 to the audio output processing unit 114. The audio output processing unit 114 outputs the processed audio Vb1 by the audio processing unit 112 to the audio device 2.

[0060] The voice recognition server 6 executes the voice recognition process based on the voice Vb, and outputs the recognition result (text information Tb1) of the voice recognition process for the voice Vb to the conference terminal 3.

[0061] When each audio device 2 acquires the audio Vb1, it outputs (plays) it from a speaker. The conference terminal 3 in the conference room R1 displays the text information Tb1 on the display device 5 (see FIG. 1).

[0062] According to the configuration of the fourth embodiment, it is possible to obtain the same effects as those of the third embodiment. Furthermore, by adopting the configurations of the third and fourth embodiments, it is possible to realize a high-quality online conference.

[0063] [Example 5] 7 is a schematic diagram illustrating an application example of the voice processing device 1 according to the fifth embodiment. In the fifth embodiment, the voice (voice from each microphone) of each voice device 2 is input to the voice recognition processing unit 113, and the voice recognition processing unit 113 executes the voice recognition process for each microphone voice of the voice device 2.

[0064] For example, the acquisition processing unit 111 acquires voices input to the microphones of each of the audio devices 2A to 2D and outputs the voices for each audio device 2. Specifically, the acquisition processing unit 111 adds identification information of the audio device 2A to the voice v1 input to the microphone of the audio device 2A and outputs the voice v1 to the voice recognition processing unit 113, adds identification information of the audio device 2A to the voice v2 input to the microphone of the audio device 2B and outputs the voice v2 to the voice recognition processing unit 113, adds identification information of the audio device 2C to the voice recognition processing unit 113, and adds identification information of the audio device 2C to the voice v3 input to the microphone of the audio device 2C and outputs the voice v4 input to the microphone of the audio device 2D to the voice recognition processing unit 113. In this way, the microphone voices for voice recognition are input to the voice recognition processing unit 113 individually for each audio device 2.

[0065] The voice recognition processing unit 113 performs voice recognition processing individually for each voice device 2 (each microphone voice). For example, the voice recognition processing unit 113 converts voice v1 to text, voice v2 to text, voice v3 to text, and voice v4 to text. The voice recognition processing unit 113 outputs the recognition results (text information Ta1) obtained by individually performing voice recognition processing to the text output processing unit 115.

[0066] According to the configuration of Example 5, it is possible to add identification information of the voice device 2 or identification information of the user of the voice device 2 and display the text information on the display device 5, or to display the text information in a different display area for each voice device 2 or each user of the voice device 2. Note that, as another embodiment of Example 5, the voice recognition processing unit 113 and the text output processing unit 115 may be arranged outside the voice processing device 1 and configured as a voice recognition server 6, similar to Example 3.

[0067] [Example 6] 8 schematically illustrates an application example of the audio processing device 1 according to the sixth embodiment. The sixth embodiment is configured to acquire audio (microphone audio) from an audio device 2A having an echo cancellation function and an audio device 2B not having an echo cancellation function.

[0068] For example, when the acquisition processing unit 111 acquires a voice v1 that is input to the microphone of the audio device 2A and has been subjected to echo cancellation, it outputs the voice v1 to the voice recognition processing unit 113. On the other hand, when the acquisition processing unit 111 acquires a voice v2 that is input to the microphone of the audio device 2B and has not been subjected to echo cancellation, it outputs the voice v2 to the echo cancellation processing unit 116. When the echo cancellation processing unit 116 acquires the voice v2, it performs echo cancellation and outputs a voice v21 that has been subjected to echo cancellation to the voice recognition processing unit 113.

[0069] The voice recognition processor 113 performs voice recognition processing on the voices v1 and v21 that have been subjected to echo cancellation.

[0070] In this way, the voice processing device 1 determines whether or not to execute processing (e.g., echo cancellation) on the voice input to the voice recognition processing unit 113 for each function (e.g., echo cancellation function) of the voice device 2. This makes it possible to use a plurality of voice devices 2 with different functions. Note that, as another embodiment of Example 6, the voice recognition processing unit 113 and the text output processing unit 115 may be arranged outside the voice processing device 1 and configured as a voice recognition server 6, similar to Example 3.

[0071] The speech processing system 100 according to the present disclosure may be configured by any one of the first to sixth embodiments described above or by combining a plurality of the embodiments.

[0072] [Voice control processing] FIG. 9 shows an example of the procedure of the voice control process executed by the control unit 11 of the voice processing device 1.

[0073] The present disclosure can be understood as a voice control method (voice processing method of the present disclosure) that executes one or more steps included in the voice control process. Furthermore, one or more steps included in the voice control process described here may be omitted as appropriate. Furthermore, the steps in the voice control process may be executed in a different order as long as the same operational effect is achieved. Furthermore, while the description here takes as an example a case where the control unit 11 executes each step in the voice control process, in other embodiments, one or more processors may execute each step in the voice control process in a distributed manner.

[0074] Here, as shown in the first embodiment (see FIG. 3), a case will be described as an example in which users in a conference room R1 hold a conference using the audio device 2.

[0075] First, in step S1, the control unit 11 acquires the voice Va input to the microphone of the audio device 2 in the conference room R1.

[0076] Next, in step S2, the control unit 11 performs routing processing to output the acquired voice Va to a predetermined output destination. Here, the control unit 11 outputs the acquired voice Va to a voice processing unit 112 that generates voice for the conference (S3), and also outputs it to a voice recognition processing unit 113 that converts the voice into text (S11). The control unit 11 performs voice processing in step S3, and voice recognition processing in step S11.

[0077] Specifically, in step S3, the control unit 11 performs the audio processing on the acquired audio Va to reproduce the audio from a speaker. For example, the control unit 11 performs at least one of echo cancellation, noise cancellation, and gain adjustment on the audio Va. After performing the audio processing on the audio Va, the control unit 11 outputs the processed audio (audio Va1) to the conference terminal 3 in the conference room R1 in the subsequent step S4. Upon acquiring the audio Va1, the conference terminal 3 reproduces the audio Va1 from the audio device 2 in the conference room R2 via, for example, the conference server 4. After step S4, the control unit 11 ends the audio control processing.

[0078] On the other hand, in step S11, the control unit 11 executes the voice recognition process based on the acquired voice Va. That is, the control unit 11 converts the voice Va, which has not been subjected to the voice processing, into text information Ta1. After executing the voice recognition process on the voice Va, the control unit 11 outputs the recognition result of the voice recognition process (text information Ta1) to the conference terminal 3 in the conference room R1 in the following step S12. Upon acquiring the text information Ta1, the conference terminal 3 displays it on the display device 5 in the conference room R1. After step S12, the control unit 11 ends the voice control process.

[0079] As described above, the control unit 11 executes the audio control process each time audio is acquired from the audio device 2. The control unit 11 executes the processes of steps S3 and S4 and steps S11 and S12 in parallel.

[0080] As described above, the speech processing system 100 according to the present disclosure acquires speech input to the microphone of the speech device 2 and performs speech processing on the acquired speech to play the speech from a speaker. The speech processing system 100 also performs speech recognition processing to convert the speech into text based on the acquired speech. The speech processing system 100 then outputs the speech after the speech processing and the recognition result of the speech recognition processing. Note that the speech recognition processing unit (speech recognition engine) that executes the speech recognition processing may be installed in the device that executes the speech processing (speech processing device 1) or in a device (speech recognition server 6) different from the speech processing device 1.

[0081] According to the above configuration, the speech input to the microphone can be processed and played back as conference speech, thereby improving conference quality (conversation quality). Also, the speech input to the microphone can be converted to text by speech recognition processing without processing the speech, thereby improving the accuracy of speech recognition.

[0082] In the above embodiment, an example has been shown in which conference rooms R1 and R2 are connected via a network to hold an online conference, but the speech processing system 100 of the present disclosure may be configured with only one conference room R1. In this case, for example, in conference room R1, the conference terminal 3 plays back speech input to a microphone of one audio device 2 from a speaker of another audio device 2, and displays text information converted from the speech on the display device 5. For example, the speech output processing unit 114 may output speech after the speech processing by the speech processing unit 112 from the speaker of the audio device 2, and the text output processing unit 115 may display text information, which is the recognition result of the speech recognition processing by the speech recognition processing unit 113, on the display device 5.

[0083] In addition, the speech recognition process may convert speech in a first language (e.g., Japanese) into text in the first language, or may convert speech in a first language (e.g., Japanese) into text in a second language (e.g., English).

[0084] [Disclosure Note] The following is a summary of the disclosure extracted from the above-described embodiment. Note that the configurations and processing functions described in the following supplementary notes can be selected and combined as desired.

[0085] <Appendix 1> an acquisition processing unit that acquires audio input to a microphone of the audio device; an audio processing unit that performs audio processing on the audio acquired by the acquisition processing unit to reproduce the audio from a speaker; a speech recognition processing unit that performs speech recognition processing to convert speech into text based on the speech acquired by the acquisition processing unit; an audio output processing unit that outputs the audio after the audio processing by the audio processing unit; a text output processing unit that outputs a recognition result of the speech recognition processing by the speech recognition processing unit; A voice processing system comprising:

[0086] <Appendix 2> the speech recognition processing unit executes the speech recognition processing based on the speech that has not been subjected to the speech processing; 10. The speech processing system of claim 1.

[0087] <Appendix 3> the audio processing unit performs at least one of echo cancellation, noise cancellation, and gain adjustment on the audio acquired by the acquisition processing unit; 3. The speech processing system according to claim 1 or 2.

[0088] <Appendix 4> the speech recognition processing unit acquires speech obtained by performing echo cancellation on the speech acquired by the acquisition processing unit, and executes the speech recognition processing. 4. A speech processing system according to any one of Supplementary notes 1 to 3.

[0089] <Appendix 5> It is possible to set whether or not to perform echo cancellation on the audio acquired by the acquisition processing unit. 5. The speech processing system of claim 4.

[0090] <Appendix 6> whether or not to perform echo cancellation on the audio acquired by the acquisition processing unit can be set for each of the audio devices. 6. The speech processing system of claim 5.

[0091] <Appendix 7> the audio output processing unit outputs the audio processed by the audio processing unit from a speaker; the text output processing unit causes a display device to display text information that is a recognition result of the speech recognition processing by the speech recognition processing unit. 7. A speech processing system according to any one of Supplementary notes 1 to 6.

[0092] <Appendix 8> the acquisition processing unit acquires audio input to a microphone of each of the plurality of audio devices; the voice recognition processing unit executes the voice recognition process for the voice for each of the voice devices; 8. A speech processing system according to any one of Supplementary Notes 1 to 7.

[0093] <Appendix 9> a voice processing device including the acquisition processing unit, the voice processing unit, the voice recognition processing unit, the voice output processing unit, and the text output processing unit; 9. A speech processing system according to any one of Supplementary notes 1 to 8. [Explanation of symbols]

[0094] 1: Audio processing device 2: Audio equipment 3: Conference terminal 4: Conference Server 5:Display device 6: Speech recognition server 11: Control section 12: Storage section 13: Communications Department 100: Audio processing system 111: Acquisition processing unit 112: Audio processing unit 113: Speech recognition processing unit 114: Audio output processing unit 115: Text output processing unit 116: Echo cancellation processing unit

Claims

1. an acquisition processing unit that acquires audio input to a microphone of the audio device; an audio processing unit that performs audio processing on the audio acquired by the acquisition processing unit to reproduce the audio from a speaker; a speech recognition processing unit that performs speech recognition processing to convert speech into text based on the speech acquired by the acquisition processing unit; an audio output processing unit that outputs the audio after the audio processing by the audio processing unit; a text output processing unit that outputs a recognition result of the speech recognition processing by the speech recognition processing unit; A voice processing system comprising:

2. the speech recognition processing unit executes the speech recognition processing based on the speech that has not been subjected to the speech processing; The audio processing system of claim 1 .

3. the audio processing unit performs at least one of echo cancellation, noise cancellation, and gain adjustment on the audio acquired by the acquisition processing unit; The audio processing system of claim 1 .

4. the speech recognition processing unit acquires speech obtained by performing echo cancellation on the speech acquired by the acquisition processing unit, and executes the speech recognition processing. The audio processing system of claim 1 .

5. whether or not to perform echo cancellation on the audio acquired by the acquisition processing unit can be set for each of the audio devices.

5. The audio processing system of claim 4.

6. the acquisition processing unit acquires audio input to a microphone of each of the plurality of audio devices; the voice recognition processing unit executes the voice recognition process for the voice for each of the voice devices; The audio processing system of claim 1 .

7. a speech processing device including the acquisition processing unit, the speech processing unit, the speech recognition processing unit, the speech output processing unit, and the text output processing unit; The audio processing system of claim 1 .

8. an acquisition processing unit that acquires audio input to a microphone of the audio device; an audio processing unit that performs audio processing on the audio acquired by the acquisition processing unit to reproduce the audio from a speaker; an output processing unit that outputs the processed voice by the voice processing unit to a voice server that outputs the voice from a speaker of another voice device, and outputs the voice acquired by the acquisition processing unit to a voice recognition server that converts the voice into text; A voice processing system comprising:

9. Acquiring audio input into a microphone of an audio device; performing audio processing on the acquired audio to play the audio from a speaker; performing a speech recognition process for converting the speech into text based on the acquired speech; outputting the processed audio; outputting a recognition result of the speech recognition processing; An audio processing method executed by one or more processors.

10. Acquiring audio input into a microphone of an audio device; performing audio processing on the acquired audio to play the audio from a speaker; performing a speech recognition process for converting the speech into text based on the acquired speech; outputting the processed audio; outputting a recognition result of the speech recognition processing; An audio processing program for causing one or more processors to execute the above.

Citation Information

Patent Citations

  • Conference system

    JP2023037813A