Voice processing system, voice processing method, and voice processing program

The speech processing system addresses delays in converting multiple user voices to text by synthesizing and individually processing voices, achieving real-time and accurate text display.

JP2026025105APending Publication Date: 2026-02-13SHARP KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024127654
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-13

Smart Images

  • Figure 2026025105000001_ABST
    Figure 2026025105000001_ABST
Patent Text Reader

Abstract

To provide a voice processing system, a voice processing method, and a voice processing program capable of performing text conversion of voice of a conversation using a plurality of voice devices in real time and improving accuracy of the text conversion.SOLUTION: The audio processing device 1 includes the acquisition processing unit 111 that acquires a plurality of input audios input to the respective microphones of the plurality of audio devices 2, the synthesis processing unit 112 that synthesizes the plurality of input audios acquired by the acquisition processing unit 111 into one synthesized audio, and the output processing unit 113 that converts the synthesized audio synthesized by the synthesis processing unit 112 into texts and outputs the plurality of input audios and the synthesized audio to the conference server 4 that individually converts each of the plurality of input audios into a plurality of texts.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technology for converting speech into text (transcription) when multiple users individually use audio devices to converse with each other. [Background technology]

[0002] Conventionally, systems have been known in which multiple users can converse using audio devices each equipped with a microphone and a speaker. For example, a system is known that includes multiple audio devices (personal call devices) and a hub device that is installed in a conference space and simultaneously connects the multiple audio devices to a local network, and that creates a group call network in which the hub device enables simultaneous mutual calls to the connected audio devices (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2023-037813 Summary of the Invention [Problem to be solved by the invention]

[0004] In a conversation in which multiple users each use a voice device, the user's speech may be converted into text (transcription) and displayed. In this case, highly accurate text conversion (voice recognition) can be achieved because the speech can be acquired individually from each voice device and converted into text. However, with conventional methods, the load of the text conversion process is high, causing delays in the conversion process, which impairs the real-time nature of displaying text during a conversation.

[0005] An object of the present disclosure is to provide a speech processing system, a speech processing method, and a speech processing program that can convert speech from conversations using multiple audio devices into text in real time and improve the accuracy of the text conversion. [Means for solving the problem]

[0006] A speech processing system according to one aspect of the present disclosure includes an acquisition processing unit, a synthesis processing unit, and an output processing unit. The acquisition processing unit acquires multiple input speeches input to microphones of multiple speech devices. The synthesis processing unit synthesizes the multiple input speeches acquired by the acquisition processing unit into a single first speech. The output processing unit converts the first speech synthesized by the synthesis processing unit into a first text, and outputs the multiple input speeches and the first speech to a conversion processing unit that individually converts each of the multiple input speeches into multiple second texts.

[0007] Another aspect of the present disclosure is a voice processing method performed by one or more processors, which includes acquiring multiple input voices input to respective microphones of multiple audio devices, synthesizing the acquired multiple input voices into a single first voice, and outputting the multiple input voices and the first voices to a conversion processing unit that converts the first voices into a first text and individually converts each of the multiple input voices into multiple second texts.

[0008] Another aspect of the present disclosure is a voice processing program for causing one or more processors to execute the following steps: acquiring multiple input voices input to respective microphones of multiple audio devices; synthesizing the acquired multiple input voices into a single first voice; and outputting the multiple input voices and the first voices to a conversion processing unit that converts the first voice into a first text and individually converts each of the multiple input voices into multiple second texts. [Effects of the Invention]

[0009] According to the present disclosure, it is possible to provide a voice processing system, a voice processing method, and a voice processing program that can convert the voice of a conversation using multiple voice devices into text in real time and improve the accuracy of the text conversion. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram illustrating an application example of a voice processing system according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram showing a configuration of a voice processing system according to an embodiment of the present disclosure. [Figure 3] FIG. 3 is a diagram illustrating an example of device information used in the speech processing system according to an embodiment of the present disclosure. [Figure 4] FIG. 4 is a diagram illustrating an example of voice information used in the voice processing system according to an embodiment of the present disclosure. [Figure 5] FIG. 5 is a diagram showing a specific example of audio input to an audio device according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a diagram illustrating a specific example of processing audio acquired from an audio device according to an embodiment of the present disclosure. [Figure 7] FIG. 7 is a diagram illustrating a specific example of processing audio acquired from an audio device according to an embodiment of the present disclosure. [Figure 8] FIG. 8 is a flowchart illustrating an example of a procedure of a voice control process executed in a voice processing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments are examples that embody the present disclosure and do not limit the technical scope of the present disclosure.

[0012] The voice processing system according to the present disclosure can be applied to a case where, for example, multiple users in the same space (e.g., a conference room) converse (conference) using audio devices each equipped with a microphone and a speaker. The voice processing system can also be applied to a case where multiple spaces are connected and users in each space converse with each other (e.g., remote conference).

[0013] Fig. 1 shows an application example of a voice processing system 100 according to this embodiment. As shown in Fig. 1, users A to D participate in a conference in a conference room R1. Users A to D converse using neckband-type voice devices 2A to 2D that can be worn around the neck, respectively.

[0014] Each audio device 2 is wirelessly connected (connected via Bluetooth (registered trademark)) to the audio processing device 1, and audio input into the microphone of each audio device 2 is input into the audio processing device 1 and output (played) from the speaker of each audio device 2.

[0015] In this way, the voice processing system 100 is a system that enables multiple users to converse in the same space (conference room R1 in FIG. 1 ) using their respective voice devices 2. The voice processing system 100 may also include a display device 5 that can be used in a conference. The display device 5 displays conference information such as camera images of conference participants and conference materials using a conference application, and also displays the results (text information) of speech-to-text conversion (transcription).

[0016] As shown in FIG. 1, the voice processing system 100 includes a voice processing device 1, a voice device 2, a user terminal 3, and a conference server 4. The voice device 2 is a wirelessly connected audio device equipped with a microphone and a speaker. The voice processing system 100 includes multiple voice devices 2 and transmits and receives audio data of user speech between the multiple voice devices 2. The voice devices 2 may be the same type of audio device or different types of audio devices. For example, the multiple voice devices 2 may include a wirelessly connected audio device and a wired connected audio device. The multiple voice devices 2 may also include a neckband-type audio device, a headset-type audio device, and a stationary audio device. The voice device 2 may also be built into the user terminal 3.

[0017] The voice processing device 1 controls the voice (input voice, output voice, etc.) of the voice device 2, and executes a process of transmitting and receiving voice to and from multiple voice devices 2 when a conference starts, for example, in a conference room. For example, the voice processing device 1 controls multiple voice devices 2 arranged in the same space. The voice processing device 1 also stores the voice acquired from the voice device 2 as audio to be recorded, and executes a process of converting the acquired voice into text (character conversion process). Note that the voice processing device 1 alone may constitute the voice processing system of the present disclosure. The voice processing device 1 may be equipped with a text conversion engine and perform the character conversion process itself, or the voice processing device 1 may not be equipped with a text conversion engine, and the voice processing device 1 may output voice data to an external server, the server may perform the character conversion process, and the voice processing device 1 may receive the converted text data from the server.

[0018] The speech processing system of the present disclosure may also have functions to provide various services, such as a conference service, a subtitling (transcription) service using speech recognition, a translation service, and a minutes service. In this embodiment, the speech processing system 100 includes a conference server 4 that provides a conference service. The conference server 4 provides a conference service in the form of a conference application, which is a type of general-purpose software. For example, the conference application is installed on a user terminal 3. By starting up the user terminal 3 and logging in, it becomes possible to hold a conference using the conference application. In this embodiment, the conference server 4 performs a process (character conversion process, transcription) of converting speech input to each speech device 2 into text and displaying it on the user terminal 3, display device 5, etc. Note that if the speech processing device 1 performs character conversion process, the conference server 4 does not need to be equipped with a character conversion function (text conversion engine).

[0019] The user terminals 3 are personal computers owned by the users who will participate in the conference, and each user can start the conference application on the user terminal 3 and view the conference screen.

[0020] [Speech processing device 1] 2, the voice processing device 1 is a device including a control unit 11, a storage unit 12, a communication unit 13, etc. For example, the voice processing device 1 is connected to a plurality of voice devices 2 and has a function of mixing or splitting voices input from the plurality of voice devices 2. The voice processing device 1 may also have a character conversion function of converting input voice into text.

[0021] The communication unit 13 is a communication unit that connects the audio processing device 1 to a communication network via a wired or wireless connection and executes data communication in accordance with a predetermined communication protocol via the communication network with external devices such as the audio device 2, the user terminal 3, the conference server 4, and the display device 5. For example, the communication unit 13 executes pairing processing using the Bluetooth method to establish a wireless connection with each audio device 2.

[0022] The storage unit 12 is a non-volatile storage unit such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory that stores various types of information. The storage unit 12 stores data such as device information D1 related to the audio device 2 and audio information D2 related to the spoken voice.

[0023] FIG. 3 shows an example of device information D1. Information such as a connection ID, a device number, and a user name is registered in device information D1. The connection ID is identification information (device information) used when connecting audio device 2, such as a Bluetooth address. The device number is identification information such as a unique number or name of audio device 2. Instead of the connection ID and the device number, identification information such as an identifier assigned by audio processing system 100 or a USB port number may be registered. The user name is the name of the user who uses audio device 2. In this way, audio device 2 and user identification information are associated and registered in device information D1. In the example shown in FIG. 3, device number "MS001" indicates audio device 2A used by user A, device number "MS002" indicates audio device 2B used by user B, device number "MS003" indicates audio device 2C used by user C, and device number "MS004" indicates audio device 2D used by user D.

[0024] FIG. 4 shows an example of the audio information D2. Information such as an audio ID, audio data, a device number, a speech time, a speaker, and text data is registered in the audio information D2. The audio ID is identification information of the audio corresponding to the content of the user's (speaker's) speech. The audio data is data of the audio (speech) acquired from the audio device 2. The device number is identification information of the audio device 2 and is stored in association with the device information D1 (see FIG. 3). Each piece of audio data is stored in association with the device number. The speech time is the time when the user spoke, and includes, for example, the speech start time and the speech end time. The speaker is the user's name or identification information (user ID). The text data is character information obtained by converting the user's speech into text. In this embodiment, the conference server 4 performs a process of converting the speech into text, and the converted text data is registered in the audio information D2. When the conference starts, the control unit 11 registers each piece of information in the audio information D2 based on the user's speech acquired from the audio device 2.

[0025] In the example shown in Figure 4, voice Va with voice ID "101" indicates voice acquired from user A's voice device 2A, and text Ta indicates data obtained by converting voice Va into text. Voice Vb with voice ID "102" indicates voice acquired from user B's voice device 2B, and text Tb indicates data obtained by converting voice Vb into text.

[0026] The storage unit 12 also stores control programs such as a voice control program (an example of a voice processing program of the present disclosure) for causing the control unit 11 to execute a voice control process (see FIG. 8 ) described below. For example, the voice control program may be non-temporarily recorded on a computer-readable recording medium such as a CD or a DVD, read by a reading device (not shown) such as a CD drive or a DVD drive provided in the voice processing device 1, and stored in the storage unit 12.

[0027] The control unit 11 has control devices such as a CPU, a ROM, and a RAM. The CPU is a processor that executes various arithmetic processes. The ROM is a non-volatile storage unit that pre-stores control programs such as a BIOS and an OS that cause the CPU to execute various arithmetic processes. The RAM is a volatile or non-volatile storage unit that stores various information and is used as a temporary storage memory (work area) for various processes executed by the CPU. The control unit 11 controls the audio processing device 1 by having the CPU execute various control programs pre-stored in the ROM or the storage unit 12.

[0028] Specifically, as shown in Fig. 2, the control unit 11 includes various processing units such as an acquisition processing unit 111, a synthesis processing unit 112, an output processing unit 113, a display processing unit 114, and a generation processing unit 115. The control unit 11 functions as the various processing units by executing various processes in accordance with the control program using the CPU. Some or all of the processing units may be configured with electronic circuits. The control program may be a program for causing multiple processors to function as the processing units.

[0029] The acquisition processing unit 111 acquires the voices uttered by the user. Specifically, the acquisition processing unit 111 acquires a plurality of input voices input to the respective microphones of a plurality of audio devices 2. For example, when a conference starts and a user speaks, the acquisition processing unit 111 acquires the voice (input voice) input to the microphone of the audio device 2 of that user. The acquisition processing unit 111 also acquires time information corresponding to the time of utterance of the user's voice. For example, the acquisition processing unit 111 acquires the time when the user's voice is input to the microphone of the audio device 2 or the time when the voice is acquired.

[0030] When the acquisition processing unit 111 acquires voice from the voice device 2, it stores the voice data in voice information D2 (see FIG. 4) in association with identification information (device number, connection ID, device name, etc.) of the voice device 2. In the example shown in FIG. 1, the acquisition processing unit 111 acquires the speech voices of users A to D in conference room R1 from voice devices 2A to 2D, respectively, and stores each input voice Va to Vd in association with the device number in voice information D2.

[0031] 5 shows an example of the audio (input audio) acquired from each audio device 2. Input audio Va indicates the audio spoken by user A from time t1 to t3, input audio Vb indicates the audio spoken by user B from time t2 to t5, input audio Vc indicates the audio spoken by user C from time t6 to t8, and input audio Vd indicates the audio spoken by user D from time t4 to t7. Note that input audio Va and input audio Vb overlap in the period from time t2 to t3, input audio Vb and input audio Vd overlap in the period from time t4 to t5, and input audio Vc and input audio Vd overlap in the period from time t6 to t7.

[0032] The synthesis processing unit 112 synthesizes multiple input voices acquired by the acquisition processing unit 111 into one voice (synthesized voice V1). The synthesized voice V1 is an example of the first voice of the present disclosure. Specifically, the synthesis processing unit 112 detects a portion where voice exists (voice section) from the voice stream acquired from each audio device 2, and synthesizes multiple voices in the detected voice section into one synthesized voice V1. For example, the synthesis processing unit 112 detects a voice section using a silent state for a predetermined period as a trigger and extracts multiple voices. In the example shown in FIG. 5, the synthesis processing unit 112 synthesizes input voices Va to Vd of users A to D corresponding to the voice section (times t1 to t8) to generate synthesized voice V1.

[0033] The output processing unit 113 outputs the multiple input voices acquired by the acquisition processing unit 111 and the synthesized voices synthesized by the synthesis processing unit 112 to a conversion processing unit (character conversion device) having a character conversion function (a function for converting voice to text). For example, the output processing unit 113 outputs synthesized voice V1, which is obtained by synthesizing input voices Va to Vd acquired from the voice devices 2A to 2D of users A to D, to the conference server 4, which performs character conversion processing (transcription).

[0034] FIG. 6 schematically shows a configuration for outputting speech from the speech processing device 1 to the conference server 4. The conference server 4 performs text conversion processing based on the synthesized speech V1 to convert the synthesized speech V1 into text T1. The conference server 4 outputs the text data (text T1) that is the result of the text conversion to the speech processing device 1. When the control unit 11 of the speech processing device 1 acquires the data of the text T1 from the conference server 4, it stores the data in the memory unit 12 in association with the synthesized speech V1.

[0035] Furthermore, the output processing unit 113 outputs each of the input speeches Va to Vd acquired from the speech devices 2A to 2D of users A to D to the conference server 4, in association with the identification information (device number) of each of the speech devices 2A to 2D or the identification information (user name) of each of the users. That is, the output processing unit 113 outputs the input speeches Va to Vd individually to the conference server 4. The conference server 4 converts each of the input speeches Va to Vd into text by performing a character conversion process. For example, the conference server 4 converts the input speech Va to text Ta, the input speech Vb to text Tb, the input speech Vc to text Tc, and the input speech Vd to text Td. Note that because the conference server 4 performs character conversion process on individual speeches, character conversion accuracy can be improved compared to performing character conversion process on synthesized speech. The conference server 4 outputs text data (text Ta to Td) that is the result of the character conversion to the speech processing device 1. When the control unit 11 acquires the data of the texts Ta to Td from the conference server 4, it stores the data in the storage unit 12 (voice information D2 (see FIG. 4)) in association with the input voices Va to Vd, respectively.

[0036] In this way, the output processing unit 113 outputs the multiple input voices acquired by the acquisition processing unit 111 and the synthesized voices synthesized by the synthesis processing unit 112 to the conversion processing unit separately. The conference server 4 is an example of a conversion processing unit of the present disclosure.

[0037] The display processing unit 114 displays the text generated by the conversion processing unit (conference server 4) on the user terminal 3 or the display device 5. For example, during a conference, the display processing unit 114 displays text T1 (see FIG. 4) obtained by converting synthesized speech V1 from the conference server 4 on the conference screen (not shown) of each of the user terminals 3A to 3D. This allows the display processing unit 114 to display the content of user utterances in text in real time during the conference.

[0038] The generation processing unit 115 generates minutes (summary) of the conference based on the text generated by the conversion processing unit (conference server 4). For example, after the conference is over, the generation processing unit 115 generates minutes of the conference based on text Ta to Td (see FIG. 4) obtained by converting input speech Va to Vd from the conference server 4. With the above configuration, minutes are generated using text obtained by converting individual speech acquired from each speech device 2, so that minutes can be generated based on highly accurate character conversion results.

[0039] Here, after the process of converting synthesized speech V1 into text T1 is completed, output processing unit 113 may output the multiple input speeches Va to Vd to conference server 4. Specifically, output processing unit 113 outputs synthesized speech V1 to conference server 4 during a predetermined period (conference period) during which acquisition processing unit 111 acquires the multiple input speeches, and outputs the multiple input speeches to conference server 4 after the predetermined period has elapsed.

[0040] In this way, control unit 11 displays the text into which synthesized speech V1 has been converted by conference server 4 while the user is speaking, and generates minutes of the user's conversation based on the multiple texts into which the multiple input speeches Va-Vd have been converted by conference server 4. This makes it possible, for example, to quickly convert synthesized speech V1 into text and display it in real time during a conference, and then, after the conference ends, to generate minutes by converting the individual input speeches Va-Vd into text with high accuracy.

[0041] The conference server 4 is equipped with a well-known character conversion function and converts the voice acquired from the voice processing device 1 into text. When a remote conference is held, the conference server 4 transmits the voice acquired from the voice processing device 1 in the conference room R1 to a voice processing device installed in another space (remote location), and outputs the voice from that voice processing device to users participating in the conference at the remote location. The conference server 4 may be a cloud server or a computer (personal computer) installed in the conference room R1. In another embodiment, the voice processing device 1 may be equipped with the character conversion function.

[0042] [Other embodiments] In the above-described embodiment, the output processing unit 113 outputs each of the multiple input voices Va to Vd individually to the conference server 4. However, in another embodiment, the output processing unit 113 may store each of the multiple input voices Va to Vd in a queue (storage area). The queue is a storage area set in the storage unit 12, and one queue may be set in the storage unit 12, or multiple queues may be set for each audio device 2.

[0043] 7, when the acquisition processing unit 111 acquires input audio Va to Vd from each of the audio devices 2A to 2D, the acquisition processing unit 111 stores the input audio Va to Vd individually in a queue. Furthermore, the output processing unit 113 detects, for each audio device 2, a portion where audio exists (a voice section) from the audio stream acquired from the audio device 2, and stores input audio containing only the detected voice section in a queue. For example, the output processing unit 113 detects a voice section from the audio stream acquired from the audio device 2A, and stores input audio Va containing only the detected voice section in a queue, and detects a voice section from the audio stream acquired from the audio device 2B, and stores input audio Vb containing only the detected voice section in a queue.

[0044] The acquisition processing unit 111 also stores the input voices Va to Vd in the queue in chronological order of acquisition so that they do not overlap with each other. For example, the acquisition processing unit 111 stores the input voice Va acquired from audio device 2A between times t1 and t3 (see FIG. 5) in the queue, the input voice Vb acquired from audio device 2B between times t2 and t5 in the queue, the input voice Vd acquired from audio device 2D between times t4 and t7 in the queue, and the input voice Vc acquired from audio device 2C between times t6 and t8 in the queue. The acquisition processing unit 111 also associates each of the multiple input voices with the identification information (device number) of the audio device 2 or the identification information (user name) of the user and stores them in the queue.

[0045] Furthermore, the output processing unit 113 outputs the voices stored in the queue to the conference server 4. Specifically, the output processing unit 113 outputs the multiple input voices to the conference server 4 so that they do not overlap with each other. For example, the output processing unit 113 outputs the input voices Va to Vd to the conference server 4 so that they do not overlap with each other. For example, as shown in FIG. 5, the input voices Va and Vb overlap in the section from time t2 to t3, the input voices Vb and Vd overlap in the section from time t4 to t5, and the input voices Vc and Vd overlap in the section from time t6 to t7. Therefore, when the input voices are synthesized according to the input time, the voices overlap in each of the above sections, as shown in the synthesized voice V1. In response to this, the output processing unit 113 outputs the input voices Va to Vd to the conference server 4 with the time intervals shifted so that the voices do not overlap in each of the above sections. Furthermore, the output processing unit 113 outputs the input audio Va to Vd as one audio (audio stream V2) to the conference server 4 (see FIG. 7).

[0046] The conference server 4 executes a character conversion process based on one audio stream V2 output from the audio processing device 1. This allows the conference server 4 to convert the input audio Va to Vd together, for example, thereby further improving the accuracy of the character conversion. For example, the conference server 4 can convert each input audio into a character while taking into account the context surrounding it by referring to other input audio, so that each input audio can be converted into a character accurately in accordance with the context.

[0047] As described above, the configuration shown in FIG. 7 can reduce the number of times of character conversion (transcription) while achieving highly accurate character conversion. Specifically, by using the input speech before synthesis, it is possible to clearly define the boundaries of the speech, thereby improving the accuracy of speech recognition. Furthermore, it is possible to accurately identify the speaker corresponding to the speech, since it is clear which audio device 2 the speech was input from. Furthermore, since only one audio stream (audio stream V2) is sent to the conference server 4, the processing load can be reduced. Furthermore, since the context (flow of conversation) can be taken into consideration, the accuracy of character conversion can be improved.

[0048] In the above configuration, the acquisition processing unit 111 may perform predetermined audio processing such as gain adjustment and noise removal on the input audio acquired from each audio device 2, detect the audio section after the audio processing, and store the input audio for that audio section in a queue. This improves the accuracy of detecting the audio section, thereby preventing unnecessary audio input to the queue and reducing the number of times the character conversion process is performed.

[0049] Furthermore, in the above configuration, the output processing unit 113 may execute a process to delete input voices stored in the queue if the same input voice of the same user is included. For example, a voice uttered by user A may be simultaneously input to the microphone of user A's voice device 2A and the microphone of user B's voice device 2B. In this case, user A's input voices input from voice device 2A and voice device 2B may be stored in the queue. In this case, user A's input voice input from voice device 2B may contain reverberation and noise, and using this voice may result in a problem of reduced character conversion accuracy. Therefore, the output processing unit 113 executes a process to delete user A's input voice acquired from voice device 2B. This further improves character conversion accuracy. Furthermore, if a silence period lasts for more than a first predetermined time, part of the silence period may be deleted to shorten the silence data to a second predetermined time, or the voice data during the silence period may be deleted and replaced with silence data of a second predetermined time shorter than the first predetermined time. This reduces the CPU load imposed by the character conversion process and shortens the character conversion process time.

[0050] [Voice control processing] FIG. 8 shows an example of the procedure of the voice control process executed by the control unit 11 of the voice processing device 1.

[0051] The present disclosure can be understood as a voice control method (voice processing method of the present disclosure) that executes one or more steps included in the voice control process. Furthermore, one or more steps included in the voice control process described here may be omitted as appropriate. Furthermore, the steps in the voice control process may be executed in a different order as long as the same operational effect is achieved. Furthermore, while the description here takes as an example a case where the control unit 11 executes each step in the voice control process, in other embodiments, one or more processors may execute each step in the voice control process in a distributed manner.

[0052] <Step S1> In step S1, the control unit 11 determines whether an operation to start a conference has been accepted. For example, user A, who is the organizer of the conference, starts a conference application on the user terminal 3A and accepts a conference start operation on the setting screen. When the control unit 11 accepts the conference start operation (S1: Yes), it shifts the processing to step S2. The control unit 11 waits until the conference start operation is accepted (S1: No).

[0053] <Step S2> In step S2, the control unit 11 starts a process of acquiring voices uttered by the users from the audio device 2. For example, when a conference starts and voice uttered by user A is input to the microphone of audio device 2A, the control unit 11 acquires the voice (input voice Va) from audio device 2A. Furthermore, when voice uttered by user B is input to the microphone of audio device 2B, the control unit 11 acquires the voice (input voice Vb) from audio device 2B. When voice uttered by user C is input to the microphone of audio device 2C, the control unit 11 acquires the voice (input voice Vc) from audio device 2C. When voice uttered by user D is input to the microphone of audio device 2D, the control unit 11 acquires the voice (input voice Vd) from audio device 2D. The control unit 11 registers information related to the acquired input voices in voice information D2 (see FIG. 4).

[0054] <Step S3> In step S3, control unit 11 determines whether or not a voice section has been detected. For example, control unit 11 detects a portion where voice exists (voice section) from the voice stream acquired from each voice device 2. In the example shown in Fig. 5, control unit 11 detects a voice section from time t1 to t8 where input voices Va to Vd exist in the voice stream acquired from voice devices 2A to 2D. When control unit 11 detects a voice section in the voice stream acquired from voice devices 2A to 2D, it causes the process to proceed to step S4.

[0055] Furthermore, the control unit 11 detects a portion where audio exists (audio section) from the acquired audio stream for each audio device 2. For example, in the example shown in Fig. 5, the control unit 11 detects an audio section from time t1 to t3 where input audio Va exists in the audio stream acquired from the audio device 2A. When the control unit 11 detects an audio section from the acquired audio stream for each audio device 2, it causes the process to proceed to step S31.

[0056] <Step S4> In step S4, the control unit 11 synthesizes a plurality of input voices corresponding to the voice section. In the example shown in Fig. 5, the control unit 11 synthesizes input voices Va to Vd extracted in the voice section from time t1 to t8 to generate one synthesized voice V1.

[0057] <Step S5> In step S5, control unit 11 outputs synthesized speech V1 to conference server 4, which executes a character conversion process. Conference server 4 recognizes synthesized speech V1 and converts it into text data (text T1).

[0058] <Step S6> In step S6, the control unit 11 determines whether or not text data (character conversion results) have been acquired from the conference server 4. If the control unit 11 has acquired the text data from the conference server 4 (S6: Yes), the control unit 11 shifts the process to step S7. On the other hand, if the control unit 11 cannot acquire the text data from the conference server 4 (S6: No), the control unit 11 shifts the process to step S3. For example, if a voice segment such as a sneeze is detected, the character conversion process may not be performed normally and no text data may exist. In such a case, the control unit 11 returns to step S3 without acquiring text data (S6: No) and re-executes the above-described process.

[0059] <Step S7> In step S7, control unit 11 outputs the text data and causes the external terminal to display the text (characters, images, etc.). For example, control unit 11 causes text T1, which is obtained by converting synthesized voice V1 into text, to be displayed on the conference screen of each of user terminals 3A to 3D. Control unit 11 also causes text T1 to be displayed on the conference screen of display device 5. After step S7, control unit 11 causes the process to proceed to step S8.

[0060] <Step S31> In step S31, the control unit 11 stores the input voice corresponding to the voice section in a queue. In the example shown in Fig. 5, the control unit 11 stores the input voice Va extracted in the voice section from time t1 to t3 in the queue in association with the identification information (device number) of audio device 2A or the identification information (user name) of user A. The control unit 11 also stores the input voice Vb extracted in the voice section from time t2 to t5 in the queue in association with the identification information of audio device 2B or the identification information of user B, stores the input voice Vd extracted in the voice section from time t4 to t7 in the queue in association with the identification information of audio device 2D or the identification information of user D, and stores the input voice Vc extracted in the voice section from time t6 to t8 in the queue in association with the identification information of audio device 2C or the identification information of user C.

[0061] Furthermore, the control unit 11 shifts the time of the input voices Va to Vd so that they do not overlap with each other, arranges them in the order of acquisition time, and stores them in the queue. After step S31, the control unit 11 causes the process to proceed to step S8.

[0062] <Step S8> In step S8, the control unit 11 determines whether an operation to end the conference has been accepted. For example, when user A wants to end the conference, he or she closes the conference application. If the control unit 11 accepts an operation to end the conference application (S8: Yes), the control unit 11 shifts the process to step S9. On the other hand, if the control unit 11 does not accept an operation to end the conference (S8: No), the control unit 11 returns the process to step S3.

[0063] Returning to step S3, the control unit 11 detects the next voice section and executes the synthesis process (S4) of the input voice and the storage process in the queue (S31). The control unit 11 repeats the above process until the conference ends. That is, the control unit 11 continues the process of displaying the text T1 of the synthesized voice V1 and storing each input voice in the queue until the conference ends.

[0064] <Step S9> In step S9, the control unit 11 outputs the input voices stored in the queue to the conference server 4. Specifically, the control unit 11 outputs the multiple input voices to the conference server 4 so that the multiple input voices do not overlap with each other. For example, the control unit 11 outputs the input voices Va to Vd to the conference server 4 with a time shift so that the voice sections of the input voices do not overlap with each other. Furthermore, the control unit 11 outputs the input voices Va to Vd to the conference server 4 as a single voice (voice stream V2) (see FIG. 7).

[0065] The conference server 4 performs a character conversion process based on one audio stream V2 output from the audio processing device 1. The conference server 4 also generates text by associating the identification information of the audio device 2 or the identification information of the user. For example, the conference server 4 converts the input audio Va into text Ta and associates it with the identification information of user A, converts the input audio Vb into text Tb and associates it with the identification information of user B, converts the input audio Vc into text Tc and associates it with the identification information of user C, and converts the input audio Vd into text Td and associates it with the identification information of user D.

[0066] <Step S10> In step S10, control unit 11 determines whether or not text data has been acquired from conference server 4. If control unit 11 has acquired the text data (S10: Yes), control unit 11 shifts the process to step S11. Control unit 11 waits until it acquires the text data from conference server 4 (S10: No). For example, control unit 11 acquires data of texts T1 to T4 from conference server 4.

[0067] <Step S11> In step S11, the control unit 11 generates minutes based on the text data acquired from the conference server 4. For example, the control unit 11 generates minutes based on texts T1 to T4. The control unit 11 also generates minutes by assigning user identification information (user name) to each text. After generating the minutes, the control unit 11 ends the voice control process.

[0068] The control unit 11 executes the voice control process as described above each time a conference starts, displays the spoken voice as text in real time during the conference, and generates minutes based on the text information of the spoken voice when the conference ends. Note that, as another embodiment, the control unit 11 may generate minutes in parallel with displaying the spoken voice as text during the conference.

[0069] As described above, the speech processing system 100 according to the present disclosure acquires a plurality of input speeches input to the respective microphones of a plurality of speech devices 2, and synthesizes the acquired plurality of input speeches into one synthesized speech V1 (first speech). The speech processing system 100 also outputs the plurality of input speeches and the synthesized speech V1 to a text conversion device (e.g., conference server 4).

[0070] This allows the synthesized speech V1 to be converted into text in real time and displayed. Also, since the input speech input to the speech device 2 can be converted into text individually, the text conversion accuracy (speech recognition accuracy) of each input speech can be improved. Therefore, it is possible to convert the speech of a conversation using multiple speech devices 2 into text in real time and improve the accuracy of the text conversion.

[0071] In the speech recognition process, speech in a first language (e.g., Japanese) may be converted into text in the first language, or speech in a first language (e.g., Japanese) may be converted into text in a second language (e.g., English).

[0072] As another embodiment of the present disclosure, the control unit 11 may switch the process of individually recognizing (transcribes) input voices between the method shown in Fig. 6 and the method shown in Fig. 7. The method shown in Fig. 6 is a method in which each input voice Va to Vd input from each voice device 2 is individually output to the conference server 4 for voice recognition processing. In contrast, the method shown in Fig. 7 is a method in which each input voice input from each voice device 2 is stored in a queue, and output from the queue to the conference server 4 as a single voice stream V2 for voice recognition processing.

[0073] Specifically, the output processing unit 113 switches between a first output mode in which each of the multiple input audios Va to Vd is output individually to the conference server 4, and a second output mode in which audio stream V2 is output to the conference server 4, based on the number of input audios stored in the queue.

[0074] For example, the output processing unit 113 sets the first output mode when the number of input voices stored in the queue is equal to or greater than a predetermined number, and sets the second output mode when the number of input voices stored in the queue is less than the predetermined number.

[0075] In another embodiment, the output processing unit 113 may be set to the first output mode in a scene where one user mainly speaks and other users do not speak (such as a presentation), and may be set to the second output mode in a scene where multiple users are engaged in a question and answer session.

[0076] The control unit 11 of the voice processing device 1 controls the entire voice processing device 1. The control unit 11 realizes various functions by reading and executing various programs stored in the storage unit 12 (for example, storage or ROM). The control unit 11 may be realized by one or more control devices / arithmetic units (CPUs (Central Processing Units), SoCs (System on a Chip)). The control unit 11 may also be configured by one or more control circuits (electronic circuits).

[0077] [Disclosure Note] The following is a summary of the disclosure extracted from the above-described embodiment. Note that the configurations and processing functions described in the following supplementary notes can be selected and combined as desired.

[0078] <Appendix 1> an acquisition processing unit that acquires a plurality of input sounds input to the respective microphones of a plurality of audio devices; a synthesis processing unit that synthesizes the plurality of input voices acquired by the acquisition processing unit into one first voice; an output processing unit that converts the first speech synthesized by the synthesis processing unit into a first text and outputs the plurality of input speeches and the first speech to a conversion processing unit that converts each of the plurality of input speeches into a plurality of second texts; A voice processing system comprising:

[0079] <Appendix 2> the output processing unit outputs the plurality of input voices to the conversion processing unit after the process of converting the first voice into the first text is completed. 10. The speech processing system of claim 1.

[0080] <Appendix 3> the output processing unit outputs the first voice to the conversion processing unit during a predetermined period in which the acquisition processing unit acquires the plurality of input voices, and outputs the plurality of input voices to the conversion processing unit after the predetermined period has elapsed. 3. The speech processing system according to claim 1 or 2.

[0081] <Appendix 4> the output processing unit outputs the plurality of input voices to the conversion processing unit so that the plurality of input voices do not overlap with each other. 4. A speech processing system according to any one of Supplementary notes 1 to 3.

[0082] <Appendix 5> the acquisition processing unit stores the plurality of input voices in a storage unit in order of acquisition time so as not to overlap with each other; the output processing unit collectively outputs the plurality of input voices stored in the storage unit to the conversion processing unit; 5. A speech processing system according to any one of Supplementary notes 1 to 4.

[0083] <Appendix 6> the acquisition processing unit associates each of the plurality of input voices with identification information of the voice device and stores the input voices in the storage unit; 6. The speech processing system of claim 5.

[0084] <Appendix 7> the acquisition processing unit performs predetermined speech processing on each of the plurality of input speeches and stores the processed speeches in the storage unit. 7. The speech processing system according to claim 5 or 6.

[0085] <Appendix 8> the output processing unit switches between a first output mode in which each of the plurality of input voices is output individually to the conversion processing unit and a second output mode in which the plurality of input voices stored in the storage unit are output collectively to the conversion processing unit, based on the number of the input voices stored in the storage unit. 8. A speech processing system according to any one of Supplementary notes 5 to 7.

[0086] <Appendix 9> displaying the first text into which the first voice has been converted by the conversion processing unit while the user is speaking; generating minutes of the user's conversation based on the plurality of second texts into which the plurality of input voices are converted by the conversion processing unit; 9. A speech processing system according to any one of Supplementary notes 1 to 8.

[0087] <Appendix 10> Acquiring a plurality of input sounds input to respective microphones of a plurality of audio devices; synthesizing the acquired input voices into one first voice; outputting the plurality of input speeches and the first speech to a conversion processor that converts the first speech into a first text and converts each of the plurality of input speeches individually into a plurality of second texts; An audio processing method executed by one or more processors.

[0088] <Appendix 11> Acquiring a plurality of input sounds input to respective microphones of a plurality of audio devices; synthesizing the acquired input voices into one first voice; outputting the plurality of input speeches and the first speech to a conversion processor that converts the first speech into a first text and converts each of the plurality of input speeches individually into a plurality of second texts; An audio processing program for causing one or more processors to execute the above. [Explanation of symbols]

[0089] 100: Audio processing system 1: Audio processing device 2: Audio equipment 3: User device 4: Conference Server 5:Display device 11: Control section 12: Storage section 13: Communications Department 111: Acquisition processing unit 112: Composition processing unit 113: Output processing section 114: Display processing unit 115: Generation processing unit D1:Device information D2: Audio information

Claims

1. an acquisition processing unit that acquires a plurality of input sounds input to the respective microphones of a plurality of audio devices; a synthesis processing unit that synthesizes the plurality of input voices acquired by the acquisition processing unit into one first voice; an output processing unit that converts the first speech synthesized by the synthesis processing unit into a first text and outputs the plurality of input speeches and the first speech to a conversion processing unit that converts each of the plurality of input speeches into a plurality of second texts; A voice processing system comprising:

2. the output processing unit outputs the plurality of input voices to the conversion processing unit after the process of converting the first voices into the first text is completed. The audio processing system of claim 1 .

3. the output processing unit outputs the first voice to the conversion processing unit during a predetermined period in which the acquisition processing unit acquires the plurality of input voices, and outputs the plurality of input voices to the conversion processing unit after the predetermined period has elapsed. The audio processing system of claim 1 .

4. the output processing unit outputs the plurality of input voices to the conversion processing unit so that the plurality of input voices do not overlap with each other. The audio processing system of claim 1 .

5. the acquisition processing unit stores the plurality of input voices in a storage unit in order of acquisition time so as not to overlap with each other; the output processing unit collectively outputs the plurality of input voices stored in the storage unit to the conversion processing unit; The audio processing system of claim 1 .

6. the acquisition processing unit associates each of the plurality of input voices with identification information of the voice device and stores the input voices in the storage unit; The audio processing system of claim 5 .

7. the output processing unit switches between a first output mode in which each of the plurality of input voices is output individually to the conversion processing unit and a second output mode in which the plurality of input voices stored in the storage unit are output collectively to the conversion processing unit, based on the number of the input voices stored in the storage unit. The audio processing system of claim 5 .

8. displaying the first text obtained by converting the first voice by the conversion processing unit while the user is speaking; generating minutes of the user's conversation based on the plurality of second texts into which the plurality of input voices are converted by the conversion processing unit; The voice processing system according to any one of claims 1 to 7.

9. Acquiring a plurality of input sounds input to respective microphones of a plurality of audio devices; synthesizing the acquired input voices into one first voice; outputting the plurality of input speeches and the first speech to a conversion processor that converts the first speech into a first text and converts each of the plurality of input speeches individually into a plurality of second texts; An audio processing method executed by one or more processors.

10. Acquiring a plurality of input sounds input to respective microphones of a plurality of audio devices; synthesizing the acquired input voices into one first voice; outputting the plurality of input speeches and the first speech to a conversion processor that converts the first speech into a first text and converts each of the plurality of input speeches individually into a plurality of second texts; An audio processing program for causing one or more processors to execute the above.

Citation Information

Patent Citations

  • Conference system

    JP2023037813A