Speech recognition system, speech recognition method, and information storage medium

The speech recognition system addresses the challenge of multilingual conversations by using a buffer and dynamic language detection to provide real-time translation and display, improving usability and accuracy in language recognition.

WO2025197100A1PCT designated stage Publication Date: 2025-09-25POCKETALK CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/011427
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Conventional translation devices struggle with speedy conversations between speakers of different languages due to difficulty in operating language-specific buttons during simultaneous speech.

Method used

A speech recognition system that includes a buffer storage mechanism to acquire partial speech data, a character string acquisition mechanism for language-independent speech recognition, and a language determination mechanism to identify and switch languages dynamically, enabling seamless translation and display of recognized speech in real-time.

Benefits of technology

Facilitates quick and accurate speech recognition and translation between multiple languages, allowing users to understand conversations in their own language without manual button operation, enhancing usability in multilingual interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024011427_25092025_PF_FP_ABST
    Figure JP2024011427_25092025_PF_FP_ABST
Patent Text Reader

Abstract

This present invention can be applied to a conversation between speakers of different languages and execute correct speech recognition in the respective languages. An input reception unit (110) sequentially acquires and stores, in a buffer (132), a plurality of pieces of partial speech data indicating a first statement and a subsequent second statement. A speech recognition unit (120) performs a speech recognition process on one or more speech packets indicating the first statement stored in the buffer. When the end part of the first statement is recognized in the speech recognition process, an end part identification unit (124) identifies speech packets related to the end part of the first statement in the buffer. A language determination unit (140): acquires, on the basis of at least a portion of speech packets that were stored in the buffer after the identified speech packets and that indicate the second statement, language determination information for determining the language related to the speech packets; and determines one language on the basis of said information. The speech recognition unit (120) performs the speech recognition process on the speech packets indicating the second statement by using the one determined language as a recognition language.
Need to check novelty before this filing date? Find Prior Art

Description

Speech recognition system, speech recognition method, and information storage medium

[0001] The present invention relates to a voice recognition system, a voice recognition method, and an information storage medium.

[0002] Patent Literature 1 discloses a two-way speech translator that generates translated text and translated speech in a second language from speech in a first language, and generates translated text and translated speech in the first language from speech in a second language. The two-way speech translator performs speech recognition on speech input while a button on the terminal is pressed, displays the results on the terminal, and generates translated text and translated speech from the speech recognition results when the button on the terminal is released.

[0003] JP 2023-084986 A

[0004] Displaying the results of speech recognition on a terminal, as in the two-way speech translator described in Patent Document 1, allows the speaker to check whether their speech has been correctly recognized, improving convenience. However, with the above-mentioned conventional translation devices, when a speaker of one language speaks immediately after a speaker of a different language, it becomes extremely difficult to operate the buttons on the terminal. This has resulted in the problem that the above-mentioned conventional translation devices are difficult to use for speedy conversations between speakers of different languages.

[0005] The present invention has been made in view of the above-mentioned problems, and its object is to provide a speech recognition system, a speech recognition method, and an information storage medium that can be used for speedy conversations between speakers of different languages ​​and that can quickly obtain correct speech recognition results in each language during the conversation.

[0006] (1) A speech recognition system according to the present invention includes a buffer storage means for sequentially acquiring a plurality of partial speech data each representing a part of a first utterance by a first speaker and a subsequent second utterance by a second speaker, and storing the acquired data in a buffer; a character string acquisition means for performing speech recognition processing on one or more of the partial speech data representing the first utterance stored in the buffer using a given language as a recognition language, and acquiring a character string representing the content of the utterance represented by the partial speech data; an end portion identification means for, when an end portion of the first utterance is recognized in the speech recognition processing, identifying the partial speech data relating to the end portion of the first utterance in the buffer; and a signal processing means for detecting the partial speech data relating to the end portion of the first utterance from a previous speech. and a language determination means for, when the partial audio data is identified in the buffer, acquiring language determination information for determining the language associated with the partial audio data based on at least a portion of one or more partial audio data indicating the second utterance that were stored in the buffer after the identified partial audio data, and determining one language based on the language determination information, wherein the string acquisition means performs voice recognition processing on the one or more partial audio data indicating the second utterance that are stored in the buffer, using the one language determined by the language determination means as the recognition language, and acquires a string in the one language that indicates the utterance content indicated by the partial audio data.

[0007] (2) The speech recognition system described in (1) may further include a language determination information acquisition means for sequentially acquiring language determination information for determining the language of the speech content indicated by the partial speech data based on a predetermined number of consecutive partial speech data acquired from one or more partial speech data stored in the buffer while shifting the start timing sequentially, and a recognition language change means for changing the recognition language based on the language determination information acquired by the language determination means and the language determination information acquired sequentially by the language determination information acquisition means.

[0008] (3) In the speech recognition system described in (2), when the recognition language change means changes the recognition language, the character string acquisition means may perform speech recognition processing on the plurality of partial speech data stored in the buffer in the changed recognition language, thereby acquiring character strings in the changed recognition language that indicate the utterance contents indicated by the partial speech data, instead of character strings in the one language.

[0009] (4) In the speech recognition system described in (2) or (3), the buffer storage means may sequentially acquire the plurality of partial speech data at a predetermined acquisition interval. The character string acquisition means may perform speech recognition processing on one or more of the partial speech data representing the first utterance at a predetermined recognition interval longer than the predetermined acquisition interval. The language determination information acquisition means may sequentially acquire the language determination information at a predetermined determination interval longer than the acquisition interval but shorter than the recognition interval.

[0010] (5) In the speech recognition system according to any one of (1) to (4), when the speech recognition process recognizes the end of the first utterance, the end portion identification means may perform speech recognition in the recognition language on the plurality of partial speech data representing the entire first utterance stored in the buffer, and may acquire a character string representing the content of the utterance represented by the partial speech data.

[0011] (6) In the speech recognition system described in any one of (1) to (5), the end portion identification means may delete the multiple partial speech data representing the first utterance stored in the buffer when the end portion of the first utterance is recognized in the speech recognition process.

[0012] (7) In the speech recognition system described in any one of (1) to (6), the end portion identification means may identify the end portion of the first utterance by detecting a predetermined delimiter symbol from the character string obtained by the speech recognition processing.

[0013] (8) A speech recognition method according to the present invention includes the steps of sequentially acquiring a plurality of partial speech data each representing a part of a first utterance by a first speaker and a subsequent second utterance by a second speaker, and storing the data in a buffer; performing speech recognition processing on one or more of the partial speech data representing the first utterance stored in the buffer using a given language as a recognition language, and acquiring a character string representing the utterance content represented by the partial speech data; when the end of the first utterance is recognized in the speech recognition processing, identifying the partial speech data relating to the end of the first utterance in the buffer; and When data is identified in the buffer, the method includes a step of acquiring language determination information for determining the language associated with the partial audio data based on at least a part of the one or more partial audio data indicating the second utterance that were stored in the buffer after the identified partial audio data, and determining one language based on the language determination information; and a step of performing speech recognition processing on the one or more partial audio data indicating the second utterance that are stored in the buffer using the determined one language as a recognition language, and acquiring a character string in the one language that indicates the utterance content indicated by the partial audio data.

[0014] (9) A program according to the present invention includes a buffer storage means for sequentially acquiring a plurality of partial voice data each representing a part of a first utterance by a first speaker and a subsequent second utterance by a second speaker, and storing the acquired data in a buffer; a character string acquisition means for performing a speech recognition process on one or more of the partial voice data representing the first utterance stored in the buffer using a given language as a recognition language, and acquiring a character string representing the content of the utterance represented by the partial voice data; an end portion identification means for, when an end portion of the first utterance is recognized in the speech recognition process, identifying the partial voice data relating to the end portion of the first utterance in the buffer; and a program for identifying the partial voice data relating to the end portion in the buffer. a program for causing a computer to function as a language determination means that, in a case where the specified partial voice data is stored in the buffer after the specified partial voice data, acquires language determination information for determining the language associated with the one or more partial voice data indicating the second utterance, based on at least a part of the partial voice data, and determines a single language based on the language determination information, wherein the character string acquisition means performs speech recognition processing on the one or more partial voice data indicating the second utterance that are stored in the buffer, using the one language determined by the language determination means as a recognition language, and acquires a character string in the one language that indicates the utterance indicated by the partial voice data. This program may be stored on a non-transitory, tangible computer-readable information storage medium.

[0015] FIG. 1 is an overall configuration diagram of a speech recognition system according to an embodiment of the present invention. FIG. 2 is a diagram showing an example of a display screen according to an embodiment of the present invention. FIG. 3 is a diagram showing an example of a display screen according to an embodiment of the present invention. FIG. 4 is a diagram showing a hardware configuration diagram of a speech recognition device according to an embodiment of the present invention. FIG. 5 is a diagram showing a hardware configuration diagram of a client device according to an embodiment of the present invention. FIG. 6 is a functional block diagram of a speech recognition system according to an embodiment of the present invention. FIG. 7 is a diagram showing a chronological correspondence between speech data stored in a buffer and speech recognition text stored in a speech recognition text storage unit. FIG. 8 is a diagram explaining a provisional language determination process when an end portion (delimiter) is recognized in speech recognition text. FIG. 9 is a diagram showing the stored contents of a language information storage unit.

[0016] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.

[0017] FIG. 1 is a diagram showing the overall configuration of a speech recognition system with a translation function according to this embodiment.

[0018] 1, a speech recognition system with a translation function 1 according to this embodiment includes a speech recognition engine 10, a language determination engine 20, a translation engine 30, a speech synthesis engine 40, a client device 50, a microphone 70, and a speech recognition device 100. The speech recognition engine 10, the language determination engine 20, the translation engine 30, the speech synthesis engine 40, the client device 50, and the speech recognition device 100 are connected to a computer network 80 such as the Internet. The speech recognition engine 10, the language determination engine 20, the translation engine 30, the speech synthesis engine 40, the client device 50, and the speech recognition device 100 are capable of communicating with each other via the computer network 80.

[0019] The speech recognition system 1 according to this embodiment is used in a situation where multiple speakers each speak a different language. For example, as shown in FIG. 1 , the speech recognition system 1 is used when a user UE who speaks in English and a user UJ who speaks in Japanese are having a conversation.

[0020] The speech recognition engine 10 is, for example, a computer system such as a server computer that performs speech recognition on utterances contained in speech input from a microphone 70. The speech recognition engine 10 may be composed of a single computer or multiple computers. The speech recognition engine 10 includes speech recognition engines for multiple candidate languages. The speech recognition engine 10 includes, for example, a Japanese speech recognition engine and an English speech recognition engine. The Japanese speech recognition engine is used to recognize utterances made by a user UJ, and the English speech recognition engine is used to recognize utterances made by a user UE.

[0021] The language determination engine 20 is a computer system such as a server computer that generates language determination information that serves as the basis for determining the language in which a utterance contained in a voice input from the microphone 70 is spoken. The language determination information is information indicating the likelihood that the voice was spoken in each of a plurality of candidate languages, and includes probability information for each candidate language. The probability information is, for example, a numerical value within a predetermined numerical range (e.g., from 0 to 1). For example, a larger numerical value indicates a higher probability that the language is that candidate language. The language determination engine 20 may be composed of a single computer or multiple computers.

[0022] The translation engine 30 is a computer system, such as a server computer, that executes a process of translating text in a first language (source text) into text in a second language (target text). The first language and the second language may be various. The translation engine 30 may include, for example, an English-Japanese translation engine in which English is the first language and Japanese is the second language, or conversely, a Japanese-English translation engine in which Japanese is the first language and English is the second language. The Japanese-English translation engine is used to translate utterances of user UJ, and the English-Japanese translation engine is used to translate utterances of user UE. The translation engine 30 may be composed of a single computer or multiple computers.

[0023] The speech synthesis engine 40 is, for example, a computer system such as a server computer that executes speech processing such as speech synthesis. The speech synthesis engine 40 may be composed of a single computer or multiple computers. The speech synthesis engine 40 also includes speech synthesis engines for various languages. For example, it may include a Japanese speech synthesis engine and an English speech synthesis engine. The synthetic Japanese speech generated by the Japanese speech synthesis engine is mainly used by user UJ, and the synthetic English speech generated by the English speech synthesis engine is mainly used by user UE.

[0024] The speech recognition engine 10, the language determination engine 20, the translation engine 30, and the speech synthesis engine 4 may be publicly known and generally available API services on the Internet, regardless of the entity that operates them.

[0025] The client device 50 according to this embodiment is a computer used by a user. For example, the client device 50 is a smartphone, a tablet terminal, or a personal computer. A microphone 70 is connected to the client device 50, and speech captured by the microphone 70 is transmitted to the speech recognition device 100 via a computer network 80.

[0026] The microphone 70 is connected to the client device 50 and generates audio data representing the speaker's voice. The audio data includes a large number of time-ordered audio packets (partial audio data) generated every short time (e.g., 20 ms). Each audio packet stores audio data for a short time (e.g., 20 ms). The microphone 70 may be separate from the client device 50 or may be built into the client device 50.

[0027] The speech recognition device 100 with translation function according to this embodiment is a computer system such as a server computer, which recognizes speech data generated by a microphone 70 in an appropriate language in parts in a situation where multiple speakers speak different languages. It also translates speech-recognized text generated by speech recognition and performs speech synthesis on the translated speech-recognized text. The speech recognition device 100 may be configured from a single computer or multiple computers.

[0028] A user of the speech recognition system 1 operates the client device 50 in advance to set the language (target language) in which the translation result will be displayed. For example, in the case of FIG. 1 , English and Japanese, which are the languages ​​used by users UE and UJ, are set as the target languages. Note that instead of manually setting the target language by the user, speech data acquired by the microphone 70 may be sent to the language determination engine 20, and the language of the user around the microphone 70 may be determined based on language determination information returned from the language determination engine 20. Then, some or all of the determined language may be used as the target language.

[0029] The client device 50 transmits speech data acquired by the microphone 70 to the speech recognition device 100. The speech acquired by the microphone 70 may include speeches in different languages. The speech recognition device 100 transmits a portion of the speech to the language determination engine 20, thereby obtaining language determination information for determining in what language the portion was spoken.

[0030] 1 , the speech acquired by the microphone 70 may include English speech uttered by the user UE and Japanese speech uttered by the user UJ. When the speech recognition device 100 transmits the speech acquired by the microphone 70 to the language determination engine 20, language determination information indicating a high probability of English for the time period when the user UE is speaking English, and language determination information indicating a high probability of Japanese for the time period when the user UJ is speaking Japanese is acquired.

[0031] The speech recognition device 100 determines the recognition language based on the language determination information thus obtained, transmits the determined recognition language to the speech recognition engine 210, and obtains speech recognition text, which is the speech recognition result. In the example of FIG. 1 , for the speech uttered by user UJ, Japanese is preferably determined as the recognition language, and the speech data is transmitted to the Japanese speech recognition engine. In this way, Japanese speech recognition text is obtained. Furthermore, for the speech uttered by user UE, English is preferably determined as the recognition language, and the speech data is transmitted to the English speech recognition engine. In this way, English speech recognition text is obtained.

[0032] Next, the speech recognition device 100 transmits the acquired speech-recognized text to the translation engine 30 and acquires translated text. As described above, the translation target language is set in advance by the user, and a language other than the language related to the speech-recognized text is selected from the set translation target languages ​​as the actual translation target language. In the example of FIG. 1 , the speech recognition device 100 transmits the English speech-recognized text to the English-Japanese translation engine 30 for translation into Japanese, which is a language other than English among the translation target languages ​​set by the user. This acquires translated English text. Furthermore, the speech recognition device 100 transmits the Japanese speech-recognized text to the Japanese-English translation engine 30 for translation into English, which is a language other than Japanese among the translation target languages ​​set by the user. This acquires translated Japanese text.

[0033] The speech recognition device 100 transmits the translated text to a speech synthesis engine 40 of the language associated with the translated text, and acquires translated speech. In the example of Fig. 1, the speech recognition device 100 transmits the translated English text to the English speech synthesis engine 40, and the translated Japanese text to the Japanese speech synthesis engine 40. In this way, translated speech in English and Japanese is acquired.

[0034] The speech recognition device 100 then acquires the speech-recognized text, the translated text, and the translated speech, and sequentially transmits them to the client device 50. The client device 50 displays the speech-recognized text and the translated text, allowing the user to check the speech recognition results and the translation results. The user can also check the translation results by listening to the translated speech using a speech output device (not shown) connected to the client device 50.

[0035] 2 shows an example of a display screen 200 according to an embodiment of the present invention. The display screen 200 includes a first language display area 202, a second language display area 204, and a speech recognition text display area 206.

[0036] The text in the first language is displayed in chronological order in the first language display area 202. In the example of Fig. 2, the first language is English, which is one of the languages ​​set as translation target languages ​​by the user, and the English speech-recognized text and translated text are displayed in chronological order in the first language display area 202.

[0037] The text in the second language is displayed in chronological order in the second language display area 204. In the example of Fig. 2, the second language is Japanese, which is another language set by the user as the translation target language, and the Japanese speech-recognized text and translated text are displayed in chronological order in the second language display area 204.

[0038] The first language display area 202 and the second language display area 204 each display the speech-recognized text and the translated text in one language. Therefore, user UE can check all utterances in their own language simply by looking at the first language display area 202, and user UJ can check all utterances in their own language simply by looking at the second language display area 204. Furthermore, on the display screen 200, only the translated text may be displayed in bold, as shown in FIG. 2. By changing the display mode of the speech-recognized text and the translated text, the utterances of user UJ and user UE can be displayed distinctly.

[0039] Speech-recognition text is displayed in the speech-recognition text display area 206. While user UE is speaking, speech-recognition text in English is displayed in the speech-recognition text display area 206. While user UJ is speaking, speech-recognition text in Japanese is displayed in the speech-recognition text display area 206.

[0040] 2 displays "Exciting, isn't it?" along with a sentence delimiter "?". When the speech recognition device 100 detects a predetermined delimiter in the speech recognition text during speech recognition, it transmits the speech recognition text before the detection of the delimiter to the translation engine 30 as the confirmed speech recognition text.

[0041] In the example of FIG. 2 , a delimiter "?" is detected from the English speech-recognized text displayed in the speech-recognized text display area 206 of FIG. 2 , and the speech-recognized text is confirmed sentence by sentence. Then, the speech recognition device 100 transmits the confirmed English speech-recognized text sentence by sentence to the English-Japanese translation engine 30, thereby obtaining translated Japanese text. The speech recognition device 100 transmits the translated Japanese text to the speech synthesis engine 40, thereby obtaining translated Japanese speech. Then, the speech recognition device 100 transmits the confirmed English speech-recognized text, translated Japanese text, and translated Japanese speech to the client device 50.

[0042] 3 shows an example of the display screen 200 that is displayed after the display screen 200 of FIG. 2 is displayed. In FIG. 3, the confirmed speech-recognition text that was displayed in the speech-recognition text display area 206 of FIG. 2 is displayed at the bottom of the first language display area 202. Also, translated Japanese text obtained by translating the confirmed speech-recognition text into Japanese is displayed at the bottom of the first language display area 202. In this way, each time a delimiter is detected in the speech-recognition text, the speech-recognition text is confirmed, and the confirmed speech-recognition text and its translated text are moved to the bottom of the first language display area 202 or the bottom of the second language area 204.

[0043] That is, when the English speech-recognition text is confirmed by detecting delimiters, the confirmed speech-recognition text is moved to the bottom of the first language display area 202, and its translation, the translated text, is displayed at the bottom of the second language display area 204. Similarly, when the Japanese speech-recognition text is confirmed by detecting delimiters, the confirmed speech-recognition text is moved to the bottom of the second language display area 204, and its translation, the translated text, is displayed at the bottom of the first language display area 204. The translated text is displayed in bold letters to make it stand out, while the confirmed speech-recognition text is displayed in normal bold letters. In FIG. 3 , the speech-recognition text display area 206 displays up to the middle of the Japanese speech-recognition text, which is the next utterance. As time passes, i.e., according to the utterance of user UJ, the number of characters in the speech-recognition text display area 206 increases.

[0044] By doing so, even when users who speak different languages ​​are conversing with each other, they can understand the content of the conversation in their own language. Furthermore, the speech-recognition text display area 206 displays speech-recognition text indicating the current utterance in real time, allowing the speaker to confirm that their utterance has been correctly recognized. By displaying the speech-recognition text in real time in the speech-recognition text display area 206, a user who is not good at listening to languages ​​other than their own but is good at reading sentences can quickly understand the content of the other person's utterance. For example, if user UJ is not good at listening to English but is good at reading English sentences, user UJ can quickly understand what user UE is saying by reading the speech-recognition text indicating user UE's utterance displayed in the speech-recognition text display area 206.

[0045] Fig. 4A is a diagram showing an example of the hardware configuration of a speech recognition device 100 according to an embodiment of the present invention. The speech recognition device 100 according to this embodiment is a computer as shown in Fig. 4A. Software installed on the computer realizes a function of determining the language of each speech input containing utterances in multiple languages ​​and recognizing the speech. As shown in Fig. 4A, the example of the speech recognition device 100 according to the embodiment includes, for example, a processor 100a, a storage unit 100b, and a communication unit 100c.

[0046] The processor 100a is, for example, a program-controlled device such as a microprocessor that operates according to a program installed in the speech recognition device 100. The storage unit 100b is, for example, a storage element such as a ROM or RAM, a solid-state drive, or a hard disk drive. The storage unit 100b stores programs to be executed by the processor 100a. The communication unit 100c is, for example, a communication interface for exchanging data between the client device 50, the speech recognition engine 10, the language determination engine 20, the translation engine 30, and the speech synthesis engine 40 via the computer network 80.

[0047] Fig. 4B is a diagram showing an example of the hardware configuration of a client device 50 according to an embodiment of the present invention. The client device 50 according to this embodiment is a computer as shown in Fig. 4A. Software installed on the computer realizes a function of determining the language of each speech input containing speech in multiple languages ​​and performing speech recognition. As shown in Fig. 4A, the example client device 50 according to the embodiment includes, for example, a processor 50a, a storage unit 50b, a communication unit 50c, an operation unit 50d, a microphone 50f, and a speaker 50g.

[0048] The processor 50a is a program-controlled device such as a microprocessor that operates according to a program installed in the client device 50. The storage unit 50b is a storage element such as a ROM or RAM, a solid-state drive, or a hard disk drive. The storage unit 50b stores programs executed by the processor 50a. The communication unit 50c is a communication interface for exchanging data with the speech recognition device 100 via the computer network 80. The operation unit 50d is a user interface such as a keyboard or mouse that accepts user input and outputs signals indicating the content of the input to the processor 50a. The display unit 50e is a display device such as a liquid crystal display that displays various images according to instructions from the processor 50a. The microphone 50f is an audio input device that converts received audio into an electrical signal. The speaker 50g is an audio output device that outputs audio.

[0049] 5 is a functional block diagram of the speech recognition device 100 and the client device 50 according to an embodiment of the present invention. Note that the speech recognition system 100 and the client device 50 according to this embodiment do not need to implement all of the functions shown in FIG. 5, and functions other than those shown in FIG. 5 may also be implemented.

[0050] 5, the speech recognition device 100 functionally includes an input receiving unit 110, a speech recognition unit 120, a buffer 132, a language information storage unit 134, a speech recognition text storage unit 136, a language determination unit 140, a translation unit 150, a speech synthesis unit 160, and a transmission unit 170. The input receiving unit 110 and the transmission unit 170 are implemented mainly in the communication unit 100c. The speech recognition unit 120, the language determination unit 140, the translation unit 150, and the speech synthesis unit 160 are implemented mainly in the processor 100a and the communication unit 100c. The buffer 132, the language information storage unit 132, and the speech recognition text storage unit 134 are implemented mainly in the storage unit 100b.

[0051] The above functions are implemented by executing a program containing instructions corresponding to the above functions on the processor 100a, which is installed in the speech recognition device 100, which is a computer. This program is supplied to the translation and distribution device 100 via a computer-readable information storage medium, such as an optical disk, a magnetic disk, a magnetic tape, a magneto-optical disk, or a flash memory, or via the Internet, for example.

[0052] 5, the client device 50 functionally includes an audio input receiving unit 52, a setting receiving unit 54, an input transmitting unit 56, an output receiving unit 58, a display control unit 60, and an audio output control unit 62. The audio input receiving unit 52 is implemented mainly using a processor 50a and a microphone 50f. The setting receiving unit 54 is implemented mainly using a processor 50a, an operation unit 50d, and a display unit 50e. The input transmitting unit 56 and the output receiving unit 58 are implemented mainly using a communication unit 50c. The display control unit 60 is implemented mainly using a processor 50a and a display unit 50e. The audio output control unit 62 is implemented mainly using a processor 50a and a speaker 50g.

[0053] The above functions are implemented by executing a program including instructions corresponding to the above functions on the processor 50a, which is installed in the client device 50, which is a computer. This program is supplied to the client device 50 via a computer-readable information storage medium such as an optical disk, a magnetic disk, a magnetic tape, a magneto-optical disk, or a flash memory, or via the Internet, for example.

[0054] In this embodiment, the voice input receiving unit 52 of the client device 50 receives voice data input by the user's speech via, for example, the microphone 50 f. As described above, the voice data includes time-series voice packets (partial voice data), and the voice input receiving unit 52 receives a large number of voice packets in chronological order.

[0055] In this embodiment, for example, the setting receiving unit 54 receives in advance the setting of the translation target language input by the user via the operation unit 50d after viewing the image displayed on the display unit 50e.

[0056] In this embodiment, for example, the input transmitting unit 56 sequentially transmits the voice packets acquired by the voice input accepting unit 52 to the voice recognition device 100. The input transmitting unit 56 also transmits target language information indicating the target language accepted by the setting accepting unit 54 to the voice recognition device 100 in advance.

[0057] In this embodiment, the input accepting unit 110 of the speech recognition device 100 receives, for example, voice packets transmitted from the input transmitting unit 56. The received voice packets are temporarily stored in a receiving buffer (not shown). The input accepting unit 110 reads one or more voice packets stored in the receiving buffer at a predetermined acquisition interval (for example, every 20 ms) and stores them in the buffer 132. As a result, the voice packets are stored in chronological order in the buffer 132. The input accepting unit 110 is an example of a buffer storage means. The input accepting unit 110 also receives target language information transmitted from the input transmitting unit 56 and stores it in the memory unit 100b.

[0058] The speech recognition unit 120 includes a speech recognition processing unit 122 and an end identification unit 124. When a speech recognition text for the immediately preceding utterance has been determined and a speech recognition text for the next utterance is to be generated, the speech recognition processing unit 122 acquires all speech packets stored in the buffer 132 at a predetermined recognition interval (e.g., every 300 ms) and transmits them to the speech recognition engine 10. In this way, the speech recognition text is acquired from the speech recognition engine 10. The processing of the speech recognition processing unit 122 is executed asynchronously with the processing of storing speech packets in the buffer 132 by the input receiving unit 110 and the periodic determination processing of the periodic determination unit 144, which will be described later.

[0059] As will be described later, when the speech recognition text for the immediately preceding utterance is finalized, the speech packets for that immediately preceding utterance are deleted from the buffer 132. Therefore, the buffer 132 stores the speech packets from the beginning of the utterance to the last speech packet stored in the buffer 132, and these speech packets are transmitted to the speech recognition engine 10. Since the number of speech packets transmitted to the speech recognition engine 10 increases over time, the speech recognition text generated by the speech recognition engine 10 becomes longer over time. Furthermore, the accuracy of speech recognition improves over time.

[0060] When transmitting a voice packet to the voice recognition engine 10, the voice recognition processing unit 122 references the latest language information stored in the language information storage unit 134. The language information includes information on the latest recognized language at that time, and the voice recognition processing unit 122 can identify the recognized language from the language information. The voice recognition processing unit 122 transmits the voice packet to the voice recognition engine 10 corresponding to the identified recognized language. The voice recognition processing unit 122 stores the voice-recognition text acquired from the voice recognition engine 10 in the voice-recognition text storage unit 136. The transmission unit 170 reads the voice-recognition text stored in the voice-recognition text storage unit 136 at regular time intervals (e.g., every 300 m) and transmits it to the client device 50. The output acceptance unit 58 of the client device 50 receives the voice-recognition text, and the display control unit 60 displays the voice-recognition text in the voice-recognition text display area 206. Furthermore, when a delimiter is detected in the voice-recognition text and a sentence of voice-recognition text is determined, the translation unit 150 causes the translation engine 30 to translate the determined voice-recognition text. The transmitting unit 170 transmits this translated text to the client device 50. The output receiving unit 58 of the client device 50 receives the translated text, and the display control unit 60 displays the translated text in either the first language display area 202 or the second language display area 204. At this time, the display control unit 60 displays the confirmed speech-recognized text in the other of the first language display area 202 or the second language display area. The speech synthesis unit 160 also transmits the translated text to the speech synthesis engine 40 and receives translated speech from the speech synthesis engine 40. This translated speech is also transmitted to the client device 50 by the transmitting unit 170. The output receiving unit 58 of the client device 50 receives the translated speech, and the speech output control unit 62 outputs it from the speaker 50g.

[0061] 6 shows the processing of the speech recognition processing unit 122, with B1 to B6 indicating the contents of the buffer 132 at each predetermined recognition interval. Specifically, the length of the white bar indicates the total size of the voice packets stored in the buffer 132, and above it, for reference, the voice content indicated by those voice packets. R1 to R6 correspond to B1 to B6, respectively, and indicate the contents of the speech recognition text storage unit 136 at each predetermined recognition interval. By sending the contents of the buffer 132 indicated by Bn to the speech recognition engine 10, the contents of the speech recognition storage unit 136 are updated as indicated by Rn (n = 1, 2, 3, 4, 5, 6, ...).

[0062] That is, the speech processing unit 122 transmits all of the contents (speech packets) of the buffer 132 indicated by Bn to the speech recognition engine 10. Thereafter, the speech recognition text is received from the speech recognition engine 10 and stored in the speech recognition text storage unit 136 as indicated by Rn. As indicated by Rn, the speech recognition text storage unit 136 stores not only the speech recognition text but also the recognition language and time code used when recognizing the speech recognition text. The recognition language is specified by the speech recognition processing unit 122 from the language information stored in the language information storage unit 134 in order to select the speech recognition engine 10. The time code is received from the speech recognition engine 10 along with the speech recognition text. The time code is information generated during the speech recognition process in the speech recognition engine 10 and includes the start time of each recognition unit (e.g., syllable) included in the speech recognition text and the end time of the speech recognition text.

[0063] For example, at the timing shown in B3, the speech recognition text "Exciting, isn't it?" is obtained from the speech recognition engine 10, along with the start time T1 of "Exciting" and the start time T2 and end time T3 of "isn't it?". This information is stored as time codes in the speech recognition text storage unit 136 as shown in R3. "English," the recognition language used in speech recognition, is also stored in the speech recognition text 136 as shown in R3.

[0064] The end identifying unit 124 monitors the speech-recognized text stored in the speech-recognized text storage unit 136 and recognizes the end of each utterance. As an example, the end identifying unit 124 determines whether the speech-recognized text contains a predetermined delimiter. One or more delimiters are predefined for each language. For example, a period "." is defined for Japanese, and a period "." and a question mark "?" are defined for English. When the speech-recognized text contains a delimiter predefined for the language associated with the speech-recognized text, the end identifying unit 124 reads the time of the delimiter from the speech-recognized text storage unit 136. In the example of FIG. 6 , the end identifying unit 124 can detect the delimiter "?" defined for English, the recognition language, in the speech-recognized text at timing R3. In this case, the end identifying unit 124 determines the current utterance as a sentence unit. Specifically, the end identification unit 124 obtains T3, which is the time of the delimiter "?" (the end time of the speech-recognition text), from the time code stored in the speech-recognition text storage unit 136. The end identification unit 124 then identifies end address A, which is the address of the buffer 132 corresponding to time T3 of the delimiter "?". Because of the round-trip communication time between the speech recognition device 100 and the speech recognition engine 10 and the processing time in the speech recognition engine 10, a certain amount of time has passed since the speech packet was transmitted at B3 when the end identification unit 124 detected the delimiter in the speech-recognition text. Therefore, as shown in FIG. 7 as an example, the contents of the buffer 132 are larger than the contents of the buffer 132 at time B3. As shown in FIG. 7, the end identification unit 124 uses end address A to delete from the buffer 132 the speech packets (the shaded portion in FIG. 7) that were stored in the buffer 132 before time T4. As a result, only the speech packets of the next utterance are stored in the buffer 132 in chronological order, starting from the beginning. Note that the buffer 132 shown in FIG. 7 is immediately after the timing of B5 in FIG. 6 and well before the timing of B6.

[0065] If the end identification unit 124 detects a delimiter in the speech-recognition text, it notifies the translation unit 150, and the translation unit 150 reads the speech-recognition text at that time from the speech-recognition text storage unit 136. (In the example of FIG. 6 , the translation unit 150 reads the speech-recognition text for one sentence stored in the speech-recognition text storage unit 136 at timing R3.) The translation unit 150 then performs translation processing using the speech-recognition text (the entire utterance (from the beginning to the delimiter)) read from the speech-recognition text storage unit 136.

[0066] The language determination unit 140 includes a temporary language determination unit 142 and a periodic determination unit 144. When the ending determination unit 124 detects a delimiter in the speech-recognition text, the temporary language determination unit 142 immediately reads out the voice packets stored in the buffer 132 and transmits them to the language determination engine 20. Figure 7 shows the contents of the buffer 132 when the temporary language determination unit 142 provisionally determines the recognition language. In the figure, the shaded area has already been deleted by the ending determination unit 124, as described above. When the ending determination unit 124 detects a delimiter in the speech-recognition text, a certain amount of time has passed since the voice packets were transmitted at B3. Therefore, when the ending determination unit 124 detects a delimiter in the speech-recognition text, a considerable number of voice packets related to the next utterance are already stored in the buffer 132.

[0067] In the figure, the portion after the termination address A corresponds to the voice packets of the next utterance. The tentative language determination unit 142 effectively utilizes the voice packets related to the next utterance stored in the buffer 132 to provisionally determine the language to be used for voice recognition, i.e., the recognition language. Specifically, if the voice length of the voice packets stored in the buffer 132 is three seconds or less, all voice packets are transmitted to the language determination engine 20. If the voice length of the voice packets stored in the buffer 132 exceeds three seconds, voice packets of three seconds' worth of voice are read from the buffer 132, going back from the most recent voice packet (the voice packet last stored in the buffer 132), and transmitted to the language determination engine 20. Note that the longer the voice length of the voice packets transmitted to the language determination engine 20, the longer the determination process in the language determination engine 20 takes. For this reason, the tentative language determination unit 142 limits the voice length of the voice packets transmitted to the language determination engine 20 to three seconds. This enables faster voice recognition of the next speech. When the temporary language determination unit 142 transmits the voice packet to the language determination engine 20, it receives language determination information from the language determination engine 20. The temporary language determination unit 142 stores the received language determination information in the language information storage unit .

[0068] 8 is a diagram showing an example of language information stored in the language information storage unit 134. The language information includes language determination information in chronological order received from the language determination engine 20. As described above, the language determination information includes multiple pieces of probability information, and in this embodiment, the language determination information includes probability information for English and probability information for Japanese. Before performing speech recognition of the next utterance, the tentative language determination unit 142 receives language determination information from the language determination engine 20 as described above and stores the language determination information in association with the sequence "0." In the example shown in the figure, "0.4" is stored as the probability information for English, and "0.6" is stored as the probability information for Japanese.

[0069] The temporary language determination unit 142 provisionally determines the recognition language. As an example, the language with the highest probability information value may be determined as the recognition language. The determined language is stored in the "effective language" and "recognition language" columns in the figure. As another example, if the probability information value is the highest and is equal to or greater than a predetermined value (e.g., 0.5), the language with the highest probability value may be determined as the recognition language. If both probability information values ​​are less than the predetermined value, the recognition language used in the speech recognition of the immediately preceding utterance may be reused. When the end identification unit 124 detects a delimiter in the speech-recognized text and the speech-recognized text for the immediately preceding utterance is determined, the speech recognition processing unit 122 acquires the recognition language provisionally determined by the temporary language determination unit 142 and transmits all voice packets stored in the buffer 132 at that time to the speech recognition engine 10 associated with the provisional recognition language. That is, at timing B6 in Fig. 6 , the contents of buffer 132 are transmitted to the Japanese speech recognition engine 10, which is the provisional recognition language. When the speech recognition processing unit 122 receives Japanese speech recognition text and the like from the Japanese speech recognition engine 10, it stores them in the speech recognition text storage unit 136, as shown by R6 in Fig. 6 . The transmission unit 170 transmits this speech recognition text to the client device 50 and causes it to be displayed on the client device 50.

[0070] In the example of FIG. 6 , since the end identification unit 124 detects a delimiter in the speech-recognition text after timing B5, the contents of the buffer 132 at timings B4 and B5 are sent to the English speech recognition engine 10, where they are subjected to speech recognition processing. Therefore, there is a possibility that the portion of the speech-recognition text following the delimiter "?" at timings R4 and R5 (the Japanese portion in the example shown in FIG. 6 ) may contain an erroneous recognition. Therefore, the speech recognition processing unit 122 may prevent the portion following the delimiter from being stored in the speech-recognition text storage unit 136. This prevents the client device 50 from displaying an erroneously recognized portion. Similarly, the portion following the delimiter at timing R3 may also be prevented from being stored in the speech-recognition text storage unit 136.

[0071] After starting to generate speech-recognized text for the next utterance, the periodic determination unit 144 periodically reviews the recognition language. That is, at a predetermined determination interval (e.g., every 100 ms), the periodic determination unit 144 reads the most recent predetermined number of consecutive voice packets (i.e., a predetermined time length, one second in this case) stored in the buffer 132 and transmits them to the language determination engine 20. That is, the periodic determination unit 144 sequentially acquires a predetermined number of consecutive voice packets while shifting the start timing and transmits them to the language determination engine 20. In this way, language determination information is received from the language determination engine 20. The periodic determination unit 144 stores the received language determination information in the language information storage unit 134. Note that if the speech length stored in the buffer 132 is less than one second, the periodic determination unit 144 does not transmit the voice packet to the language determination engine 20. Therefore, after the provisional recognition language is determined by the provisional language determination unit 142, the recognition language is not reviewed until a voice packet with a speech length of one second or more is accumulated in the buffer 132. In the example of FIG. 8, the language determination information acquired by the periodic determination unit 144 is associated with the order "2," "3," "4," . . . and stored in the language information storage unit 134 in chronological order.

[0072] The periodic determination unit 144 changes the recognition language based on the language determination information stored in the language information storage unit 134, i.e., the language determination information stored by the temporary language determination unit 142 and the language determination information stored by the periodic determination unit 144 itself. When the recognition language is changed by the periodic determination unit 144, the speech recognition processing unit 122 performs speech recognition processing on the speech packets stored in the buffer 132 using the changed recognition language, and the transmission unit 70 transmits the resulting speech recognition text to the client device 50. This causes the client device 50 to display the speech recognition text in the updated recognition language. For example, each time language determination information is stored in the language information storage unit 134, the periodic determination unit 144 sets the language with the highest probability information value as the "effective language." If a recognition language other than the one determined by the temporary language determination unit 142 is set as the "effective language" multiple times (here, twice) in a row, the periodic determination unit 144 sets that language as the new recognition language. The periodic determination unit 144 reviews the recognition language at a determination interval (for example, 100 ms) that is shorter than the recognition interval (for example, 300 ms) by the speech recognition processing unit 122, so that the speech recognition processing unit 122 can perform speech recognition using a recognition language that is likely to be correct, in accordance with the most recent determination by the language determination unit 140. Furthermore, the periodic determination unit 144 transmits speech packets with a fixed length of 1 second to the language determination engine 20, so that responses from the language determination engine 20 can be prevented from becoming too long. This allows the periodic determination unit 144 to review the recognition language approximately periodically.

[0073] According to the speech recognition system with translation function 1 described above, even at the transition between utterances of user UJ and user UE, voice packets are continuously acquired at short acquisition intervals, such as every 20 ms, without the need for button or other operations, and are sequentially stored in the buffer 132. When the termination identification unit 124 detects a delimiter in the speech-recognized text acquired by the speech recognition processing unit 122, a considerable number of voice packets related to the next utterance are already stored in the buffer 132. The provisional language determination unit 142 can effectively utilize the voice packets related to the next utterance stored in the buffer 132 in this way to provisionally determine the language to be used in speech recognition of the next utterance, i.e., the recognition language. This allows speech recognition of the next utterance to begin promptly. Furthermore, the periodic determination unit 144 reviews the recognition language using voice packets of a certain voice length at a predetermined determination interval (e.g., 100 ms) that is longer than the voice packet acquisition interval (e.g., 20 ms) and shorter than the recognition interval (e.g., 300 ms) that is the interval between voice recognition processes. Therefore, voice recognition can be performed using a likely recognition language in each periodic voice recognition process. As a result, even in cases where user UJ speaks immediately after user UE's speech, the voice-recognized text of user UJ's speech can be displayed immediately. Furthermore, the recognition language during voice recognition can be set correctly.

Claims

1. A system including: a buffer storage means for sequentially acquiring a plurality of partial audio data each representing a part of a first utterance by a first speaker and a subsequent second utterance by a second speaker, and storing the data in a buffer; a character string acquisition means for performing a speech recognition process on one or more of the partial audio data representing the first utterance stored in the buffer using a given language as a recognition language, and acquiring a character string representing the utterance content represented by the partial audio data; an end portion identification means for identifying the partial audio data relating to the end portion of the first utterance in the buffer when the end portion of the first utterance is recognized in the speech recognition process; and a language determination means for acquiring language determination information for determining the language of the partial audio data based on at least a part of the one or more partial audio data representing the second utterance stored in the buffer after the identified partial audio data, when the partial audio data relating to the end portion is identified in the buffer, and determining a language based on the language determination information. the character string acquisition means performs a speech recognition process on one or more partial voice data representing the second utterance stored in the buffer, using the one language determined by the language determination means as a recognition language, and acquires a character string in the one language indicating the utterance content indicated by the partial voice data.

2. A speech recognition system as claimed in claim 1, further comprising: a language determination information acquisition means for sequentially acquiring language determination information for determining the language of the speech content indicated by the partial speech data, based on a predetermined number of consecutive partial speech data acquired from one or more partial speech data stored in the buffer while shifting the start timing in sequence; and a recognition language change means for changing the recognition language based on the language determination information acquired by the language determination means and the language determination information acquired in sequence by the language determination information acquisition means.

3. A speech recognition system as claimed in claim 2, wherein, when the recognition language is changed by the recognition language change means, the character string acquisition means performs speech recognition processing on the plurality of partial speech data stored in the buffer in the changed recognition language, and acquires, instead of the character string in the one language, a character string in the changed recognition language that indicates the utterance indicated by the partial speech data.

4. A speech recognition system as claimed in claim 2 or 3, wherein the buffer storage means sequentially acquires the plurality of partial speech data at a predetermined acquisition interval; the character string acquisition means performs speech recognition processing on the one or more partial speech data representing the first utterance at a predetermined recognition interval that is longer than the predetermined acquisition interval; and the language determination information acquisition means sequentially acquires the language determination information at a predetermined determination interval that is longer than the acquisition interval and shorter than the recognition interval.

5. A speech recognition system as claimed in claim 1, wherein, when the speech recognition process recognizes the end of the first utterance, the end part identification means performs speech recognition in the recognition language on the plurality of partial speech data representing the entire first utterance stored in the buffer by the character string acquisition means, and acquires a character string representing the content of the utterance represented by the partial speech data.

6. A speech recognition system according to claim 5, wherein the end portion identification means deletes the plurality of partial speech data representing the first utterance stored in the buffer when the end portion of the first utterance is recognized in the speech recognition processing.

7. A speech recognition system according to claim 1, wherein the end portion identification means identifies the end portion of the first utterance by detecting a predetermined delimiter from the character string obtained by the speech recognition processing.

8. A step of sequentially acquiring a plurality of partial audio data each indicating a part of a first utterance by a first speaker and a subsequent second utterance by a second speaker, and storing the data in a buffer; a step of performing speech recognition processing on one or more of the partial audio data indicating the first utterance stored in the buffer using a given language as a recognition language, and acquiring a character string indicating the utterance content indicated by the partial audio data; a step of identifying the partial audio data relating to the end of the first utterance in the buffer when the end of the first utterance is recognized in the speech recognition processing; a step of acquiring language determination information for determining the language associated with the partial audio data based on at least a part of the one or more partial audio data indicating the second utterance stored in the buffer after the identified partial audio data when the partial audio data relating to the end is identified in the buffer, and determining a language based on the language determination information; a step of performing speech recognition processing on one or more of the partial audio data indicating the second utterance stored in the buffer using the determined language as a recognition language, and acquiring a character string in the one language indicating the utterance content indicated by the partial audio data. A speech recognition method comprising:

9. An information storage medium containing a program for causing a computer to function as: buffer storage means for sequentially acquiring a plurality of partial voice data each indicating a part of a first utterance by a first speaker and a subsequent second utterance by a second speaker, and storing the acquired data in a buffer; character string acquisition means for performing voice recognition processing on one or more of the partial voice data indicating the first utterance stored in the buffer using a given language as a recognition language, and acquiring a character string indicating the content of the utterance indicated by the partial voice data; end point identification means for identifying the partial voice data relating to the end of the first utterance in the buffer when the end of the first utterance is recognized in the voice recognition processing; and language determination means for acquiring language determination information for determining the language associated with the partial voice data based on at least a part of one or more of the partial voice data indicating the second utterance stored in the buffer after the identified partial voice data, when the partial voice data relating to the end of the first utterance is identified in the buffer, and determining a language based on the language determination information, the character string acquisition means performs a voice recognition process on one or more partial voice data representing the second utterance stored in the buffer, using the one language determined by the language determination means as a recognition language, and acquires a character string in the one language indicating the content of the utterance indicated by the partial voice data.

Citation Information

Patent Citations

  • Automatic language identification method and system

    JP2004347732A

  • Voice translation apparatus and its method

    JP2007322523A

  • Voice input device, translation device, voice input method, and voice input program

    WO2018034059A1