Translation distribution system, translation distribution method, and program
The translation distribution system dynamically manages participant languages and provides targeted translations using speech recognition and synthesis, addressing resource wastage and processing load by adapting to real-time participant presence and language needs.
Patent Information
- Application Number
- PCT/JP2024/000071
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-10
AI Technical Summary
Existing translation systems face challenges in efficiently providing translations in necessary languages to participants with varying native languages, leading to resource wastage and increased processing load due to the uncertainty of participant languages and the need for continuous translations even when participants leave.
A translation distribution system that dynamically updates and manages participant languages by associating terminal identification information with user languages, selectively generating and transmitting translations based on real-time participant presence and language needs, using speech recognition, translation engines, and voice synthesis to provide targeted translations.
Reduces processing load and conserves resources by providing translations only in required languages, ensuring efficient use of resources and accurate language support for participants as they arrive and depart.
Smart Images

Figure JP2024000071_10072025_PF_FP_ABST
Abstract
Description
Translation distribution system, translation distribution method, and program
[0001] The present invention relates to a translation and distribution system, a translation and distribution method, and a program.
[0002] Patent document 1 describes a translation device that, when it detects input of input speech in a first language from a source device, translates the input speech into output speech in multiple target languages different from the first language and transmits the output speech to multiple destination devices.
[0003] Patent No. 6710818
[0004] The inventors of the present application are considering distributing translation information in which speeches of speakers and announcements are translated into various languages, for example, at conferences attended by participants with various native languages, or at airports and other facilities where participants with various native languages gather. However, at conference halls, airports, and other facilities where translations are provided, the languages of the participants for whom translations are provided may not be known in advance. Anticipating and preparing all the languages needed in advance imposes a large processing load and is difficult, and resources are wasted if there are no participants who speak the languages prepared. Furthermore, it is inconvenient not to be able to provide translations for languages that are not prepared in advance. Furthermore, if a participant who needs a translation in a specific language leaves midway, continuing the translation in that language is a waste of resources.
[0005] The present invention has been made in consideration of the above-mentioned problems, and one of its objectives is to provide a translation distribution system, translation distribution method, and program that can distribute translations in the required languages according to the needs of participants while reducing the processing load and saving resources.
[0006] (1) A translation delivery system according to the present invention includes a terminal identification information storage means for storing identification information for identifying terminals of one or more users who use translation of a speaker's voice at a service location where the speech of the speaker is provided, in association with the language used by the user; a first update means for acquiring the identification information of the terminal and the language used by the user who uses the terminal based on communication with the terminal of a user who visits the service location, and updating the contents of the terminal identification information storage means so that, when at least the language used by the user is different from the language used by the speaker, the terminal identification information of the user is stored in association with the language used by the user; and a first update means for repeatedly determining whether communication with a terminal identified by each of the identification information stored in the terminal identification information storage means has ended, and deleting or invalidating the identification information when communication has ended. a second updating means for updating the contents of the terminal identification information storage means; a voice acquiring means for acquiring the speaker's voice portion by portion in sequence; a translation language determining means for determining, at a given timing, one or more translation target languages, which are languages different from the language used by the speaker and which are respectively associated with one or more terminals based on the contents of the terminal identification information storage means; a translation means for generating, based on the portions of the speaker's voice acquired in sequence by the voice acquiring means, one or more pieces of translation information which are partial translations of the speaker's voice into each of the one or more translation target languages by text or voice; and a transmission means for transmitting, to a terminal identified by each of the identification information stored in the terminal identification information storage means, the translation information in the language associated with the identification information.
[0007] (2) The translation delivery system described in (1) above may further include a translation information storage means for storing at least one of the last-generated pieces of translation information generated by the translation means based on a portion of the speaker's speech sequentially acquired by the speech acquisition means, in association with the target language corresponding to the translation information. The transmission means may transmit the translation information stored in the translation information storage means to a terminal identified by each of the identification information stored in the terminal identification information storage means, in association with the target language, which is the working language associated with the identification information, among the one or more pieces of translation information.
[0008] (3) In the translation delivery system described in (2) above, the translation information storage means may sequentially store one or more pieces of translation information generated by the translation means based on a portion of the speaker's voice sequentially acquired by the voice acquisition means, in association with the target language corresponding to the translation information. The transmission means may transmit, to the terminal that has transmitted a predetermined transmission request, the past translation information stored in the translation information storage means in association with the target language, which is a working language associated with identification information of the terminal.
[0009] (4) The translation delivery system described in (3) above may further include a translation information detection means for detecting whether requested translation information, which is one or more pieces of past translation information that the user requests to be sent in response to the specified transmission request from the terminal, is included in the past translation information stored in the translation information storage means in association with the target language, which is the language used and associated with the identification information of the terminal.
[0010] (5) In the translation delivery system described in (4) above, the translation information detection means may cause the translation means to generate the past translation information corresponding to the request translation information when the request translation information is not included in the past translation information stored in the translation information storage means in association with the target language, which is the used language associated with the identification information of the terminal.
[0011] (6) In the translation and distribution system described in any one of (1) to (5) above, a language determination means may be further included that determines the language of the speaker by inputting a portion of the speaker's speech into a trained machine model.
[0012] (7) In the translation delivery system described in any one of (1) to (6) above, speech recognition text that is speech-recognized based on a portion of the speaker’s voice may be transmitted to a terminal identified by each of the identification information stored in the terminal identification information storage means, together with the translation information in the language used that is associated with the identification information.
[0013] (8) In the translation and delivery system according to any one of (1) to (7), the one or more pieces of translation information may include one or more translated texts, which are texts generated based on a portion of the speech acquired by the speech acquisition means. The translation means may generate the one or more translated texts by having a translation engine of each of the one or more target languages translate a speech-recognition text that is speech-recognized based on the portion of the speaker's speech.
[0014] (9) In the translation delivery system described in (8) above, the one or more pieces of translation information may include one or more translated speeches, which are speeches generated based on a portion of the speech acquired by the speech acquisition means. The translation means may generate the one or more translated speeches by causing speech synthesis engines of the one or more target languages to perform speech synthesis of the one or more translated texts.
[0015] (10) A translation delivery method according to the present invention includes a terminal identification information storage step of storing identification information identifying each terminal of one or more users who use translation of the speaker's voice at a service location where the speaker's voice is provided, in association with the language used by the user; a first update step of acquiring the identification information of the terminal and the language used by the user who uses the terminal based on communication with the terminal of a user who visits the service location, and updating the contents of the terminal identification information storage means so that, when at least the language used by the user is different from the language used by the speaker, the terminal identification information of the user is stored in association with the language used by the user; and a second update step of repeatedly determining whether communication with a terminal identified by each of the identification information stored in the terminal identification information storage means has ended, and deleting or invalidating the identification information when the communication has ended. The method includes a second updating step of updating the contents of the terminal identification information storage means, a voice acquisition step of acquiring the speaker's voice portion by portion, a translation language determination step of determining, at a given timing, one or more translation target languages associated with one or more terminals, the languages being different from the language used by the speaker, based on the contents of the terminal identification information storage means, a translation step of generating one or more pieces of translation information which are partial text or voice translations of the speaker's voice into each of the one or more translation target languages, based on the portions of the speaker's voice acquired sequentially by the voice acquisition means, and a transmission step of transmitting the one or more pieces of translation information in the language used associated with the identification information to a terminal identified by each of the identification information stored in the terminal identification information storage means.
[0016] (11) A translation delivery program according to the present invention includes a terminal identification information storage means for storing identification information for identifying terminals of one or more users who use translation of the speaker's voice at a service location where the speech of the speaker is provided, in association with the language used by the user; a first update means for acquiring the identification information of the terminal and the language used by the user who uses the terminal based on communication with the terminal of the user who visits the service location, and updating the contents of the terminal identification information storage means so that, when at least the language used by the user is different from the language used by the speaker, the terminal identification information of the user is stored in association with the language used by the user; and a second update means for repeatedly determining whether communication with a terminal identified by each of the identification information stored in the terminal identification information storage means has ended, and deleting or invalidating the identification information when the communication has ended. The computer functions as a second updating means for updating the contents of the terminal identification information storage means, a voice acquisition means for sequentially acquiring portions of the speaker's voice, a translation language determination means for determining, at a given timing, one or more translation languages associated with one or more terminals, the languages being different from the language used by the speaker, based on the contents of the terminal identification information storage means, a translation means for generating one or more pieces of translation information that are partial text or voice translations of the speaker's voice into each of the one or more translation languages, based on portions of the speaker's voice sequentially acquired by the voice acquisition means, and a transmission means for transmitting, to a terminal identified by each of the identification information stored in the terminal identification information storage means, the translation information of the one or more pieces of translation information in the language used associated with the identification information.
[0017] 5 is a diagram showing an overall configuration of a translation and delivery system according to an embodiment of the present invention. FIG. 6 is a diagram showing another situation in the overall configuration diagram of FIG. 1. FIG. 7 is a diagram showing yet another situation in the overall configuration diagram of FIG. 1. FIG. 8 is a diagram showing an example of a hardware configuration of a translation and delivery device according to an embodiment of the present invention. FIG. 9 is a functional block diagram of a translation and delivery device according to an embodiment of the present invention. FIG. 10 is a functional block diagram showing in detail the speech recognition unit, translation unit, and translation information storage unit in FIG. 5. FIG. 11 is a diagram showing an example of speech recognition text data. FIG. 12 is a diagram showing an example of translated text data. FIG. 13 is a functional block diagram showing in detail the translation and delivery information management unit and delivery unit in FIG. 5. FIG. 14 is a diagram showing an example of terminal identification information and translation target language information in FIG. 1. FIG. 15 is a diagram showing an example of terminal identification information and translation target language information in FIG. 2. FIG. 16 is a flow diagram showing an example of the flow of a process of generating translated text according to the present embodiment. FIG. 17 is a diagram showing an example of a screen of a translation page displayed on a participant terminal. FIG. 18 is a diagram showing an example of a screen of a translation page displayed on a participant terminal. FIG. 19 is a diagram showing an example of translated text data. FIG. 19 is a flow diagram showing an example of a process of transmitting past translation information according to an embodiment of the present invention. FIG. 19 is a diagram showing a modified example of a translation page according to an embodiment of the present invention.
[0018] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0019] Fig. 1 is a diagram showing the overall configuration of a translation distribution system 1 according to this embodiment. The translation distribution system 1 is used at a translation provision location A, such as an international conference center. At the translation provision location A, a speaker gives a speech in a specific language, such as their native language. While Fig. 1 shows a speaker SJ giving a speech in Japanese as an example, the speaker may also give a speech in English or another language, as will be described later (see Fig. 3).
[0020] Furthermore, various conference participants come and go from the translation provision location A, and they use a variety of languages. The figure shows, as an example, a situation in which not only a participant PJ who uses Japanese, the same language as the speaker, and understands Japanese, but also a participant PE who uses English and understands English (but does not necessarily understand Japanese) are present at the translation provision location A. In such a case, the translation distribution system 1 provides translation information of the speaker's speech to participants who use a language different from the speaker. The translation information includes not only text but also speech. For example, in the example shown in the figure, the speech of speaker SJ in Japanese is translated into English and provided to each participant PE as English text, and the same English speech is also provided to each participant PE.
[0021] In order to realize the above-mentioned function of providing translation information from multiple languages to multiple languages, the translation distribution system 1 includes a translation distribution device 100, multiple speech recognition engines 2, multiple translation engines 3, multiple speech synthesis engines 4, and multiple participant terminals 6 (participant terminals 6E, 6J, ...), which are capable of communicating with each other via a computer network 5.
[0022] Each of the speech recognition engines 2 is a server computer that outputs the contents of speech data in a specific language as text in that language. A speech recognition engine 2 is provided for each planned language (speaker language), and speech data in any of the planned speaker languages can be output as text in that language by using any of the speech recognition engines 2. For example, if N languages are planned as speaker languages, N speech recognition engines 2 are available via the computer network 5.
[0023] Each of the translation engines 3 is a server computer that translates text in a specific source language into text in a specific target language. The translation engines 3 are used to translate text in a speaker's language into text in a working language (participant's language). For example, if N languages are planned as speaker languages and M languages are planned as participant languages, N x M translation engines 3 are available via the computer network 5.
[0024] Each of the speech synthesis engines 4 is a server computer that converts text in a specific language into speech data in that language. A speech synthesis engine 4 is provided for each of the planned participant languages, and by using any of the speech synthesis engines 4, text in any of the planned participant languages can be output as speech data in that language. If M languages are planned as participant languages, M speech synthesis engines 4 are available via the computer network 5.
[0025] A participant terminal 6 is a terminal carried by a participant, and is, for example, a portable computer such as a PC, tablet, or smartphone. Here, the participant terminal carried by participant PE, whose language of use is English, is referred to as participant terminal 6E, and the participant terminal carried by participant PJ, whose language of use is Japanese, is referred to as participant terminal 6J. When referring to a participant terminal without specifying the language of use, it is simply referred to as participant terminal 6.
[0026] A participant visiting translation provision location A uses a participant terminal 6 to read a two-dimensional barcode 103 displayed within translation provision location A. This allows the participant terminal 6 to begin communication with a translation delivery device 100 having the address specified by the two-dimensional barcode 103, and to receive translation information from the translation delivery device 100. The translation information includes text and audio data in the participant's language; the text is displayed on a display included in the participant terminal 6, and the audio data is output by an audio output device 7, such as headphones, connected to the participant terminal 6 by wire or wirelessly. This allows participants who speak a language different from the speaker to understand what is being said by the speaker.
[0027] The translation distribution device 100 is a computer installed, for example, at the translation provision site A, and is connected to a microphone 101 used by a speaker such as speaker SJ, and is also connected to a computer network 5. The translation distribution device 100 acquires speech data representing the speaker's speech acquired by the microphone 101, and generates translation information in one or more participant languages using a speech recognition engine 2, a translation engine 3, and a speech synthesis engine 4. The generated translation information is then distributed to participant terminals 6.
[0028] An operator (not shown) of translation provision site A uses an input device (not shown) to set in advance the speaker language (Japanese in the figure) of a speaker (speaker SJ in the figure) in the translation distribution device 100. Furthermore, the translation distribution device 100 manages the target language (only English in the figure), which is the participant language of a participant (participant PE in the figure) who needs translation at translation provision site A, using a method described below. This allows the translation distribution device 100 to select the required engine from among multiple speech recognition engines 2, multiple translation engines 3, and multiple speech synthesis engines 4.
[0029] For example, in the situation shown in Figure 1, the operator of translation provider A has pre-set Japanese as the speaker language of speaker SJ in the translation distribution device 100, so the translation distribution device 100 selects the Japanese speech recognition engine 2J from among the multiple speech recognition engines 2.
[0030] Furthermore, because the translation delivery device 100 detects that a participant PE is present at translation provision location A, it selects English as the target language. From among the multiple translation engines 3, the translation delivery device 100 selects a Japanese-English translation engine 3JE to translate from Japanese, the speaker language of speaker SJ, into English, the participant language of participant PE. Furthermore, from among the multiple speech synthesis engines 4, the translation delivery device 100 selects an English speech synthesis engine 4E to provide a speech translation to participant PE.
[0031] The translation and distribution device 100 acquires the speaker SJ's speech portion by portion (for example, at each segment of speech, such as silence) via a wireless or wired microphone 101 installed at the translation provision location A, and transmits speech data representing the speech to a Japanese speech recognition engine 2J. This generates Japanese speech-recognition text, which is the recognition result in Japanese. Next, the Japanese speech-recognition text is transmitted to a Japanese-to-English translation engine 3JE. This generates English text, which is a translation of the speech-recognition text into English. Furthermore, the translation and distribution device 100 transmits the generated English text to an English speech synthesis engine 4E, which generates English speech with the same content as the English text.
[0032] The translation and distribution device 100 transmits the Japanese speech-recognition text and the generated translation information, that is, the English text and English speech, to the participant terminal 6E of the participant PE. The speech-recognition text and English text are displayed on the display of the participant terminal 6E held by the participant PE. The participant PE can also use the audio output device 7 to listen to the English speech received at the participant terminal 6E. This allows the participant PJ to understand in English what the speaker SJ is saying in Japanese.
[0033] Next, Fig. 2 is a diagram showing another situation in the overall configuration diagram of Fig. 1. The diagram shows, as an example, a situation in which not only a participant PJ who understands Japanese and a participant PE whose participant language is English, but also a participant PC who understands Chinese (but does not necessarily understand Japanese) are present at translation provision location A. In this case, in the example of the diagram, the speech of speaker SJ in Japanese is provided with English translation information to participant PE as in Fig. 1, and Chinese translation information is also provided to participant PC.
[0034] In the situation shown in Figure 2, the operator of translation provider A has pre-set Japanese as the speaker language of speaker SJ in the translation delivery device 100, so the translation delivery device 100 selects a Japanese speech recognition engine 2J from among multiple speech recognition engines 2.
[0035] Furthermore, because the translation delivery device 100 detects that the participant PE and participant PC are present at the translation provision location A, it selects English and Chinese as the translation target languages. From among the multiple translation engines 3, the translation delivery device 100 selects the same Japanese-English translation engine 3JE as in FIG. 1 , as well as a Japanese-Chinese translation engine 3JC for translating from Japanese, the participant language of speaker SJ, to Chinese, the participant language of the participant PC. Furthermore, to provide translations via speech to the participant PE and participant PC, the translation delivery device 100 selects the same English speech synthesis engine 4E as in FIG. 1 , as well as a Chinese speech synthesis engine 4C from among the multiple speech synthesis engines 4. Furthermore, to generate translated speech in English and Chinese, the translation delivery device 100 selects the same English speech synthesis engine 4E as in FIG. 1 , as well as a Chinese speech synthesis engine 4C from among the multiple speech synthesis engines 4.
[0036] In addition to generating English translation information as shown in Figure 1, translation and distribution device 100 also generates Chinese translation information. Translation and distribution device 100 generates Chinese text, which is a translation of the speech recognition text into Chinese, by sending Japanese speech recognition text, which is the generated Japanese recognition result, to Japanese-Chinese translation engine 3JC. Furthermore, translation and distribution device 100 sends the Chinese text to Chinese speech synthesis engine 4C, similar to generating English speech in Figure 1, and generates Chinese speech with the same content as the Chinese text.
[0037] The translation and distribution device 100 transmits Japanese speech-recognition text, English text, and English audio to the participant terminal 6E of the participant PE, as in FIG. 1, and transmits Japanese speech-recognition text and the generated translation information, Chinese text and Chinese audio, to the participant terminal 6C of the participant PC. As in FIG. 1, the participant PE displays the speech-recognition text and English text on the display of the participant terminal 6E that he or she owns, and can listen to the English audio received at the participant terminal 6E using the audio output device 7. The speech-recognition text and Chinese text are also displayed on the display of the participant terminal 6C that the participant PC owns. The participant PC can also listen to the Chinese audio received at the participant terminal 6C using the audio output device 7. This allows the participants PE and PC to understand what the speaker SJ is saying in Japanese in their own participant languages, English or Chinese.
[0038] Next, Figure 3 is a diagram showing yet another situation in the overall configuration diagram of Figure 1. Unlike Figures 1 and 2, this figure shows, as an example, a situation in which the speaker has changed to a speaker SE who speaks in English. The figure shows, as an example, a situation in which a participant PJ who understands Japanese (but does not necessarily understand English), a participant PE who understands English, and a participant PC who understands Chinese (but does not necessarily understand English) are visiting a translation service location A.
[0039] When the speaker changes to speaker SE, the operator of translation provision site A sets English as the speaker language of speaker SE in the translation distribution device 100. Then, the translation distribution device 100 selects an English speech recognition engine 2E from among the multiple speech recognition engines 2.
[0040] Furthermore, because the translation distribution device 100 detects that there are participants PJ and PC at the translation provision location A who speak different languages from the speaker SJ, it determines Japanese and Chinese as the translation target languages. The translation distribution device 100 selects, from among the multiple translation engines 3, an English-Japanese translation engine 3EJ and an English-Chinese translation engine 3EC to translate from English, the speaker language of speaker SE, into Japanese and Chinese, the participant languages of the participants at the translation provision location A, respectively. Furthermore, the translation distribution device 100 selects, from among the multiple speech synthesis engines 4, a Japanese speech synthesis engine 4J and a Chinese speech synthesis engine 4C to generate speech in Japanese and Chinese.
[0041] The translation and distribution device 100 sequentially acquires portions of the speech of the speaker SJ via a wireless or wired microphone 101 installed at the translation provision location A and transmits speech data representing the speech to the English speech recognition engine 2E. This generates English speech-recognition text, which is the recognition result in English. The English speech-recognition text is then transmitted to an English-Japanese translation engine 3EJ and an English-Chinese translation engine 3EC, respectively, to generate Japanese text, which is a translation of the speech-recognition text into Japanese, and Chinese text, which is a translation of the speech-recognition text into Chinese. Furthermore, the translation and distribution device 100 transmits the Japanese text to a Japanese speech synthesis engine 4J and the Chinese text to a Chinese speech synthesis engine 4C, to generate Japanese speech with the same content as the Japanese text and Chinese speech with the same content as the Chinese text.
[0042] The translation distribution device 100 transmits the English speech recognition text and the generated Japanese text and Japanese audio, which are the translation information, to the participant terminal 6J of the participant PJ, and transmits the English speech recognition text and the generated Chinese text and Chinese audio, which are the translation information, to the participant terminal 6C of the participant PC. The speech recognition text and Japanese text are displayed on the display of the participant terminal 6J held by the participant PJ. The participant PJ can also use the audio output device 7 to listen to the Japanese audio received at the participant terminal 6J. The speech recognition text and Chinese text are also displayed on the display of the participant terminal 6C held by the participant PC. The participant PC can also use the audio output device 7 to listen to the Chinese audio received at the participant terminal 6C.
[0043] In an embodiment of the present invention, the translation distribution device 100 manages the participant languages of participants visiting the translation provision location A, and generates translation information according to the combination of the speaker language and the participant language, thereby reducing the processing load and saving resources, and delivering translations in the required languages according to the needs of the participants.
[0044] Fig. 4 is a diagram showing an example of the hardware configuration of a translation and distribution device 100 according to an embodiment of the present invention. The translation and distribution device 100 according to this embodiment is a computer as shown in Fig. 4. Software installed on the computer realizes a function of translating an input speaker's speech into one or more different languages and distributing the translated speech. As shown in Fig. 4, the example of the translation and distribution device 100 according to the embodiment includes, for example, a processor 100a, a storage unit 100b, and a communication unit 100c.
[0045] The processor 100a is, for example, a program-controlled device such as a microprocessor that operates according to a program installed in the translation and distribution device 100. The storage unit 100b is, for example, a storage element such as a ROM or RAM, a solid-state drive, or a hard disk drive. The storage unit 100b stores programs and the like to be executed by the processor 100a. The communication unit 100c is, for example, a communication interface for exchanging data with the participant terminals 6 via the computer network 5.
[0046] 5 is a functional block diagram of a translation and distribution device 100 according to an embodiment of the present invention. As shown in FIG. 5, the translation and distribution system 1 functionally includes a speech recognition unit 10, a translation unit 20, a translation information storage unit 30, a translation and distribution information management unit 40, and a distribution unit 50.
[0047] The above functions are implemented by executing a program containing instructions corresponding to the above functions on the processor 100a, which is installed in the translation and distribution device 100, which is a computer. This program is supplied to the translation and distribution device 100 via a computer-readable information storage medium such as an optical disk, magnetic disk, magnetic tape, magneto-optical disk, or flash memory, or via the Internet, for example.
[0048] The following describes the function of providing translation information from multiple languages to multiple languages.
[0049] FIG. 6 is a functional block diagram showing in detail the speech recognition unit 10, translation unit 20, and translation information storage unit 30 shown in FIG.
[0050] The speech recognition unit 10 acquires speech data representing a speaker's speech and generates speech-recognized text by performing speech recognition on the speech data. The speech recognition unit 10 includes a speech input acceptance unit 11, a language determination unit 13, and a speech-recognized text acquisition unit 15. The speech input acceptance unit 11, the language determination unit 13, and the speech-recognized text acquisition unit 15 are implemented mainly by the processor 100a and the communication unit 100c.
[0051] The translation unit 20 generates translation information from the speech-recognition text transmitted from the speech-recognition text acquisition unit 15. The translation unit 20 includes a translation engine determination unit 21, a translated text acquisition unit 23, a previous translated text acquisition unit 23a, and a translated speech acquisition unit 25. The translation engine determination unit 21, the translated text acquisition unit 23, the previous translated text acquisition unit 23a, and the translated speech acquisition unit 25 are implemented mainly by the processor 100a and the communication unit 100c.
[0052] The translation information storage unit 30 stores the speech-recognition text data Dvr transmitted from the speech-recognition text acquisition unit 15 and the translated text data Dptt transmitted from the translated text acquisition unit 23. The translation information storage unit 30 includes a speech-recognition text storage unit 31 and a translated text storage unit 33. The speech-recognition text storage unit 31 and the translated text storage unit 33 are implemented mainly in the storage unit 100b.
[0053] The voice input receiving unit 11 sequentially receives, one portion at a time, voice data representing the speaker's voice input from a microphone 101 installed at the translation providing location A. The voice input receiving unit 11 transmits the received voice data to the language determination unit 13 and the voice recognition text acquisition unit 15.
[0054] The voice input receiving unit 11 may detect silent intervals in the speaker's voice and receive the voice by dividing the voice into silent intervals. The voice input receiving unit 11 may also receive the voice by dividing the voice into predetermined time intervals without dividing the voice into silent intervals.
[0055] The language determination unit 13 determines the speaker language, which indicates the language spoken by the speaker. The language determination unit 13 determines the speaker language based on settings input by the operator of translation providing site A. The speaker language determination may be performed each time the speaker changes, or may be performed periodically at predetermined timing. The language determination unit 13 notifies the speech-recognition text acquisition unit 15 and the translation target language management unit 47, which will be described later, of the speaker language. For example, in the situation shown in FIG. 2 , the language determination unit 13 determines the speaker language to be Japanese based on the Japanese setting input by the operator of translation providing site A.
[0056] The language determination unit 13 may also determine the language of a speaker by inputting the speaker's voice data into a trained machine learning model. By inputting voice data in various languages and performing training that outputs the language of the voice, it is possible to create a machine learning model that, when input, determines the language in which the voice is spoken.
[0057] The speech-recognition text acquisition unit 15 determines the speech recognition engine 2 based on the speaker language notified by the language determination unit 13. The speech-recognition text acquisition unit 15 then transmits a portion of the speech data transmitted from the speech input receiving unit 11 to the speech recognition engine 2 to acquire speech-recognition text. In other words, the speech-recognition text may be text obtained by partially recognizing speech. Furthermore, the speech-recognition text acquisition unit 15 assigns text identification information that enables identification of the speech-recognition text. The speech-recognition text acquisition unit 15 then associates the speech-recognition text data Dvr with the assigned text identification information and the speaker language information of the speech-recognition text, and sequentially transmits the data to the translated text acquisition unit 23 and the speech-recognition text storage unit 31. The timing at which the speech-recognition text acquisition unit 15 transmits the data is not limited. The data may be transmitted each time a piece of speech-recognition text data Dvr is generated, or each time any number of pieces of speech-recognition text data Dvr are generated.
[0058] The speech-recognition text storage unit 31 stores speech-recognition text. The speech-recognition text storage unit 31 stores speech-recognition text data Dvr, text identification information, and speaker language information of the speech-recognition text transmitted by the speech-recognition text acquisition unit 15. The speech-recognition text storage unit 31 creates a new record and stores the latest speech-recognition text data Dvr transmitted by the speech-recognition text acquisition unit 15.
[0059] Fig. 7 is a diagram showing an example of data storage of the speech recognition text data group 31D stored in the speech recognition text storage unit 31 in the situation shown in Fig. 2. In the example shown in Fig. 7, four records are stored in the speech recognition text data group 31D, and each column of the record stores speech recognition text data Dvr, text identification information, and speaker language information of the speech recognition text.
[0060] The text ID is text identification information for identifying the speech-recognition text. The text identification information may be the time when the speech-recognition text was recognized or an identification number sequentially assigned when the speech-recognition text is recognized by the speech-recognition text acquisition unit 15. In this embodiment, the identification number assigned by the speech-recognition text acquisition unit 15 is stored as the value of the text ID.
[0061] The speech recognition text is text information acquired by the speech recognition text acquisition unit 15 by transmitting speech data to the speech recognition engine 2 for speech recognition. As an example of the situation in Fig. 2, Fig. 7 shows Japanese text information obtained by speech recognition of Japanese speech data by the Japanese speech recognition engine 2J, which is stored as the value of the speech recognition text.
[0062] The speaker language information is information indicating the language spoken in the speech-recognized text. In Fig. 7, as an example in the situation of Fig. 2, "Japanese", which is the speaker language of speaker SJ determined by the language determination unit 13, is stored as the value of the speaker language information.
[0063] In this embodiment, when a speaker makes a new utterance, speech recognition text data Dvr corresponding to the new utterance is generated and stored in the speech recognition text storage unit 31. For example, when a Japanese speaker SJ makes a new utterance as shown in Figure 2, the speech recognition text storage unit 31 creates a new record and stores a new identification number assigned by the speech recognition text acquisition unit 15, Japanese text information (utterance content), and "Japanese" as the speaker language information. In this way, the speech recognition text storage unit 31 stores the results of speech recognition of the speaker's utterance content in order.
[0064] The translation engine determination unit 21 determines a translation engine to translate the speech-recognized text. The translation engine determination unit 21 determines a translation engine 3 based on the speaker language transmitted from the speech-recognized text acquisition unit 15 and one or more translation-target languages transmitted from the translation distribution information management unit 40. The translation distribution information management unit 40 monitors one or more participant terminals 6 at the translation providing location A and manages the participant languages to determine one or more translation-target languages. The translation engine determination unit 21 may determine a translation engine 3 each time speech-recognized text data Dvr and speaker language information are transmitted from the speech-recognized text acquisition unit 15, or may determine a translation engine 3 periodically regardless of the transmission frequency. The translation engine determination unit 21 notifies the translated text acquisition unit 23 of the determined translation engine 3.
[0065] When there are multiple target languages, the translation engine determination unit 21 determines, for each target language, a translation engine that translates from the speaker language to the target language. For example, in the case of Figure 1, the target language is English only, so the translation engine determination unit 21 determines the Japanese-English translation engine 3JE. When the target languages are English and Chinese as in Figure 2, the translation engine determination unit 21 determines the Japanese-English translation engine 3JE and the Japanese-Chinese translation engine 3JC.
[0066] The translated text acquisition unit 23 transmits the speech-recognition text data Dvr to the translation engine 3 to acquire the translated text data Dptt translated into the target language. The translated text acquisition unit 23 transmits the speech-recognition text data Dvr to the translation engine 3 determined by the translation engine determination unit 21 to acquire the translated text data Dptt translated into the target language. By translating the speech-recognition text data Dvr into the target language, it is possible to acquire the translated text data Dptt, which is a partial translation based on a portion of the speaker's speech. Furthermore, the translated text acquisition unit 23 associates the acquired translated text data Dptt with text identification information assigned to the speech-recognition text data Dvr corresponding to the translated text data Dptt and target language information of the translated text, and transmits them to the translated speech acquisition unit 25 and the translated text storage unit 33.
[0067] The timing at which the translated text acquisition unit 23 transmits the data is not limited. The data may be transmitted to the translated speech acquisition unit 25 and the translated text storage unit 33 each time translated text data Dptt is generated, or multiple translated text data Dptt may be transmitted collectively. The translated text data Dptt may be transmitted each time one piece of translated text data Dptt is generated, or each time any number of translated text data Dptt are generated.
[0068] When there are multiple target languages, the translation engine determination unit 21 notifies the translation engines of multiple engines, and the translated text acquisition unit 23 transmits the speech-recognition text data Dvr to each of the multiple translation engines to acquire multiple pieces of translated text data Dptt. For example, in the case shown in FIG. 2 , the translated text acquisition unit 23 is notified by the translation engine determination unit 21 of the Japanese-English translation engine 3JE and the Japanese-Chinese translation engine 3JC, and transmits the speech-recognition text data Dvr to each of the translation engines 3JE and 3JC to acquire English text data Det translated into English and Chinese text data Dct translated into Chinese. The translated text acquisition unit 23 then transmits the text ID, the English text data Det, and the target language information "English," as well as the same text ID, the Chinese text data Dct, and the target language information "Chinese," to the translated text storage unit 33 and the translated speech acquisition unit 25.
[0069] The previous translated text acquisition unit 23a acquires previous translated text when it receives a translation generation notification from the log request processing unit 53. The function of acquiring previous translated text data Dptt will be described later.
[0070] The translated text storage unit 33 stores the translated text data Dptt. The translated text storage unit 33 associates and stores the translated text data Dptt, text identification information, and translation target language information transmitted from the translated text acquisition unit 23. The translated text storage unit 33 creates a new record and stores the latest translated text data Dptt transmitted from the translated text acquisition unit 23.
[0071] The translated text storage unit 33 has a column for storing translated text data Dptt for each language scheduled as a participant language. The column for storing the translated text data Dptt for each language is initially blank. The translated text data Dptt sent from the translated text acquisition unit 23 is stored in the column for the target language of the translated text data Dptt.
[0072] 8A is a diagram showing an example of data storage of the translated text data group 33D stored in the translated text storage unit 33 when the translation target language is only English in the situation shown in Fig. 2. In the example shown in Fig. 8A, the translated text data group 33D has columns for English, Chinese, and other languages as participant languages.
[0073] The translated text data group 33D shown in Figure 8A stores three records, with the "English Text" column storing the translated English text data Dptt storing English text data Det and the "Text ID" column storing text identification information.
[0074] Next, Fig. 8B is a diagram showing an example of data storage of translated text data group 33D stored in translated text storage unit 33 when the translation target languages are English and Chinese in the situation shown in Fig. 2. Fig. 8A shows translated text data group 33D stored in translated text storage unit 33 after an utterance corresponding to text ID "4" is made.
[0075] 2 , if speaker SJ makes an utterance corresponding to text ID "4" and the translation target languages are English and Chinese, the translated text acquisition unit 23 sends the speech-recognition text data Dvr to each of translation engines 3JE and 3JC, and acquires the English text data Det and Chinese text data Dct associated with text ID "4." The translated text acquisition unit 23 then sends the English text data Det and Chinese text data Dct associated with text ID "4" to the translated text storage unit 33.
[0076] 8B, a new record with a text ID of "4" is added to the translated text group 33D shown in Fig. 8A. The "English Text" column that stores the translated English text data Dptt stores English text data Det, and the "Chinese Text" column that stores the translated Chinese text data Dptt stores Chinese text data Dct.
[0077] In this way, the translated text storage unit 33 can store the translated text data Dptt acquired for each translation target language.
[0078] The translated speech acquisition unit 25 acquires speech in the target language. The translated speech acquisition unit 25 acquires translated speech data Dptv by having the speech synthesis engine 4 perform speech synthesis of the translated text data Dptt. The translated speech acquisition unit 25 receives the target language information and the translated text data Dptt transmitted from the translated text acquisition unit 23. The translated speech acquisition unit 25 transmits the translated text data Dptt to the speech synthesis engine 4 of the target language determined by the target language information received from the translated text acquisition unit 23, and acquires the translated speech data Dptv in the target language. The translated speech acquisition unit 25 then transmits the text identification information, the target language information, and the translated speech data Dptv to the distribution translation information generation unit 51 of the distribution unit 50.
[0079] For example, in the situation shown in FIG. 2 , if the target languages are English and Chinese, the translated speech acquisition unit 25 receives from the translated text acquisition unit 23 a text ID, English text data Det, and target language information "English," as well as the same text ID, Chinese text data Dct, and target language information "Chinese." The translated speech acquisition unit 25 then determines the English speech synthesis engine 4E and the Chinese speech synthesis engine 4C based on the target language information. Next, the translated speech acquisition unit 25 transmits the English text data Det to the English speech synthesis engine 4E for speech synthesis, thereby obtaining English speech Dev, which is translated English speech data Dvr. The translated speech acquisition unit 25 also transmits the Chinese text data Dct to the Chinese speech synthesis engine 4C for speech synthesis, thereby obtaining Chinese speech Dcv, which is translated Chinese speech data Dptv. Then, the translated speech acquisition unit 25 transmits the target language information “English” and the English speech Dev, and the target language information “Chinese” and the Chinese speech Dcv to the distribution translation information generation unit 51 .
[0080] Next, a description will be given of the function of the translation distribution device 100 to manage the identification information and participant languages of the participant terminals 6 at the translation providing location A. Fig. 9 is a functional block diagram showing in detail the translation distribution information management unit 40 and distribution unit 50 in Fig. 5 .
[0081] The translation distribution information management unit 40 manages information of the participant terminals 6 that distribute translation information. The translation distribution information management unit 40 includes a language information storage unit 43, an individual language information acquisition unit 41, an individual communication status acquisition unit 45, and a translation target language management unit 47. The individual language information acquisition unit 41, the individual communication status acquisition unit 45, and the translation target language management unit 47 are implemented mainly in the processor 100a and the communication unit 100c. The language information storage unit 43 is implemented mainly in the storage unit 100b.
[0082] The distribution unit 50 distributes the translation information to the participant terminal 6. The distribution unit 50 includes a distribution translation information generation unit 51 and a log request processing unit 53. The distribution translation information generation unit 51 and the log request processing unit 53 are implemented mainly by the processor 100a and the communication unit 100c.
[0083] The distribution translation information generation unit 51 generates translation information data Dpt for each target language. The distribution translation information generation unit 51 generates translation information data Dpt for each target language by combining the translated text data Dptt and the translated information obtained by accessing the translated text storage unit 33, and the translated voice data Dptv transmitted by the translated voice acquisition unit 25. The distribution translation information generation unit 51 may also obtain the latest translated text data Dptt stored in the translated text storage unit 33.
[0084] For example, translated text data Dptt and translated speech data Dptv that are associated with the same target language and text identification information may be combined into the same translation information data Dpt. Note that the translation information data Dpt may include at least either the translated text data Dptt or the translated speech data Dptv.
[0085] The distribution translation information generation unit 51 then transmits the speech recognition text data Dvr and the translation information data Dpt to the participant terminal 6. The distribution translation information generation unit 51 accesses the language information storage unit 43, which will be described later, to identify the translation target language of the participant terminal 6. The distribution translation information generation unit 51 then transmits the translation information data Dpt and speech recognition text data Dvr in the identified translation target language to the participant terminal 6.
[0086] Furthermore, when a notification of transmission of past translated text is received from the log request processing unit 53 (described later), the past translated text data Dptt stored in the translated text storage unit 33 may be transmitted.
[0087] Here, the translation and delivery system 1 has a log transmission function that accepts transmission requests from the participant terminal 6 and transmits past translated text. The log request processing unit 53 accepts a predetermined transmission request from the participant terminal 6 and identifies the past translated text requested by the participant terminal 6. Then, it sends a transmission request to the delivery translation information generation unit 51 to send the identified past translated text. The log transmission function will be described in detail later.
[0088] The individual language information acquisition unit 41 initiates communication with the participant terminal 6 of a participant who has visited the translation provision location A, and acquires the terminal identification information of the participant terminal 6 and the participant language of the participant who possesses the participant terminal 6.
[0089] Here, for example, in the situation shown in Figure 1, a participant visiting translation provider location A reads a two-dimensional barcode 103 displayed within translation provider location A with the participant terminal 6 in order to use a translation, thereby accessing the translation distribution device 100 having the address specified by the two-dimensional barcode 103 and starting communication with the individual language information acquisition unit 41.
[0090] The terminal identification information of the participant terminal 6 acquired by the individual language information acquisition unit 41 need only be information that can identify the participant terminal 6, and may be a unique number assigned to the participant terminal 6, or a number assigned by the individual language information acquisition unit 41 to manage the participant terminal 6 may be used as the terminal identification information.
[0091] The participant language is primarily the participant's native language, but may also be the language that the participant normally uses on their own participant terminal 6. The individual language information acquisition unit 41 may acquire the participant language, for example, by displaying a language selection screen on the browser and having the participant select a language. Alternatively, the participant language may be acquired by displaying an input screen on the browser and having the participant input the type of language. Furthermore, the participant language may be acquired by reading the language setting of the operating system or browser of the participant terminal 6.
[0092] The individual language information acquisition unit 41 then stores the terminal identification information associated with the acquired participant language in the language information storage unit 43. Based on communication with the participant terminal 6, the individual language information acquisition unit 41 stores the participant's terminal identification information in association with the participant language, at least when the participant language and the speaker language are different. In this way, it is possible to manage the participant languages of participants who visit the translation providing location A. Furthermore, even when the participant language and the speaker language are the same, the terminal identification information may be stored in the language information storage unit 43.
[0093] The language information storage unit 43 stores terminal identification information that identifies each participant terminal 6 owned by one or more participants visiting the translation providing location A, in association with the participant language. The language information storage unit 43 may store a combination of terminal identification information and participant language. Furthermore, for example, the language information storage unit 43 may prepare a database for each language and store the terminal identification information of the participant terminal 6 in association with that language.
[0094] The individual communication status acquisition unit 45 repeatedly determines whether communication has ended with each participant terminal 6 having terminal identification information stored in the language information storage unit 43. The individual communication status acquisition unit 45 may determine whether communication has ended by, for example, periodically sending a ping signal to the participant terminal 6.
[0095] When communication with a participant terminal 6 is terminated, the individual communication status acquisition unit 45 deletes or invalidates the terminal identification information stored in the language information storage unit 43. Then, when the individual communication status acquisition unit 45 confirms that communication has been terminated, it deletes or invalidates the terminal identification information held by that participant terminal 6.
[0096] Deleting terminal identification information means deleting the terminal identification information stored in the language information storage unit 43. Invalidating terminal identification information means setting a flag for communication termination when communication ends in a situation where a flag for communication termination is managed in association with terminal identification information in the language information storage unit 43. In this way, the participant terminal 6 for which communication has ended will no longer be counted in the number of participants by language information described below.
[0097] The target language management unit 47 manages one or more target languages into which a speaker language is translated. The target language management unit 47 includes a language-specific number-of-people management unit 47a and a target language determination unit 47b.
[0098] The number-of-participants-by-language management unit 47a acquires number-of-participants-by-language information, which is the number of participant terminals 6 identified by each of the terminal identification information stored in the language information storage unit 43 for each participant language, based on the contents of the language information storage unit 43.
[0099] The target language determination unit 47b determines, at a given timing, a participant language different from the speaker language associated with one or more participant terminals 6 as one or more target languages based on the contents of the language information storage unit 43. The target language determination unit 47b may determine, as the target language, a language other than the speaker language notified by the language determination unit 13 and for which the language-specific number-of-participants information acquired by the language-specific number-of-participants management unit 47a indicates one or more participant languages. The target language may also be determined to be one or more participant languages associated with terminal identification information stored in the language information storage unit 43. Note that the operator of the translation providing site A (not shown) may set any language as the target language using an input device (not shown).
[0100] FIG. 10( a) shows an example of the terminal identification information of multiple participant terminals 6 stored in the language information storage unit 43 in the situation shown in FIG. 1. In the situation shown in FIG. 1, the translation provision location A includes a speaker SJ whose speaker language is Japanese, three participants PJ whose participant language is Japanese, and two participants PE whose participant language is English. Each participant performs a predetermined operation to initiate communication with the individual language information acquisition unit 41. The individual language information acquisition unit 41 then assigns a terminal ID, which is terminal identification information, to each participant terminal 6 held by the participant and acquires each participant language. The individual language information acquisition unit then associates the terminal ID of each participant terminal 6 with the participant language and stores the associated ID in the language information storage unit 43. In FIG. 10( a), the participant terminal 6E, which has the terminal ID "T001" assigned as its terminal identification information, is stored in association with the participant language, which is English.
[0101] Fig. 10(b) is an example of the number-of-people information by language stored in the language information storage unit 43 in the situation of Fig. 1. Based on the terminal identification information (see Fig. 10(a)), the number-of-people-by-language management unit 47a obtains, as the number-of-people information by language, the number of Japanese participants, "3", the number of English participants, and the number of other language participants, "0", and stores these in the language information storage unit 43.
[0102] As shown in FIG. 10B, six languages may be planned as participant languages. Furthermore, when a new, unplanned language is detected based on the contents of the language information storage unit 43, the new language may be added. Participant languages may be assigned language IDs for management. In this embodiment, for example, the participant language "Japanese" is assigned the language ID "L001."
[0103] The target language determination unit 47b sets English, which is a language with one or more participants and is other than Japanese, which is the speaker language notified to the language determination unit 13, as the target language. As a result, in the situation of Figure 1, the speech of speaker SJ is translated into English and distributed.
[0104] FIG. 11(a) is an example of the terminal identification information of multiple participant terminals 6 stored in the language information storage unit 43 in the situation of FIG. 2. The difference between FIG. 2 and FIG. 1 is that a participant PC whose participant language is Chinese is present at translation provision location A. When the participant PC performs a predetermined operation, communication with the individual language information acquisition unit 41 is initiated. Then, the individual language information acquisition unit 41 assigns the terminal ID "T006" as the terminal identification information of the participant terminal 6C possessed by the participant PC, and acquires Chinese, which is the participant language. The individual language information acquisition unit 41 associates the terminal ID "T006" with Chinese and stores it in the language information storage unit 43.
[0105] Fig. 11(b) is an example of the number-of-people information by language stored in the language information storage unit 43 in the situation of Fig. 2. Based on the terminal identification information (see Fig. 11(a)), the number-of-people-by-language management unit 47a obtains the number of Japanese participants "3", the number of English participants "2", the number of Chinese participants "1", and the number of other language participants "0" as the number of people-by-language information, and stores these in the language information storage unit 43.
[0106] The target language determination unit 47b sets English and Chinese as target languages, which are languages with one or more participants other than Japanese, which is the speaker language notified to the language determination unit 13. As a result, in the situation of Figure 2, the Japanese speech of speaker SJ is translated into English and Chinese and distributed.
[0107] Next, the determination of the translation target language when the participant PC leaves the translation provision location A in the situation of Fig. 2 will be explained. The individual communication status acquisition unit 45 detects that communication with the participant terminal 6C possessed by the participant PC has ended, and deletes the row for the terminal ID "T006," which is the terminal identification information of the participant terminal 6C stored in the language information storage unit 43. As a result, when the participant PC leaves the translation provision location A in the situation of Fig. 2, the terminal identification information and language-specific number-of-person information shown in Figs. 10(a) and 10(b) are stored in the language information storage unit 43.
[0108] Based on the terminal identification information (see FIG. 10(a)), the language-specific number-of-participants management unit 47a obtains the number of Japanese participants, "3," the number of English participants, "0," the number of Chinese participants, and the number of other language participants, "0," as language-specific number-of-participants information, and stores these in the language information storage unit 43. The translation-target language determination unit 47b then sets English, which is a language other than Japanese, the speaker language notified to the language determination unit 13, and which has one or more participants, as the translation-target language. As a result, if the participant PC leaves the translation-providing location A in the situation of FIG. 2, the translation into Chinese will be stopped, and the Japanese speech of speaker SJ will be translated into English and distributed.
[0109] In this way, by managing the current participant languages of one or more participants visiting translation provider location A and determining the translation target language, translation into participant languages that currently require translation can be performed while translating into unnecessary participant languages can be stopped, thereby saving resources.
[0110] 11(a) and 11(b) are also examples of terminal identification information and language-specific number-of-participants information stored in the language information storage unit 43 in the situation of FIG. 3. The difference between FIG. 3 and FIG. 2 is that the speaker is the speaker SE. In the situation of FIG. 3, the language determination unit 13 determines that the speaker language is English based on the setting input by the operator of the translation provision site A. Then, the language determination unit 13 notifies the translation target language management unit 47 that the speaker language is English. The translation target language determination unit 47b sets Japanese and Chinese as the translation target languages, which are languages other than English, the speaker language notified to the language determination unit 13, and which have one or more participants.
[0111] In this way, even if the speaker's language changes, the current participant language of one or more participants visiting translation provision location A is managed and the translation target language is determined, so that the translation can be delivered in the participant language required by the participant.
[0112] Next, the process in which the translation and delivery system 1 generates one or more translated texts after the target language has been determined will be described with reference to the flow diagram shown in FIG.
[0113] The target language management unit 47 determines the participant language associated with one or more participant terminals 6 as the target language based on the contents stored in the language information storage unit 43 (S101).
[0114] The voice input receiving unit 11 acquires a voice-recognized text by having the voice recognition engine 2 recognize the voice based on a part of the speaker's voice (S102).
[0115] The target language management unit 47 monitors whether the contents of the language information storage unit 43 have been updated (S103). If the contents of the language information storage unit 43 have not been updated (N in S103), the target language management unit 47 does not update the target language and maintains the current target language (S104). If the contents of the language information storage unit 43 have been updated (Y in S103), the target language management unit 47 updates the target language based on the contents of the language information storage unit 43 (S105).
[0116] The translated text acquisition unit 23 transmits the speech recognition text acquired in the processing of S102 to the translation engine 3 of the target language determined in the processing of S104 or S105, acquires the translated text (S106), and the processing shown in this processing example is terminated.
[0117] The speech-recognition text data Dvr and the translation information data Dpt sent by the distribution translation information generation unit 51 are displayed as a translation page 60 on the display of the participant terminal 6 using a browser application or the like.
[0118] The translation page 60 is a page displayed on the display of the participant terminal 6, and participants check the translation information sequentially transmitted via the translation page 60. As shown in FIG. 13 , the translation page 60 includes a speech recognition text display area 61, a translated text display area 63, a translation target language display button 65, and a translated voice switch button 67.
[0119] The speech-recognized text display area 61 displays speech-recognized text that has been speech-recognized based on a portion of the speaker's speech. The speech-recognized text display area 61 displays speech-recognized text that is sequentially transmitted from the translation and distribution system 1. The latest speech-recognized text is displayed at the bottom of the speech-recognized text display area 61.
[0120] The translated text display area 63 displays text translated in the participant language of the participant who has the participant terminal 6, based on a portion of the speaker's speech. The translated text display area 63 displays translated text that is sequentially transmitted from the translation distribution system 1. The translated text display area 63 displays the most recent translated text at the top, with translated text transmitted earlier displayed below. Therefore, participants can scroll down the translated text display area 63 to go back and check the translated text that has been transmitted.
[0121] The translation target language display button 65 displays the language in which the translated text displayed on the translation page 60 has been translated. By tapping the translation target language display button 65, it is possible to select another language.
[0122] The translated audio switch button 67 displays an icon such as a speaker, allowing the user to select whether to play the translated audio. The participant PE can listen to the translated audio by using the audio output device 7.
[0123] 13 and 14 are diagrams showing examples of the screen of the translation page 60 displayed on the participant terminal 6E possessed by the participant PE and the participant terminal 6C possessed by the participant PC.
[0124] In the situation of Figure 2, if a participant PE understands in English the speech content of a speaker SJ whose speaker language is Japanese, the Japanese speech recognition text is displayed in the speech recognition text display area 61, and the translated English text is displayed in the translated text display area 63, as shown in Figure 13. The participant PE can confirm that the speech has been translated into English by pressing the translation target language display button 65, and can select whether or not to listen to the translated English audio by pressing the translated audio switch button 67.
[0125] In the situation of Figure 2, if the participant PC understands the speech content of speaker SJ, whose speaker language is Japanese, in Chinese, the Japanese speech recognition text is displayed in the speech recognition text display area 61, and the translated Chinese text is displayed in the translated text display area 63, as shown in Figure 13. The participant PC can confirm that the translation has been made in Chinese by pressing the translation target language display button 65, and can select whether or not to listen to the translated Chinese audio by pressing the translated audio switch button 67.
[0126] Here, the log transmission function mentioned above will be explained.
[0127] 15A shows an example of a translation page 60 displayed on a participant terminal 6E owned by a participant PE. The translation page 60 in the figure displays the latest speech-recognized text in a speech-recognized text display area 61 and the latest English text data Det in a translated text display area 63.
[0128] Next, when the participant PE performs a log checking operation such as scrolling back in the translated text display area 63, the English text that is the past translation information is displayed below the English text in the translated text display area 63, as shown in Figure 15(b). This allows the participant PE to understand what the speaker SJ has said in the past in English.
[0129] After the past translation information is displayed, further past translation information may be displayed by performing a log confirmation operation. In this embodiment, the past translation information transmitted by a single log confirmation operation may be one translated text or multiple translated texts. Furthermore, when a predetermined log confirmation operation is performed, all past translation information stored in the translated text storage unit 33 may be transmitted as past translation information.
[0130] The log transmission operation for transmitting past translation information in response to a request from a participant who visits translation providing location A will be described.
[0131] The log request processing unit 53 (see FIG. 8) of the distribution unit 50 receives a predetermined transmission request from the participant terminal 6 whose terminal identification information is stored in the language information storage unit 43 .
[0132] A transmission request is a request for past translation information from the participant terminal 6. The transmission request includes the oldest text identification information included in the translation information data Dpt received by the participant terminal 6, the terminal identification information of the participant terminal 6, and the participant language of the participant terminal 6. Because the participant terminal 6 has transmitted translated text data Dptt in which the text identification information and the translated text are associated, or translated audio data Dptv in which the text identification information and the translated audio are associated, it is possible to identify the oldest translated text or translated audio among the translated texts or translated audio received by the participant terminal 6.
[0133] When the log request processing unit 53 receives a transmission request, it identifies the past translation information for which the transmission request was received from the text identification information included in the transmission request. The log request processing unit 53 identifies the past translation information for which the transmission request was received as the past translation information that is older than the text identification information received as the transmission request. For example, if the transmission request includes text ID "n", the log request processing unit 53 identifies the translation information associated with text ID "n-1", which is older than the text ID "n", as the past translation information for which the transmission request was received.
[0134] Then, the log request processing unit 53 accesses the translated text storage unit 33 and detects whether there is any translated text that has been requested in the past.
[0135] If the translated text storage unit 33 contains previously requested translated text, the log request processing unit 53 sends a notification of transmission of the past translated information to the distribution translation information generation unit 51. The notification of transmission of the past translated information includes text identification information of the past translated information, the target language, and the terminal identification information of the participant terminal 6 that sent the transmission request. Then, based on the transmission notification, the distribution translation information generation unit 51 transmits the past translated text data Dptt in the target language acquired from the translated text storage unit 33 to the participant terminal 6 that sent the transmission request.
[0136] Here, the past translated text sent in response to a single transmission request may be one translated text or multiple translated texts. Furthermore, for example, when a specific log confirmation operation is performed, the distribution translation information generation unit 51 may send all translated texts stored in the translated text storage unit 33.
[0137] If there is no previous translated text for the requested information, the log request processing unit 53 sends a translation generation notification to the previous translated text acquisition unit 23a (see FIG. 6 ), which translates the speech-recognition text corresponding to the previous translation information for the requested information, and generates the previous translated text. The translation generation notification includes the participant language of the participant terminal 6 that sent the request and text identification information for the previous translation information.
[0138] The previous translated text acquisition unit 23a accesses the speech-recognition text storage unit 31 and acquires speech-recognition text data Dvr corresponding to the requested previous translated text, identified from the previous translation information text identification information included in the translation generation notification. The speech-recognition text data Dvr corresponding to the previous translated text is, for example, speech-recognition text data Dvr associated with the same text identification information. The translation engine determination unit 21 then determines a translation engine 3 that will translate the acquired speech-recognition text data Dvr from the speaker language to the participant language included in the translation generation notification. The translated text acquisition unit 23 then transmits the speech-recognition text of the acquired speech-recognition text data Dvr to the translation engine determined by the translation engine determination unit 21, generating the previous translated text. The translated text acquisition unit 23 transmits the generated previous translated text to the translated text storage unit 33 for storage.
[0139] Then, the log request processing unit 53 transmits a transmission notification of the past translation information to the distribution translation information generation unit 51. Then, based on the transmission notification, the distribution translation information generation unit 51 transmits the past translated text data Dptt in the target language acquired from the translated text storage unit 33 to the participant terminal 6 that transmitted the transmission request.
[0140] In this way, participants who visit translation providing location A can check past translated text. Even if there is no past translated text in the translated text storage unit 33, past translated text is generated and stored in the translated text storage unit 33, eliminating the need to translate every time a transmission request is received and saving resources.
[0141] 15, one translated text is sent for one transmission request. A case will be described in which a participant PE having a participant terminal 6E with a terminal ID of "T004" makes a predetermined transmission request.
[0142] In this situation, the translated text storage unit 33 stores a translated text group 33D shown in FIG. 8B. The translated text group D33 includes English text data Det and Dct. Because the presence of participant PE at translation provision location A was detected before speaker SJ spoke, the "English text data" column includes English text data Det associated with text IDs "1" to "4." Meanwhile, because the presence of participant PC at translation provision location A was detected when the content of text ID "4" was spoken, the "Chinese text data" column does not include Chinese text associated with text IDs "1" to "3," but only includes Chinese text associated with text ID "4."
[0143] When a participant PE performs a log confirmation operation, such as scrolling back, in the translated text display area 63 of the translation page 60 displayed on the participant terminal 6E, a predetermined transmission request is sent to the log request processing unit 53. In this embodiment, English text associated with text ID "4" is sent to the participant terminal 6E as translated text data Dptt (see FIG. 16(a)). Therefore, the text ID of the oldest translated text held by the participant terminal 6E is "4." Therefore, the transmission request from the participant terminal 6E includes the text ID "4," which is the oldest text identification information, the terminal ID "T004," which is terminal identification information, and English, which is the participant language.
[0144] The log request processor 53 receives a transmission request from the participant terminal 6 E. Because the transmission request includes information that the participant language is English and that the oldest text ID is “4,” the log request processor 53 detects whether the translated English text with the text ID “3” is in the translated text storage unit 33.
[0145] 15, because the translated text storage unit 33 stores English texts with text IDs "1" to "4," the log request processing unit 53 detects that an English text with text ID "3" is in the translated text storage unit 33. Because there is a previous translated text for which a request has been received, the log request processing unit 53 notifies the distribution translation information generation unit 51 of the text ID "3," which is the text identification information of the previous translated text, the translation target language (English), and the terminal ID "T004," which is the terminal identification information of the participant terminal 6E that sent the transmission request.
[0146] Upon receiving notification of the transmission of past translated text from the log request processing unit 53, the distribution translation information generation unit 51 accesses the translated text storage unit 33 and transmits the English text data Det of the notified English text with text ID "3" to the participant terminal 6E with terminal ID "T004." Upon receiving the translation information data Dpt, the participant terminal 6E displays the English text below the English text in the translated text display area 63 of the translation page 60. This allows the participant PE to understand the content of the speaker SJ's past utterances in English.
[0147] Next, we will explain what happens when a participant PC having a participant terminal 6C with a terminal ID of "T006" makes a specific transmission request. When the participant PC performs a log confirmation operation, such as scrolling back, in the translated text display area 63 of the translation page 60 displayed on the participant terminal 6C, the specific transmission request is sent to the log request processing unit 53.
[0148] Chinese text associated with text ID "4" has been sent to participant terminal 6C as translated text data Dptt. Therefore, the text ID of the oldest translated text held by participant terminal 6C is "4." The send request from participant terminal 6C includes the text ID "4" which is the oldest text identification information, terminal ID "T006" which is terminal identification information, and Chinese which is the participant language.
[0149] The log request processor 53 receives a transmission request from the participant terminal 6 E. Because the transmission request includes information that the participant language is Chinese and that the oldest text ID is “4,” the log request processor 53 detects whether the translated Chinese text with the text ID “3” is in the translated text storage unit 33.
[0150] 8B , there is no Chinese text with text ID "3" in the translated text storage unit 33, so the log request processing unit 53 detects that there is no Chinese text with text ID "3" in the translated text storage unit 33. Because there is no previous translated text that has been requested, the log request processing unit 53 instructs the previous translated text acquisition unit 23a to generate Chinese text with text ID "3".
[0151] The translation engine determination unit 21 accesses the speech-recognition text storage unit 31 and acquires the speech-recognition text data Dvr of text ID "3" and the speaker language information "Japanese." The translation engine determination unit 21 determines a Japanese-Chinese translation engine 3JC for translating from Japanese, the speaker language, to Chinese, the participant language, and notifies the translated text acquisition unit 23. The translated text acquisition unit 23 then transmits the speech-recognition text data Dvr of text ID "3" to the translation engine 3JC and acquires the Chinese text. The translated text acquisition unit 23 transmits the acquired Chinese text data Dct and text ID "3" to the translated text storage unit 33 for storage.
[0152] The translated text storage unit 33 stores the transmitted Chinese text data Dct in the "Chinese text" column of text ID "3." As shown in Figure 16, the Chinese text data Dct of text ID "3," which is the requested past translation information, is stored in the translated text storage unit 33.
[0153] Then, the log request processing unit 53 notifies the distribution translation information generation unit 51 of the text ID "3" which is the text identification information of the past translated text, the target language (Chinese), and the terminal ID "003" which is the terminal identification information of the participant terminal 6C that sent the transmission request.
[0154] Upon receiving notification of the transmission of past translated text from the log request processing unit 53, the distributed translation information generation unit 51 accesses the translated text storage unit 33 and transmits the Chinese text data Dct of the notified Chinese text with text ID "3" to the participant terminal 6C with terminal ID "003." On the participant terminal 6C that receives the translated text data Dptt, the Chinese text is displayed below the Chinese text in the translated text display area 63 of the translation page 60. This allows the participant PC to understand the past translation information in Chinese.
[0155] Next, the process in which the translation and distribution system 1 transmits past translation information after receiving a transmission request from the participant terminal 6 will be described with reference to the flowchart shown in FIG.
[0156] The log request processing unit 53 receives a transmission request from a participant terminal 6 whose terminal identification information is stored in the language information storage unit 43 (S201).
[0157] The log request processing unit 53 identifies the participant language of the participant terminal 6 and the requested past translation information based on the content included in the transmission request received in S201 (S202).
[0158] The log request processing unit 53 detects whether or not past translation information in the participant language of the participant terminal 6 identified in S202 is stored in the translated text storage unit 33 (S203). If past translation information in the participant language of the participant terminal 6 is not stored in the translated text storage unit 33 (S203: N), the log request processing unit 53 has the past translated text acquisition unit 23a translate the speech recognition text corresponding to the past translation information, thereby generating past translation information in the participant language of the participant terminal 6 and storing it in the translated text storage unit 33 (S204).
[0159] The log request processing unit 53 causes the distribution translation information generation unit 51 to transmit the past translation information in the participant language of the participant terminal 6 stored in the translated text storage unit 33 (S205), and the processing shown in this processing example is terminated.
[0160] In this embodiment, a translation page 60 in Fig. 18 is shown as a modified example of the translation page 60 in Fig. 13 . A speech-recognition text display area 61 of the translation page 60 in Fig. 18 displays Japanese speech-recognition text and English speech-recognition text. Furthermore, a translated text display area 63 of the translation page 60 displays Chinese text obtained by translating the Japanese speech-recognition text and the English speech-recognition text in the speech-recognition text display area 61 into Chinese. When a speaker speaks a mixture of multiple languages as in Fig. 18 , the language determination unit 13 can distinguish between parts of the speaker's speech spoken in Japanese and parts spoken in English and perform speech recognition in each of the speaker's languages by, for example, determining the speaker language for each silent section of the speaker's speech.
[0161] Even when the speaker's language changes frequently, the translation and distribution device 100 selects an appropriate speech recognition engine 2 based on the speaker's language determined by the language determination unit 13. Furthermore, the translation and distribution device 100 selects an appropriate translation engine 3 according to the combination of the speaker's language and the participant's language, and therefore generates translation information in an appropriate participant language according to the demand of the translation provision site A and transmits it to the participant terminal 6. In this way, even when the speaker's language changes frequently, participants can understand content spoken in various languages in their own participant language.
[0162] In this embodiment, it is assumed that the speaker is at translation service location A, but the speaker may also transmit voice data representing the speaker's voice from a remote location via the Internet or a telephone line to translation service location A. In this case, the voice data representing the speaker's voice may be directly input to translation and distribution device 100.
[0163] Furthermore, at the translation providing location A, speech-recognition text may be sent as translated text to the participant terminal 6 of a participant whose participant language is the same as the speaker's language. In this case, the speech-recognition text in the speaker's language is displayed in the translated text display area 63 of the translation page 60 displayed by the participant terminal 6. When the participant language is the same as the speaker's language, the participant can check the speaker's spoken content in text.
[0164] The present invention is not limited to the above-described embodiment. In this embodiment, an example of translating Japanese into English or Chinese has been described, but translation information can be similarly generated and distributed for other language combinations.
[0165] Furthermore, the specific text and numerical values above and in the drawings are examples, and the present invention is not limited to these text and numerical values.
Claims
1. At a location where the speaker's voice is provided, terminal identification information storage means for storing, in association with the language used by each user, identification information for identifying each terminal of one or more users who each utilize the translation of the voice; first update means for obtaining, based on communication with the terminal of a user who has visited the location, the identification information of the terminal and the language used by the user who uses the terminal, and updating the content of the terminal identification information storage means so as to store the terminal identification information of the user in association with the language used by the user, at least when the language used by the user is different from the language used by the speaker; second update means for repeatedly determining whether communication with the terminal identified by each of the identification information stored in the terminal identification information storage means has ended, and when the communication has ended, updating the content of the terminal identification information storage means so as to delete or invalidate the identification information; voice acquisition means for sequentially acquiring the speaker's voice in portions; translation destination language determination means for determining, at a given timing, one or more translation destination languages that are languages used by the users different from the language used by the speaker and are associated with one or more terminals based on the content of the terminal identification information storage means; translation means for generating one or more pieces of translation information that are partial translations in text or voice of the speaker's voice into each of the one or more translation destination languages based on a portion of the speaker's voice sequentially acquired by the voice acquisition means; and transmission means for transmitting, to the terminal identified by each of the identification information stored in the terminal identification information storage means, the translation information in the language associated with the identification information among the one or more pieces of translation information. A translation distribution system including the above components.
2. The translation distribution system according to claim 1, further comprising translation information storage means for associating and storing, based on a part of the speaker's voice sequentially acquired by the voice acquisition means, at least the last generated one or more pieces of translation information among the one or more pieces of translation information generated by the translation means, with the target language of translation corresponding to the translation information; The transmission means transmits, to each terminal identified by each of the identification information stored in the terminal identification information storage means, the translation information stored in the translation information storage means in association with the target language of translation which is the language of use associated with the identification information, among the one or more pieces of translation information. Translation distribution system.
3. The translation distribution system according to claim 2, wherein The translation information storage means stores, in order, one or more pieces of translation information generated by the translation means, in association with the target language of translation corresponding to the translation information, based on a part of the speaker's voice sequentially acquired by the voice acquisition means; The transmission means transmits, to the terminal that has transmitted a predetermined transmission request, the past translation information stored in the translation information storage means in association with the target language of translation which is the language of use associated with the identification information of the terminal. Translation distribution system.
4. The translation distribution system according to claim 3, further comprising translation information detection means for detecting whether or not request translation information, which is one or more pieces of past translation information requested by the user to be transmitted, in response to the predetermined transmission request from the terminal, is included in the past translation information stored in the translation information storage means in association with the target language of translation which is the language of use associated with the identification information of the terminal. Translation distribution system.
5. The translation distribution system according to claim 4, wherein When the translation information detection means determines that the request translation information is not included in the past translation information stored in the translation information storage means in association with the target language of translation which is the language of use associated with the identification information of the terminal, the translation information detection means causes the translation means to generate the past translation information corresponding to the request translation information. Translation distribution system.
6. The translation distribution system according to claim 1, further comprising language determination means for determining the language used by the speaker by inputting a part of the speaker's voice into a learned machine model.
7. The translation distribution system according to claim 1, wherein for each terminal identified by each of the identification information stored in the terminal identification information storage means, together with the translation information in the language used associated with the identification information, a speech recognition text speech-recognized based on a part of the speaker's voice is transmitted.
8. The translation distribution system according to claim 1, wherein the one or more translation information includes one or more translated texts that are texts generated based on a part of the voice acquired by the voice acquisition means, and the translation means generates the one or more translated texts by causing the speech recognition text obtained by speech recognition based on a part of the speaker's voice to be translated by each of the translation engines in the one or more target languages.
9. The translation distribution system according to claim 8, wherein the one or more translation information includes one or more translated voices that are voices generated based on a part of the voice acquired by the voice acquisition means, and the translation means generates the one or more translated voices by causing the one or more translated texts to be speech-synthesized by each of the speech synthesis engines in the one or more target languages.
10. A terminal identification information storage step of associating and storing identification information for identifying each of one or more users' terminals that use the translation of the voice, respectively, with the language used by the user at the location where the speaker's voice is provided; A first update step of obtaining the identification information of the terminal and the language used by the user using the terminal based on communication with the terminal of the user who has visited the location, and updating the content of the terminal identification information storage means so as to store the terminal identification information of the user in association with the language used by the user, at least when the language used by the user is different from the language used by the speaker; A second update step of repeatedly determining whether communication with the terminal identified by each of the identification information stored in the terminal identification information storage means has ended, and if the communication has ended, updating the content of the terminal identification information storage means so as to delete or invalidate the identification information; A voice acquisition step of sequentially acquiring the speaker's voice in part; A translated language determination step of determining, at a given timing, one or more translated languages that are languages used different from the language used by the speaker, each associated with one or more terminals, based on the content of the terminal identification information storage means; A translation step of generating one or more translation information that are partial translations in text or voice to each of the one or more translated languages of the speaker's voice based on a part of the speaker's voice sequentially acquired by the voice acquisition means; A transmission step of transmitting, to the terminal identified by each of the identification information stored in the terminal identification information storage means, the translation information in the language associated with the identification information among the one or more translation information. A translation distribution method comprising the above steps.
11. Terminal identification information storage means for storing, in association with the language used by each user, identification information for identifying each terminal of one or more users who use the translation of the voice at the location where the speaker's voice is provided; First update means for obtaining, based on communication with the terminal of a user who has visited the location, the identification information of the terminal and the language used by the user who uses the terminal, and updating the content of the terminal identification information storage means so as to store the terminal identification information of the user in association with the language used by the user, at least when the language used by the user is different from the language used by the speaker; Second update means for repeatedly determining whether communication with the terminal identified by each of the identification information stored in the terminal identification information storage means has ended, and when the communication has ended, updating the content of the terminal identification information storage means so as to delete or invalidate the identification information; Voice acquisition means for sequentially acquiring the voice of the speaker in parts; Translation destination language determination means for determining, at a given timing, one or more translation destination languages that are languages used by the user different from the language used by the speaker and are associated with one or more terminals based on the content of the terminal identification information storage means; Translation means for generating one or more pieces of translation information that are partial translations of the voice of the speaker into text or voice in each of the one or more translation destination languages based on a part of the voice of the speaker sequentially acquired by the voice acquisition means; and Transmission means for transmitting, to the terminal identified by each of the identification information stored in the terminal identification information storage means, the translation information in the language associated with the identification information among the one or more pieces of translation information. A program for causing a computer to function as described above.
Citation Information
Patent Citations
Translation management system
JP2021086264A
Translation device, translation method, and recording medium
WO2021149267A1
Communication system
WO2022038928A1