Translation management system

The translation management system addresses redundant processing in speech translation devices by grouping terminals by language and collectively performing translation, improving efficiency and reducing processing load.

JP2026012242AActive Publication Date: 2026-01-23TOPPAN HOLDINGS INC
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2025179846
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-01-23
Estimated Expiration
2039-11-26

AI Technical Summary

Technical Problem

Existing speech translation devices require redundant processing when translating into multiple languages, leading to inefficiencies and increased processing load due to duplicated speech recognition, translation, and speech synthesis tasks for participants speaking the same language.

Method used

A translation management system that groups terminals based on language distribution, collectively performs translation processing, and distributes results to terminals within the same language group, reducing duplication and processing load.

Benefits of technology

The system enhances translation efficiency by minimizing redundant processing, optimizing data exchange, and simplifying terminal configurations, making it suitable for temporary activities like sightseeing tours.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012242000001_ABST
    Figure 2026012242000001_ABST
Patent Text Reader

Abstract

To provide a translation management system capable of improving efficiency of processing related to translation into a plurality of languages.SOLUTION: A translation management system 50 sets a translation group to which a first terminal and a plurality of second terminals that are distribution targets of a translation result of first information input to the first terminal belong, and sets a translation target language of the first information in the translation group in association with a type of a distribution language set for each second terminal as a language used for distribution. Further, when receiving the first information from the first terminal, the translation management system 50 instructs translation of the first information into the translation target language of the translation group to which the first terminal belongs, and distributes the translation result of each translation target language to the second terminal in which the translation target language is set as the distribution language.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a translation management system for managing translations into multiple languages. [Background technology]

[0002] In recent years, as communication between people who speak different languages ​​has become more active, portable speech translation devices have been developed. A speech translation device recognizes speech in a first language, converts it into text, and translates the text in the first language into text in a second language. The speech translation device then synthesizes and outputs speech in the second language from the text in the second language (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2018-124695 Summary of the Invention [Problem to be solved by the invention]

[0004] There are an increasing number of occasions where it is desirable to translate the content of a speaker's speech into multiple languages. For example, with the diversification of sightseeing tours, it is becoming more common for speakers of various languages ​​to participate in a single tour. Tour participants may include multiple participants who speak the same language as each other, as well as multiple participants who speak different languages.

[0005] In such tours, the tour guide's speech needs to be translated into multiple languages ​​depending on the participants' languages. While such translation is possible if each participant carries a speech translation device, performing speech recognition, translation, and speech synthesis on each speech translation device results in redundant processing, resulting in a lot of waste. For example, speech recognition processing is duplicated on the speech translation devices of all participants. Furthermore, translation processing and speech synthesis processing are duplicated on the speech translation devices of participants who speak the same language.

[0006] An object of the present invention is to provide a translation management system that can improve the efficiency of processing related to translation into multiple languages. [Means for solving the problem]

[0007] A translation management system that solves the above problem includes a group management unit that sets a translation group to which a first terminal and a plurality of second terminals to which the translation results of first information input to the first terminal belong, and sets the translation target language of the first information in the translation group to correspond to the type of distribution language set for each second terminal as the language to be used for distribution, and a translation management unit that, when receiving the first information from the first terminal, instructs the translation of the first information into the translation target language of the translation group to which the first terminal belongs, and distributes the translation results of each translation target language to the second terminal for which the translation target language is set as the distribution language.

[0008] According to the above configuration, translation processing for distribution to the second terminal is performed collectively for each distribution language, and common data is distributed to terminals of the same distribution language. Therefore, compared to when translation processing is performed for each terminal and different data is generated, duplication of processing is reduced.

[0009] In the above configuration, a period during which the second terminal can be added to the translation group may be limited to a predetermined period. With the above configuration, participants in a temporary activity can form a translation group using the devices they carry and use the translation results distributed by that translation group. This realizes a translation management system suitable for use in temporary activities such as sightseeing tours.

[0010] In the above configuration, the group management unit may, upon receiving a request from the first terminal to create the translation group, set up a new translation group to which the first terminal belongs and send identification information of the translation group to the first terminal, and, upon receiving a request from the second terminal that has acquired the identification information to be added to the translation group together with the identification information, add the second terminal to the translation group corresponding to the identification information.

[0011] According to the above configuration, the second terminal to be added to the translation group, i.e., the terminal to which the translation result of the first information input to the first terminal is delivered, is determined through the exchange of identification information between the first terminal and the second terminal. Therefore, it is possible to accurately select the second terminal to be added to the translation group, and by using the identification information, it is possible to accurately identify the translation group to which the terminal will be added.

[0012] In the above configuration, the group management unit may set a plurality of translation groups and set the target language for translation for each translation group, and the plurality of translation groups may include a plurality of translation groups to which the same first terminal belongs, and when the translation management unit receives the first information from the first terminal, it may instruct the translation of the first information into the target language for translation of the translation group selected by the first terminal, and distribute the translation result to the second terminal belonging to the translation group.

[0013] According to the above configuration, it is possible to set multiple translation groups to which one first terminal belongs, each with different terminals to which translation results are distributed, and the user of the first terminal can use multiple translation groups depending on the purpose of use.

[0014] In the above configuration, when the group management unit receives a request from the second terminal to add the second terminal to the translation group, the group management unit may add the second terminal to the translation group on the condition that the second terminal is located within a predetermined range from the location of the first terminal.

[0015] According to the above configuration, the second terminals added to the translation group can be limited to terminals located within a predetermined range from the first terminal. Therefore, when it is desired to form a translation group from a first terminal and a second terminal located in a common space, the second terminal to be added to the translation group can be accurately selected.

[0016] In the above configuration, a memory unit is provided that stores auxiliary information data that associates text with an image related to the text, and the translation management unit may deliver the image that is associated with the text corresponding to the first information in the auxiliary information data to the second terminal together with the translation result.

[0017] According to the above configuration, an image related to the translation content is displayed on the second terminal together with the translation result, which allows the user of the second terminal to deepen their understanding of the translation content. In the above configuration, the first information may be input to the first terminal as speech, and the translation result delivered to the second terminal may include speech.

[0018] The processing load required for speech translation is greater than the processing load required for text translation. With the above configuration, a translation management system is used for speech translation, which reduces duplication of processing and significantly reduces the processing load.

[0019] In the above configuration, the information processing device may further include a translation unit that translates the first information into the translation target language designated by the translation management unit. According to the above configuration, the translation management system is equipped with a translation unit, which makes it possible to improve the efficiency of the processing required for data exchange and optimize data such as vocabulary books used for translation, compared to when an external translation engine is used. [Effects of the Invention]

[0020] According to the present invention, it is possible to improve the efficiency of processing related to translation into multiple languages. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a diagram showing the configuration of a multilingual translation system including a translation management system according to a first embodiment of the translation management system. [Figure 2] FIG. 2 is a diagram showing the concept of a translation group set in the multilingual translation system of the first embodiment. [Figure 3] FIG. 3 is a sequence diagram showing the procedure for generating a new translation group in the multilingual translation system of the first embodiment. [Figure 4] FIG. 2 is a diagram showing an example of a screen of an input terminal in the multilingual translation system of the first embodiment. [Figure 5] FIG. 2 is a diagram showing an example of a screen of an input terminal in the multilingual translation system of the first embodiment. [Figure 6] FIG. 4 is a sequence diagram showing the procedure for adding an output terminal to a translation group in the multilingual translation system of the first embodiment. [Figure 7] FIG. 2 is a sequence diagram showing the procedure of speech translation in the multilingual translation system of the first embodiment. [Figure 8] FIG. 2 is a diagram showing an example of a screen of an input terminal in the multilingual translation system of the first embodiment. [Figure 9] FIG. 2A is a diagram showing an example of a screen of an input terminal in the multilingual translation system of the first embodiment, and FIG. 2B is a diagram showing an example of a screen of an output terminal in the multilingual translation system of the first embodiment. [Figure 10] FIG. 2 is a diagram showing a schematic diagram of the flow of data in the multilingual translation system of the first embodiment. [Figure 11] FIG. 10 is a sequence diagram showing the procedure for adding an output terminal to a translation group in a multilingual translation system according to a second embodiment of the translation management system. [Figure 12]FIG. 10 is a diagram showing the configuration of a multilingual translation system including a translation management system according to a third embodiment of the translation management system. [Figure 13] FIG. 11 is a diagram showing an example of association between text and images in auxiliary information data included in the translation management system of the third embodiment. [Figure 14] FIG. 11 is a sequence diagram showing the procedure of speech translation in the multilingual translation system of the third embodiment. [Figure 15] 10A is a diagram showing an example of a screen of an input terminal in the multilingual translation system of the third embodiment, and FIG. 10B is a diagram showing an example of a screen of an output terminal in the multilingual translation system of the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0022] (First embodiment) A first embodiment of a translation management system will be described with reference to FIGS. [Overall configuration of translation management system] 1, the multilingual translation system 100 includes an input terminal 10, an output terminal 20, a language management server 30, and a speech translation system 40. Of these, the language management server 30 and the speech translation system 40 constitute a translation management system 50.

[0023] The input terminal 10 and the output terminal 20 transmit and receive data to and from the language management server 30 via a network. The language management server 30 also transmits and receives data to and from the speech translation system 40 via a network. The network used for communication between the terminals 10, 20 and the language management server 30, and the network used for communication between the language management server 30 and the speech translation system 40, may each be a general-purpose communication line such as the Internet, or a dedicated communication line for communication between the respective devices.

[0024] The input terminal 10 is an example of a first terminal and has a communication function and a voice input function. The input terminal 10 receives voice input of the speaker's speech. The output terminal 20 is an example of a second terminal and has a communication function and a voice output function. The output terminal 20 outputs a translated voice of the speech. The multilingual translation system 100 includes a plurality of output terminals 20, and the plurality of output terminals 20 includes terminals used by users of different languages. For example, when the multilingual translation system 100 is used in a sightseeing tour, the input terminal 10 is carried by the tour guide and used, and the output terminal 20 is carried by the tour participants and used.

[0025] Furthermore, the input terminal 10 and the output terminal 20 have a display function, a photographing function, etc. The input terminal 10 and the output terminal 20 are, for example, general-purpose terminals such as smartphones and tablet terminals.

[0026] The language management server 30 manages combinations of the input terminal 10 and a plurality of output terminals 20 to which translation results of utterances input to the input terminal 10 are delivered, as well as languages ​​to be translated.

[0027] The language management server 30 includes a communication unit 31 , a control unit 32 , and a storage unit 33 . The communication unit 31 executes connection processing between the language management server 30, the terminals 10 and 20, and the speech translation system 40 via the network, and transmits and receives data between the connected devices.

[0028] The control unit 32 includes a CPU and volatile memory such as RAM. Based on the programs and data stored in the storage unit 33, the control unit 32 controls the various units of the language management server 30, such as controlling processing by the communication unit 31, reading and writing information from and to the storage unit 33, and performing various types of calculation processing.

[0029] In the multilingual translation system 100, the control unit 32 functions as a group management unit 32a and a translation management unit 32b. The group management unit 32a manages translation groups Gp. Each translation group Gp includes one input terminal 10 and multiple output terminals 20 to which the translation results of the utterances input to the input terminal 10 are delivered. In the translation group Gp, a delivery language is set for each output terminal 20, which specifies the language used for delivery, i.e., the language used in the delivered translation results. For each translation group Gp, a translation target language, into which the utterances are translated, is set according to the delivery language of the output terminals 20 belonging to the translation group Gp. A validity period is also set for each translation group Gp. The validity period is the period during which addition of output terminals 20 to the translation group Gp is ​​permitted.

[0030] The group management unit 32a sets up multiple translation groups Gp. The setting process for each translation group Gp includes registering the input terminals 10 and output terminals 20 that belong to the translation group Gp, setting the delivery language for the output terminal 20, setting the translation target language for the translation group Gp, and setting the validity period for the translation group Gp.

[0031] The function of transmitting and receiving data in the group management unit 32a is performed by the communication unit 31. When the translation management unit 32b receives speech data of the utterance content from the input terminal 10, it instructs the speech translation system 40 to translate the speech into the translation target language of the translation group Gp to which the input terminal 10 belongs. Then, the translation management unit 32b transmits the translation result into the language corresponding to the distribution language of the output terminal 20 to each output terminal 20 belonging to the translation group Gp. The data transmission and reception function of the translation management unit 32b is performed by the communication unit 31.

[0032] The storage unit 33 includes a non-volatile memory and stores programs and data necessary for the processing executed by the control unit 32. As part of this data, the storage unit 33 stores group data 33a and translation history data 33b. The group data 33a is data related to the translation group Gp. The group data 33a includes, for each translation group Gp, identification information of the input terminals 10 and output terminals 20 belonging to the translation group Gp, the distribution language of the output terminal 20, the language to be translated, and the validity period. The translation history data 33b includes the translation history for each translation group Gp. The translation history data 33b includes, for example, text data that is the recognition result of the speech input to the input terminal 10, and text data that is the translation result of the text data into the language to be translated.

[0033] The functions of the group management unit 32a and translation management unit 32b in the control unit 32 may be realized by various hardware components such as multiple CPUs and memories such as RAM, and software that makes these components function, or may be realized by software that gives multiple functions to a single common piece of hardware.

[0034] The speech translation system 40 functions as a translation unit and performs the following processes: recognizing speech and converting it into text, translating the text into a specified language, and synthesizing speech from the translated text. The speech translation system 40 is composed of one or more servers and has the programs and data required for each of the above processes. For example, the speech translation system 40 is composed of a speech recognition server 41 that performs speech recognition processing, a translation server 42 that performs translation processing, and a speech synthesis server 43 that performs speech synthesis processing.

[0035] [Translation group configuration] The translation group Gp will be described in detail with reference to Fig. 2. Fig. 2 conceptually shows an example of a translation group Gp established in the language management server 30.

[0036] A plurality of translation groups Gp are set in the language management server 30. Fig. 2 shows an example in which three translation groups Gp are set. Each translation group Gp includes one input terminal 10 and multiple output terminals 20. For example, translation group GpA includes an input terminal 10a and six output terminals 20a to 20f. The same input terminal 10 may belong to different translation groups Gp, and the same output terminal 20 may belong to different translation groups Gp. In other words, the same input terminal 10 may belong to multiple translation groups Gp, and the same output terminal 20 may belong to multiple translation groups Gp.

[0037] As described above, a translation group Gp is ​​a group consisting of an input terminal 10 and an output terminal 20 to which the translation results of the utterances input to the input terminal 10 are distributed. For example, if the multilingual translation system 100 is used for a sightseeing tour, one translation group Gp is ​​made up of the input terminal 10 carried by the tour guide of one sightseeing tour and the output terminals 20 carried by the participants of that sightseeing tour. For example, if one tour guide guides multiple sightseeing tours, a translation group Gp is ​​made up for each sightseeing tour, consisting of the tour guide's input terminal 10 and the participants' output terminals 20.

[0038] In each translation group Gp, a distribution language is set for each output terminal 20. Then, languages ​​to be translated in that translation group Gp are set corresponding to the types of distribution languages ​​of the output terminals 20 that belong to that translation group Gp. In other words, the distribution languages ​​of the multiple output terminals 20 that belong to the translation group Gp are categorized, and languages ​​to be translated in that translation group Gp are set corresponding to the languages ​​of this category.

[0039] For example, in translation group GpA, English is set as the delivery language for output terminals 20a, 20b, and 20c, Chinese is set as the delivery language for output terminal 20d, and Korean is set as the delivery language for output terminals 20e and 20f. Thus, while six output terminals 20a to 20f belong to translation group GpA, the delivery languages ​​of these output terminals 20 are English, Chinese, and Korean. That is, the delivery languages ​​of output terminals 20a to 20f belonging to translation group GpA are categorized into three. Languages ​​corresponding to these categories are then set as translation target languages. That is, three languages, English, Chinese, and Korean, are set as translation target languages ​​in translation group GpA.

[0040] The translation target languages ​​are set for each translation group Gp to correspond to the type of distribution language of the output terminal 20. In other words, the translation target languages ​​of a translation group Gp are limited to the distribution languages ​​of the output terminals 20 belonging to the translation group Gp, among the languages ​​that can be translated by the speech translation system 40. Therefore, the number and breakdown of translation target languages ​​in different translation groups Gp may or may not match. For example, the translation target languages ​​of translation group GpA are three: English, Chinese, and Korean, while the translation target languages ​​of translation group GpB are two: English and French.

[0041] The language of the speech content input to the input terminal 10, i.e., the language before translation, may be uniformly set to a predetermined language, or may be set for each translation group Gp. In the example shown in Fig. 2, it is assumed that the input language to the input terminal 10 is uniformly set to Japanese.

[0042] A validity period is set for each translation group Gp. The validity period may be a period from a specific date to a specific date, or a period from a specific time on a specific date to a specific time. As described above, the validity period is the period during which addition of an output terminal 20 to the translation group Gp is ​​permitted; in other words, it is the period during which speech translation is permitted in the translation group Gp. For example, if the multilingual translation system 100 is used for a sightseeing tour, the period during which the sightseeing tour is scheduled to take place is set as the validity period of the translation group Gp corresponding to the sightseeing tour.

[0043] [Translation Management System Operation] 3 to 9, the operation of the multilingual translation system 100 including the translation management system 50 will be described. Note that the processing of the input terminal 10 and the output terminal 20 in the operation of the multilingual translation system 100 may be performed based on a web application, or may be performed based on application software installed on these terminals.

[0044] For example, since the processing of the input terminal 10 tends to be more numerous and complex than the processing of the output terminal 20, executing the processing of the input terminal 10 based on application software installed on the input terminal 10 is highly effective in smoothly progressing the processing of the input terminal 10. On the other hand, for the output terminal 20, which has more users than the input terminal 10, using a web application reduces the burden of setting up the output terminal 20, etc., required to use the multilingual translation system 100, thereby greatly improving convenience. If a web application is used, for example, participants in a sightseeing tour can easily use their own mobile devices as output terminals 20 while traveling, making the multilingual translation system 100 easier to use.

[0045] <Translation group settings> First, the procedure for setting up a translation group Gp will be described. 3, when a predetermined operation is performed on the input terminal 10, a request to generate a new translation group Gp is ​​sent from the input terminal 10 to the language management server 30 (step S10). For example, as shown in FIG. 4, when an area Ra instructing the generation of a new translation group Gp is ​​selected on a screen Pa displayed upon the start of application software on the input terminal 10, the request to generate the new translation group Gp is ​​sent to the language management server 30.

[0046] 3, upon receiving a request to create a new translation group Gp, the group management unit 32a of the language management server 30 performs processing to create a new translation group Gp (step S11). Specifically, the group management unit 32a adds data of the new translation group Gp to the group data 33a. In the data of the new translation group Gp, the input terminal 10 that sent the request to create the new translation group Gp in step S10 is set as the input terminal 10 belonging to the translation group Gp. The data of the new translation group Gp also sets the validity period of the translation group Gp. An arbitrary period may be specified in the input terminal 10, and the validity period may be set to the specified period based on notification of the specified period to the language management server 30, or may be set to a period predetermined by the language management server 30.

[0047] Next, the group management unit 32a transmits information including the identification information Id of the generated translation group Gp to the input terminal 10 (step S12). Upon receiving the information, the input terminal 10 displays the information including the identification information Id on the display unit of the input terminal 10 (step S13).

[0048] The identification information Id of the translation group Gp is ​​displayed on the input terminal 10 in a form that can be mechanically read by the output terminal 20 or in a form that can be manually communicated to the user of the output terminal 20. For example, the identification information Id of the translation group Gp is ​​preferably displayed on the input terminal 10 as a two-dimensional code and read by the output terminal 20. The two-dimensional code contains, for example, the identification information Id and holds, as information, a URL that indicates the access destination of the language management server 30.

[0049] 5 shows an example of the input terminal 10 on which a two-dimensional code Cd including the identification information Id of a translation group Gp is ​​displayed. After a new translation group Gp is ​​generated, it is preferable that the identification information Id of the translation group Gp can be displayed on the input terminal 10 at any time. For example, as shown in FIG. 4 above, an area Rb for selecting a generated translation group Gp is ​​included in the screen Pa and displayed on the input terminal 10. When the area Rb is selected, information including the identification information Id of the translation group Gp corresponding to the area Rb is transmitted from the language management server 30 to the input terminal 10 in response to a request from the input terminal 10, and the information is displayed on the input terminal 10.

[0050] The two-dimensional code including the identification information Id of the translation group Gp is ​​not limited to being displayed on the input terminal 10, but may be provided to the user of the output terminal 20 in a state where it is printed on paper based on the printout on the screen of the input terminal 10. Similarly, even if the identification information Id of the translation group Gp is ​​displayed on the input terminal 10 in a form other than a two-dimensional code, the identification information Id of the translation group Gp may be provided to the user of the output terminal 20 through paper or verbal communication, and the user may input the identification information Id into the output terminal 20. The key is that the output terminal 20 should be able to acquire the identification information Id of the translation group Gp.

[0051] As shown in FIG. 5, the screen Pb displayed on the input terminal 10 also includes an area Rc for instructing the start of speech translation for the selected translation group Gp. 6, the output terminal 20 acquires the identification information Id of the translation group Gp and then transmits a request to add the translation group Gp to the language management server 30 (step S20). For example, if the identification information Id of the translation group Gp is ​​included in a two-dimensional code, the output terminal 20 reads the two-dimensional code and accesses the URL acquired thereby. As a result, the identification information Id of the translation group Gp is ​​notified from the output terminal 20 to the language management server 30, which functions as a request to add the translation group Gp.

[0052] When a request to add to the translation group Gp is ​​received, the group management unit 32a of the language management server 30 requests the output terminal 20 to transmit language information (step S21). The language information is information that specifies the distribution language.

[0053] Upon receiving the request from the language management server 30, the output terminal 20 transmits language information to the language management server 30 (step S22). For example, based on a language specified by a user via the screen of the output terminal 20, language information indicating the specified language is transmitted from the output terminal 20 to the language management server 30. Alternatively, the language used in the output terminal 20, i.e., information indicating the language used to display information on the output terminal 20, may be transmitted from the output terminal 20 to the language management server 30 as language information.

[0054] Upon receiving the language information, the group management unit 32a of the language management server 30 performs processing to add the output terminal 20 to the translation group Gp corresponding to the identification information Id acquired in response to the addition request (step S23).

[0055] Specifically, the group management unit 32a assigns the output terminal 20 to the translation group Gp corresponding to the identification information Id. That is, in the data for the translation group Gp corresponding to the identification information Id in the group data 33a, the output terminal 20 that sent the request to be added to the translation group Gp in step S20 is set as the output terminal 20 belonging to the translation group Gp. Furthermore, the group management unit 32a sets the language indicated by the language information received from the output terminal 20 as the distribution language of the added output terminal 20 in the group data 33a.

[0056] In setting the translation group Gp, for example, access from the output terminal 20 to the language management server 30 based on the acquired identification information Id is permitted only during the validity period of the translation group Gp corresponding to the identification information Id, thereby limiting the period during which the output terminal 20 can be added to the translation group Gp to the validity period. For example, a time-limited URL that can be accessed only during the validity period may be used as the URL held by the two-dimensional code containing the identification information Id.

[0057] Alternatively, when a request to add an output terminal 20 to a translation group Gp is ​​received from the output terminal 20, the group management unit 32a of the language management server 30 may check whether the time is within the validity period of the target translation group Gp, and add the output terminal 20 to the translation group Gp only if the time is within the validity period. Furthermore, when transmitting the identification information Id of the generated translation group Gp to the input terminal 10, the group management unit 32a may check whether the time is within the validity period of the translation group Gp, and send the identification information Id to the input terminal 10 only if the time is within the validity period. With these configurations, the period during which an output terminal 20 can be added to a translation group Gp is ​​also limited to the validity period of the translation group Gp.

[0058] The translation target languages ​​of a translation group Gp are set to correspond to the types of delivery languages ​​of the output terminals 20 that belong to the translation group Gp at that time. That is, if the addition of an output terminal 20 to the translation group Gp increases the number of delivery languages ​​of the output terminals 20 that belong to the translation group Gp, in other words, if the delivery language of the added output terminal 20 is different from any of the delivery languages ​​of the output terminals 20 that already belong to the translation group Gp, the delivery language of the added output terminal 20 is added to the translation target languages ​​of the translation group Gp. That is, the translation target languages ​​are updated in the data of the translation group Gp in the group data 33a, and the delivery language of the added output terminal 20 is added to the translation target languages.

[0059] In this way, the translation target languages ​​of the translation group Gp change in accordance with changes in the output terminals 20 belonging to the translation group Gp and changes in the types of distribution languages. It is also preferable that the delivery language of output terminal 20 can be changed arbitrarily. Changing the delivery language is performed, for example, in the following manner. That is, the screen displayed on output terminal 20 includes an area for instructing a change of the delivery language, and when that area is selected and a new language is specified, a request to change the delivery language together with the new language information is sent from output terminal 20 to language management server 30. Upon receiving the request to change the delivery language, group management unit 32a of language management server 30 changes group data 33a so as to change the delivery language of that output terminal 20 to the language indicated by the new language information.

[0060] If the type of delivery language in the translation group Gp to which the output terminal 20 belongs changes in association with a change in the delivery language of the output terminal 20, the translation target language of the translation group Gp will change in accordance with this change.

[0061] In the above configuration, for example, if the multilingual translation system 100 is used for a sightseeing tour, at the start of the sightseeing tour, the tour guide displays a two-dimensional code containing the identification information Id of the translation group Gp on the input terminal 10, and the tour participants read the two-dimensional code using the output terminal 20. This adds the output terminal 20 of the tour participant to the translation group Gp, making it possible to proceed with the sightseeing tour using the speech translation of that translation group Gp. If the identification information Id of the generated translation group Gp can be displayed on the input terminal 10 at any time, it is possible to add the output terminal 20 of a new tour participant to the translation group Gp during the sightseeing tour, thereby increasing the degree of freedom in configuring the translation group Gp.

[0062] <Speech translation> Next, the procedure for speech translation will be described. As shown in Fig. 7, first, with a translation group Gp selected, speech indicating the content of the utterance is input to the input terminal 10 (step S30). Specifying a translation group Gp that has already been created, or creating a new translation group, functions as selecting a translation group Gp. For example, as shown in Fig. 8, when a translation group Gp is ​​selected and an instruction to start speech translation is given, a screen Pc for speech input is displayed on the input terminal 10, and the user speaks into the input terminal 10, whereupon the speech is input to the input terminal 10.

[0063] First voice data S1 representing the voice input to the input terminal 10 is transmitted from the input terminal 10 to the language management server 30 (step S31). The first voice data S1 is an example of first information. Furthermore, when or before transmitting the first voice data S1, information representing the selected translation group Gp is ​​also transmitted from the input terminal 10 to the language management server 30.

[0064] Upon receiving the first speech data S1, the translation management unit 32b of the language management server 30 sends a request for speech translation to the speech translation system 40 (step S32). In detail, the translation management unit 32b refers to the group data 33a and identifies the translation target language of the translation group Gp selected on the input terminal 10. Then, the translation management unit 32b instructs the speech translation system 40 to translate into the identified translation target language. The first speech data S1 is sent from the language management server 30 to the speech translation system 40.

[0065] Upon receiving a request from the language management server 30, the speech translation system 40 performs a speech translation process (step S33). That is, the speech translation system 40 first performs a speech recognition process on the first speech data S1 to generate first text data T1, which is text data in a language corresponding to the first speech data S1. Next, the speech translation system 40 translates the first text data T1 into a target language to generate second text data T2, which is text data in the target language. When there are multiple target languages, the speech translation system 40 translates the first text data T1 into each of the target languages ​​to generate second text data T2 in each language. Then, the speech translation system 40 synthesizes speech based on the second text data T2 to generate second speech data S2, which is speech data in a language corresponding to the second text data T2. In this way, second speech data S2 is generated in a number and variety of languages ​​corresponding to the number and variety of target languages.

[0066] When the speech translation process is completed, the speech translation system 40 transmits the processing result to the language management server 30 (step S34). The processing result includes the second speech data S2, the first text data T1, and the second text data T2.

[0067] Upon receiving the processing result, the translation management unit 32b of the language management server 30 transmits the second voice data S2 in a language corresponding to the set delivery language to the output terminals 20 belonging to the translation group Gp selected on the input terminal 10 (step S35). At this time, the translation management unit 32b preferably also transmits the second text data T2 corresponding to the second voice data S2 to the output terminal 20. In addition, the translation management unit 32b preferably transmits the first text data T1 to the input terminal 10.

[0068] Furthermore, the translation management unit 32b stores at least the first text data T1 and the second text data T2 of the received processing results in the storage unit 33, including them in the data of the corresponding translation group Gp in the translation history data 33b (step S36).

[0069] Upon receiving the second voice data S2, the output terminal 20 outputs voice based on the second voice data S2 (step S37). As a result, the voice obtained by translating the utterance content input to the input terminal 10 is output from the output terminal 20.

[0070] When the output terminal 20 receives the second text data T2 together with the second voice data S2, the output terminal 20 displays a character string based on the second text data T2 on the display unit of the output terminal 20. When the input terminal 10 receives the first text data T1, the input terminal 10 displays a character string based on the first text data T1 on the display unit of the input terminal 10.

[0071] FIG. 9(a) shows an example of a screen Pd displayed on the input terminal 10 as a result of the above-described processing, and FIG. 9(b) shows an example of a screen Pe displayed on the output terminal 20 as a result of the above-described processing. The screen Pd of the input terminal 10 includes an area Rd showing a character string based on the first text data T1, i.e., a speech recognition result of the utterance input to the input terminal 10. By viewing the screen Pd, the user of the input terminal 10 can confirm whether the utterance has been correctly recognized. The screen Pe of the output terminal 20 includes an area Re showing a character string based on the second text data T2, i.e., a sentence corresponding to the translation result of the utterance input to the input terminal 10. By viewing the screen Pe, the user of the output terminal 20 can confirm the translation result in text in addition to audio, making it easier to understand the translation result.

[0072] Alternatively, the second text data T2 in each translation target language may be transmitted from the language management server 30 to the input terminal 10, and a character string based on the second text data T2 may be displayed on the input terminal 10. With this configuration, the user of the input terminal 10 can check the translation result of the utterance content, and therefore can understand whether the translation result is appropriate.

[0073] Furthermore, the input terminal 10 is capable of checking the translation history data 33b. Checking the translation history data 33b is performed, for example, by the following procedure. That is, with a translation group Gp selected, the input terminal 10 sends a request to check the translation history data 33b to the language management server 30. The request to check the translation history data 33b is made, for example, by selecting an area on the screen of the input terminal 10 that instructs checking the translation history. Upon receiving the request to check the translation history data 33b, the translation management unit 32b of the language management server 30 sends to the input terminal 10 the data included in the translation history data 33b for the translation group Gp selected on the input terminal 10. The input terminal 10 displays the received data. This allows the user of the input terminal 10 to check the history of speech translation for the selected translation group Gp, i.e., what translation results were sent to the output terminal 20 for what spoken content, for each translation target language.

[0074] The operation of the multilingual translation system 100 will be described with reference to Figure 10. Figure 10 conceptually illustrates speech translation in the translation group GpA shown in Figure 2. In the translation group GpA, the target languages ​​for translation are three languages: English, Chinese, and Korean, and Japanese speech is input to the input terminal 10a.

[0075] When Japanese speech data is transmitted from the input terminal 10a to the translation management system 50, the translation management system 50 performs speech translation of the speech data into the target language. That is, Japanese is translated into English to generate English speech data, Japanese is translated into Chinese to generate Chinese speech data, and Japanese is translated into Korean to generate Korean speech data.

[0076] Then, the English audio data is distributed to three output terminals 20a to 20c whose distribution language is set to English, the Chinese audio data is distributed to one output terminal 20d whose distribution language is set to Chinese, and the Korean audio data is distributed to two output terminals 20e and 20f whose distribution language is set to Korean.

[0077] In this way, even when there are multiple output terminals 20 that use the same distribution language, only one piece of voice data is generated based on translation into a language corresponding to the distribution language, and the same voice data is distributed to the multiple output terminals 20. In other words, translation and voice synthesis processing for distribution to the multiple output terminals 20 is performed collectively for each distribution language. Therefore, overlapping processing is reduced compared to when translation and voice synthesis processing are performed for each output terminal 20 and different voice data are generated.

[0078] Furthermore, for text data generated by speech recognition of pre-translation speech data, one data can be commonly used for translation into each language. In other words, by performing speech recognition processing collectively for distribution to multiple output terminals 20, duplication of processing can be reduced compared to when speech recognition processing is performed separately for each output terminal 20.

[0079] Furthermore, because the speech recognition, translation, and speech synthesis processes are all performed together as described above, the processing load on the translation management system 50 is reduced compared to when these processes are repeatedly performed on a server for each target terminal. This enables rapid speech translation and delivery of translation results, shortening the time required from inputting the spoken content to delivering the translation results.

[0080] Furthermore, because the input terminal 10 and output terminal 20 carried by the user do not need to have translation functions, the terminal configuration is simplified, and it becomes easy to use general-purpose terminals as the input terminal 10 and output terminal 20. Therefore, compared to using dedicated terminals, the labor required for preparing and distributing terminals is reduced, and compared to speech translation performed through direct communication between dedicated terminals, the burden required for establishing a communication environment is also reduced. This increases the convenience of using the multilingual translation system 100.

[0081] As described above, according to the first embodiment, the following effects can be obtained. (1) For each distribution language, translation and speech synthesis processing for distribution to the output terminal 20 is performed collectively, and common data is distributed to output terminals 20 of the same distribution language. Therefore, overlapping processing is reduced compared to when translation and speech synthesis processing is performed for each output terminal 20 and different data is generated. Furthermore, speech recognition processing for distribution to multiple output terminals 20 is performed collectively, and common data is used for translation into each translation target language. Therefore, overlapping processing is reduced compared to when speech recognition processing is performed for each output terminal 20 and different data is generated. This makes it possible to improve the efficiency of processing related to translation into multiple languages.

[0082] (2) By setting a validity period, adding an output terminal 20 to a translation group Gp is ​​limited to a specific period. Therefore, participants in a temporary activity can form a translation group Gp using the terminals they carry and use the distribution of translation results from that translation group Gp. This realizes a system suitable for use in temporary activities, such as sightseeing tours.

[0083] (3) In response to a request to create a translation group Gp from the input terminal 10, the language management server 30 creates a new translation group Gp and transmits the identification information Id of the translation group Gp to the input terminal 10. Then, based on receiving a request to add the output terminal 20 to the translation group Gp from the output terminal 20 that has acquired the identification information Id, the language management server 30 adds the output terminal 20 to the translation group Gp corresponding to the identification information Id. In this way, the output terminal 20 to be added to the translation group Gp, i.e., the terminal to which the translation result will be distributed, is determined through the exchange of the identification information Id between the input terminal 10 and the output terminal 20. Therefore, it is possible to accurately select the output terminal 20 to be added to the translation group Gp, and by using the identification information Id, it is possible to accurately identify the translation group Gp to which the output terminal 20 will be added.

[0084] (4) The language management server 30 sets multiple translation groups Gp, and these translation groups Gp include multiple translation groups Gp to which the same input terminal 10 belongs. When the language management server 30 receives the first voice data S1, it instructs the translation group Gp selected by the input terminal 10 to translate the data into the target language for translation, and distributes the translation results to the output terminals 20 that belong to that translation group Gp. With this configuration, it is possible to set multiple translation groups Gp to which one input terminal 10 belongs, each with different target terminals for distributing the translation results, and the user of the input terminal 10 can use multiple translation groups Gp depending on the purpose of use.

[0085] (5) Because the translation management system 50 is used for speech translation, duplication of processing is reduced in speech translation, which has a greater processing load than text translation. Therefore, the effect of reducing the processing load is greatly achieved.

[0086] (6) The translation management system 50 includes a speech translation system 40 that functions as a translation unit. This makes it possible to improve the efficiency of the processing required for data exchange and to optimize data such as vocabulary books used for translation, compared to when an external translation engine is used.

[0087] (Second embodiment) A second embodiment of the translation management system will be described with reference to Figure 11. The translation management system of the second embodiment differs from the first embodiment in that the location of the terminal is confirmed when an output terminal is added to a translation group. The following description will focus on the differences between the second embodiment and the first embodiment, and components similar to those in the first embodiment will be assigned the same reference numerals and will not be described again.

[0088] For example, when the multilingual translation system 100 is used for a sightseeing tour, the user of the input terminal 10 and the user of the output terminal 20 are located in the same space. In the second embodiment, the output terminal 20 is added to the translation group Gp on the condition that the output terminal 20 is located within a predetermined range from the input terminal 10.

[0089] In the second embodiment, each of the input terminal 10 and the output terminal 20 has a function of acquiring location information indicating the location of the terminal. The location information may be information that can identify the absolute location of each terminal or the relative location of each terminal with respect to a predetermined reference, and may be information that can be used to determine whether the output terminal 20 is located within a predetermined range centered on the input terminal 10.

[0090] For example, the location information may be information consisting of latitude and longitude obtained using a GPS (Global Positioning System), or information that enables relative location to be identified using short-range wireless communication such as Wi-Fi (registered trademark) or Bluetooth (registered trademark), sound waves, ultrasound, light, etc. If necessary, transmitters and receivers of radio waves or the like for acquiring location information may be installed in locations where the input terminal 10 and the output terminal 20 may be located, for example, in tourist facilities where sightseeing tours are held.

[0091] As shown in FIG. 11, similarly to the first embodiment, the output terminal 20 acquires the identification information Id of the translation group Gp and then transmits a request to add the translation group Gp to the language management server 30 (step S40).

[0092] When a request to add to the translation group Gp is ​​received, the group management unit 32a of the language management server 30 requests the output terminal 20 to transmit language information (step S41). The group management unit 32a also requests the output terminal 20 and the input terminal 10 belonging to the translation group Gp corresponding to the identification information Id to transmit location information (step S42).

[0093] Upon receiving a request from the language management server 30, the output terminal 20 transmits the language information and the location information of the output terminal 20 to the language management server 30 (step S43). Also, upon receiving a request from the language management server 30, the input terminal 10 transmits the location information of the input terminal 10 to the language management server 30 (step S44).

[0094] When receiving the location information from each of the input terminal 10 and the output terminal 20, the group management unit 32a of the language management server 30 determines whether the output terminal 20 is located within a predetermined range from the input terminal 10 based on the received location information (step S45).

[0095] When it is confirmed that the output terminal 20 is located within a predetermined range from the input terminal 10, the group management unit 32a performs a process of adding the output terminal 20 to the translation group Gp corresponding to the above-mentioned identification information Id (step S46). When it is determined that the output terminal 20 is not located within a predetermined range from the input terminal 10, the process of adding the output terminal 20 to the translation group Gp is ​​not performed, and the output terminal 20 is notified that it cannot be added to the translation group Gp.

[0096] In the second embodiment, speech translation processing in the translation group Gp is ​​performed in the same manner as in the first embodiment. In the second embodiment, an output terminal 20 is added to the translation group Gp on the condition that the output terminal 20 is located within a predetermined range from the position of the input terminal 10, so that the output terminal 20 to be added to the translation group Gp can be accurately identified. This is particularly effective when it is desired to form a translation group Gp with input terminals 10 and output terminals 20 located in a common space.

[0097] The configuration of the second embodiment also contributes to improving security regarding the distribution of translation results. For example, if the content of an utterance input to the input terminal 10 contains information that should not be leaked to a third party, even if a third party illegally obtains identification information Id and attempts to join the translation group Gp from a remote location, the addition of the terminal owned by the third party to the translation group Gp is ​​denied, thereby preventing the translation results of the utterance from being leaked to a third party.

[0098] In addition to adding the output terminal 20 to the translation group Gp, the delivery of the translation result to the output terminal 20 may also be conditional on the output terminal 20 being located within a predetermined range from the location of the input terminal 10. That is, before sending the second voice data S2 to the output terminal 20, the translation management unit 32b of the language management server 30 acquires location information from each terminal 10, 20, and sends the second voice data S2 to the output terminal 20 if the output terminal 20 is located within a predetermined range from the location of the input terminal 10. This configuration further enhances the security of the delivery of the translation result.

[0099] Furthermore, in the above configuration, if there is an output terminal 20 that belongs to the translation group Gp that is not located within a predetermined range from the position of the input terminal 10 when the translation result is distributed, the input terminal 10 may be notified that there is an output terminal 20 that is outside the predetermined range. With this configuration, for example, if the multilingual translation system 100 is used for a sightseeing tour, and a participant strays from the tour group, the input terminal 10 will be notified that there is an output terminal 20 that is outside the predetermined range. Therefore, a tour guide using the input terminal 10 can know whether any of the participants are lost, etc.

[0100] In the second embodiment, the location information of the input terminal 10 and the output terminal 20 does not necessarily have to be acquired from the terminals themselves, and the language management server 30 may acquire the location information of the terminals based on the output of a device such as a radio wave receiver installed near the input terminal 10 and the output terminal 20. Furthermore, if the locations of the input terminal 10 and the output terminal 20 are limited to a specific range, the condition for adding the output terminal 20 to the translation group Gp may be that the output terminal 20 is located within the specific range, instead of the location relative to the input terminal 10.

[0101] As described above, according to the second embodiment, in addition to the effects (1) to (6) of the first embodiment, the following effects can be obtained. (7) When the language management server 30 receives a request from an output terminal 20 to add an output terminal 20 to the translation group Gp, the language management server 30 adds the output terminal 20 to the translation group Gp on the condition that the output terminal 20 is located within a predetermined range from the position of the input terminal 10. Therefore, when it is desired to form a translation group Gp from an input terminal 10 and an output terminal 20 located in a common space, the output terminal 20 to be added to the translation group Gp can be accurately selected. This also improves security regarding the distribution of translation results.

[0102] (Third embodiment) A third embodiment of the translation management system will be described with reference to Figures 12 to 15. The translation management system of the third embodiment differs from the first embodiment in that, in addition to the translation result, an image related to the utterance content is also delivered to the output terminal. The following description will focus on the differences between the third embodiment and the first embodiment, and the same components as in the first embodiment will be assigned the same reference numerals and their description will be omitted.

[0103] 12, the language management server 30 of the third embodiment stores auxiliary information data 33c in the storage unit 33. In the auxiliary information data 33c, text such as words is associated with images related to the text. For example, if the multilingual translation system 100 is used for a sightseeing tour, key words for the tour are associated with images related to the words.

[0104] FIG. 13 shows an example of correspondence between text and images in the auxiliary information data 33c. In FIG. 13, the text "Pine Tree Corridor" is associated with an image of a painting depicting a sword fight in the Pine Tree Corridor of Edo Castle, i.e., a scene from the Ako Incident. The relationship between text and images is not particularly limited. For example, a picture or photograph of an object, person, or place indicated by the text may be associated with the text, or an image of a scene from a historical event or occurrence indicated by the text may be associated with the text.

[0105] In the third embodiment, the translation group Gp is ​​set in the same manner as in the first embodiment. The procedure for speech translation in the third embodiment will be described with reference to Figure 14. As Figure 14 shows, the processes of steps S50 to S53 are the same as the processes of steps S30 to S33 in the first embodiment. That is, when speech indicating the content of an utterance is input to the input terminal 10, first speech data S1 indicating the speech is transmitted from the input terminal 10 to the language management server 30. The language management server 30 instructs the speech translation system 40 to translate the first speech data S1 into the translation target language of the translation group Gp selected on the input terminal 10. The speech translation system 40 performs speech translation processing and transmits the first text data T1 and the second text data T2 and second speech data S2 in each translation target language to the language management server 30.

[0106] Upon receiving the processing result including the first text data T1, the second text data T2, and the second audio data S2, the translation management unit 32b of the language management server 30 identifies an image associated with the text included in the first text data T1 in the auxiliary information data 33c (step S55).

[0107] If the image is identified, the translation management unit 32b transmits image data Di representing the identified image together with the second audio data S2 in a language corresponding to the distribution language to the output terminal 20 (step S56). If there is no image associated with the text included in the first text data T1 in the auxiliary information data 33c, the image data Di is not transmitted, and the same process as in the first embodiment is carried out.

[0108] Furthermore, similarly to the first embodiment, the translation management unit 32b stores the first text data T1 and the second text data T2 as translation history data 33b in the storage unit 33 (step S57). The translation history data 33b may also include image data Di.

[0109] Upon receiving the second voice data S2 and the image data Di, the output terminal 20 outputs voice based on the second voice data S2 and displays an image based on the image data Di on the display unit (step S58). As a result, voice obtained by translating the utterance content input to the input terminal 10 is output at the output terminal 20, and an image related to the utterance content is displayed at the output terminal 20.

[0110] It is preferable that the second text data T2 is transmitted to the output terminal 20 and the first text data T1 is transmitted to the input terminal 10 in the same manner as in the first embodiment. FIG. 15(a) shows an example of a screen Pf displayed on the input terminal 10 as a result of the above-described processing when the text and image are associated in the auxiliary information data 33c as shown in FIG. 13, and FIG. 15(b) shows an example of a screen Pg displayed on the output terminal 20 as a result of the above-described processing. The screen Pf of the input terminal 10 includes an area Rd displaying a character string based on the first text data T1. The screen Pg of the output terminal 20 includes an area Re displaying a character string based on the second text data T2 and an area Rf displaying an image based on the image data Di. In the example shown in FIG. 15, the speech input to the input terminal 10 includes the word "pine corridor," and as a result, an image associated with this word in the auxiliary information data 33c is displayed on the output terminal 20. This allows the user of the output terminal 20 to view images related to the translation result, thereby deepening their understanding of the translation result.

[0111] The auxiliary information data 33c may be stored in the speech translation system 40 instead of the language management server 30. In this case, after generating the first text data T1, the speech translation system 40 identifies an image associated in the auxiliary information data 33c with the text included in the first text data T1, and transmits image data Di, which is data for that image, together with the first text data T1, the second text data T2, and the second audio data S2, as a processing result to the language management server 30. The translation management unit 32b of the language management server 30 then transmits the image data Di to the output terminal 20 together with the second audio data S2 in a language corresponding to the distribution language.

[0112] If the speech translation system 40 stores auxiliary information data 33c, the auxiliary information data 33c may constitute flashcard data for translating the first text data T1 into the second text data T2. That is, the flashcard data may be formed by associating text in a language before translation, text in a language after translation corresponding to the original text, and images related to these texts, and images may be identified when translating the first text data T1 into the second text data T2 using the flashcard data. With this configuration, for example, by registering technical terms used by guides on sightseeing tours in the flashcard data, accurate translation of the technical terms becomes possible.

[0113] In short, the auxiliary information data 33c is stored in the translation management system 50, and the image associated with the text corresponding to the utterance content input to the input terminal 10 in the auxiliary information data 33c is delivered to the output terminal 20 together with the translation result of the utterance content.

[0114] Furthermore, the auxiliary information data 33c may associate text such as words with images related to the text and explanatory text in various languages ​​about the text and images. When identifying an image associated with text included in the first text data T1, the explanatory text may also be identified, and data of the explanatory text in the language corresponding to the second voice data S2 may also be transmitted to the output terminal 20 in addition to the second voice data S2 and image data Di. The explanatory text is displayed on the output terminal 20 together with an image based on the image data Di. This configuration allows the user of the output terminal 20 to view the explanatory text associated with the image in addition to the image related to the translation result, thereby further deepening their understanding of the translation result.

[0115] Note that the user of the output terminal 20 may be able to select whether or not to display an image or an explanatory text on the output terminal 20. For example, by selecting a predetermined area included in the screen of the output terminal 20, it may be possible to switch between displaying and not displaying an image.

[0116] Furthermore, it may be possible to set auxiliary information data 33c to be used for speech translation in each translation group Gp. For example, auxiliary information data 33c corresponding to the content of a sightseeing tour, i.e., auxiliary information data 33c in which words that may be included in the content explained by the tour guide on the sightseeing tour are associated with images related to those words, is generated for each sightseeing tour and stored in the language management server 30. Then, when a new translation group Gp is ​​generated in response to a request from the input terminal 10, the type of sightseeing tour is selected on the input terminal 10, and the auxiliary information data 33c corresponding to the selected sightseeing tour is selected from the multiple auxiliary information data 33c and set as the auxiliary information data 33c for the translation group Gp. With this configuration, images that are more suitable for the purpose of using speech translation in the translation group Gp are displayed on the output terminal 20.

[0117] The content associated with the text in the auxiliary information data 33c is not limited to still images, but may be video or 3DCG. Furthermore, the content associated with the text in the auxiliary information data 33c may be AR (Augmented Reality) content. That is, the still images, video, 3DCG, etc. delivered together with the translation result may be displayed on the display unit of the output terminal 20 superimposed on the real environment around the output terminal 20.

[0118] In particular, when the multilingual translation system 100 is used for sightseeing tours, the tour guide's speech often includes cultural and historical elements, as well as highly specialized topics. Therefore, displaying video, 3DCG, or AR content on the output terminal 20 along with the translation results helps participants understand the translation results. Furthermore, since the tour guide's speech often relates to the location of the participants and surrounding exhibits, displaying AR content superimposed on the real-world environment surrounding the participants along with the translation results makes it easier to understand the translation results and increases participants' interest in the translation results. As a result, communication between the tour guide and participants proceeds smoothly. Furthermore, since multiple participants can share such content through the output terminal 20, communication between participants is also promoted.

[0119] As described above, according to the third embodiment, in addition to the effects (1) to (6) of the first embodiment, the following effects can be obtained. (8) The language management server 30 distributes content such as images associated in the auxiliary information data 33c with the text corresponding to the first voice data S1, together with the translation result, to the output terminal 20. As a result, content related to the translation content is displayed on the output terminal 20 along with the translation result, thereby deepening the user of the output terminal 20's understanding of the translation content.

[0120] (Variation) The above-described embodiments can be modified as follows: The above-described embodiments and the following modifications may be combined with each other.

[0121] A translation group Gp may include multiple input terminals 10. Furthermore, one terminal may function as an input terminal 10, and also function as an output terminal 20 that receives the translation results of the utterances input to the other input terminals 10. In other words, in one translation group Gp, the translation results of the utterances input to each terminal may be distributed to terminals other than that terminal, and two-way communication between users of the terminals may be possible.

[0122] Processing using the translation group Gp may be stopped in response to a request from the input terminal 10. Processing of the translation group Gp is ​​stopped, for example, by the following procedure. That is, the screen displayed on the input terminal 10 includes an area for instructing the stopping of processing of the translation group Gp, and when that area is selected, a request to stop processing of the translation group Gp is ​​sent from the input terminal 10 to the language management server 30. Upon receiving this request, the group management unit 32a of the language management server 30 prohibits the addition of output terminals 20 to the translation group Gp selected on the input terminal 10 and the processing of speech translation in that translation group Gp. With this configuration, speech translation using the translation group Gp can be stopped at the will of the user of the input terminal 10, even during the validity period of the translation group Gp.

[0123] In response to a request from the output terminal 20, it may be possible to remove the output terminal 20 from the translation group Gp. Removal of the output terminal 20 is performed, for example, by the following procedure. That is, the screen displayed on the output terminal 20 includes an area for instructing removal from the translation group Gp, and by selecting this area, a request for removal from the translation group Gp is ​​sent from the output terminal 20 to the language management server 30. Upon receiving this request, the group management unit 32a of the language management server 30 changes the group data 33a so as to remove the output terminal 20 from the translation group Gp selected by the output terminal 20. If the removal of the output terminal 20 from the translation group Gp changes the type of distribution language of the output terminal 20 belonging to the translation group Gp, the translation target language is changed in accordance with this change.

[0124] The conditions for adding an output terminal 20 to a translation group Gp may include the establishment of authentication using a password or biometric information. Authentication is performed by comparing authentication information, such as a password or biometric information, entered into the output terminal 20 with information registered in advance in the language management server 30. This configuration further enhances security regarding the addition of an output terminal 20 to a translation group Gp.

[0125] In the above embodiments, the speech translation system 40 performs speech recognition, translation, and speech synthesis consecutively. Alternatively, the processing results may be returned to the language management server 30 after each of the speech recognition, translation, and speech synthesis processes, as described below. That is, when the language management server 30 transmits the first speech data S1 to the speech translation system 40, the speech recognition result, i.e., the first text data T1, is transmitted from the speech translation system 40 to the language management server 30. In response, the language management server 30 requests the speech translation system 40 to translate the first text data T1 into each target language. The speech translation system 40 translates the first text data T1 into the specified language to generate second text data T2 and transmits the second text data T2 to the language management server 30. Next, the language management server 30 requests the speech translation system 40 to perform speech synthesis based on the second text data T2. The speech translation system 40 performs speech synthesis on the second text data T2 and transmits the generated second speech data S2 to the language management server 30. As a result, the language management server 30 acquires the first text data T1, the second text data T2, and the second speech data S2, and transmits these data to the terminals 10 and 20.

[0126] The speech translation method is not limited to a method that sequentially performs speech recognition, translation, and speech synthesis. It is sufficient if translated second speech data S2 can be generated from first speech data S1. The translation content may be editable on the input terminal 10.

[0127] For example, when the first text data T1 is displayed on the input terminal 10 as a result of speech translation, the first text data T1 may be edited on the input terminal 10. In this case, the edited text data is sent to the language management server 30, and the translation management unit 32b instructs the speech translation system 40 to translate the edited text data into each target language, and the translation result is delivered to the output terminal 20. With this configuration, if the user of the input terminal 10 finds that their utterance is not correctly recognized in the speech recognition process, they can edit the recognition result by editing the first text data T1 and have the edit reflected in the translation result. This makes it possible to deliver appropriate translation results. Note that, as described above, when the speech translation system 40 returns the processing results of speech recognition, translation, and speech synthesis to the language management server 30, the first text data T1 may be sent to the input terminal 10 and edited before translation into the target language, and the translation and speech synthesis processes may be carried out using the edited text data.

[0128] Furthermore, for example, when second text data T2 in each translation target language is displayed on the input terminal 10 as a result of speech translation, the second text data T2 may be edited on the input terminal 10. In this case, the edited text data is sent to the language management server 30, and the translation management unit 32b instructs the speech translation system 40 to perform speech synthesis of the edited text data, and the processing result is delivered to the output terminal 20. With this configuration, the user of the input terminal 10 can directly edit the translation result by editing the second text data T2. This makes it possible to deliver appropriate translation results. Note that, as described above, when the processing results of each of the speech recognition, translation, and speech synthesis processes in the speech translation system 40 are returned to the language management server 30, the second text data T2 may be sent to the input terminal 10 and edited before the speech synthesis process, and the edited text data may be used to proceed with the speech synthesis process.

[0129] The language management server 30 may acquire the status of transmission of the second voice data S2 and the second text data T2 to the output terminal 20 and the status of output of the second voice data S2 and the second text data T2 at the output terminal 20, and the acquired information may be transmitted to the input terminal 10 and confirmed at the input terminal 10. With this configuration, the user of the input terminal 10 can ascertain whether the translation result has been conveyed to the user of the output terminal 20. For example, if there is a problem with the transmission of each piece of data to the output terminal 20, it may be possible for the input terminal 10 to issue an instruction to resend each piece of data.

[0130] The first text data T1 and second text data T2 received by the input terminal 10 may be stored in the input terminal 10. Similarly, the second text data T2 and image data Di received by the output terminal 20 may be stored in the output terminal 20. With this configuration, the contents of each of the above data can be confirmed at each terminal 10, 20 even after the validity period of the translation group Gp has expired. For example, after the end of a sightseeing tour, the user of the output terminal 20 can look back on the sightseeing tour by checking the image represented by the image data Di.

[0131] The translation group Gp may include an output terminal 20 to which the speech input to the input terminal 10 is delivered without translation. The output terminal 20 is a terminal whose delivery language is set to the language of the speech input to the input terminal 10, i.e., the same language as the original language. In this case, the language management server 30 transmits first text data T1, which is the speech recognition result of the speech, to the output terminal 20 to which the untranslated speech is delivered. A character string based on the first text data T1 is displayed on the output terminal 20. This allows the user of the output terminal 20 to confirm the speech of the user of the input terminal 10 in text. With this configuration, for example, if participants in a sightseeing tour speak the same language as the tour guide, they can use the output terminal 20 to confirm the tour guide's speech in text, thereby helping them understand the speech. Furthermore, trainee tour guides can use the output terminal 20 to easily understand the guide's content when learning about the content.

[0132] The validity period of the translation group Gp does not have to be set. Alternatively, the validity period may be set as a period during which the translation result of the speech content input to the input terminal 10 is permitted to be delivered to the output terminal 20, and the validity period may be checked when performing speech translation. For example, before sending data such as a screen PC for speech input to the input terminal 10 when starting speech translation, when the language management server 30 receives the first speech data S1, before sending the second speech data S2 to the output terminal 20, etc., the translation management unit 32b of the language management server 30 checks whether the current time is within the validity period of the translation group Gp. If the current time is within the validity period, the translation management unit 32b proceeds with the processing, and if the current time is outside the validity period, the processing is stopped.

[0133] The translation history data 33b does not have to be stored in the language management server 30. Furthermore, as a result of the speech translation, the first text data T1 does not have to be sent to the input terminal 10, and the second text data T2 does not have to be sent to the output terminal 20. In other words, only the second speech data S2 may be sent to the output terminal 20. If the first text data T1 and the second text data T2 are not sent to the terminals 10 and 20 or stored in the language management server 30, then the data does not have to be sent from the speech translation system 40 to the language management server 30.

[0134] The translation management system 50 may use an external general-purpose translation engine instead of the speech translation system 40. That is, the language management server 30 may instruct the external translation engine to translate the first speech data S1 into each target language, and deliver the obtained translation results to the output terminal 20.

[0135] The information to be translated may be input to the input terminal 10 as text. That is, the first information to be translated may be not only speech but also text. In this case, speech recognition processing of the first information is not required. Furthermore, the translation result at the output terminal 20 may be output only as text. That is, speech synthesis may not be performed, and second text data T2 may be delivered to the output terminal 20 as the translation result.

[0136] The input terminal 10 and the output terminal 20 are not limited to smartphones or tablet terminals, but may be dedicated terminals used in the multilingual translation system 100. For example, the output terminal 20 may be a device consisting of earphones and a main body that is connected to the earphones via a wired or wireless connection and communicates data.

[0137] Furthermore, the input terminal 10 and the output terminal 20 may be wearable, such as glasses-type wearable terminals. In the third embodiment, if the output terminal 20 is a glasses-type terminal capable of displaying AR content, when text and AR content are associated in the auxiliary information data 33c, the AR content delivered to the output terminal 20 together with the translation result is displayed on the lenses of the glasses. Therefore, the user of the output terminal 20 can view the AR content superimposed on the real environment, and can enjoy the AR content in a suitable manner.

[0138] There are no particular limitations on the uses of the multilingual translation system 100. The multilingual translation system 100 is not limited to sightseeing tours, and may also be used for various types of guidance and meetings that are not related to sightseeing. [Explanation of symbols]

[0139] Gp...translation group, S1...first voice data, S2...second voice data, T1...first text data, T2...second text data, 10...input terminal, 20...output terminal, 30...language management server, 31...communication unit, 32...control unit, 32a...group management unit, 32b...translation management unit, 33...memory unit, 33a...group data, 33b...translation history data, 33c...attached information data, 40...speech translation system, 41...speech recognition server, 42...translation server, 43...speech synthesis server, 50...translation management system, 100...multilingual translation system.

Claims

1. a group management unit that sets a translation group to which a first terminal and a plurality of second terminals to which translation results of first information input to the first terminal belong, and that sets the translation target language of the first information in the translation group in correspondence with the type of distribution language set for each of the second terminals as a language used for distribution; a translation management unit that, when receiving the first information from the first terminal, instructs the translation of the first information into the translation target languages ​​of the translation group to which the first terminal belongs, and distributes the translation results of each translation target language to the second terminal that has the translation target language set as the distribution language; the group management unit sets a plurality of translation groups to which the same first terminal belongs, and the plurality of translation groups include a plurality of translation groups in which a plurality of target languages ​​for translation are set corresponding to the types of distribution languages ​​of the second terminal, and the plurality of translation groups include a plurality of translation groups in which the combinations of the target languages ​​for translation set are different from each other; When the translation management unit receives the first information from the first terminal, it instructs the translation of the first information into the translation target language of the translation group selected by the first terminal, and distributes the translation result to the second terminal belonging to the translation group. Translation management system.

2. the first terminal is a terminal operated by a tour guide of a sightseeing tour, the second terminal is a terminal operated by a participant of the sightseeing tour, The group management unit sets the translation group for each sightseeing tour as the plurality of translation groups to which the same first terminal belongs. The translation management system of claim 1 .

3. The period during which the second terminal can be added to the translation group is limited to a predetermined period. The translation management system according to claim 1 or 2.

4. The group management unit Upon receiving a request to create the translation group from the first terminal, a new translation group is created to which the first terminal belongs, and identification information of the translation group is transmitted to the first terminal; Adding the second terminal to the translation group corresponding to the identification information based on a request for addition to the translation group together with the identification information received from the second terminal that has acquired the identification information. The translation management system according to any one of claims 1 to 3.

5. When the group management unit receives a request from the second terminal to add the second terminal to the translation group, the group management unit adds the second terminal to the translation group on the condition that the second terminal is located within a predetermined range from the location of the first terminal. The translation management system according to any one of claims 1 to 4.

6. a storage unit for storing auxiliary information data in which text and an image related to the text are associated with each other; The translation management unit delivers an image associated with the text corresponding to the first information in the auxiliary information data to the second terminal together with the translation result. The translation management system according to any one of claims 1 to 5.

7. The first information is input to the first terminal as speech, and the translation result delivered to the second terminal includes speech. The translation management system according to any one of claims 1 to 6.

8. a translation unit that translates the first information into the target language designated by the translation management unit; The translation management system according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Information processor, server device, group generation system, group generation method and program

    JP2012124753A

  • Facility management device, facility management method, program, and facility management system

    JP2018156374A

  • Signal processor, communication system, method to be performed by signal processor, program to be executed by signal processor, method to be performed by communication terminal, program to be executed by communication terminal

    JP2019004392A

  • Information process system, control method of the same, and program

    JP2019121812A

  • Lesson support system, information processing device, lesson support method, and program

    JP2019200233A