Translation Management System
The translation management system optimizes voice translation by grouping terminals by language, reducing redundancy and improving efficiency through collective processing and flexible group management.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TOPPAN HOLDINGS INC
- Filing Date
- 2025-10-24
- Publication Date
- 2026-07-22
AI Technical Summary
Existing voice translation devices require redundant speech recognition, translation, and speech synthesis processes for each participant in a group with different languages, leading to inefficient translation management in multilingual settings.
A translation management system that groups terminals by language distribution, collectively performs translation and speech synthesis for each distribution language, and manages translation groups with expiration periods and identification information for accurate terminal addition.
Reduces processing duplication and load, enhances translation efficiency, and facilitates flexible group management suitable for temporary activities like sightseeing tours.
Smart Images

Figure 0007893359000001 
Figure 0007893359000002 
Figure 0007893359000003
Abstract
Description
Technical Field
[0001] The present invention relates to a translation management system for managing translations into multiple languages.
Background Art
[0002] In recent years, with the increasing communication between people speaking different languages, portable voice translation devices have been developed. A voice translation device recognizes the voice of a first language, converts it into text, and translates the text of the first language into text of a second language. Then, the voice translation device synthesizes and outputs the voice of the second language from the text of the second language (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, the opportunity to translate the utterance content of a speaker into multiple languages is increasing. For example, with the diversification of tourism, it is becoming more common for a single tour to include participants who speak various languages. The participants in a tour include a plurality of participants whose languages match each other, as well as a plurality of participants whose languages are different from each other.
[0005] In these types of tours, the tour guide's speech needs to be translated into multiple languages depending on the participants' languages. While this translation is possible if each participant carries a voice translation device, the process of speech recognition, translation, and speech synthesis performed on each device is redundant and inefficient. For example, speech recognition is duplicated in all participants' voice translation devices. Similarly, translation and speech synthesis are duplicated in the voice translation devices of participants whose languages match.
[0006] The present invention aims to provide a translation management system that can improve the efficiency of processing related to translation into multiple languages. [Means for solving the problem]
[0007] A translation management system that solves the above problems includes: a first terminal and a group management unit that sets a translation group to which a plurality of second terminals to which the translation results of first information input into the first terminal belong, and sets the target languages for translation of the first information in the translation group in accordance with the type of distribution language set for each second terminal as the language used for distribution; and a translation management unit that, when it receives the first information from the first terminal, instructs the first terminal to translate the first information into the target languages of the translation group to which the first terminal belongs, and distributes the translation results of each target language to the second terminals to which that target language is set as the distribution language.
[0008] With the above configuration, translation processing for distribution to a second terminal is performed collectively for each distribution language, and common data is distributed to terminals of the same distribution language. Therefore, compared to cases where translation processing is performed for each terminal and separate data is generated, processing duplication is reduced.
[0009] In the above configuration, the period during which the second terminal can be added to the translation group may be limited to a predetermined period. According to the above configuration, participants in a temporary activity can form a translation group using their mobile devices and utilize the distribution of translation results within that group. Therefore, a translation management system suitable for use in temporary activities, such as sightseeing tours, is realized.
[0010] In the above configuration, the group management unit may, upon receiving a request from the first terminal to generate the translation group, set up a new translation group to which the first terminal belongs and transmit the identification information of the translation group to the first terminal. The group management unit may also, upon receiving a request from the second terminal, which has acquired the identification information, to add the second terminal to the translation group corresponding to the identification information.
[0011] According to the above configuration, the exchange of identification information between the first terminal and the second terminal determines which second terminal will be added to the translation group, that is, which terminal will receive the translation results of the first information input to the first terminal. Therefore, accurate selection of the second terminal to be added to the translation group is possible, and the use of identification information allows for accurate identification of the translation group to which the terminal will be added.
[0012] In the above configuration, the group management unit may set up a plurality of translation groups and set the target language for translation for each translation group, and the plurality of translation groups may include a plurality of translation groups to which the same first terminal belongs, and when the translation management unit receives the first information from the first terminal, it may instruct the first terminal to translate the first information into the target language for translation of the translation group selected by the first terminal, and distribute the translation results to the second terminals belonging to that translation group.
[0013] According to the above configuration, a single first terminal can be configured as a translation group to which multiple translation groups are set up, each with different target terminals for receiving translation results. Users of the first terminal can then use multiple translation groups according to their purpose.
[0014] In the above configuration, when the group management unit receives a request from the second terminal to add the second terminal to the translation group, it may add the second terminal to the translation group on the condition that the second terminal is located within a predetermined range from the location of the first terminal.
[0015] According to the above configuration, the second terminal added to the translation group can be restricted to a terminal located within a predetermined range from the first terminal. Therefore, when it is desired to form a translation group from a first terminal and a second terminal located in a common space, the second terminal to be added to the translation group can be accurately selected.
[0016] In the above configuration, the system includes a storage unit that stores supplementary information data associating text with images related to the text, and the translation management unit may deliver the images associated with the text corresponding to the first information, along with the translation results, to the second terminal.
[0017] With the above configuration, images related to the translation content are displayed on the second device along with the translation result, thereby deepening the user's understanding of the translation content. In the above configuration, the first information is input to the first terminal as audio, and the translation result delivered to the second terminal may include audio.
[0018] The processing load required for speech translation is greater than that required for text translation. With the above configuration, since a translation management system is used for speech translation, the redundancy of processing is minimized, resulting in a significant reduction in the processing load.
[0019] In the above configuration, a translation unit may be provided that translates the first information into the target language specified by the translation management unit. With the above configuration, since the translation management system includes a translation unit, it is possible to improve the efficiency of data exchange and optimize data such as vocabulary lists used for translation, compared to cases where an external translation engine is used.
Advantages of the Invention
[0020] According to the present invention, the efficiency of processing related to translation into multiple languages can be improved.
Brief Description of the Drawings
[0021] [Figure 1] A diagram showing the configuration of a multilingual translation system including a translation management system, regarding the first embodiment of the translation management system. [Figure 2] A diagram showing the concept of a translation group set in the multilingual translation system of the first embodiment. [Figure 3] A sequence diagram showing the procedure for generating a new translation group in the multilingual translation system of the first embodiment. [Figure 4] A diagram showing an example of the screen of an input terminal in the multilingual translation system of the first embodiment. [Figure 5] A diagram showing an example of the screen of an input terminal in the multilingual translation system of the first embodiment. [Figure 6] A sequence diagram showing the procedure for adding an output terminal to a translation group in the multilingual translation system of the first embodiment. [Figure 7] A sequence diagram showing the procedure for voice translation in the multilingual translation system of the first embodiment. [Figure 8] A diagram showing an example of the screen of an input terminal in the multilingual translation system of the first embodiment. [Figure 9] (a) is a diagram showing an example of the screen of an input terminal in the multilingual translation system of the first embodiment, and (b) is a diagram showing an example of the screen of an output terminal in the multilingual translation system of the first embodiment. [Figure 10] A diagram schematically showing the data flow in the multilingual translation system of the first embodiment. <9000106>Regarding the second embodiment of the translation management system, a sequence diagram showing the procedure for adding an output terminal to a translation group in the multilingual translation system of the second embodiment. [Figure 12]A diagram showing the configuration of a multilingual translation system including a translation management system, representing a third embodiment of the translation management system. [Figure 13] A diagram showing an example of the correspondence between text and images in the supplementary information data provided by the translation management system of the third embodiment. [Figure 14] A sequence diagram showing the procedure for speech translation in the multilingual translation system of the third embodiment. [Figure 15] (a) is a diagram showing an example of the input terminal screen in the multilingual translation system of the third embodiment, and (b) is a diagram showing an example of the output terminal screen in the multilingual translation system of the third embodiment. [Modes for carrying out the invention]
[0022] (First Embodiment) A first embodiment of the translation management system will be described with reference to Figures 1 to 10. [Overall structure of the translation management system] As shown in Figure 1, the multilingual translation system 100 includes an input terminal 10, an output terminal 20, a language management server 30, and a speech translation system 40. Of these, the language management server 30 and the speech translation system 40 constitute the translation management system 50.
[0023] Each of the input terminals 10 and output terminal 20 and the language management server 30 transmit and receive data from each other via the network. The language management server 30 and the speech translation system 40 also transmit and receive data from each other via the network. The network used for communication between each terminal 10, 20 and the language management server 30, and the network used for communication between the language management server 30 and the speech translation system 40, may be a general-purpose communication line such as the Internet, or a dedicated communication line for each device.
[0024] Input terminal 10 is an example of a first terminal and has communication and voice input functions. The content of the speaker's speech is input to input terminal 10 as voice. Output terminal 20 is an example of a second terminal and has communication and voice output functions. The translated voice of the above speech content is output from output terminal 20. The multilingual translation system 100 includes multiple output terminals 20, and the multiple output terminals 20 include terminals used by users of different languages. For example, if the multilingual translation system 100 is used for a sightseeing tour, input terminal 10 is carried and used by the tour guide, and output terminals 20 are carried and used by the tour participants.
[0025] Furthermore, the input terminal 10 and output terminal 20 have display functions, shooting functions, etc. The input terminal 10 and output terminal 20 are general-purpose terminals such as smartphones and tablet devices.
[0026] The language management server 30 manages the combination of the input terminal 10 and the multiple output terminals 20 to which the translation results of the utterances entered into the input terminal 10 are delivered, as well as the languages to be translated.
[0027] The language management server 30 comprises a communication unit 31, a control unit 32, and a storage unit 33. The communications unit 31 performs connection processing between the language management server 30, each terminal 10, 20, and the voice translation system 40 via the network, and transmits and receives data between the connected devices.
[0028] The control unit 32 includes a CPU and volatile memory such as RAM. Based on the programs and data stored in the storage unit 33, the control unit 32 controls the processing performed by the communication unit 31, reads and writes information to the storage unit 33, and performs various arithmetic operations, as well as controlling each part of the language management server 30.
[0029] In the multilingual translation system 100, the control unit 32 functions as a group management unit 32a and a translation management unit 32b. The group management unit 32a manages translation groups Gp. Each translation group Gp consists of one input terminal 10 and multiple output terminals 20 that receive the translation results of the utterances entered into the input terminal 10. In each translation group Gp, a distribution language is set for each output terminal 20, which defines the language used for distribution, i.e., the language in which the distributed translation results will be used. Furthermore, for each translation group Gp, a target translation language is set according to the distribution language of the output terminals 20 belonging to the translation group Gp, which is the language in which the utterances are translated. In addition, each translation group Gp has an expiration period. The expiration period is the period during which the addition of output terminals 20 to the translation group Gp is permitted.
[0030] The group management unit 32a sets up multiple translation groups Gp. The setup process for each translation group Gp includes registering the input terminals 10 and output terminals 20 belonging to the translation group Gp, setting the distribution language for the output terminals 20, setting the target languages for translation in the translation group Gp, and setting the validity period of the translation group Gp.
[0031] Furthermore, the data transmission and reception functions in the group management unit 32a are handled by the communication unit 31. When the translation management unit 32b receives audio data of spoken content from the input terminal 10, it instructs the speech translation system 40 to perform speech translation into the target language of the translation group Gp to which the input terminal 10 belongs. The translation management unit 32b then transmits the translation result into the language corresponding to the distribution language of each output terminal 20 belonging to the translation group Gp. The data transmission and reception functions of the translation management unit 32b are handled by the communication unit 31.
[0032] The storage unit 33 includes non-volatile memory and stores programs and data necessary for processing performed by the control unit 32. As part of this data, the storage unit 33 stores group data 33a and translation history data 33b. The group data 33a is data relating to translation groups Gp. For example, the group data 33a includes identification information of the input terminals 10 and output terminals 20 belonging to each translation group Gp, the distribution language of the output terminal 20, the target language for translation, and the validity period. The translation history data 33b includes the translation history for each translation group Gp. For example, the translation history data 33b includes text data which is the recognition result of speech input to the input terminal 10, and text data which is the translation result of said text data into the target language.
[0033] Furthermore, the functions of the group management unit 32a and the translation management unit 32b in the control unit 32 may be implemented by various hardware components such as multiple CPUs and RAM, and software that enables them to function, or by software that provides multiple functions to a single common piece of hardware.
[0034] The speech translation system 40 functions as a translation unit and performs the following processes: recognizing speech and converting it into text, translating the text into a specified language, and synthesizing speech from the translated text. The speech translation system 40 consists of one or more servers and has the programs and data required for each of the above processes. For example, the speech translation system 40 consists of a speech recognition server 41 that performs speech recognition processing, a translation server 42 that performs translation processing, and a speech synthesis server 43 that performs speech synthesis processing.
[0035] [Composition of the translation group] Refer to Figure 2 for details on the translation group Gp. Figure 2 conceptually shows an example of a translation group Gp built on the language management server 30.
[0036] In the language management server 30, multiple translation groups Gp are configured. Figure 2 shows an example where three translation groups Gp are configured. Each translation group Gp includes one input terminal 10 and multiple output terminals 20. For example, translation group GpA includes input terminal 10a and six output terminals 20a to 20f. The same input terminal 10 may belong to different translation groups Gp, and the same output terminal 20 may belong to different translation groups Gp. In other words, the same input terminal 10 may belong to multiple translation groups Gp, and the same output terminal 20 may belong to multiple translation groups Gp.
[0037] As described above, a translation group Gp is a group consisting of an input terminal 10 and an output terminal 20 that delivers the translation result of the spoken content entered into the input terminal 10. For example, when the multilingual translation system 100 is used for a sightseeing tour, one translation group Gp is constructed from the input terminal 10 carried by the tour guide of one sightseeing tour and the output terminals 20 carried by the participants of that sightseeing tour. For example, if one tour guide guides multiple sightseeing tours, a translation group Gp is constructed for each sightseeing tour from the tour guide's input terminal 10 and the participants' output terminals 20.
[0038] In each translation group Gp, the distribution language for each output terminal 20 is set. Then, the target languages for translation in that translation group Gp are set according to the type of distribution language of the output terminal 20 belonging to that translation group Gp. In other words, the distribution languages of the multiple output terminals 20 belonging to the translation group Gp are classified, and the target languages for translation in that translation group Gp are set according to the languages of this classification.
[0039] For example, in translation group GpA, English is set as the distribution language for output terminals 20a, 20b, and 20c, Chinese is set as the distribution language for output terminal 20d, and Korean is set as the distribution language for output terminals 20e and 20f. Thus, while translation group GpA has six output terminals 20a to 20f, the distribution languages for these output terminals 20 are English, Chinese, and Korean. In other words, the distribution languages of output terminals 20a to 20f belonging to translation group GpA are classified into three categories. The languages corresponding to these classifications are then set as the target languages for translation. That is, translation group GpA has three target languages set: English, Chinese, and Korean.
[0040] The target languages for translation are set for each translation group Gp to correspond to the types of languages distributed by the output terminal 20. In other words, the target languages for translation group Gp are limited to the languages that can be translated by the speech translation system 40 and are distributed by the output terminal 20 belonging to translation group Gp. Therefore, the number and breakdown of target languages in different translation groups Gp may or may not be the same. For example, translation group GpA has three target languages: English, Chinese, and Korean, while translation group GpB has two target languages: English and French.
[0041] The language of the spoken content entered into the input terminal 10, i.e., the language before translation, may be uniformly set to a predetermined language, or it may be set for each translation group Gp. In the example shown in Figure 2, it is assumed that the input language to the input terminal 10 is uniformly set to Japanese.
[0042] Each translation group Gp has an expiration period. The expiration period may be from a specified date to a specified date, or from a specified time on a specified date to a specified time. As described above, the expiration period is the period during which the addition of output terminals 20 to the translation group Gp is permitted, or in other words, the period during which voice translation is permitted in the translation group Gp. For example, if the multilingual translation system 100 is used for a tourist tour, the period during which the tourist tour is scheduled to be conducted is set as the expiration period of the translation group Gp corresponding to the tourist tour.
[0043] [Translation Management System Operation] The operation of the multilingual translation system 100, including the translation management system 50, will be explained with reference to Figures 3 to 9. Note that the processing of the input terminal 10 and output terminal 20 in the operation of the multilingual translation system 100 may be performed based on a web application or based on application software installed on these terminals.
[0044] For example, the processing of the input terminal 10 tends to be more extensive and complex than that of the output terminal 20. Therefore, executing the processing of the input terminal 10 based on application software installed on the input terminal 10 is highly effective in facilitating the smooth operation of the input terminal 10. On the other hand, for the output terminal 20, which has more users than the input terminal 10, the use of a web application reduces the burden of setting up the output terminal 20 for use with the multilingual translation system 100, thus greatly increasing its convenience. If a web application is used, for example, participants in a tourist tour can easily use their own mobile devices as output terminals 20 while traveling, making the multilingual translation system 100 easier to use.
[0045] <Translation group settings> First, let's explain the procedure for setting up a translation group (Gp). As shown in Figure 3, based on a predetermined operation performed on the input terminal 10, a request for the creation of a new translation group Gp is sent from the input terminal 10 to the language management server 30 (step S10). For example, as shown in Figure 4, based on the selection of an area Ra that instructs the creation of a new translation group Gp on the screen Pa displayed when application software is launched on the input terminal 10, the above creation request is sent to the language management server 30.
[0046] As shown in Figure 3 above, when a request for the creation of a new translation group Gp is received, the group management unit 32a of the language management server 30 performs the process of creating a new translation group Gp (step S11). Specifically, the group management unit 32a adds the data of the new translation group Gp to the group data 33a. In the data of the new translation group Gp, the input terminal 10 that sent the request for the creation of a new translation group Gp in step S10 is set as the input terminal 10 belonging to the translation group Gp. In addition, the validity period of the translation group Gp is also set in the data of the new translation group Gp. The validity period may be set to an arbitrary period specified by the input terminal 10 and notified to the language management server 30, or it may be set to a period predetermined by the language management server 30.
[0047] Next, the group management unit 32a transmits information including the identification information Id of the generated translation group Gp to the input terminal 10 (step S12). Upon receiving this information, the input terminal 10 displays the information including the identification information Id on its display unit (step S13).
[0048] The identification information ID of the translation group Gp is displayed on the input terminal 10 in a form that can be mechanically read by the output terminal 20, or in a form that can be manually transmitted to the user of the output terminal 20. For example, it is preferable that the identification information ID of the translation group Gp is displayed on the input terminal 10 as a two-dimensional code and read by the output terminal 20. The two-dimensional code contains, for example, the identification information ID and a URL indicating the access destination of the language management server 30.
[0049] Figure 5 shows an example of an input terminal 10 displaying a two-dimensional code Cd containing the identification information Id of a translation group Gp. It is preferable that, after the generation of a new translation group Gp, the identification information Id of that translation group Gp be made available for display on the input terminal 10 at any time. For example, as shown in Figure 4, an area Rb for selecting an already generated translation group Gp is included on screen Pa and displayed on the input terminal 10. When area Rb is selected, in response to a request from the input terminal 10, information including the identification information Id of the translation group Gp corresponding to area Rb is sent from the language management server 30 to the input terminal 10, and this information is displayed on the input terminal 10.
[0050] Furthermore, the two-dimensional code containing the identification information ID of the translation group Gp may be provided to the user of the output terminal 20 not only when displayed on the input terminal 10, but also when printed on paper based on a printout of the input terminal 10's screen. Similarly, if the identification information ID of the translation group Gp is displayed on the input terminal 10 in a form other than a two-dimensional code, the identification information ID of the translation group Gp may be provided to the user of the output terminal 20 through written or oral communication, and the user may input the identification information ID into the output terminal 20. In short, it is sufficient that the output terminal 20 can obtain the identification information ID of the translation group Gp.
[0051] Furthermore, as shown in Figure 5, the screen Pb displayed on the input terminal 10 also includes an area Rc for instructing the start of speech translation in the selected translation group Gp. As shown in Figure 6, the output terminal 20 obtains the identification information ID of the translation group Gp and sends an additional request to the translation group Gp to the language management server 30 (step S20). For example, if the identification information ID of the translation group Gp is included in a two-dimensional code, the output terminal 20 reads the two-dimensional code and accesses the URL obtained therefrom. This notifies the language management server 30 of the identification information ID of the translation group Gp from the output terminal 20, and functions as an additional request to the translation group Gp.
[0052] Upon receiving an additional request to the translation group Gp, the group management unit 32a of the language management server 30 requests the output terminal 20 to send language information (step S21). The language information specifies the distribution language.
[0053] Upon receiving a request from the language management server 30, the output terminal 20 transmits language information to the language management server 30 (step S22). For example, based on the user specifying a language through the screen of the output terminal 20, language information indicating the specified language is transmitted from the output terminal 20 to the language management server 30. Alternatively, information indicating the language used on the output terminal 20, that is, the language used to display information on the output terminal 20, may be transmitted from the output terminal 20 to the language management server 30 as language information.
[0054] Upon receiving language information, the group management unit 32a of the language management server 30 performs the process of adding the output terminal 20 to the translation group Gp corresponding to the identification information Id obtained in connection with the above additional request (step S23).
[0055] Specifically, the group management unit 32a assigns the output terminal 20 to the translation group Gp corresponding to the identification information Id. That is, in the data for the translation group Gp corresponding to the identification information Id in the group data 33a, the output terminal 20 that was the source of the additional request to the translation group Gp in step S20 is set as the output terminal 20 belonging to the translation group Gp. Furthermore, in the group data 33a, the group management unit 32a sets the language indicated by the language information received from the output terminal 20 as the distribution language for the added output terminal 20.
[0056] In the above-described configuration of the translation group Gp, for example, access from the output terminal 20 to the language management server 30 based on the acquisition of identification information Id is permitted only within the validity period of the translation group Gp corresponding to the identification information Id, thereby limiting the period during which the output terminal 20 can be added to the translation group Gp to the above validity period. For example, a time-limited URL that is only accessible within the above validity period can be used as the URL held by the two-dimensional code containing the identification information Id.
[0057] Alternatively, when the language management server 30 receives a request from the output terminal 20 to add to the translation group Gp, the group management unit 32a of the language management server 30 may check whether the current time is within the validity period of the target translation group Gp, and only add the output terminal 20 to the translation group Gp if the current time is within the validity period. Furthermore, when the group management unit 32a sends the generated identification information Id of the translation group Gp to the input terminal 10, it may check whether the current time is within the validity period of the translation group Gp, and only send the identification information Id to the input terminal 10 if the current time is within the validity period. These configurations also limit the period during which the output terminal 20 can be added to the translation group Gp to within the validity period of the translation group Gp.
[0058] The target languages for translation group Gp are set to correspond to the types of distribution languages of the output terminals 20 belonging to that translation group Gp at that time. In other words, if the number of distribution languages of the output terminals 20 belonging to that translation group Gp increases due to the addition of an output terminal 20 to that translation group Gp, or in other words, if the distribution language of the added output terminal 20 is different from any of the distribution languages of the output terminals 20 already belonging to that translation group Gp, then the distribution language of the added output terminal 20 is added to the target languages of that translation group Gp. That is, the target languages in the data of the translation group Gp in group data 33a are updated, and the distribution language of the added output terminal 20 is added to the target languages.
[0059] Thus, the target languages for translation by translation group Gp change in accordance with the changes in the output terminals 20 belonging to translation group Gp, and in accordance with the changes in the types of languages to be delivered. Furthermore, it is preferable that the distribution language of the output terminal 20 can be changed at will. Changing the distribution language is performed, for example, by the following procedure. That is, the screen displayed on the output terminal 20 includes an area for instructing a change in the distribution language, and when this area is selected and a new language is specified, a request to change the distribution language, along with new language information, is sent from the output terminal 20 to the language management server 30. Upon receiving a request to change the distribution language, the group management unit 32a of the language management server 30 modifies the group data 33a to change the distribution language of the output terminal 20 to the language indicated by the new language information.
[0060] If a change in the distribution language of output terminal 20 results in a change in the distribution language type of the translation group Gp to which output terminal 20 belongs, the target languages for translation in the translation group Gp will be changed accordingly.
[0061] In the above configuration, for example, if the multilingual translation system 100 is used for a sightseeing tour, at the start of the tour, the tour guide displays a two-dimensional code containing the identification information Id of the translation group Gp on the input terminal 10, and the tour participants read the two-dimensional code with the output terminal 20. As a result, the output terminal 20 of the tour participants is added to the translation group Gp, and the sightseeing tour can proceed using the voice translation of the translation group Gp. If the generated identification information Id of the translation group Gp can be displayed on the input terminal 10 at any time, it is possible to add a new tour participant's output terminal 20 to the translation group Gp in the middle of the sightseeing tour, thus increasing the flexibility of the configuration of the translation group Gp.
[0062] <Voice Translation> Next, I will explain the procedure for voice translation. As shown in Figure 7, first, with a translation group Gp selected, audio representing the spoken content is input to the input terminal 10 (step S30). Specifying an existing translation group Gp or creating a new translation group functions as selecting a translation group Gp. For example, as shown in Figure 8, when a translation group Gp is selected and the voice translation is instructed to start, a screen PC for voice input is displayed on the input terminal 10, and the voice is input to the input terminal 10 when the user speaks towards the input terminal 10.
[0063] The first audio data S1, which represents the voice input into the input terminal 10, is transmitted from the input terminal 10 to the language management server 30 (step S31). The first audio data S1 is an example of the first information. Also, at or before the transmission of the first audio data S1, information indicating the selected translation group Gp is transmitted from the input terminal 10 to the language management server 30.
[0064] Upon receiving the first voice data S1, the translation management unit 32b of the language management server 30 sends a voice translation request to the voice translation system 40 (step S32). Specifically, the translation management unit 32b refers to the group data 33a and identifies the target language for translation of the translation group Gp selected on the input terminal 10. Then, the translation management unit 32b instructs the voice translation system 40 to translate to the identified target language. The first voice data S1 is transmitted from the language management server 30 to the voice translation system 40.
[0065] Upon receiving a request from the language management server 30, the speech translation system 40 performs speech translation processing (step S33). Specifically, the speech translation system 40 first performs speech recognition processing on the first speech data S1 and generates first text data T1, which is text data in the language corresponding to the first speech data S1. Next, the speech translation system 40 translates the first text data T1 into the target language and generates second text data T2, which is text data in the target language. When there are multiple target languages, the speech translation system 40 translates the first text data T1 into each of the target languages and generates second text data T2 for each language. Then, the speech translation system 40 synthesizes speech based on the second text data T2 and generates second speech data S2, which is speech data in the language corresponding to the second text data T2. As a result, second speech data S2 is generated in a number of languages corresponding to the number of target languages.
[0066] Once the speech translation process is complete, the speech translation system 40 sends the processing results to the language management server 30 (step S34). The processing results include the second speech data S2 and the first text data T1 and the second text data T2.
[0067] Upon receiving the processing result, the translation management unit 32b of the language management server 30 transmits the second audio data S2 in the language corresponding to the set distribution language to the output terminal 20 belonging to the translation group Gp selected by the input terminal 10 (step S35). At this time, it is preferable that the translation management unit 32b also transmits the second text data T2 corresponding to the second audio data S2 to the output terminal 20. It is also preferable that the translation management unit 32b transmits the first text data T1 to the input terminal 10.
[0068] Furthermore, the translation management unit 32b stores at least the first text data T1 and the second text data T2 from the received processing results in the storage unit 33, along with the data of the corresponding translation group Gp in the translation history data 33b (step S36).
[0069] Upon receiving the second audio data S2, the output terminal 20 outputs audio based on the second audio data S2 (step S37). As a result, the audio translated from the utterance input to the input terminal 10 is output by the output terminal 20.
[0070] When the second audio data S2 is received along with the second text data T2, the output terminal 20 displays a string based on the second text data T2 on its display unit. When the first text data T1 is received, the input terminal 10 displays a string based on the first text data T1 on its display unit.
[0071] Figure 9(a) shows an example of screen Pd displayed on the input terminal 10 as a result of the above-described process, and Figure 9(b) shows an example of screen Pe displayed on the output terminal 20 as a result of the above-described process. Screen Pd on the input terminal 10 includes an area Rd that shows a string based on the first text data T1, i.e., the speech recognition result of the utterance entered into the input terminal 10. By looking at screen Pd, the user of the input terminal 10 can confirm whether the utterance has been correctly recognized. Screen Pe on the output terminal 20 includes an area Re that shows a string based on the second text data T2, i.e., the sentence corresponding to the translation result of the utterance entered into the input terminal 10. By looking at screen Pe, the user of the output terminal 20 can confirm the translation result in text in addition to audio, making it easier to understand the translation result.
[0072] Furthermore, the second text data T2 for each target language may be sent from the language management server 30 to the input terminal 10, and a string based on the second text data T2 may be displayed on the input terminal 10. With this configuration, the user of the input terminal 10 can check the translation result of the spoken content and determine whether the translation result is appropriate or not.
[0073] Furthermore, the input terminal 10 is capable of checking the translation history data 33b. Checking the translation history data 33b is performed, for example, in the following procedure. That is, with a translation group Gp selected, a request to check the translation history data 33b is sent from the input terminal 10 to the language management server 30. The request to check the translation history data 33b is made, for example, based on the selection of an area on the screen of the input terminal 10 that instructs the user to check the translation history. Upon receiving the request to check the translation history data 33b, the translation management unit 32b of the language management server 30 sends the data of the translation group Gp selected by the input terminal 10 from the data contained in the translation history data 33b to the input terminal 10. The input terminal 10 displays the received data. This allows the user of the input terminal 10 to check the history of speech translation in the selected translation group Gp, that is, what kind of translation result was sent to the output terminal 20 for what kind of utterance for each target language.
[0074] The operation of the multilingual translation system 100 will be explained with reference to Figure 10. Figure 10 conceptually illustrates speech translation in translation group GpA shown in Figure 2. In translation group GpA, the target languages for translation are English, Chinese, and Korean, and Japanese speech is input to the input terminal 10a.
[0075] When Japanese voice data is transmitted from the input terminal 10a to the translation management system 50, the translation management system 50 performs voice translation of the voice data into the target language. Specifically, translation from Japanese to English is performed to generate English voice data, translation from Japanese to Chinese is performed to generate Chinese voice data, and translation from Japanese to Korean is performed to generate Korean voice data.
[0076] Then, the English audio data is distributed to three output terminals 20a-20c whose distribution language is set to English, the Chinese audio data is distributed to one output terminal 20d whose distribution language is set to Chinese, and the Korean audio data is distributed to two output terminals 20e and 20f whose distribution language is set to Korean.
[0077] Thus, even when there are multiple output terminals 20 with the same distribution language, only one audio data file is generated based on the translation into the language corresponding to that distribution language, and the same audio data is distributed to multiple output terminals 20. In other words, the translation and speech synthesis processing for distribution to multiple output terminals 20 is performed collectively for each distribution language. Therefore, compared to the case where translation and speech synthesis processing are performed for each output terminal 20 and separate audio data is generated, processing duplication is reduced.
[0078] Furthermore, the text data generated by speech recognition on the audio data before translation can be used as a single, common data set for translation into each language. In other words, by performing the speech recognition processing for distribution to multiple output terminals 20 in a single batch, processing duplication is reduced compared to the case where speech recognition processing is performed for each output terminal 20.
[0079] Furthermore, since speech recognition, translation, and speech synthesis are performed together as described above, the processing load on the translation management system 50 is reduced compared to when these processes are repeatedly performed on a server for each target terminal. Therefore, rapid speech translation and delivery of translation results become possible, meaning that the time required from inputting spoken content to delivering translation results can be shortened.
[0080] Furthermore, since the input terminal 10 and output terminal 20 carried by the user do not need to have translation functions, the terminal configuration is simplified, and it becomes easy to use general-purpose terminals as input terminal 10 and output terminal 20. Therefore, compared to using dedicated terminals, the effort required to prepare and distribute terminals is reduced, and compared to a system that performs voice translation through direct communication between dedicated terminals, the burden required to set up the communication environment is reduced. Thus, the convenience of using the multilingual translation system 100 is enhanced.
[0081] As described above, the first embodiment provides the following effects. (1) For each distribution language, translation and speech synthesis processing for distribution to the output terminal 20 is performed collectively, and common data is distributed to the output terminal 20 of the same distribution language. Therefore, processing duplication is reduced compared to the case where translation and speech synthesis processing is performed for each output terminal 20 and separate data is generated. In addition, speech recognition processing for distribution to multiple output terminals 20 is performed collectively, and common data is used for translation to each target language. Therefore, processing duplication is reduced compared to the case where speech recognition processing is performed for each output terminal 20 and separate data is generated. This makes it possible to improve the efficiency of processing related to translation to multiple languages.
[0082] (2) The setting of an effective period restricts the addition of output terminals 20 to the translation group Gp to within a predetermined period. Therefore, participants in temporary activities can configure the translation group Gp from their own devices and utilize the distribution of translation results from that translation group Gp. Thus, a system suitable for use in temporary activities such as sightseeing tours is realized.
[0083] (3) In response to a request from the input terminal 10 to generate a translation group Gp, the language management server 30 generates a new translation group Gp and sends the identification information Id of the translation group Gp to the input terminal 10. Then, based on the language management server 30 receiving a request from the output terminal 20, which has acquired the identification information Id, to add the output terminal 20 to the translation group Gp corresponding to the identification information Id. In this way, the output terminal 20 to be added to the translation group Gp, i.e., the terminal to which the translation results will be distributed, is determined through the exchange of identification information Id between the input terminal 10 and the output terminal 20. Therefore, it is possible to accurately select the output terminal 20 to be added to the translation group Gp, and by using the identification information Id, it is possible to accurately identify the translation group Gp to which the output terminal 20 is added.
[0084] (4) The language management server 30 sets up multiple translation groups Gp, and these translation groups Gp include multiple translation groups Gp to which the same input terminal 10 belongs. When the language management server 30 receives the first voice data S1, it instructs the translation to be performed in the target language of the translation group Gp selected by the input terminal 10, and delivers the translation result to the output terminal 20 belonging to that translation group Gp. With this configuration, it is possible to set up multiple translation groups Gp to which a single input terminal 10 belongs, each with different terminals to which the translation result is delivered, and users of the input terminal 10 can use multiple translation groups Gp according to their purpose of use.
[0085] (5) Since the translation management system 50 is used for speech translation, redundant processing is reduced in speech translation, which has a higher processing load than text translation. Therefore, a significant reduction in processing load is obtained.
[0086] (6) The translation management system 50 includes a voice translation system 40 that functions as a translation unit. Therefore, compared to the case where an external translation engine is used, it is possible to improve the efficiency of the processing required for data exchange and to optimize the data such as vocabulary lists used for translation.
[0087] (Second Embodiment) Referring to Figure 11, a second embodiment of the translation management system will be described. The translation management system of the second embodiment differs from the first embodiment in that it verifies the location of the terminal when adding an output terminal to a translation group. In the following, the differences between the second embodiment and the first embodiment will be described in detail, and components similar to those in the first embodiment will be denoted by the same reference numerals and their descriptions will be omitted.
[0088] For example, when the multilingual translation system 100 is used in a tourist tour, the user of the input terminal 10 and the user of the output terminal 20 are located in the same space. In the second embodiment, the output terminal 20 is added to the translation group Gp on the condition that the output terminal 20 is located within a predetermined range from the input terminal 10.
[0089] In the second embodiment, each of the input terminal 10 and the output terminal 20 has a function for acquiring location information indicating the location of the terminal. The location information can be any information that can identify the absolute position of each terminal or the relative position of each terminal with respect to a predetermined reference, and can be used to determine whether or not the output terminal 20 is located within a predetermined range centered on the input terminal 10.
[0090] For example, location information may include latitude and longitude data obtained using GPS (Global Positioning System), or information that allows for the determination of relative location using short-range wireless communication such as Wi-Fi® or Bluetooth®, sound waves, ultrasound, or light. If necessary, transmitters and receivers for acquiring location information may be installed in locations where the input terminal 10 and output terminal 20 may be located, such as within tourist facilities where tours are conducted.
[0091] As shown in Figure 11, similar to the first embodiment, the output terminal 20 obtains the identification information Id of the translation group Gp and sends an additional request to the translation group Gp to the language management server 30 (step S40).
[0092] Upon receiving an additional request to the translation group Gp, the group management unit 32a of the language management server 30 requests the output terminal 20 to transmit language information (step S41). The group management unit 32a also requests the output terminal 20 and the input terminal 10 belonging to the translation group Gp corresponding to the above identification information Id to transmit location information (step S42).
[0093] Upon receiving a request from the language management server 30, the output terminal 20 transmits language information and the location information of the output terminal 20 to the language management server 30 (step S43). Also, upon receiving a request from the language management server 30, the input terminal 10 transmits the location information of the input terminal 10 to the language management server 30 (step S44).
[0094] Upon receiving location information from both the input terminal 10 and the output terminal 20, the group management unit 32a of the language management server 30 determines, based on the received location information, whether or not the output terminal 20 is located within a predetermined range from the input terminal 10 (step S45).
[0095] When it is confirmed that the output terminal 20 is located within a predetermined range from the input terminal 10, the group management unit 32a performs the process of adding the output terminal 20 to the translation group Gp corresponding to the identification information Id (step S46). If it is determined that the output terminal 20 is not located within a predetermined range from the input terminal 10, the process of adding the output terminal 20 to the translation group Gp is not performed, and the output terminal 20 is notified that it cannot be added to the translation group Gp.
[0096] In the second embodiment, the speech translation processing in the translation group Gp is performed in the same manner as in the first embodiment. In the second embodiment, the output terminal 20 is added to the translation group Gp on the condition that the output terminal 20 is located within a predetermined range from the input terminal 10, so that the output terminal 20 to be added to the translation group Gp is accurately identified. This is particularly effective when it is desired to configure the translation group Gp with an input terminal 10 and an output terminal 20 located in a common space.
[0097] Furthermore, the configuration of the second embodiment also contributes to improving security regarding the distribution of translation results. For example, if the spoken content entered into the input terminal 10 contains information that should not be leaked to a third party, even if a third party illegally obtains the identification information Id and tries to join the translation group Gp from a remote location, the addition of the third party's terminal to the translation group Gp will be refused, thus preventing the translation results of the spoken content from being leaked to a third party.
[0098] In addition to adding the output terminal 20 to the translation group Gp, the distribution of translation results to the output terminal 20 may also be conditional on the output terminal 20 being located within a predetermined range from the input terminal 10. That is, before transmitting the second audio data S2 to the output terminal 20, the translation management unit 32b of the language management server 30 obtains location information from each terminal 10, 20, and transmits the second audio data S2 to the output terminal 20 if the output terminal 20 is located within a predetermined range from the input terminal 10. This configuration further enhances the security of the distribution of translation results.
[0099] Furthermore, in the above configuration, if there is an output terminal 20 belonging to the translation group Gp that is not located within a predetermined range from the input terminal 10 when distributing the translation results, the input terminal 10 may be notified that there is an output terminal 20 outside the predetermined range. With this configuration, for example, if the multilingual translation system 100 is used for a sightseeing tour, if there is a participant who has become separated from the tour group, the input terminal 10 will be notified that there is an output terminal 20 outside the predetermined range. Therefore, the tour guide using the input terminal 10 can find out if there are any lost participants, etc.
[0100] In the second embodiment, the location information of the input terminal 10 and the output terminal 20 does not necessarily have to be obtained from the terminals themselves. The language management server 30 may obtain the location information of the terminals based on the output of a device such as a radio wave receiver located near the input terminal 10 and the output terminal 20. Furthermore, if the locations of the input terminal 10 and the output terminal 20 are limited to a specific range, the location of the output terminal 20 within that specific range may be a condition for adding the output terminal 20 to the translation group Gp, instead of the location of the input terminal 10.
[0101] As described above, according to the second embodiment, in addition to the effects (1) to (6) of the first embodiment, the following effects can be obtained. (7) When the language management server 30 receives a request from the output terminal 20 to add it to the translation group Gp, it adds the output terminal 20 to the translation group Gp on the condition that the output terminal 20 is located within a predetermined range from the input terminal 10. Therefore, when it is desired to configure the translation group Gp from the input terminal 10 and the output terminal 20 located in a common space, the output terminal 20 to be added to the translation group Gp is accurately selected. Furthermore, security regarding the distribution of translation results is also enhanced.
[0102] (Third embodiment) A third embodiment of the translation management system will be described with reference to Figures 12 to 15. The third embodiment differs from the first embodiment in that, in addition to the translation results, images related to the spoken content are delivered to the output terminal. In the following, the differences between the third embodiment and the first embodiment will be described in detail, and components similar to those in the first embodiment will be denoted by the same reference numerals and their descriptions will be omitted.
[0103] As shown in Figure 12, the language management server 30 of the third embodiment stores supplemental information data 33c in the storage unit 33. In the supplemental information data 33c, text such as words is associated with images related to that text. For example, if the multilingual translation system 100 is used for a tourist tour, key words related to the tourist are associated with images related to those words.
[0104] Figure 13 shows an example of the correspondence between text and images in the supplementary information data 33c. In Figure 13, the text "Matsuno Ōrōka" is associated with an image of a painting depicting a scene from the Akō Incident, specifically the sword fight that occurred in the Matsuno Ōrōka of Edo Castle. The relationship between text and images is not particularly limited. For example, an image of a picture or photograph of an object, person, or place mentioned in the text may be associated with that text, or an image of a scene from a historical event or occurrence mentioned in the text may be associated with that text.
[0105] In the third embodiment, the translation group Gp is set up in the same manner as in the first embodiment. Referring to Figure 14, the procedure for speech translation in the third embodiment will be described. As shown in Figure 14, the processing in steps S50 to S53 is the same as the processing in steps S30 to S33 of the first embodiment. That is, when speech representing the content of a speech is input to the input terminal 10, first speech data S1 representing the speech is sent from the input terminal 10 to the language management server 30. The language management server 30 instructs the speech translation system 40 to translate the first speech data S1 into the target language of the translation group Gp selected by the input terminal 10. The speech translation system 40 performs speech translation processing and sends first text data T1, second text data T2 and second speech data S2 for each target language to the language management server 30.
[0106] Upon receiving the processing result including the first text data T1, the second text data T2, and the second audio data S2, the translation management unit 32b of the language management server 30 identifies the image associated with the text contained in the first text data T1 using the accompanying information data 33c (step S55).
[0107] If the above image is identified, the translation management unit 32b transmits image data Di indicating the identified image to the output terminal 20 along with the second audio data S2 in the language corresponding to the distribution language (step S56). If there is no image associated with the text contained in the first text data T1 in the accompanying information data 33c, the image data Di is not transmitted, and the same processing as in the first embodiment proceeds.
[0108] Furthermore, the translation management unit 32b stores the first text data T1 and the second text data T2 in the storage unit 33 as translation history data 33b, similar to the first embodiment (step S57). The translation history data 33b may also include image data Di.
[0109] Upon receiving the second audio data S2 and image data Di, the output terminal 20 outputs audio based on the second audio data S2 and displays an image based on the image data Di on the display unit (step S58). As a result, the audio translated from the utterance input to the input terminal 10 is output by the output terminal 20, and an image related to the utterance is displayed on the output terminal 20.
[0110] Furthermore, it is preferable that the transmission of the second text data T2 to the output terminal 20 and the transmission of the first text data T1 to the input terminal 10 be carried out in the same manner as in the first embodiment. Figure 15(a) shows an example of screen Pf displayed on the input terminal 10 as a result of the above-described process, when the accompanying information data 33c has been associated with the text and image as illustrated in Figure 13. Similarly, Figure 15(b) shows an example of screen Pg displayed on the output terminal 20 as a result of the above-described process. Screen Pf on the input terminal 10 includes an area Rd showing a string based on the first text data T1. Screen Pg on the output terminal 20 includes an area Re showing a string based on the second text data T2 and an area Rf showing an image based on image data Di. In the example shown in Figure 15, the utterance entered into the input terminal 10 includes the word "Matsuno Ōrōka" (Pine Great Corridor), and as a result, the image associated with that word in the accompanying information data 33c is displayed on the output terminal 20. This allows the user of the output terminal 20 to see an image related to the translation result, thereby deepening their understanding of the translation result.
[0111] The supplementary information data 33c may be stored in the speech translation system 40 instead of the language management server 30. In this case, after generating the first text data T1, the speech translation system 40 identifies the image associated with the text contained in the first text data T1 using the supplementary information data 33c, and sends the image data Di, which is the data of that image, along with the first text data T1, the second text data T2, and the second speech data S2, as part of the processing result to the language management server 30. The translation management unit 32b of the language management server 30 then sends the image data Di along with the second speech data S2 in the language corresponding to the distribution language to the output terminal 20.
[0112] If the voice translation system 40 stores supplemental information data 33c, the supplemental information data 33c may constitute vocabulary list data for translation from first text data T1 to second text data T2. That is, vocabulary list data is formed by associating the text of the language to be translated, the text of the language to be translated corresponding to that text, and images associated with these texts, and when translating from first text data T1 to second text data T2 using the vocabulary list data, the images may be identified. With such a configuration, for example, by registering specialized terms used by guides in tourist tours in the vocabulary list data, accurate translation of specialized terms becomes possible.
[0113] In short, the translation management system 50 stores the supplementary information data 33c, and the images associated with the text corresponding to the spoken content entered into the input terminal 10, along with the translation result of the spoken content, are delivered to the output terminal 20.
[0114] Furthermore, in the supplemental information data 33c, text such as words may be associated with images related to the text and explanatory texts in each language. When the image associated with the text contained in the first text data T1 is identified, the explanatory text is also identified, and in addition to the second audio data S2 and image data Di, the output terminal 20 may also receive the data of the explanatory text in the language corresponding to the second audio data S2. The output terminal 20 displays the explanatory text along with the image based on the image data Di. With this configuration, the user of the output terminal 20 can check the explanation related to the image in addition to the image related to the translation result, thereby deepening their understanding of the translation result.
[0115] Furthermore, the user of the output terminal 20 may choose whether or not to display images and explanatory text on the output terminal 20. For example, the display and hiding of images may be switched by selecting a predetermined area on the screen of the output terminal 20.
[0116] Furthermore, it may be possible to set supplemental information data 33c used for speech translation in each translation group Gp. For example, supplemental information data 33c corresponding to the content of a sightseeing tour, that is, supplemental information data 33c that associates words that may be included in the content explained by the tour guide during the sightseeing tour with images related to those words, is generated for each sightseeing tour and stored in the language management server 30. Then, when a new translation group Gp is generated in response to a request from the input terminal 10, based on the selection of the type of sightseeing tour at the input terminal 10, supplemental information data 33c corresponding to the selected sightseeing tour is selected from among the multiple supplemental information data 33c and set as the supplemental information data 33c for that translation group Gp. With this configuration, images more suitable for the purpose of using speech translation in the translation group Gp are displayed on the output terminal 20.
[0117] Furthermore, the content associated with the text in the supplementary information data 33c is not limited to still images; it may also be video or 3DCG. In addition, the content associated with the text in the supplementary information data 33c may also be AR (Augmented Reality) content. That is, the still images, videos, 3DCG, etc., delivered along with the translation results on the output terminal 20 may be displayed on the display unit superimposed on the real environment surrounding the output terminal 20.
[0118] In particular, when the multilingual translation system 100 is used in tourist tours, the tour guide's speech often includes cultural and historical elements, as well as highly specialized topics. Therefore, displaying videos, 3DCG, or AR content along with the translation results on the output terminal 20 helps in understanding the translation. Furthermore, since the tour guide's speech often relates to the location of the participants or the surrounding exhibits, displaying AR content overlaid on the real environment surrounding the participants, along with the translation results, makes it easier to understand the translation and increases participants' interest in the translation. As a result, communication between the tour guide and participants can proceed smoothly. In addition, since multiple participants can share this content through the output terminal 20, communication among participants is also promoted.
[0119] As described above, according to the third embodiment, in addition to the effects (1) to (6) of the first embodiment, the following effects can be obtained. (8) The language management server 30 delivers to the output terminal 20, along with the translation result, content such as images associated with the text corresponding to the first audio data S1 in the accompanying information data 33c. As a result, the content related to the translation is displayed on the output terminal 20 along with the translation result, thereby deepening the user's understanding of the translation.
[0120] (modified version) Each of the above embodiments can be implemented with the following modifications. Furthermore, each of the above embodiments and each of the following modifications may be implemented in combination with each other.
[0121] A translation group Gp may include multiple input terminals 10. Furthermore, one terminal may function as both an input terminal 10 and an output terminal 20 that receives translation results of speech input to other input terminals 10. In other words, within a single translation group Gp, translation results of speech input to each terminal may be distributed to terminals other than that terminal, and two-way communication between terminal users may be possible.
[0122] - Processing using translation group Gp may be stopped in response to a request from input terminal 10. Stopping processing of translation group Gp is performed, for example, by the following procedure. That is, the screen displayed on input terminal 10 includes an area for instructing to stop processing of translation group Gp, and by selecting this area, a request to stop processing of translation group Gp is sent from input terminal 10 to language management server 30. Upon receiving this request, the group management unit 32a of language management server 30 prohibits the addition of output terminal 20 to translation group Gp selected on input terminal 10, and prohibits speech translation processing in said translation group Gp. With this configuration, even within the validity period of translation group Gp, speech translation using translation group Gp can be stopped at the discretion of the user of input terminal 10.
[0123] - In response to a request from the output terminal 20, it may be possible to remove the output terminal 20 from the translation group Gp. The removal of the output terminal 20 is performed, for example, by the following procedure. That is, the screen displayed on the output terminal 20 includes an area for instructing removal from the translation group Gp, and by selecting this area, a request for removal from the translation group Gp is sent from the output terminal 20 to the language management server 30. Upon receiving this request, the group management unit 32a of the language management server 30 modifies the group data 33a to remove the output terminal 20 from the translation group Gp selected by the output terminal 20. If the removal of the output terminal 20 from the translation group Gp results in a change in the types of languages distributed by the output terminal 20 belonging to the translation group Gp, the target languages for translation are changed accordingly.
[0124] The conditions for adding the output terminal 20 to the translation group Gp may include successful authentication using a password or biometric information. Authentication is performed by comparing the authentication information, such as the password or biometric information entered into the output terminal 20, with information pre-registered in the language management server 30. This configuration enhances the security of adding the output terminal 20 to the translation group Gp.
[0125] In the embodiments described above, the speech translation system 40 is shown to perform speech recognition, translation, and speech synthesis in a continuous sequence. Alternatively, the processing results may be returned to the language management server 30 after each stage of speech recognition, translation, and speech synthesis, as described below. That is, when the first speech data S1 is sent from the language management server 30 to the speech translation system 40, the speech recognition result, i.e., the first text data T1, is sent from the speech translation system 40 to the language management server 30. Upon receiving this, the language management server 30 requests the speech translation system 40 to translate the first text data T1 into each target language. The speech translation system 40 translates the first text data T1 into the specified language to generate second text data T2 and sends the second text data T2 to the language management server 30. Subsequently, the language management server 30 requests the speech translation system 40 to perform speech synthesis based on the second text data T2. The speech translation system 40 performs speech synthesis on the second text data T2 and sends the generated second speech data S2 to the language management server 30. The language management server 30 then acquires the first text data T1, the second text data T2, and the second speech data S2, and sends these data to the terminals 10 and 20.
[0126] The method of speech translation is not limited to performing speech recognition, translation, and speech synthesis in sequence. It is sufficient as long as the translated second speech data S2 can be generated from the first speech data S1. • The translation content on the input terminal 10 may be editable.
[0127] For example, if the first text data T1 is displayed on the input terminal 10 as a result of speech translation, the input terminal 10 may be able to modify the first text data T1. In this case, the modified text data is sent to the language management server 30, and the translation management unit 32b instructs the speech translation system 40 to translate the modified text data into each target language, and delivers the translation results to the output terminal 20. With this configuration, if the user of the input terminal 10's speech is not correctly recognized by the speech recognition process, they can correct the recognition result by modifying the first text data T1 and reflect the correction in the translation result. Therefore, it becomes possible to deliver appropriate translation results. In addition, as described above, if the processing results are returned to the language management server 30 for each process of speech recognition, translation, and speech synthesis in the speech translation system 40, the first text data T1 may be sent to the input terminal 10 and modified before the translation process into the target language, and the translation and speech synthesis processes may proceed using the modified text data.
[0128] For example, if the second text data T2 for each target language is displayed on the input terminal 10 as a result of the speech translation, the input terminal 10 may be able to modify the second text data T2. In this case, the modified text data is sent to the language management server 30, and the translation management unit 32b instructs the speech translation system 40 to synthesize the modified text data into speech, and distributes the processing result to the output terminal 20. With this configuration, the user of the input terminal 10 can directly modify the translation result by modifying the second text data T2. Therefore, it becomes possible to distribute appropriate translation results. In addition, as described above, if the processing result is returned to the language management server 30 for each process of speech recognition, translation, and speech synthesis in the speech translation system 40, the second text data T2 may be sent to the input terminal 10 and modified before the speech synthesis process, and the speech synthesis process may proceed using the modified text data.
[0129] The language management server 30 may acquire the transmission status of the second audio data S2 and second text data T2 to the output terminal 20, and the output status of the second audio data S2 and second text data T2 at the output terminal 20. The acquired information may be sent to the input terminal 10, where it can be checked. With this configuration, the user of the input terminal 10 can understand whether the translation result has been conveyed to the user of the output terminal 20. Furthermore, for example, if there is a problem with the transmission of each data to the output terminal 20, the input terminal 10 may be able to instruct the retransmission of each data.
[0130] The first text data T1 and the second text data T2 received by the input terminal 10 may be stored in the input terminal 10. Similarly, the second text data T2 and image data Di received by the output terminal 20 may be stored in the output terminal 20. With this configuration, even after the validity period of the translation group Gp has expired, the contents of each of the above data can be checked at each terminal 10, 20. For example, after the sightseeing tour has ended, the user of the output terminal 20 can review the sightseeing tour by checking the image shown by the image data Di.
[0131] The translation group Gp may include an output terminal 20 to which the spoken content entered into the input terminal 10 is delivered without translation. This output terminal 20 is a terminal set to the same language as the spoken content entered into the input terminal 10, i.e., the language before translation, as the delivery language. In this case, the language management server 30 sends the first text data T1, which is the speech recognition result of the spoken content, to the output terminal 20 to which the untranslated spoken content is delivered, and a string based on the first text data T1 is displayed on the output terminal 20. This allows the user of the output terminal 20 to confirm the spoken content of the user of the input terminal 10 in text. With this configuration, for example, if a tourist tour includes participants who speak the same language as the tour guide, these participants can use the output terminal 20 to confirm the tour guide's spoken content in text, thus assisting in understanding the spoken content. Also, when tour guide trainees are learning the content of the guide, they can use the output terminal 20 to make it easier to grasp the content of the guide.
[0132] The validity period of the translation group Gp does not need to be set. On the other hand, a validity period may be set for the period during which the translation result of the utterance entered into the input terminal 10 is permitted to be delivered to the output terminal 20, and the validity period may be checked when performing voice translation. For example, before sending data such as a screen PC for voice input to the input terminal 10 when starting voice translation, when the language management server 30 receives the first voice data S1, before sending the second voice data S2 to the output terminal 20, etc., the translation management unit 32b of the language management server 30 checks whether that time is within the validity period of the translation group Gp. If the time is within the validity period, the translation management unit 32b proceeds with processing, and if the time is outside the validity period, it stops processing.
[0133] The translation history data 33b does not have to be stored in the language management server 30. Also, as a result of speech translation, the first text data T1 does not have to be sent to the input terminal 10, and the second text data T2 does not have to be sent to the output terminal 20. In other words, only the second speech data S2 may be sent to the output terminal 20. If the first text data T1 and the second text data T2 are not sent to terminals 10 and 20 or stored in the language management server 30, then the data does not have to be sent from the speech translation system 40 to the language management server 30.
[0134] The translation management system 50 does not include the speech translation system 40, and an external general-purpose translation engine may be used. That is, the language management server 30 may instruct the external translation engine to translate the first speech data S1 into each target language, and distribute the obtained translation results to the output terminal 20.
[0135] The input of information to be translated at the input terminal 10 may be in text format. That is, the first information to be translated may be text, not just speech. In this case, speech recognition processing of the first information is unnecessary. Furthermore, the output of the translation result at the output terminal 20 may be in text format only. That is, speech synthesis may not be performed, and the second text data T2 may be delivered to the output terminal 20 as the translation result.
[0136] The input terminal 10 and output terminal 20 are not limited to smartphones or tablet devices, but may also be dedicated terminals used in the multilingual translation system 100. For example, the output terminal 20 may be a device consisting of earphones and a main unit that is connected to the earphones by wire or wirelessly to perform data communication.
[0137] Furthermore, the input terminal 10 and output terminal 20 may be wearable devices, such as eyeglasses. In the third embodiment, if the output terminal 20 is an eyeglasses-type terminal that supports the display of AR content, when text and AR content are associated in the accompanying information data 33c, the AR content delivered to the output terminal 20 along with the translation result is displayed on the lenses of the eyeglasses. Therefore, the user of the output terminal 20 can view the AR content superimposed on the real environment and enjoy the AR content to the fullest.
[0138] The uses of the multilingual translation system 100 are not particularly limited. The multilingual translation system 100 may be used not only for tourist tours, but also for various types of guidance and meetings unrelated to tourism. [Explanation of symbols]
[0139] Gp...Translation group, S1...First audio data, S2...Second audio data, T1...First text data, T2...Second text data, 10...Input terminal, 20...Output terminal, 30...Language management server, 31...Communication unit, 32...Control unit, 32a...Group management unit, 32b...Translation management unit, 33...Storage unit, 33a...Group data, 33b...Translation history data, 33c...Additional information data, 40...Speech translation system, 41...Speech recognition server, 42...Translation server, 43...Speech synthesis server, 50...Translation management system, 100...Multilingual translation system.
Claims
1. A group management unit sets up a translation group to which a first terminal and a plurality of second terminals to which the translation results of the first information input into the first terminal belong, and sets the target language for translation of the first information in the translation group in accordance with the type of distribution language set for each second terminal as the language used for distribution. The system includes a translation management unit that, upon receiving the first information from the first terminal, instructs the translation of the first information into the target language of the translation group to which the first terminal belongs, and distributes the translation results for each target language to the second terminal, where that target language is set as the distribution language, The group management unit sets up a plurality of translation groups to which the same first terminal belongs, and the plurality of translation groups include a plurality of translation groups to which a plurality of target translation languages are set corresponding to the type of distribution language of the second terminal, and the combination of the set target translation languages is different from that of the plurality of translation groups. When the translation management unit receives the first information from the first terminal, it instructs the first terminal to translate the first information into the target language of the translation group selected by the first terminal, and performs a distribution process to distribute the translation result to the second terminal belonging to that translation group. The period during which the second terminal can be added to the translation group is limited to a predetermined first period, and when the first terminal requests the cessation of processing for the translation group within the first period, the group management unit prohibits the addition of the second terminal to the translation group selected by the first terminal, and prohibits the processing of translations in that translation group. A second period is set, which is the period during which the translation results are permitted to be delivered to the second terminal. The translation management unit proceeds with the delivery process when it is within the second period and stops the delivery process when it is outside the second period. Translation management system.
2. The first terminal is a terminal operated by a tour guide on a sightseeing tour, The second terminal is a terminal operated by the participants of the sightseeing tour, The group management unit sets up the translation group for each of the tourist tours as the multiple translation groups to which the same first terminal belongs. The translation management system according to claim 1.
3. The aforementioned Group Management Department, Based on receiving a request from the first terminal to generate the translation group, a new translation group to which the first terminal belongs is set up and identification information of the translation group is transmitted to the first terminal. Based on receiving a request from the second terminal, which acquired the aforementioned identification information, to add the identification information to the translation group along with the identification information, the second terminal is added to the translation group corresponding to the said identification information. The translation management system according to claim 1 or 2.
4. When the group management unit receives a request from the second terminal to add it to the translation group, it adds the second terminal to the translation group, provided that the second terminal is located within a predetermined range from the location of the first terminal. A translation management system according to any one of claims 1 to 3.
5. It includes a storage unit that stores supplementary information data that associates text with images related to that text, The translation management unit distributes the image associated with the text corresponding to the first information, along with the translation result, to the second terminal. A translation management system according to any one of claims 1 to 4.
6. The first information is input to the first terminal as audio, and the translation result, which is distributed to the second terminal, includes audio. A translation management system according to any one of claims 1 to 5.
7. The system includes a translation unit that translates the first information into the target language specified by the translation management unit. A translation management system according to any one of claims 1 to 6.