Communication support system, computer program, and communication support method
The communication support system addresses the inflexibility of existing language translation devices by using a terminal device with voice recognition activated by a two-dimensional code, enabling immediate and effective communication support for individuals with different languages.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ONE MINUTE CO LTD
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-29
AI Technical Summary
Existing communication devices for facilitating language translation between individuals with different languages are not flexible or easily accessible, making them unsuitable for urgent communication needs in settings like hospitals and construction sites.
A communication support system that uses a terminal device with voice recognition functionality activated by reading a two-dimensional code, allowing users to select languages and display translated texts side by side in a predetermined order, enabling immediate and flexible communication support.
Facilitates simple, instantaneous, and smooth communication between individuals with different languages, enhancing understanding of the context and improving usability through simultaneous display of translated texts.
Smart Images

Figure 2026123278000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a communication support system, a computer program, and a communication support method for facilitating communication between users using different languages.
Background Art
[0002] In recent years, the number of countries facing labor shortages has been increasing. Many of these countries accept foreigners as workers. For example, foreigners who migrate to other countries under such circumstances do not necessarily have sufficient proficiency in the language of the host country. Therefore, such foreigners and the residents of the host country may sometimes have difficulty communicating with each other. Such situations are seen, for example, in local governments, hospitals, construction sites, etc.
[0003] As a means for facilitating communication between users using different languages, machine translation tools are useful. For example, Patent Document 1 discloses a conversation assistance device assumed to be installed at a local government window or the like.
[0004] In the device disclosed in Patent Document 1, when a person speaks from one screen side of a display, a translated text corresponding to the uttered voice is displayed on the other screen of the display. Also, when another person speaks from the other screen side of the display, a translated text corresponding to the voice of the other person is displayed on one screen of the display.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] The device described in Patent Document 1 offers excellent visibility when reviewing conversation content because it displays the conversation content of two people on two opposing screens of a display. However, because this device requires a special display, it cannot be used easily or immediately. Furthermore, its use is limited and it cannot be said to be flexible.
[0007] For example, in hospitals and construction sites, situations may arise where it is necessary to understand what a foreigner is saying as quickly as possible, or to give appropriate instructions to a foreigner. In such cases, the device described in Patent Document 1 is not suitable.
[0008] This invention was conceived against the above background, and aims to provide a communication support system, computer program, and communication support method that can easily and immediately facilitate communication between people who use different languages. [Means for solving the problem]
[0009] The present invention relates to the following aspects.
[0010] An interface providing unit that provides a user interface with voice recognition functionality to the terminal device in response to the reading of a two-dimensional code by the terminal device, The terminal device includes a language setting unit that sets a language selected from among multiple languages as the language to be recognized by the user interface, A translation generation unit generates a translation in the default language when speech in the language setting unit is recognized via the user interface, and generates a translation in the default language when speech in the default language is recognized via the user interface. A communication support system comprising: a translation display processing unit that displays the translated texts generated by the translation generation unit on the user interface in a predetermined direction in the order in which the speech recognition was performed.
[0011] An interface provision step provides a user interface having a voice recognition function to the terminal device in response to the terminal device reading a two-dimensional code. A language setting step in which the terminal device selects a language from among multiple languages and sets it as the language to be recognized by the user interface, A translation generation step which generates a translation in the default language when speech in the setting language set in the language setting step is recognized via the user interface, and generates a translation in the setting language when speech in the default language is recognized via the user interface, A computer program that causes a computer to perform a translated text display step, which displays the translated texts generated in the translation generation step on the user interface in a predetermined direction in the order in which the speech recognition was performed.
[0012] An interface provision step provides a user interface having a voice recognition function to the terminal device in response to the terminal device reading a two-dimensional code. A language setting step in which the terminal device selects a language from among multiple languages and sets it as the language to be recognized by the user interface, A translation generation step which generates a translation in the default language when speech in the setting language set in the language setting step is recognized via the user interface, and generates a translation in the setting language when speech in the default language is recognized via the user interface, A communication support method comprising: a translation display step, which displays the translated text generated in the translation generation step on the user interface in a predetermined direction in the order in which the speech recognition was performed. [Effects of the Invention]
[0013] According to this invention, communication between people who speak different languages can be made simple, instantaneous, and smooth. In particular, since a user interface for communication support can be used by reading a two-dimensional code, it requires no effort or time to use and offers a high degree of flexibility in terms of location. Furthermore, in communication between people who speak different languages, it is often difficult to properly understand the other person's intentions without considering the context of what is being said. In contrast, with this invention, the translated text corresponding to what is being said is displayed side by side, making it easier to recognize the context and enabling effective communication with the other person. [Brief explanation of the drawing]
[0014] [Figure 1] This diagram schematically shows the configuration of a communication support system according to one embodiment. [Figure 2] Figure 1 is a block diagram showing the functional configuration of each component that makes up the communication support system. [Figure 3] Figure 1 is a flowchart illustrating the processes performed by the communication support system shown. [Figure 4] This figure shows an example of a user interface screen displayed on a terminal device that constitutes the communication support system shown in Figure 1. [Figure 5] This figure shows an example of a user interface screen displayed on a terminal device that constitutes the communication support system shown in Figure 1. [Figure 6] This figure shows an example of a user interface screen displayed on a terminal device that constitutes the communication support system shown in Figure 1. [Figure 7] This figure shows an example of a user interface screen displayed on a terminal device that constitutes the communication support system shown in Figure 1. [Figure 8] This figure shows an example of a user interface screen displayed on a terminal device that constitutes the communication support system shown in Figure 1. [Figure 9]It is a diagram showing an example of a screen of a user interface displayed on a terminal device constituting a communication support system shown in FIG. 1. [Figure 10] It is a diagram explaining the relationship between the user's set language set in the communication support system shown in FIG. 1 and the additional information displayed on the user interface. [Figure 11] It is a diagram explaining the user interface according to a modification example. [Figure 12] It is a diagram showing the hardware configuration of a computer that can constitute a first server and a second server constituting the communication support system according to an embodiment. <于
Mode for Carrying Out the Invention
[0015] Hereinafter, an embodiment will be described.
[0016] <Schematic Configuration of Communication Support System> FIG. 1 is a diagram schematically showing the configuration of a communication support system S according to an embodiment. The communication support system S according to the present embodiment shown in FIG. 1 includes a terminal device 1, a first server 10, and a second server 20.
[0017] The terminal device 1, the first server 10, and the second server 20 are connected to each other via a network 30 so as to be able to transmit and receive information. The network 30 is a communication network including, for example, the Internet, WAN (Wide Area Network), LAN (Local Area Network), etc. The terminal device 1, the first server 10, and the second server 20 are connected to the network 30 via wireless or wired and can communicate with external devices.
[0018] The terminal device 1 is a smartphone, a tablet terminal, etc. used by the user U. The first server 10 and the second server 20 are servers that provide services for supporting communication between the user U of the terminal device 1 and a conversation partner C who uses a different language from the user U.
[0019] In Figure 1, reference numeral 40 indicates a medium to which a two-dimensional code 41 is attached. The medium 40 may be a physical medium such as a poster, leaflet, or card, or it may be a display that shows the two-dimensional code 41 as electronic information. In this example, a URL (Uniform Resource Locator) is embedded in the two-dimensional code 41 as information for accessing the first server 10.
[0020] In the communication support system S according to this embodiment, when user U attempts to use the services provided by the first server 10 and the second server 20, user U first reads the two-dimensional code 41 using the terminal device 1. As a result, the terminal device 1 is connected to the first server 10, and the first server 10 provides a user interface to the terminal device 1.
[0021] The user interface provided by the first server 10 translates and presents the voice spoken by user U in the language used by conversation partner C, and also translates and presents the voice spoken by conversation partner C in the language used by user U. The second server 20 is a server that stores the translated text provided by the first server 10 and provides a function to allow users to check it afterward.
[0022] While there are no particular limitations on how the communication support system S can be used, this embodiment assumes that it will be used in typical settings such as municipal offices, hospitals, and construction sites. In this case, the media 40 will be installed or attached to appropriate locations at municipal offices, hospitals, and construction sites. At municipal offices, hospitals, and construction sites, there is a relatively large amount of interaction between foreigners and Japanese (local) people, and communication often requires urgency. As will become clear in the following explanation, the communication support system S according to this embodiment provides a service that enables simple and immediate communication support. This allows it to function beneficially in the above-mentioned usage scenarios.
[0023] In this embodiment, it is assumed that a QR code (registered trademark) will be used as the two-dimensional code 41 described above. However, the format of the two-dimensional code 41 is not particularly limited, and for example, a Data Matrix, Maxi Code, Aztec Code, or a so-called barcode may be used. Furthermore, the information embedded in the two-dimensional code 41 is not limited to a URL, but may also be an IP address, a domain name, or the like.
[0024] <Functional Configuration of Each Part of the Communication Support System> The following describes the functional configuration of each component of the communication support system S. Figure 2 is a block diagram showing the functional configuration of each component (terminal device 1, first server 10, and second server 20) of the communication support system S.
[0025] (Terminal device) First, as described above, terminal device 1 is a smartphone, tablet, or the like used by user U, and comprises a code reading unit 2, an information transmission / reception unit 3, a display control unit 4, and an instruction receiving unit 5. Terminal device 1 may also be composed of a computer having a processor such as a CPU and a storage device such as ROM or RAM. In this case, terminal device 1 can perform the processing of each of the above parts by, for example, loading a computer program stored in ROM into RAM and executing it with the processor.
[0026] The code reading unit 2 is a means for reading the two-dimensional code 41. A known application for reading two-dimensional codes may be used as the code reading unit 2. When the code reading unit 2 reads the two-dimensional code 41, the URL embedded in the two-dimensional code 41 is identified. As a result, the code reading unit 2 accesses the first server 10 via the URL and connects the terminal device 1 and the first server 10.
[0027] The information transmission / reception unit 3 performs the process of transmitting information to the first server 10 and / or the second server 20, and the process of receiving information from the first server 10 and / or the second server 20. The information transmitted by the information transmission / reception unit 3 includes requests for the provision of a user interface to the first server 10, which are generated in response to the reading of the two-dimensional code 41 by the code reading unit 2. The information received by the information transmission / reception unit 3 includes information for displaying the user interface provided by the first server 10.
[0028] The display control unit 4 controls the display of information on the terminal device 1 when the information transmission / reception unit 3 receives information from the first server 10 and / or the second server 20. For example, when the information transmission / reception unit 3 receives information from the first server 10 for displaying a user interface, the display control unit 4 displays the user interface on the terminal device 1's display based on the information received by the information transmission / reception unit 3. The display control unit 4 also updates the display state of the user interface when the first server 10 updates the information to be displayed on the user interface.
[0029] The instruction receiving unit 5 receives various instructions from user U. The user interface receives selection instructions and the like from user U through operations such as tapping. The instruction receiving unit 5 may also detect operations (instructions) from user U on the user interface via a touch panel or the like. Instructions from user U detected by the instruction receiving unit 5 on the user interface are transmitted to the first server 10, etc., via the information transmission / reception unit 3, and the first server 10, etc., performs control according to the instructions.
[0030] (Server 1) The first server 10 is a server that communicates with the terminal device 1 and includes an interface provisioning unit 101, a language setting unit 102, a translation generation unit 103, and a translated text display processing unit 104. The first server 10 may be composed of a computer having a processor such as a CPU and a storage device such as ROM and RAM. In this case, the first server 10 can perform the processing of each of the above units by, for example, loading a computer program stored in ROM into RAM and executing it with the processor.
[0031] The interface provider 101 receives a user interface provision request transmitted from the terminal device 1 in response to the reading of the two-dimensional code 41 by the terminal device 1. In response, the interface provider 101 provides the terminal device 1 with a user interface having a voice recognition function. The voice recognition function in the user interface has the function of performing voice recognition via the user interface using the microphone of the terminal device 1, converting the voice-recognized voice into data, and transmitting it from the terminal device 1 to the first server 10.
[0032] There are no particular limitations on the method for implementing the speech recognition function, but for example, the speech recognition function may be implemented in the user interface by incorporating a speech recognition API (Application Programming Interface) into the user interface. The speech recognition function of the user interface is capable of recognizing speech in many languages, but when a setting language is made by the language setting unit 102 described below, the system is configured to perform speech recognition only in the set language and a pre-set default language.
[0033] After the user interface is displayed on the terminal device 1, the language setting unit 102 receives information about the language selected by user U from among several languages on the terminal device 1, and sets the provided language as the language to be recognized for speech recognition in the user interface. As will be described in detail later, in the user interface of this embodiment, an image part is displayed to allow user U to select the language they will use. The user interface displayed on the terminal device 1 accepts the language selection from user U via the image part.
[0034] The translation generation unit 103 generates a translation in the default language when speech in the set language is recognized via the user interface after the language setting unit 102 has set the set language, and generates a translation in the set language when speech in the default language is recognized via the user interface. Here, the default language is set in advance on the first server 10 as described above, and is specifically set to the language of the conversation partner C. In this embodiment, as described above, it is assumed that the communication support system S will be used at local government offices, hospitals, construction sites, etc. In this case, the default language will be set to the language used by the staff at local government offices, hospitals, construction sites, etc. (Japanese, etc.).
[0035] The translation generation unit 103 may generate the translated text corresponding to the speech recognized by the user interface either internally within the first server 10 or externally within the first server 10. If the translated text is generated externally, the translation generation unit 103 may request a translation of the text corresponding to the speech recognized by the speech from a generation AI on an external server. In this embodiment, the translation generation unit 103 also performs the transcription of the speech recognized by the user interface into text. Specifically, the terminal device 1 transmits the speech recognized by the user interface as data to the translation generation unit 103. At this time, the translation generation unit 103 transcribes the received speech into text. This transcription may be performed internally within the first server 10 or externally within the first server 10, similar to the translation generation process. Subsequently, the translation generation unit 103 generates a translated text for the text corresponding to the speech.
[0036] The translated text display processing unit 104 processes the translation generated by the translation generation unit 103 to display on the user interface on the terminal device 1. Specifically, the translated text display processing unit 104 displays the translated text in the set language and the translated text in the default language in a predetermined direction in the order in which the speech recognition was performed. In practice, the translated text display processing unit 104 processes the display of the user interface that user U and conversation partner C see on the terminal device 1 to display the new translated text each time a new translated text is generated. As a result, user U and conversation partner C can sequentially grasp the content of what the other person is saying on the terminal device 1.
[0037] In this embodiment, the translated text display processing unit 104 displays the translated text generated by the translation generation unit 103 on the user interface, along with the text produced when the corresponding voice is recognized by speech recognition. Furthermore, the translated text display processing unit 104 displays the translated text in the default language and the translated text in the set language in different display modes. Specifically, the translated text display processing unit 104 displays the translated text in the default language and the translated text in the set language shifted in a direction orthogonal to the predetermined direction. In addition, the translated text display processing unit 104 displays the translated text in the default language and the translated text in the set language with different display modes for the background portion of each translated text. Details of these will be described later with reference to the user interface screen examples shown in Figures 4 to 9.
[0038] (Server 20) The second server 20 is a server that provides services to user U together with the first server 10, and comprises a conversation text storage unit 201, a URL issuing unit 202, a conversation text redisplay processing unit 203, a summary generation unit 204, and an additional information presentation processing unit 205. The second server 20 may be composed of a computer having a processor such as a CPU and a storage device such as ROM and RAM. In this case, the second server 20 can perform the processing of each of the above units by, for example, loading a computer program stored in ROM into RAM and executing it with the processor.
[0039] The conversation text storage unit 201 stores and stores information of the translated text displayed on the user interface as processed by the translated text display processing unit 104 of the first server 10, in accordance with instructions from the user. As will be described later, when user U finishes using the user interface displayed on terminal device 1, user U issues a termination instruction. In response to this termination instruction, the conversation text storage unit 201 stores the information of the translated text displayed on the user interface. The conversation text storage unit 201 may also store this translated text information in ROM or the like.
[0040] The URL issuing unit 202 issues a URL associated with a webpage that displays the translated text information stored in the conversation text storage unit 201.
[0041] The conversation text redisplay processing unit 203, in response to user instructions, redisplays the translated text displayed by the translated text display processing unit 104 on the terminal device 1 based on the translated text information stored in the conversation text storage unit 201.
[0042] As will be explained in more detail later, when user U gives an instruction to end the use of the user interface, an image part is then displayed on the user interface that accepts the option of whether or not to redisplay the translated text, and a URL issued by the URL issuing unit 202 is associated with this image part. The conversation text redisplay processing unit 203 detects access to the second server 20 via the URL of terminal device 1, in response to user U's operation on the image part. As a result, the conversation text redisplay processing unit 203 retrieves the translated text displayed by the translated text display processing unit 104 from the conversation text storage unit 201 and redisplays it on terminal device 1.
[0043] The summary generation unit 204 performs a process to summarize the content of the translated text displayed on the user interface by the translation display processing unit 104 of the first server 10. In this embodiment, the summary generation unit 204 summarizes the information of the translated text (conversation content) displayed on the user interface that is stored in the conversation storage unit 201. The summary generation unit 204 may generate the conversation content summary internally within the first server 10 or externally. If the summary is generated externally, a summary request may be made to a generation AI on an external server. The summary generated by the summary generation unit 204 may be stored in association with the corresponding translated text information stored in the conversation storage unit 201. The summary generated by the summary generation unit 204 may be displayed together with the entire translated text by the conversation redisplay processing unit 203. The summary generation unit 204 may also convert and store the generated summary as vectorized features, discretely represented features, distributedly represented features, etc. In this embodiment, the summary generation unit 204 stores the generated summary as a feature quantity converted using Word2Vec, BERT, Bag-of-Words, etc., for use by the additional information presentation processing unit 205 described below.
[0044] The additional information presentation processing unit 205 performs processing to present (display) additional information, such as advertisements, questionnaires, and coupons, on the user interface displayed on the terminal device 1. In this embodiment, the additional information that the additional information presentation processing unit 205 can present is stored in the second server 20, but it may also be stored on other servers. Prior to presenting the additional information, the additional information presentation processing unit 205 obtains the setting language information set by the language setting unit 102 from the first server 10 and presents the additional information in the setting language set by the language setting unit 102. That is, for example, if the setting language is set to English, the questionnaire will be displayed in English, and if the setting language is set to Persian, the questionnaire will be displayed in Persian. As described above, in this embodiment, when user U gives an instruction to end the use of the user interface, an image part is then displayed on the user interface that accepts the selection of whether or not to redisplay the translated text. In this embodiment, at this time, the additional information presentation processing unit 205 performs processing so that the additional information is displayed.
[0045] In this embodiment, as an example, additional information is associated with each language that can be set by the language setting unit 102 in the first server 10. The additional information presentation processing unit 205 then selects and presents additional information from among multiple pieces of additional information that is associated with the set language set by the language setting unit 102. That is, for users who use English, the additional information presentation processing unit 205 displays advertisements, surveys, coupons, etc. that tend to be of interest to users in English-speaking regions, and for users who use Chinese, it displays advertisements, surveys, coupons, etc. that tend to be of interest to users in Chinese-speaking regions.
[0046] In this embodiment, the additional information presentation processing unit 205 selects additional information from among multiple pieces of additional information based on the content of the translated text displayed by the translated text display processing unit 104. Specifically, in this embodiment, since the additional information is associated with the set language, the additional information presentation processing unit 205 selects additional information from among the additional information associated with the set language set by the language setting unit 102, based on the content of the translated text displayed by the translated text display processing unit 104. However, the additional information presentation processing unit 205 may ignore the association with the set language and select additional information from all multiple pieces of additional information based on the content of the translated text.
[0047] More specifically, the additional information presentation processing unit 205 selects additional information from among multiple pieces of additional information that has a high similarity to the summary result of the translated text produced by the summary generation unit 204. In this case, the additional information presentation processing unit 205 in this embodiment may compare the feature quantities obtained by vectorizing the summary generated by the summary generation unit 204 with the feature quantities of the additional information identified using the same method, and select additional information that has a high similarity to the summary result. The additional information that has a high similarity to the summary result may be the additional information with the highest similarity, or it may be one randomly selected from multiple pieces of additional information whose similarity is above a certain value. By performing such a selection process for additional information that takes summarization into consideration, it becomes possible to efficiently present effective additional information to the user.
[0048] <Example of processing> Below, an example of processing performed by the communication support system S having the above configuration will be explained with reference to Figure 3. In Figure 3, the left side shows a flowchart explaining the processing performed by terminal device 1, and the right side shows a flowchart explaining the processing performed by first server 10.
[0049] The processing of terminal device 1 begins when the code reading unit 2 is activated to read the two-dimensional code 41. When the two-dimensional code 41 is read in step S11, terminal device 1 accesses the first server 10, and terminal device 1 and the first server 10 are connected. When terminal device 1 accesses the first server 10, a request for provision of a user interface is sent from terminal device 1 to the first server 10.
[0050] The first server 10 is basically always running, and in step S21, it monitors whether or not the two-dimensional code 41 has been read by the terminal device 1. In this example, the first server 10 detects the reading of the two-dimensional code 41 by the terminal device 1 when it receives a request from the terminal device 1 to provide a user interface. If the reading of the two-dimensional code 41 is detected in step S21 (YES in step S21), the process proceeds to step S22. If the reading of the two-dimensional code 41 is not confirmed (NO in step S21), the monitoring is repeated.
[0051] When the process moves from step S21 to step S22, the interface provider 101 provides the terminal device 1 with a user interface having a voice recognition function. In response, in step S12, the terminal device 1 displays the user interface on its display. In this embodiment, the interface provider 101 provides the user interface as a web page that can be displayed by the web browser of the terminal device 1. In this case, it is advantageous for the terminal device 1 in that the user interface can be displayed immediately without any effort.
[0052] When the user interface is displayed on terminal device 1 in step S12, the user interface first prompts user U to select the language to use. Once user U has selected a language, terminal device 1 transmits information about the language selected by user U (e.g., English) to the first server 10 in step S13.
[0053] When the first server 10 receives language information selected by user U from the terminal device 1, in step S23, the first server 10 sets the language selected by user U in step S13 as the set language using the language setting unit 102, and retains it as the language of the translated text generated when the default language is recognized by speech.
[0054] After the language is selected by user U in step S13, in step S14, terminal device 1 monitors whether or not voice input has occurred to the user interface. If voice input is confirmed in step S14, in step S15, terminal device 1 sends the input voice as data to the first server 10.
[0055] In step S15, when terminal device 1 transmits voice to first server 10, in step S24, first server 10 generates a translated text corresponding to the received voice using translation generation unit 103. Here, when voice in the set language (English, etc.) is recognized via the user interface, translation generation unit 103 generates a translated text corresponding to the voice in the set language in the default language (Japanese, etc.), and when voice in the default language (Japanese, etc.) is recognized via the user interface, translation generation unit 103 generates a translated text corresponding to the voice in the default language in the set language (English, etc.).
[0056] Then, in step S25, the translation display processing unit 104 of the first server 10 performs display processing for the translated text generated in step S24. Here, the translation display processing unit 104 displays the translated text in the set language and the translated text in the default language generated by the translation generation unit 103 on the user interface displayed on the terminal device 1, arranged in a predetermined direction in the order in which the speech recognition was performed. As described above, the translation display processing unit 104 updates the display of the user interface that user U and conversation partner C see on the terminal device 1 each time a translated text is generated. As a result, on the terminal device 1 side, the translated text is displayed sequentially on the user interface in step S16.
[0057] After the translated text is displayed in step S16, terminal device 1 monitors whether user U has given an instruction to end the conversation (use of the user interface) in step S17. If no end instruction is confirmed in step S17 (NO in step S17), the process returns to step S14. In this case, the monitoring of voice input, generation of a translated text according to the voice input, and display of the translated text are repeated (steps S14-S17, S24-S25).
[0058] On the other hand, if the termination instruction is confirmed in step S17 (YES in step S17), in step S18, terminal device 1 sends the termination instruction and the translation storage instruction to the first server 10 and the second server 20 via the user interface. After that, terminal device 1's processing ends. At this point, the second server 20 receives the translation storage instruction from terminal device 1 and also receives the information of the translated text displayed on the user interface by terminal device 1, which is stored in the conversation text storage unit 201. Then, the URL issuing unit 202 of the second server 20 issues a URL associated with a web page that displays the translated text information stored in the conversation text storage unit 201. As a result, the second server 20 becomes capable of displaying the corresponding translated text information when accessed via the URL.
[0059] Furthermore, after the translation is displayed in step S25, the first server 10 monitors for the reception of a termination instruction from the terminal device 1 in step S26. If the reception of a termination instruction is not confirmed (NO in step S26), the first server 10 repeats the process of generating a translation in response to the received audio information. On the other hand, if the reception of a termination instruction is confirmed in step S26 (YES in step S26), the first server 10 returns to step S21 and monitors for the next reading of the two-dimensional code from the terminal device 1. Through the above process, user U and conversation partner C can confirm the content of the conversation with the other party through the translation displayed on the user interface.
[0060] <Example of a user interface screen> The process described in Figure 3 will be explained in more detail below, referring to an example screen of the user interface displayed on terminal device 1.
[0061] Figure 4 shows the process of reading the two-dimensional code 41 with terminal device 1, and an example of the user interface screen displayed on terminal device 1 in response to the reading of the two-dimensional code 41. The example screen shown in Figure 4 is the initial screen 50 of the user interface that is displayed in response to the reading of the two-dimensional code 41.
[0062] The initial screen 50 includes a default language display area 51, a language selection area 52, and a start button 53. The initial screen 50 is displayed by the processing of steps S11, S12 and steps S21, S22 in Figure 3.
[0063] The default language display area 51 is an area that shows the default language that has been set in advance. In the illustrated example, Japanese is set as the default language. Typically, the default language is the language used by the person who presents the two-dimensional code 41 (in this example, conversation partner C) in order to allow user U to use the services provided by the first server 10. The default language may also be set by the person who presents the two-dimensional code 41 accessing the first server 10, and may be changed as needed.
[0064] The language selection area 52 is an image component that allows user U to select the language they wish to use. As shown in Figure 5, the language selection area 52 presents multiple languages (English, Chinese, Portuguese (Brazil), etc.) for selection. User U can select from the presented languages the language (set language) to translate the voice spoken by conversation partner C.
[0065] The Start button 53 is a button for initiating the process of speech recognition of the voice spoken by user U, etc., and the process of generating a translated text corresponding to the recognized voice. When user U selects a language in the language selection area 52 and taps the Start button 53, the user interface transitions to a speech recognition-enabled state. At this time, along with the tap of the Start button 53, the language information selected by user U is sent to the first server 10, and the first server 10 sets the language selected by user U as the language (setting language) for generating the translated text. The above process using the language selection area 52 and the Start button 53 corresponds to the processes in steps S13 and S23 in Figure 3. After the language selected by user U is set as the setting language in the first server 10, buttons and explanatory texts regarding operations, etc., displayed on the user interface thereafter will be displayed in the setting language.
[0066] Figure 5 shows an example of the voice reception screen 60, which is displayed after the start button 53 in Figure 4 is tapped. When the voice reception screen 60 is displayed, voice recognition via the user interface becomes possible.
[0067] Figure 6 shows an example of the translated text display screen 70, which appears after the voice reception screen 60 of Figure 5 is displayed, when the voices spoken by user U and conversation partner C are recognized and the corresponding translated texts are displayed. In the translated text display screen 70, codes Ua to Uc indicate user-side speech frames that display the translated text corresponding to the voice spoken by user U, and codes Ca to Cc indicate conversation partner C's speech frames that display the translated text corresponding to the voice spoken by conversation partner C.
[0068] The user's speech boxes Ua~Uc and the other party's speech boxes Ca~Cc are displayed vertically from top to bottom in the order in which speech recognition was performed. In other words, on the translated text display screen 70, the translated text in the default language corresponding to the speech spoken by user U and the translated text in the set language corresponding to the speech spoken by conversation partner C are displayed in chronological order in the order in which speech recognition was performed.
[0069] Furthermore, in the example screen shown in Figure 6, on the user interface (translation display screen 70), the text (in English) obtained when the voice spoken by user U is recognized, and the translated text in the default language (Japanese) corresponding to the voice spoken by user U, are displayed together in the user's speech frames Ua to Uc. The text (in Japanese) obtained when the voice spoken by conversation partner C is recognized, and the translated text in the setting language (English) corresponding to the voice spoken by conversation partner C, are displayed together in the other party's speech frames Ca to Cc.
[0070] As described above, when the speech-recognized text and its corresponding translation are displayed together, convenience is improved because the speaker can check the speech recognition results of their own speech and recognize whether their speech and its corresponding translation are appropriate. In particular, the display order of the speech-recognized text and translation in each set is arranged vertically in the order of default language followed by the set language, regardless of whether the speech was spoken by user U or conversation partner C. In this case, user U and conversation partner C can narrow down the parts they should focus on (for example, the English text for user U), making it easier to recognize their own speech and the speech of the other party.
[0071] Furthermore, the display order of the speech-recognized text and the translated text in each of the above sets may be arranged vertically in the order of the set language followed by the default language. Also, the speech-recognized text and the corresponding translated text do not necessarily have to be displayed together; a configuration in which only the translated text is displayed may be adopted.
[0072] Furthermore, in the example screen shown in Figure 6, the translations in the default language (Japanese) corresponding to the voice spoken by user U (translations within Ua~Uc) and the translations in the set language (English) corresponding to the voice spoken by conversation partner C (translations within Ca~Cc) are displayed shifted in a direction perpendicular to the predetermined direction (vertical direction) (horizontal direction). In addition, the translations in the default language (Japanese) corresponding to the voice spoken by user U and the translations in the set language corresponding to the voice spoken by conversation partner C are displayed with different background patterns for each translation (backgrounds for Ua~Uc and Ca~Cc). Specifically, in this example, the translations in the default language (Japanese) and the translations in the set language (English) are displayed with different background colors and / or patterns (colors in this example).
[0073] As described above, when the translation in the default language and the translation in the set language are displayed in different ways, it becomes easier to visually recognize the speaker corresponding to the translation on the user interface, thereby making it easier to understand the content of the speech and improving usability. The above translation display process is performed by repeating steps S14-S17 and S24-S26 in Figure 3.
[0074] On the other hand, as shown in Figure 6, the translated text display screen 70 includes an end command button 71. User U can stop speech recognition via the user interface by tapping the end command button 71.
[0075] Figure 7 shows an example of the termination confirmation screen 80, which is displayed when the termination instruction button 71 is tapped in Figure 6. The termination confirmation screen 80 includes a modal 81 containing a continue button 82 and a termination confirmation button 83. When the continue button 82 is tapped, the display state returns from the termination confirmation screen 80 in Figure 7 to the translated text display screen 70 in Figure 6. On the other hand, when the termination confirmation button 83 is tapped, the termination process is performed.
[0076] Figure 8 shows an example of the termination screen 90, which is displayed when the termination confirmation button 83 is tapped. When user U issues a termination command, the command is sent to the second server 20, and the conversation text storage unit 201 of the second server 20 stores the information of the translated text displayed on the user interface. The URL issuing unit 202 of the second server 20 also issues a URL associated with a web page that displays the translated text information stored in the conversation text storage unit 201.
[0077] The ending screen 90 shown in Figure 8 includes a button 91 that instructs the user to redisplay the conversation text. The button 91 is associated with a URL issued by the URL issuing unit 202 of the second server 20. In Figure 9, the symbol 92 conceptually represents the URL issued by the URL issuing unit 202. When the button 91 is tapped, the translated text display screen is redisplayed, as shown by the symbol 95 in Figure 8. This redisplay process is performed by the conversation redisplay processing unit 203 of the second server 20.
[0078] Furthermore, the termination screen 90 shown in Figure 8 includes an additional information display area 300. In the additional information display area 300, additional information such as advertisements, questionnaires, and coupons is displayed by the processing of the additional information presentation processing unit 205 of the second server 20 described above. The additional information displayed in the additional information display area 300 is displayed in the setting language set in the language setting unit 102. Figure 9 shows the termination screen 90 with additional information displayed in the additional information display area 300. In the termination screen 90 shown, since the setting language set in the language setting unit 102 is English, a questionnaire explained in English is displayed. However, as shown by the arrow in the figure, if the setting language is another language (Chinese, Korean, Vietnamese, etc.), the same questionnaire will be displayed in the other language.
[0079] In this embodiment, the additional information presentation processing unit 205 selects additional information associated with the set language set by the language setting unit 102 from among multiple pieces of additional information, and in doing so, it also considers the content of the translated text displayed by the translated text display processing unit 104 when selecting the additional information. Figure 10 is a diagram illustrating the relationship between the set language and the additional information, namely advertisements and coupons, associated with the set language. Figure 10 shows an example of a data structure in which multiple pieces of additional information are associated with each set language. In Figure 10, for example, English is associated with advertisements and coupons for Japanese food, hot springs, and WIFI rentals, Chinese is associated with advertisements and coupons for sushi, and Persian is associated with advertisements and coupons for halal shops, etc. The additional information presentation processing unit 205 extracts additional information associated with the set language, and selects and presents additional information from among them based on the conversation content of the translated text.
[0080] It goes without saying that the user interface provided to terminal device 1 may display screens different from those shown in Figures 4 to 9 described above. For example, Figure 11 is a diagram illustrating a modified user interface, specifically showing a modified version of the translated text display screen 70 shown in Figure 6. In the modified version, the same reference numerals are used for components that are the same as those described in the above-described embodiment, and redundant explanations are omitted.
[0081] In the example screen shown in Figure 11, the first speaker display area 96 and the second speaker display area 97 are displayed at the top of the area where the vertically arranged translated texts are displayed. The first speaker display area 96 is the area that displays information about user U, and the second speaker display area 97 is the area that displays information about conversation partner C. The first speaker display area 96 and the second speaker display area 97 are displayed side by side. The first speaker display area 96 is displayed on the side of the display of the translation in the default language corresponding to the voice spoken by user U and the translation in the set language corresponding to the voice spoken by conversation partner C, i.e., on the Ua-Uc side, while the second speaker display area 97 is displayed on the side of the display of the translation in the set language, i.e., on the Ca-Cc side.
[0082] The first speaker display area 96 contains information 96a of the user U's name (Carlos in this example) and an image 96b of user U's face or icon. This information 96a and image 96b may be automatically acquired from the terminal device 1 by the first server 10 when user U reads the two-dimensional code 41 and the user interface is displayed on the terminal device 1, and reflected in the user interface. On the other hand, the second speaker display area 97 contains information 97a of the conversation partner C, in this example the information "Staffer," but may also contain an image of C's face or icon. The face image information of user U or conversation partner C may be acquired when the user interface activates the camera of the terminal device 1.
[0083] Furthermore, in this embodiment, the user U is asked to select a language on the initial screen 50. However, instead, in response to the user U reading the two-dimensional code 41, the first server 10 may automatically obtain information on the language set by the user U in the terminal device 1 (OS language setting, browser language, etc.), and the obtained language information may be automatically set as the setting language in the first server 10. In this case, the user U does not have to select a language, improving convenience. When the two-dimensional code 41 is read by the terminal device 1 and the terminal device 1 accesses the first server 10, access information (information including IP address, browser language, etc.) is provided from the terminal device 1 to the first server 10. The first server 10 may automatically obtain language information from the terminal device 1 based on such access information. In this case, when the first server 10 provides a user interface, the user interface described in the obtained language is provided. In other words, if the acquired language is English, for example, the voice reception screen 60 explained in English as shown in Figure 5 may be displayed directly, and if the acquired language is Chinese, the voice reception screen 60 explained in Chinese as shown in Figure 5 may be displayed directly. Also, when the start button 53 on the initial screen 50 shown in Figure 4 is tapped, the request information may include language-specific information such as the browser language. The first server 10 may automatically acquire language information from the terminal device 1 based on such request information.
[0084] The communication support system S according to this embodiment has been described above. The communication support system S comprises an interface providing unit 101 that provides a user interface with a voice recognition function to the terminal device 1 in response to the reading of a two-dimensional code 41 by the terminal device 1; a language setting unit 102 that sets a language selected from among a plurality of languages in the terminal device 1 as the language to be voice-recognized in the user interface; a translation generation unit 103 that generates a translated sentence in the default language corresponding to the voice in the setting language when voice in the setting language set by the language setting unit 102 is voice-recognized via the user interface, and generates a translated sentence in the setting language corresponding to the voice in the default language when voice in the default language is voice-recognized via the user interface; and a translated sentence display processing unit 104 that displays the translated sentences generated by the translation generation unit 103 on the user interface in a predetermined direction in the order in which the voice recognition occurred.
[0085] This technology facilitates communication between speakers of different languages in a simple, instantaneous, and smooth manner. In particular, by reading a QR code, a user interface for communication support can be accessed, requiring minimal effort and time to use, and offering a high degree of flexibility in usage location. Furthermore, in communication between speakers of different languages, it is often difficult to properly understand the other person's intentions without considering the context of what is being said. In contrast, this embodiment displays the translated text corresponding to the spoken content side by side, making it easier to recognize the context and enabling more effective communication.
[0086] Furthermore, the communication support system S according to this embodiment further includes a conversation storage unit 201 that stores the information of the translated text displayed by the translated text display processing unit 104 in response to instructions from the user U. This makes it possible to review the content of the translated text afterward and utilize it as stored data.
[0087] Furthermore, the communication support system S according to this embodiment further includes a conversation redisplay processing unit 203 that, in response to instructions from the user U, redisplays the translated text displayed by the translated text display processing unit 104 on the terminal device 1 based on the translated text information stored in the conversation text storage unit 201. This makes it easy to confirm the content of the translated text afterward.
[0088] Furthermore, the communication support system S according to this embodiment further includes a URL issuing unit 202 that issues a URL associated with a web page that displays the translated text information stored by the conversation text storage unit 201. This makes it extremely easy to check the content of the translated text afterward.
[0089] Furthermore, in this embodiment, the translated text display processing unit 104 displays the translated text generated by the translation generation unit 103 on the user interface along with the text obtained when the corresponding speech was recognized. In this case, the person who spoke can check the speech recognition result of their own speech and recognize whether their speech and the corresponding translated text are appropriate or not. This allows for smoother communication, for example, by rephrasing if the speech is inappropriate.
[0090] Furthermore, the translated text display processing unit 104 displays the translated text in the default language and the translated text in the set language in different display modes. This makes it easier to visually recognize the speaker corresponding to the translated text on the user interface, thereby making it easier to understand the content of the speech and improving usability.
[0091] In detail, the translated text display processing unit 104 displays the translated text in the default language and the translated text in the set language, shifted in a direction orthogonal to a predetermined direction. In this case, the simple difference in display mode makes it easier to visually recognize the speaker corresponding to the translated text on the user interface. This configuration is particularly beneficial in improving the visibility of the translated text when the terminal device 1 is a small device such as a smartphone and the size of the display area is limited.
[0092] Furthermore, the translated text display processing unit 104 displays the translated text in the default language and the translated text in the set language with different background display modes for each translated text. This also makes it easier to visually recognize the speaker corresponding to the translated text on the user interface through simple differences in display modes.
[0093] Furthermore, the communication support system S according to this embodiment further includes an additional information presentation processing unit 205 that presents additional information in the user interface. The additional information presentation processing unit 205 presents additional information such as advertisements, questionnaires, or coupons in the language set by the language setting unit 102. This enables effective information provision to the user U of the terminal device 1.
[0094] In this embodiment in particular, the additional information presentation processing unit 205 selects and presents additional information associated with the language setting set by the language setting unit 102 from among a plurality of additional information. More specifically, the additional information presentation processing unit 205 selects and presents additional information from among a plurality of additional information based on the content of the translated text displayed by the translated text display processing unit 104. This makes it possible to present additional information that is likely to be of interest to user U of terminal device 1, thereby enabling effective information provision.
[0095] <Example of computer hardware configuration> The following describes a computer 400 that can constitute the first server 10 and the second server 20. Figure 12 shows the hardware configuration of a computer 400 that can constitute the first server 10 and the second server 20. The computer 400 includes a processor 401, a memory device 402, a storage device 403, a communication interface 404, an operating device 405, and a display device 406, which are interconnected by a bus.
[0096] The processor 401 may be a CPU, a GPU, or a combination of a CPU and a GPU. The processor 401 executes a computer program stored in the ROM constituting the memory device 402 on the RAM constituting the memory device 402. The computer program is, for example, a program that can realize the above-described functional configurations of the first server 10 and the second server 20. For example, each part of the first server 10, such as the interface providing unit 101, can be realized by executing a computer program. The computer program may not be a single program, but may be realized by a combination of multiple programs or scripts.
[0097] The memory device 402 includes ROM and RAM. ROM is a non-temporary computer-readable storage medium that stores computer programs, data necessary for the execution of computer programs, and data generated by the execution of computer programs. These computer programs and data are read into RAM when the computer programs are executed. ROM includes, but is not limited to, hard disks, optical disks, flash memory, and magnetic tape.
[0098] The computer program may be pre-installed on the computer 400, or it may be stored on a non-temporary computer-readable storage medium such as a CD-ROM. Furthermore, the computer program may be uploaded to the internet. RAM is, for example, DRAM or SRAM, but is not limited to these.
[0099] The storage device 403 is a storage device capable of recording and retaining information. The storage device 403 may be, for example, an SSD or an HDD. The conversation storage unit 201 of the second server 20 may store the translated text information in the storage device 403.
[0100] The communication interface 404 is a circuit for the first server 10 and the second server 20 to communicate with external devices wirelessly or via wired connections. The first server 10 and the second server 20 can transmit various information to the terminal device 1 via the communication interface 404 and can also receive information from the terminal device 1. The operating device 405 is a circuit for inputting operation signals from input devices such as keyboards, mice, and touch panels to the first server 10 and the second server 20. The display device 406 is, for example, the display of the first server 10 and the second server 20. The display device 406 may be, for example, an LCD (liquid crystal display), an organic electroluminescent display, etc.
[0101] Furthermore, the computer 400 may consist of a single device or it may be configured as a system consisting of multiple interconnected computers.
[0102] The embodiments described above are merely examples of how the present invention can be implemented, and it is possible to carry out the present invention in various other forms. For example, various modifications, substitutions, omissions, or combinations thereof are possible without departing from the spirit of the present invention. Such modified, substituted, or omission forms are also included within the scope of the present invention, as well as within the scope of the invention and its equivalents as described in the claims. [Explanation of Symbols]
[0103] S...Communication support system 1…Terminal device 2... Code reading unit 3…Information transmission and reception unit 4…Display Control Unit 5…Instruction Reception Department 10…Server 1 101…Interface Provisioning Department 102...Language settings section 103…Translation Generation Department 104...Translation display processing unit 20…Second Server 201...Conversation text storage section 202…URL Issuance Department 203...Conversation text redisplay processing unit 204…Summary generator 205... Additional Information Presentation Processing Unit 30…Network 400... Computer 401… Processor 402...Memory device 403…Storage device 404…Communication Interface 405...Operating device 406…Display device
Claims
1. An interface providing unit that provides a user interface with voice recognition functionality to the terminal device in response to the reading of a two-dimensional code by the terminal device, The terminal device includes a language setting unit that sets a language selected from among multiple languages as the first language, A translation generation unit generates a translation in the first language corresponding to the audio in the second language when audio in a second language, which is a language different from the first language set in the language setting unit and is set in advance, is recognized via the user interface. The system includes a translation display processing unit that displays the translated text generated by the translation generation unit on the user interface, The translated text display processing unit is a communication support system that displays, in the user interface, a translated text in the first language generated by the translation generation unit and a text in the second language when the voice corresponding to the translated text is recognized by speech recognition, arranged vertically as a set.
2. The translation generation unit generates a translated sentence in the second language corresponding to the speech in the first language when the speech in the first language is recognized via the user interface. The communication support system according to claim 1, wherein the translated text display processing unit displays the translated text generated by the translation generation unit and the set of text when the voice corresponding to the translated text is recognized on the user interface in the order in which the voice recognition was performed.
3. The communication support system according to claim 2, wherein the translated text display processing unit displays the translated text generated by the translation generation unit and the set of text when the corresponding voice is speech-recognized, arranged vertically in the order in which the speech recognition was performed.
4. An interface provision step provides a user interface with voice recognition functionality to a terminal device in response to the reading of a two-dimensional code by the terminal device. A language setting step in which a language selected from among multiple languages is set as the first language in the terminal device, A translation generation step in which, when speech in a pre-configured second language, different from the first language set in the language setting step, is recognized via the user interface, a translation corresponding to the speech in the second language is generated in the first language. The computer is instructed to perform a translation display processing step, which involves displaying the translated text generated in the translation generation step on the user interface, A computer program that, in the translated text display processing step, displays the translated text in the first language generated in the translation generation step and the text in the second language when the speech corresponding to the translated text is recognized, arranged vertically in the user interface.
5. A communication support method that is performed using a server that communicates with a terminal device, An interface provision step is provided to the terminal device in response to the reading of a two-dimensional code by the terminal device, providing the terminal device with a user interface having a voice recognition function. A language setting step in which a language selected from among multiple languages is set as the first language in the terminal device, A translation generation step in which, when speech in a pre-configured second language, different from the first language set in the language setting step, is recognized via the user interface, a translation corresponding to the speech in the second language is generated in the first language. The server is instructed to perform a translation display processing step, which involves displaying the translated text generated in the translation generation step on the user interface. A communication support method in which, in the translated text display processing step, the translated text in the first language generated in the translation generation step and the text in the second language when the voice corresponding to the translated text is recognized are displayed together vertically in the user interface.