Speech translation system, speech translation method, and program
The speech translation system allows speakers to identify and correct translation errors in real-time, enhancing communication clarity by displaying translated and reverse-translated words with interactive buttons.
Patent Information
- Application Number
- JP2024184476
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-29
- Filing Date
- 2024-10-18
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2044-10-18
AI Technical Summary
Existing speech translation systems fail to allow speakers to correct mistranslations promptly, leading to potential miscommunication as speakers only notice errors when listeners point them out, delaying correction.
A speech translation system that displays translated and reverse-translated words on terminals, enabling speakers to identify and correct errors directly, with features like 'Respeak' and 'Cancel' buttons to manage mistranslations.
Enables immediate correction of translation errors by speakers, ensuring smooth communication without listener confusion and reducing the risk of misinterpretation.
Smart Images

Figure 2026021220000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech translation system and a speech translation method for translating speech that has been speech-recognized into another language, and also to a program used in the speech translation system. [Background technology]
[0002] Conventionally, systems capable of simultaneously interpreting conversations in different languages have been known (see, for example, Patent Document 1). In the system of Patent Document 1, for example, a Japanese speaker and an English speaker each have a terminal, and speech recognized by the Japanese speaker is translated into English on a first terminal, and the translation result is displayed on a second terminal of the English speaker. At this time, the English speaker can specify any mistranslations or unclear parts in the translation result displayed on their second terminal. If the English speaker specifies a mistranslation, the specified part of the mistranslation is displayed on the first terminal of the Japanese speaker, and the Japanese speaker can correct the mistranslation on their first terminal. It is also disclosed that at this time, the reverse translation of the speaker's own speech can be displayed on the first terminal of the Japanese speaker. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2024-031617 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in the system of Patent Document 1, the speaker (the Japanese speaker in the example above) is only given an opportunity to correct the mistranslation once the listener (the English speaker in the example above) points out that there is an error in the translation result. Therefore, even if the mistranslation is obvious to the speaker, the speaker will not notice the mistranslation unless the listener recognizes it as an error, and will not be able to correct it. Therefore, in the system of Patent Document 1, there is a risk that the speaker's intention may be conveyed to the listener in an incorrect way, and that neither the speaker nor the listener may realize this for a long time.
[0005] Therefore, the main object of the present invention is to promote smooth communication between the parties by enabling the speaker to inform the listener of an error in the speech recognition or translation results of the speaker's utterance. [Means for solving the problem]
[0006] Therefore, the inventor of the present invention has intensively studied means for solving the problems of the above-mentioned conventional inventions, and has come up with a system in which a translation of a word uttered by a speaker is displayed on a terminal of the listener, and a reverse-translated word / uttered word of the translated word is displayed on a terminal of the speaker, and if the speaker specifies that the reverse-translated word / uttered word contains an error, a message to that effect is displayed on the terminal of the listener. This effectively prevents the intention of the speaker's words from being conveyed to the listener in an incorrect way. Based on this knowledge, the inventor has conceived the idea that the problems of the conventional inventions can be solved, and has completed the present invention. Specifically, the present invention has the following configuration or steps.
[0007] A first aspect of the present invention relates to a speech translation system. The speech translation system according to the present invention includes a plurality of terminals connected via a communication line. Each terminal may be a PC, a smartphone, or a head-mounted computer for VR or AR. The speech translation system may further include a server device connected to the plurality of terminals via a communication line. The server device may be a web server connected to each terminal via the Internet, or a local server connected to each terminal via a local area network.
[0008] The speech translation system according to the present invention includes a speech recognition unit, a translation unit, a reverse translation unit, a first display control unit, and a second display control unit. These functional units may be included in any of multiple terminals or in another server device. In other words, a speech translation system can be constructed using only multiple terminals, or a speech translation system can be constructed using multiple terminals and a server device. The speech recognition unit generates spoken words by converting an utterance in a first language input to a first terminal into text. The translation unit generates translated words by translating the spoken words into a second language different from the first language. The reverse translation unit generates reverse-translated words by reverse-translating the translated words into the first language. The first display control unit displays the spoken words and / or the reverse-translated words on a display screen of the first terminal. The second display control unit displays the translated words on a display screen of the second terminal.
[0009] The "first terminal" here refers to a speaker's terminal that accepts a speech input from the speaker. The "second terminal" refers to a listener's terminal that displays the translated version of the speaker's utterance. The speech translation system according to the present invention is basically assumed to involve multiple users having a conversation, each using one terminal. Because a user may be both a speaker and a listener during a conversation, any of the multiple terminals included in the system can be both a "first terminal" and a "second terminal." In addition, in the system, at a given time, there is basically only one "first terminal," but multiple "second terminals" may be included. For example, in a situation where users 1, 2, and 3, who each use different languages, are having a conversation, and each user is using terminal A, terminal B, and terminal C, it is assumed that terminal A accepts user 1's speech input, and the translated version of that utterance is displayed on user 2's terminal B and user 3's terminal C, respectively. In this case, terminal A is the "first terminal," and terminals B and C are the "second terminals."
[0010] Here, the first display control unit displays, on the display screen of the first terminal, a screen element for specifying that all or part of the spoken word and / or backward-translated word is incorrect, along with the spoken word and / or backward-translated word. Typical examples of the screen element are icons or buttons that the user can select arbitrarily. Examples of icons and buttons include a "Respeak" button, a "Cancel" button, and a "Retranslate" button. When a screen element is selected via the first terminal, the second display control unit displays a predetermined display on the display screen of the second terminal together with the translated word. The "predetermined display" here refers to a display that is displayed on the second terminal on the listener side due to the selection of the screen element by the user (speaker) of the first terminal. Since the screen element is selected when all or part of the spoken word and / or backward-translated word is incorrect, the predetermined display on the second terminal may be any display mode that can inform the listener that all or part of the spoken word and / or backward-translated word is incorrect. For example, the predetermined display on the second terminal may simply display, using text or an icon, that the translated word displayed on the second terminal has an "error," or may display, using text or an icon, that the translated word is being "corrected" by the speaker. For example, when a "Respeak" button is selected on the first terminal, the speaker re-enters the utterance into the first terminal, so the second terminal may display that the utterance is being re-executed. Also, when a "Cancel" button is selected on the first terminal, the utterance input into the first terminal is canceled, so the second terminal may display that the translated word has been canceled. Also, when a "Retranslate" button is selected on the first terminal, the utterance input into the first terminal will be translated again, so the second terminal may display that the translated word is being corrected.
[0011] As described above, by displaying at least one of the spoken words, which are the speech recognition results, and the reverse-translated words, which are the reverse-translation results, on the first terminal on the speaker's side, the speaker can easily check whether there are any errors in the translated words displayed on the second terminal. When the speaker recognizes that there is an error in the speech recognition results or the reverse-translation results, the speaker can select a screen element on the first terminal, and a message to that effect will be displayed on the second terminal. This allows the listener to immediately confirm that there is an error in the translated words displayed on the second terminal. This allows for smooth communication between the speaker and listener without any misunderstandings.
[0012] In the speech translation system according to the present invention, when a screen element is selected via the first terminal, the second display control unit preferably erases the translated word already displayed on the display screen of the second terminal or displays a similar display for the translated word. Here, a display similar to erasing the translated word refers to a display that makes the listener perceive the translated word as hidden, such as by striking through the translated word, blacking out the translated word, or graying out the translated word. If there is an error in the speech recognition result or reverse translation result, leaving the translated word displayed on the listener's second terminal may cause the listener to misunderstand the speaker's intention. Therefore, it is preferable to erase or display a similar display so that the listener immediately understands that the translated word contains an error. Note that if a translated word is being generated while it is being deleted, the generation of the translated word may be stopped, or the generation of the translated word may continue and the translated word generated after the deletion may not be displayed on the second terminal.
[0013] In the speech translation system according to the present invention, the first display control unit preferably includes, as screen elements, an element for specifying that the speech be stopped (a cancel button) and an element for specifying that the speech be rephrased (a re-speech button). Here, when the element for specifying that the speech be stopped is selected via the first terminal, the second display control unit erases the translated word already displayed on the display screen of the second terminal or displays a similar message indicating that the translated word be deleted. On the other hand, when the element for specifying that the speech be rephrased is selected via the first terminal, the speech recognition unit generates a re-translated word by converting the re-translated speech input to the first terminal into text, the translation unit generates a re-translated word by translating the re-translated word into a language different from the re-translated speech, and the second display control unit displays the re-translated word on the display screen of the second terminal. Note that, when the element for specifying that the speech be rephrased is selected, the second display control unit may, in addition to the above, erase the translated word already displayed on the display screen of the second terminal or display a similar message indicating that the translated word be deleted. In this way, by simply providing a cancel button for deleting translated words from the second terminal and a respeak button for repeating utterances that contain errors in translated words, etc., on the speaker's first terminal, the speaker can decide how to correct their own utterances, taking into account the flow of the conversation, etc.
[0014] In the speech translation system according to the present invention, when an element (re-speak button) specifying paraphrasing is selected via the first terminal, the second display control unit preferably displays a message in two languages on the second terminal indicating that speech should not be made for a predetermined time. If the listener continues speaking while the speaker is paraphrasing, the speaker's paraphrasing may not follow the flow of the conversation. For this reason, by displaying a message on the second terminal that inhibits the listener's speech while the speaker is paraphrasing, the speaker can paraphrase without disrupting the flow of the conversation.
[0015] In the speech translation system according to the present invention, when an element specifying paraphrasing via the first terminal is selected, the second terminal may disable input of speech for a predetermined time. In this way, while the speaker is paraphrasing to the first terminal, input of the listener's speech may be forcibly disabled at the second terminal.
[0016] The speech translation system according to the present invention may be configured such that, when an element specifying paraphrasing is selected via the first terminal, a spoken word designated as erroneous is associated with a re-translated word translated by a translation unit based on the re-translated spoken word subsequently converted into text by a speech recognition unit, and the associated re-translated word is stored in a storage unit. For a spoken word designated as erroneous, the subsequent re-translated word can be used as a so-called correct label. Therefore, by associating the spoken word with the re-translated word as a dataset, the dataset can be used for retraining a translation model (trained model) for generating translated words.
[0017] In the speech translation system according to the present invention, the screen element may be used to specify that a portion of the spoken word and / or the backward-translated word contains an error. In this case, the speech translation system preferably further includes a modification word estimation unit that, when the screen element is selected via the first terminal, estimates a modification word candidate to replace the portion of the spoken word or backward-translated word specified as containing an error. In this case, the translation unit regenerates the translated word by referring to the modification word candidate estimated by the modification word estimation unit. In this case, if there are multiple modification word candidates, the translation unit may automatically identify the most likely modification word and regenerate the translated word using that modification word. Alternatively, if there are multiple modification word candidates, the first display control unit may display the modification word candidate obtained by the modification word estimation unit on the display screen of the first terminal. Then, when a specific modification word is selected from the modification word candidates via the first terminal, the translation unit may regenerate the translated word by referring to the selected specific modification word. Furthermore, the second display control unit displays the regenerated translated word on the display screen of the second terminal. In this way, by allowing the speaker to specify a part of the uttered words or back-translated words that contains an error, the speaker does not have to re-utter the entire sentence when there is an error in speech recognition or translation, which can promote smooth communication. Such correction of a part of the uttered words or back-translated words is particularly effective when the speaker's utterance is a long sentence.
[0018] The speech translation system according to the present invention may further include a server device. The server device is connected to a plurality of terminals via a communication line (including the Internet and a local area network). In this case, it is preferable that one or more of a speech recognition unit, a translation unit, and a reverse translation unit be included in the server device. Because the processing by the speech recognition unit, translation unit, and reverse translation unit places a heavy load on the server device, it is preferable that the processing be performed by a server device equipped with a processor with high processing performance.
[0019] A second aspect of the present invention relates to a computer program. The program according to the present invention causes a mobile terminal to function as a terminal equipped with the first display control unit and / or the second display control unit in the speech translation system according to the first aspect. The program may be downloadable and installable in the mobile terminal via the Internet, or may be recorded on a non-transitory recording medium such as a CD-ROM.
[0020] A third aspect of the present invention relates to a speech translation method. The speech translation method according to the present invention is executed by a system including a plurality of terminals connected via a communication line. In the speech translation method according to the present invention, the system generates a spoken word by converting an utterance in a first language input to a first terminal into text (a speech recognition step). The system also generates a translated word by translating the spoken word into a second language different from the first language (a translation step). The system also generates a backward-translated word by back-translating the translated word into the first language (a backward-translation step). The system also displays the spoken word and / or the backward-translated word on a display screen of the first terminal (a first display step). The system also displays the translated word on a display screen of a second terminal (a second display step). Here, in the first display step, the system displays, on the display screen of the first terminal, a screen element for specifying that the spoken word and / or the backward-translated word contains an error, along with the spoken word and / or the backward-translated word. Furthermore, when a screen element is selected via the first terminal, in the second display step, the system displays a predetermined display together with the translated word on the display screen of the second terminal. [Effects of the Invention]
[0021] According to the present invention, if there is an error in the speech recognition result or translation result of the speaker's utterance, the speaker can notify the listener of the error, thereby promoting smooth communication between the speakers. [Brief explanation of the drawings]
[0022] [Figure 1] FIG. 1 is an overall view showing a speech translation system according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing components of a speech translation system according to a first embodiment of the present invention. [Figure 3] FIG. 3 is a flow diagram showing a speech translation method according to the first embodiment of the present invention. [Figure 4] FIG. 4 shows an example of the display screen of each terminal. [Figure 5] FIG. 5 shows an example of the display screen of each terminal. [Figure 6] FIG. 6 shows an example of the display screen of each terminal. [Figure 7] FIG. 7 shows an example of the display screen of each terminal. [Figure 8] FIG. 8 shows an example of the display screen of each terminal. [Figure 9] FIG. 9 shows an example of the display screen of each terminal. [Figure 10] FIG. 10 shows an example of the display screen of each terminal. [Figure 11] FIG. 11 is a block diagram showing the components of a speech translation system according to the second embodiment of the present invention. [Figure 12] FIG. 12 is a flowchart showing a speech translation method according to the second embodiment of the present invention. [Figure 13] FIG. 13 shows an example of the display screen of each terminal. [Figure 14] FIG. 14 shows an example of the display screen of each terminal. [Figure 15] FIG. 15 shows an example of the display screen of each terminal. DETAILED DESCRIPTION OF THE INVENTION
[0023] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The present invention is not limited to the embodiments described below, and includes appropriate modifications of the embodiments described below within the scope obvious to those skilled in the art.
[0024] FIG. 1 shows a schematic diagram of the overall configuration of a speech translation system 100 according to a first embodiment of the present invention. As shown in FIG. 1, the speech translation system 100 includes three terminals A, B, and C, and one server device 30 connected to each terminal via a communication line. In the situation shown in FIG. 1, of the three terminals A, B, and C, terminal A used by Japanese speaker 1 is the first terminal 10 on the speaker side, while terminal B used by English speaker 2 and terminal C used by French speaker 3 are the second terminals 20 on the listener side. The speech translation system 100 is basically configured to translate an utterance input into the first terminal 10 on the speaker side into another language in real time and display the translated words on the second terminal 20 on the listener side. Specifically, in the example shown in FIG. 1, when an utterance is input in Japanese to terminal A, the utterance is translated into English and French, respectively, and the English translation is displayed on terminal B, while the French translation is displayed on terminal C. In the situation shown in Figure 1, terminal A functions as the first terminal 10 on the speaker side, but naturally, in a conversation, an utterance may be input to terminal B or terminal C. In this case, terminal B or terminal C may become the first terminal 10 on the speaker side, and terminal A may become the second terminal 20 on the listener side. There is no particular limit to the number of terminals included in the speech translation system 100, as long as there are at least two or more.
[0025] In the embodiment shown in FIG. 1, the terminals 10 and 20 are connected to the server device 30 via short-range wireless communication. Known standards such as Wi-Fi (registered trademark) and Bluetooth (registered trademark) may be used as the short-range wireless communication standard. In this manner, the terminals 10 and 20 and the server device 30 can be connected to each other via a local area network (LAN). Note that the method of connecting the terminals 10 and 20 and the server device 30 is not limited to the method shown in FIG. 1. For example, a web server may be used as the server device 30, and the terminals 10 and 20 may be connected to the web server via the Internet.
[0026] The block diagram in FIG. 2 shows an example of the components of the speech translation system 100. There are no particular limitations on the type of the terminals 10 and 20, and mobile terminals such as smartphones, tablet terminals, and laptop computers can be used. Alternatively, wearable devices such as head-mounted computers, wristwatch computers, and eyeglasses computers can also be used as the terminals 10 and 20. Furthermore, a laptop computer or desktop computer can be used as the server device 30. In particular, in this embodiment, since the processing load is likely to be concentrated on the server device 30, it is preferable to use a laptop or desktop computer as the server device 30, which can achieve higher processing performance than a smartphone or tablet terminal.
[0027] As shown in FIG. 2, the first terminal 10 includes a control unit 11, a memory unit 12, a communication unit 13, a sound collection unit 14, an operation unit 15, a display unit 16, and a sound emission unit 17. The control unit 11 of the first terminal 10 executes predetermined arithmetic processing and performs overall control of the other elements 12 to 17. The control unit 11 is configured with a known processor, which includes, for example, a CPU (Central Processing Unit) and a main memory consisting of a volatile memory such as RAM or DRAM. The memory unit 12 is a storage consisting of a nonvolatile memory such as an HDD or SSD, and stores programs for the terminal. The control unit 11 reads a predetermined program from the memory unit 12, loads it into the main memory, and executes processing in accordance with the program. The control unit 11 also allocates memory areas in the main memory required to execute each process in accordance with each program. The terminal program may be an application program dedicated to this system or a web browser program. In the former case, an application program needs to be installed in advance on the first terminal 10 and the second terminal 20. In the latter case, it is sufficient for each of the terminals 10 and 20 to have a general-purpose browser program installed, and the server device 30 provides services to each of the terminals 10 and 20 through this browser program.
[0028] When the first terminal 10 is a mobile terminal such as a smartphone or a tablet terminal, the communication unit 13 of the first terminal 10 may include a communication device for wirelessly communicating with the server device 30. As described above, in this embodiment, the first terminal 10 is connected to the server device 30 via short-range wireless communication. In this case, the communication unit 13 may employ a communication module for wireless communication using a wireless LAN system such as Wi-Fi (registered trademark). Note that when the server device 30 is a web server, the communication unit 13 may employ a communication module for wireless communication using a known mobile communication standard such as 3G (W-CDMA), 4G (LTE / LTE-Advanced), or 5G in order to communicate with the server device 30 via the Internet.
[0029] The sound collection unit 14 of the first terminal 10 is an element for inputting the user's speech as an audio signal to the control unit 11. The sound collection unit 14 may be a known microphone such as a dynamic microphone, a condenser microphone, a MEMS (Micro Electrical Mechanical Systems) microphone, a headset that can be connected wired or wirelessly, or a wired skin conduction microphone. The sound collection unit 14 may be an omnidirectional (non-directional) microphone or a directional microphone. The sound collection unit 14 converts audio into an electrical signal, amplifies the electrical signal using an amplifier circuit, converts the signal into digital information using an A / D conversion circuit, and outputs the digital information to the control unit 11. The audio signal acquired by the sound collection unit 14 is transmitted to the control unit 11 provided in the housing.
[0030] The operation unit 15 of the first terminal 10 is an element for inputting operation information by the user to the control unit 11. Known devices such as a touch panel, a mouse, a keyboard, a trackpad, a stylus pen, and a pen tablet can be used as the operation unit 15. It is also possible to configure a touch panel display by overlaying a touch panel on a display screen (display).
[0031] The display unit 16 of the first terminal 10 is an element for displaying images. As the display unit 16, a known display such as a liquid crystal display or an organic EL display can be used.
[0032] The sound emitting unit 17 of the first terminal 10 is an element for outputting an audio signal as sound to the outside. An example of the sound emitting unit 17 is a general speaker that transmits sound to the user by air vibration.
[0033] Although not shown in the drawings, the second terminal 20 may have the same configuration as the first terminal 10 described above. That is, as described above, a terminal such as a smartphone sometimes serves as the first terminal 10 on the speaker side, and sometimes serves as the second terminal 20 on the listener side. Therefore, the second terminal 20 and the first terminal 10 have basically the same configuration.
[0034] As shown in FIG. 1, the first terminal 10 has a first display control unit 11a as a functional configuration, and the second terminal 20 also has a second display control unit 21a as a functional configuration. These first display control unit 11a and second display control unit 21a are functions realized by a processor constituting the control unit 11. That is, when a terminal such as a smartphone functions as the first terminal 10 on the speaker side, the control unit 11 executes the function of the first display control unit 11a in accordance with a program or the like, while when this terminal functions as the second terminal 20 on the listener side, the control unit 11 executes the function of the second display control unit 21a in accordance with a program or the like. Each of the display control units 11a and 21a basically has the function of controlling the display unit 16 of the first terminal 10 and the second terminal 20 and adjusting the display mode on the display screen. Details of the function of each of these display control units 11a and 21a will be described later with reference to the flow chart of FIG. 3 and specific examples shown in FIGS. 4 to 10.
[0035] As shown in FIG. 2, the server device 30 includes a central control unit 31, a storage unit 32, and a communication unit 33. The central control unit 31 of the server device 30 executes predetermined arithmetic processing and is responsible for overall control of the other elements 32-33 and the terminals 10 and 20. The central control unit 31 is configured with a known processor. This processor includes, for example, a central processing unit (CPU) and / or a graphics processing unit (GPU), and a main memory consisting of a volatile memory such as a RAM and / or a DRAM. The storage unit 32 of the server device 30 is a storage consisting of a nonvolatile memory such as an HDD or SSD, and this storage unit 32 stores programs for the server device 30. The central control unit 31 reads a predetermined program from the storage unit 32, loads it into the main memory, and executes processing in accordance with the program. The central control unit 31 also allocates memory areas in the main memory required for executing each process in accordance with each program.
[0036] The storage unit 32 of the server device 30 preferably stores a trained speech recognition model 32a and a trained translation model 32b. The speech recognition model 32a is a model for recognizing a character string of spoken content from an input speech signal. The speech recognition model 32a is a speech recognition engine trained using a machine learning method such as deep learning, and is capable of recognizing a spoken character string from an input speech signal with high accuracy. A known model can be used as the speech recognition model 32a, but it is also possible to employ a hybrid model combining a Gaussian mixture model and an n-gram language model, or a deep learning model such as an end-to-end forward neural network. The translation model 32b is a model for translating an input character string from a given language into another language. The translation model 32b is a machine translation engine trained using a machine learning method such as deep learning, and is capable of translating a character string into another language with high accuracy. A known model can be used as the translation model 32b, but it is also possible to use a sequential generation model such as Transformer or Convolutional Seq2Seq, or a model that is fine-tuned from a pre-trained language model such as BERT (Bidirectional Encoder Representations from Transformers).
[0037] As shown in FIG. 2, the central control unit 31 executes the functions of the speech recognition unit 31a, the translation unit 31b, and the reverse translation unit 31c in accordance with a program. Specific operations of the functional units 31a to 31c of the central control unit 31 will be described later with reference to the flowchart of FIG. 3 and specific examples shown in FIGS. 4 to 10, and will therefore only be briefly described here. The speech recognition unit 31a performs speech recognition processing on a speech signal of an utterance received from the first terminal 10, and generates spoken words by converting the speech signal into text. At this time, the speech recognition unit 31a may use a speech recognition model 32a stored in the storage unit 32. The translation unit 31b performs translation processing on the spoken words generated by the speech recognition unit 31a, and generates translated words by translating the spoken words in the first language into the second language. At this time, the translation unit 31b may use a translation model 32b stored in the storage unit 32. Furthermore, the reverse translation unit 31c performs a reverse translation process on the translated terms generated by the translation unit 31b to generate reverse-translated terms by converting the translated terms of the second language back into the first language. At this time, the reverse translation unit 31c may use the translation model 32b stored in the storage unit 32.
[0038] Next, the processing flow of the speech translation system 100 according to this embodiment will be described with reference to the flow diagram of Fig. 3. As shown in Fig. 3, first, a user's speech is input to the first terminal 10 (step S1). The first terminal 10 converts the user's speech into a voice signal using the sound collection unit 14 and inputs the voice signal to the control unit 11. Next, the control unit 11 of the first terminal 10 transmits the voice signal related to the user's speech to the server device 30 via a communication line such as Wi-Fi using the communication unit 13 (step S2). The server device 30 receives the voice signal from the first terminal 10 using the communication unit 33 (step S3).
[0039] Next, the server device 30 performs a speech recognition process on the speech signal from the first terminal 10 using the speech recognition unit 31a of the central control unit 31 (step S4). This speech recognition process converts the speech signal into spoken words, which are text information. This speech recognition process is preferably performed in real time. That is, the speech recognition unit 31a preferably generates spoken words by sequentially converting the speech signals sequentially transmitted from the first terminal 10 to the server device 30 into text. The spoken words generated by the speech recognition unit 31a are transmitted to the first terminal 10 on the speaker's side. The first display control unit 11a of the first terminal 10 sequentially displays the spoken words received from the server device 30 on the display unit 16 of the first terminal 10 (step S5). This allows the user (speaker) of the first terminal 10 to immediately check whether the spoken words recognized by the server device 30 are correct.
[0040] Next, the server device 30 causes the translation unit 31b of the central control unit 31 to perform a translation process on the spoken words generated by the speech recognition unit 31a (step S6). This translation process converts the spoken words into translated words in another language. The translation unit 31b generates translated words in the language set by the second terminal 20 on the listener side. For example, an application program on the second terminal 20 allows the user to set a desired language, and the server device 30 receives information on the language set by the user from the second terminal 20. For example, if the user of the second terminal 20 has set English as the language, the translation unit 31b converts the spoken words into English. The translated words generated by the translation unit 31b are transmitted to the second terminal 20 on the listener side. The second display control unit 21a of the second terminal 20 causes the translated words received from the server device 30 to be displayed on the display unit 16 of the second terminal 20 (step S7). This allows the user (listener) of the second terminal 20 to read the content of the utterance made by the user (speaker) of the first terminal 10 in the language that the user has set.
[0041] Next, the server device 30 causes the reverse translation unit 31c of the central control unit 31 to perform a reverse translation process on the translated words generated by the translation unit 31b (step S8). This reverse translation process converts the translated words into reverse-translated words in the original language. The reverse translation unit 31c generates reverse-translated words in the language set by the first terminal 10 on the speaker's side. For example, an application program on the first terminal 10 allows the user to set a desired language, and the server device 30 receives information about the language set by the user from the first terminal 10. For example, if the user of the first terminal 10 has set Japanese as the language, the reverse translation unit 31c converts the translated words into Japanese. The reverse-translated words generated by the reverse translation unit 31c are transmitted to the first terminal 10 on the speaker's side. The first display control unit 11a of the first terminal 10 causes the display unit 16 of the first terminal 10 to display the reverse-translated words received from the server device 30 (step S9). This allows the user (speaker) of the first terminal 10 to read the reverse-translated words displayed on the display unit 16 and check in the language that he or she has set to use how the content of his or her utterance has been translated and conveyed to the user (listener) of the second terminal 20.
[0042] The processing of steps S1 to S9 up to this point constitutes the basic function of the speech translation system 100. FIG. 4 shows a specific example of the basic function of this speech translation system 100. FIG. 4 shows examples of the display screen of terminal A of Japanese speaker 1 and the display screen of terminal B of English speaker 2. For example, when Japanese speaker 1 speaks in Japanese to terminal A, "Please tell me what the weather will be like tomorrow," the server device 30 performs speech recognition processing, and the uttered words, "Please tell me what the weather will be like tomorrow," are displayed on terminal A as a result of the speech recognition. Meanwhile, the uttered Japanese words are translated by server device 30, and the translated words, "Please let me know the weather for tomorrow," are displayed on terminal B of English speaker 2 as a result of the translation. Furthermore, the translated English words are reverse-translated by server device 30, and the reverse-translated words, "Please let me know what the weather will be like tomorrow," are displayed on terminal A of Japanese speaker 1 as a result of the reverse translation. In this case, it is advisable to display the uttered words and the reverse-translated words together in the same speech bubble on the display screen of terminal A, for example. This allows Japanese speaker 1 to check whether his or her utterance has been correctly recognized or translated by reading the spoken words and back-translated words displayed on his or her own terminal A. In this sequence, terminal A of Japanese speaker 1 functions as the first terminal 10 on the speaker side, and terminal B of English speaker 2 functions as the second terminal 20.
[0043] Similarly, when English speaker 2 replies to terminal B in English, "Tomorrow will be sunny," server device 30 performs speech recognition processing, and terminal B displays the spoken words, "Tomorrow will be sunny," converted into text as a result of the speech recognition. Meanwhile, server device 30 performs translation processing on the English spoken words, and terminal A of Japanese speaker 1 displays the translated words, "Tomorrow will be sunny," as a result of the translation. Furthermore, server device 30 performs reverse translation processing on the translated Japanese words, and terminal B of English speaker 2 displays the reverse-translated words, "Tomorrow will be fine," as a result of the reverse translation. In this case, too, the spoken words and the reverse-translated words are displayed together in the same speech bubble on the display screen of terminal B. This allows English speaker 2 to check whether his or her utterance has been correctly recognized and translated by reading the spoken words and the reverse-translated words displayed on his or her own terminal B. In this sequence, terminal B of English speaker 2 functions as the first terminal 10 on the speaker side, and terminal A of Japanese speaker 1 functions as the second terminal 20.
[0044] Next, with reference to FIG. 5, an example in which a speaker's utterance is mistranslated will be described. First, as shown in FIG. 5(a), a speech input to terminal A of Japanese speaker 1 is speech-recognized in real time by server device 30, and the speech recognition results, that is, the speech recognition results, are sequentially displayed on terminal A by the function of first display control unit 11a. Furthermore, when Japanese speaker 1 starts speaking, server device 30 transmits information to terminal B of English speaker 2 informing that the speech is being input to terminal A during the speech recognition process. In this case, terminal B of English speaker 2 notifies English speaker 2 that Japanese speaker 1 is speaking by displaying a screen element such as "..." by the function of second display control unit 21a. This prevents the speech of English speaker 2 from overlapping with the speech of Japanese speaker 1.
[0045] Next, as shown in Figure 5(b), the translated phrase "Please let me know the weather in ASHITA" is displayed on terminal B of English speaker 2 by the function of second display control unit 21a. The part "in ASHITA" is a mistranslation, and the correct translation is "for tomorrow." Note that Japanese speaker 1 can immediately notice that "in ASHITA" is a mistranslation, but it is difficult for English speaker 2 to immediately determine whether "in ASHITA" is a mistranslation or whether Japanese speaker 1 intended to say it that way.
[0046] Next, as shown in FIG. 5(c), the function of the first display control unit 11a displays the backward-translated phrase "Please tell me the weather in Ashida." on terminal A of Japanese speaker 1. This is the result of backward translating the incorrectly translated phrase shown in FIG. 5(b). By displaying the backward-translated phrase on terminal A in this way, Japanese speaker 1 can notice that what he or she has said differs from the backward-translated phrase. In this case, Japanese speaker 1 can notice that what he or she has said has not been translated correctly and that an incorrectly translated phrase may be displayed on terminal B of English speaker 2.
[0047] Next, as shown in FIG. 5(d), if the Japanese speaker 1 notices an error in the backward-translated word, he or she operates terminal A to specify the incorrect backward-translated word. In the example of FIG. 5(d), the mistranslated portion can be specified by touching and sliding the speech bubble displaying both the erroneous spoken word and the backward-translated word horizontally. When the Japanese speaker 1 specifies the speech bubble containing the spoken word and the backward-translated word in this manner, terminal A, by the function of the first display control unit 11a, displays screen elements G1 and G2 for specifying that the spoken word or the backward-translated word has an error. In this embodiment, a respeak button G1 and a cancel button G2 are displayed as screen elements for the speech bubble containing the specified spoken word and the backward-translated word. The respeak button G1 is a screen element for canceling the translation of the second terminal 20 (terminal B) on the listener's side and then speaking again. The cancel button G2 is a screen element for simply canceling the translation of the second terminal 20 (terminal B) on the listener's side.
[0048] Next, a processing flow when the re-speak button G1 is selected will be described with reference to steps S10 to S16 in the flowchart of FIG. 3 and a specific example shown in FIG. 6. First, as shown in FIG. 3, an operation to select the re-speak button G1 is input to the first terminal 10 on the speaker side (step S10). In this case, the first terminal 10 transmits an input signal related to the re-speak button G1 to the server device 30 (step S11), and the server device 30 receives this signal (step S12). Upon receiving the input signal of the re-speak button G1, the server device 30 first executes a process to cancel the translated word (step S13). Specifically, the server device 30 transmits a cancellation signal to the second terminal 20 on the listener side to cancel the translated word corresponding to the spoken word / back-translated word specified to be re-spoken on the first terminal 10. Upon receiving the cancellation signal, the second terminal 20 deletes the translated word identified by the cancellation signal or displays a similar indication for the translated word (step S14).
[0049] A more specific explanation will be given with reference to the example of FIG. 6. As shown in FIG. 6(a), Japanese speaker 1, who is the speaker, operates terminal A, designates a speech bubble containing the spoken word "Please tell me the weather tomorrow" and the back-translated word "Please tell me the weather in Ashida," and selects the re-speak button G1. At this time, the translated word "Please let me know the weather in ASHITA" is already displayed on terminal B of English speaker 2, who is the listener, but this translated word is recognized as an incorrect translation by Japanese speaker 1. When the re-speak button G1 is selected on terminal A, a line is drawn through the translated word already displayed on terminal B of English speaker 2, who is the listener, as shown in FIG. 6(b). The crossed-out translated word corresponds to the spoken word and back-translated word in the speech bubble for which re-speak button G1 was selected on terminal A. This strike-through is an example of a display similar to erasure. This allows English speaker 2, who is the listener, to realize that the translated word was an incorrect translation. 6, the translated words are simply crossed out, and the mistranslated words themselves remain as a record of the conversation on the display screen of terminal B. It is possible to erase the mistranslated words from the listener's terminal B, but in that case the translated words that were once displayed on the display screen of terminal B will suddenly disappear, which may confuse the listener, English speaker 2. For this reason, to avoid confusing the listener, it is preferable to cross out or gray out the translated words that are deemed to be mistranslated and leave them as a record of the conversation on the display screen of terminal B.
[0050] Returning to the flow diagram of FIG. 3, the server device 30 performs the re-speech process after the translated word cancellation process (step S15). In the re-speech process, the server device 30 returns to step S1 and receives a speech input again from the first terminal 10 on the speaker side. Thereafter, steps S2 to S9 are repeated again. As a result, the spoken word, which is the result of speech recognition, is displayed again on the first terminal 10 on the speaker side, the translated word of the spoken word is displayed again on the second terminal 20 on the listener side, and the backward-translated word of the translated word is displayed again on the first terminal 10 on the speaker side. The speaker reads the spoken word and backward-translated word displayed again on the first terminal 10 and confirms whether their utterance has been correctly recognized and translated. If there is an error in the re-displayed spoken word or backward-translated word, the speaker can again input an operation to re-speak into the first terminal 10. In this case, steps S10 to S16 are repeated again.
[0051] Furthermore, in the re-speech processing (step S15), the server device 30 transmits a signal to the second terminal 20 on the listener side to prohibit input of an utterance until the speaker has completed re-speech. Upon receiving this signal prohibiting input of an utterance, the second terminal 20 prohibits input of an utterance from the user for a predetermined time (step S16). For example, the second terminal 20 may turn off the sound collection unit 14 (microphone) during the speech prohibition period to prevent input of an utterance, or may not transmit an audio signal to the server device 30 even if an utterance is input to the sound collection unit 14. In this way, by prohibiting the listener from speaking while the speaker is re-speaking, it is possible to avoid a situation in which the listener misunderstands the intention of the speaker's utterance due to, for example, a mistranslation and the conversation continues.
[0052] In the specific example shown in FIG. 6(b), the screen of terminal B of English speaker 2, who is the listener, displays a message indicating that speaking in the language used by terminal B (i.e., English) is prohibited during the Japanese speaker 1's respeaking, by the function of second display control unit 21a. That is, in the example of FIG. 6(b), the message "No speaking allowed" is displayed during the Japanese speaker 1's respeaking. While this message is displayed, terminal B turns off the function of sound collection unit 14, forcibly preventing the speech of English speaker 2 from being input to terminal B. During this time, Japanese speaker 1, who is the speaker, respeaks to first terminal 10. For example, in the example shown in FIG. 6(b), Japanese speaker 1 rephrases his respeaking as "Please tell me the weather tomorrow." After that, as shown in Figure 6(c), when Japanese speaker 1 finishes re-speaking, the Japanese spoken words "Please tell me the weather tomorrow" are displayed on terminal A, the English translation "Please let me know the weather for tomorrow" is displayed on terminal B, and the Japanese back-translation "Please tell me the weather tomorrow" is displayed on terminal A. In this example, since the spoken words and the back-translation words have roughly the same meaning, Japanese speaker 1 can understand that the correct translation is being displayed on terminal B.
[0053] Furthermore, when the re-speech button G1 is selected and the speaker paraphrases the utterance, the server device 30 may collect the utterance designated as having an error in translation as an error example. Specifically, the utterance designated as having an error in translation (error example) may be associated with a re-translated word (correct answer label) translated by the translation unit 31b based on the re-speech word subsequently converted into text by the speech recognition unit 31a, and stored in the storage unit 32. This allows information on what kind of mistranslation (error) occurred relative to the original correct answer to be accumulated. For example, in the example shown in FIG. 6, the utterance "Please tell me the weather for tomorrow" (error example) initially recognized by speech recognition of Japanese speaker 1 is translated into "Please let me know the weather for tomorrow," and the translated word "Please let me know the weather for tomorrow" into which the re-speech of Japanese speaker 1 is finally translated is determined by Japanese speaker 1 to be error-free and can therefore be used as a correct answer label for machine learning. Therefore, the data set of the initially spoken words and the final translated words is stored in the memory unit 32 of the server device 30, and once a sufficient number of such data sets have been accumulated, the translation model 32b stored in the memory unit 32 can be retrained, thereby improving the accuracy of machine translation.
[0054] While the illustrated example shows a case where the translated word output from the translation model 32b contains an error, it is also possible that the spoken word output by the speech recognition model 32a contains an error, causing the re-speech button G1 to be selected. In this case, the spoken word designated as having an error may be associated with the re-spoken word subsequently converted into text by the speech recognition unit 31a, and stored in the storage unit 32. In this way, the server device 30 may collect the spoken word designated as having an error in speech recognition as an error example.
[0055] Furthermore, as described above, when the repetition button G1 is selected and the speaker paraphrases the utterance, buttons (screen elements) labeled "Repetition due to translation error" and "Repetition due to speech recognition error" may be displayed on the first terminal 10. This makes it possible to distinguish whether the cause of the speaker's repetition is a translation error or a speech recognition error. When distinguishing the cause of the error, it is preferable to collect and store four data sets for machine learning: the speech recognition result of the initial utterance, its translation result, the speech recognition result of the repetition, and its translation result. In other words, if the speech recognition result of the initial utterance and the speech recognition result of the repetition are the same, it can be assumed to be a translation error, and if the speech recognition result of the initial utterance and the speech recognition result of the repetition are different, it can be assumed to be a speech recognition error.
[0056] FIG. 7 shows an example of a screen display when the Japanese speaker 1 (speaker) selects the cancel button G2 displayed on terminal A (first terminal 10). As shown in FIG. 7(a), if the Japanese speaker 1 notices that a back-translated word is incorrect, he or she selects the cancel button G2 after specifying the back-translated word. Then, as shown in FIG. 7(b), on terminal B (second terminal 20) of the English speaker 2 (listener), a strikethrough is drawn, for example, on the translated word corresponding to the back-translated word for which the cancel button G2 was selected. The processing when the cancel button G2 is selected is the same as steps S10 to S14 shown in the flowchart of FIG. 3. However, when the cancel button G2 is selected, unlike when the re-speech button G1 is selected, the re-speech processing (steps S15 to S16) is not performed. Therefore, as shown in FIG. 7(b), the English speaker 2 can speak without waiting for the Japanese speaker 1 to re-speak. In this way, if the Japanese speaker 1 notices an error in the recognition or translation of the utterance, and needs to rephrase the utterance, he or she can select the respeak button G1, or if there is no need to rephrase the utterance, he or she can select the cancel button G2. In this way, having two types of buttons available allows the conversation to proceed smoothly.
[0057] In the examples shown in FIGS. 6 and 7, the Japanese speaker 1 (speaker) selects the respeak button G1 or the cancel button G2 after the spoken words and retranslated words have been displayed on the speaker-side terminal A (first terminal 10). However, as shown in the example of FIG. 5(a), if the Japanese speaker 1 discovers an error in the speech recognition before all of the spoken words in the speech recognition result have been displayed on terminal A, the Japanese speaker 1 can select the respeak button G1 or the cancel button G2 at that time. In this case, it is preferable that the server device 30 halts the speech recognition process, translation process, and retranslation process when either button G1 or G2 is selected. However, the server device 30 can also continue the speech recognition process, translation process, and retranslation process even when either button G1 or G2 is selected. In this case, the spoken words, translated words, and retranslated words are displayed on each of terminals A and B. However, as in the example described above, at least the translated words displayed on terminal B on the listener side are erased or displayed in a similar manner.
[0058] FIG. 8 shows a function for inputting utterances as text. For example, a software keyboard is displayed on the display unit 16 of terminal A (first terminal 10) by the function of the first display control unit 11a. The Japanese speaker 1 (speaker) operates the operation unit 15, such as a touch panel, to input the utterances he or she wishes to convey to the English speaker 2 (listener) as text via the software keyboard. The utterances thus input as text are translated into English by the server device 30, and the translated words are displayed on terminal B (second terminal 20) of the English speaker 2 (listener). The translated words are also reverse-translated into Japanese by the server device 30, and the reverse-translated words are displayed on terminal A of the Japanese speaker 1. If there is an error in the reverse translation, the Japanese speaker 1 can select the respeak button G1 or the cancel button G2 displayed on terminal A using the same procedure as described above.
[0059] 8, instead of inputting text via a software keyboard, a plurality of fixed phrases may be prepared in advance, and the Japanese speaker 1 (speaker) may select any one of the fixed phrases. The fixed phrase selected by the Japanese speaker 1 is translated in the server device 30, and the translated phrase is displayed on the terminal B of the English speaker 2 (listener). For fixed phrases, correct translations may be stored in advance in the storage unit 32 of the server device 30. In this case, when a fixed phrase is selected, the server device 30 only needs to read out the translation corresponding to the fixed phrase from the storage unit 32, and therefore the translation process and reverse translation process using the translation model 32b may be omitted.
[0060] FIG. 9 shows a function that allows approval of the content of an utterance. Specifically, on the speaker's terminal A (first terminal 10), a screen element for approving the spoken word and the backward-translated word is displayed near the display location of the spoken word and the backward-translated word by the function of the first display control unit 11a. This screen element is referred to as the approval button G3. The Japanese speaker 1, who is the speaker, reads the spoken word and the backward-translated word displayed on his / her own terminal A and selects the approval button G3 if he / she confirms that the content matches what he / she uttered. A selection signal for the approval button G3 is transmitted from terminal A to the server device 30, and further transmitted from the server device 30 to terminal B (second terminal 20) on the listener's side. Terminal B then displays a screen element indicating that the translated word corresponding to the spoken word and the backward-translated word for which the approval button G3 was selected, near the translated word, is displayed by the function of the second display control unit 21a. This screen element is referred to as the approval icon G4. In this way, when the Japanese speaker 1 (speaker) approves that the spoken word and the back-translated word are correct, this is communicated to the English speaker 2 (listener). This enables the English speaker 2 to understand that the translated word marked with the approval icon G4 is the correct translation and will not be respoken or retracted later. In this way, providing an approval function for translated words and back-translated words allows the conversation to proceed smoothly.
[0061] Figure 10 shows an example of a case where the system is used by three people who use different languages. Specifically, Figure 10 assumes a case where Japanese speaker 1 uses terminal A, English speaker 2 uses terminal B, and French speaker 3 uses terminal C, and they are conversing via the system. In this way, even when there are three or more people in a conversation and three or more languages are used, the system basically operates in the same way as in the example described above.
[0062] Specifically, when an utterance by Japanese speaker 1, who is the speaker, is input to terminal A (first terminal 10), speech recognition processing is performed in server device 30, and the spoken words converted into text are displayed on terminal A. The uttered words are also translated in server device 30, and the translated English words are displayed on terminal B (second terminal 20) of English speaker 2, and the translated French words are displayed on terminal C (second terminal 20) of French speaker 3. Furthermore, the translated English words and the translated French words are each subjected to reverse translation processing in server device 30, and the reverse-translated words translated from English back to Japanese and the reverse-translated words translated from French back to Japanese are each displayed on terminal A. At this time, the spoken Japanese words, the reverse-translated words translated from English back to Japanese, and the reverse-translated words translated from French back to Japanese are displayed together in the same speech bubble. This allows Japanese speaker 1, who is the speaker, to check whether the content of his or her utterance has been correctly translated into English and French by reading each reverse-translated word.
[0063] Furthermore, if the Japanese speaker 1 notices an error in the spoken Japanese word, the back-translated English word, or the back-translated French word, he or she can select the speech bubble containing the error and select the re-speak button G1 or the cancel button G2 (see, for example, Figure 5(d)). Regardless of whether the re-speak button G1 or the cancel button G2 is selected, the translated word displayed on the English speaker 2's terminal B and the translation displayed on the French speaker 3's terminal C will be erased or a similar display will be displayed. Note that a situation can also be assumed in which the multiple back-translated words displayed on terminal A include both correct and incorrect back-translated words. Even in this case, if the speech bubble containing both the spoken word and multiple back-translated words is selected and the re-speak button G1 or the cancel button G2 is selected, it is preferable to erase the translated words on all terminals or a similar display will be displayed. It is possible to select only the incorrectly translated words among the multiple backward-translated words displayed on terminal A and select the respeak button G1 or the cancel button G2, but in that case, the flow of conversation involving three or more people becomes complicated, especially when respeaking. For this reason, if even one incorrectly translated word is included among the multiple backward-translated words displayed on terminal A on the speaker's side, it is preferable to uniformly erase all translated words (including the correct translated word) displayed on all terminals B and C on the listener's side.
[0064] 10, Japanese speaker 1, English speaker 2, and French speaker 3 are conversing via this system, but the conversation may also include speakers of other languages, or two or more speakers of the same language. In this way, this system can handle conversations between multiple people in multiple languages.
[0065] Next, a speech translation system 100 according to a second embodiment of the present invention will be described with reference to Figures 11 to 15. In the second embodiment, the same elements and functions as those in the first embodiment will not be described, and differences from the first embodiment will be mainly described. The first and second embodiments are not alternatives, and elements and functions of both embodiments can be used together.
[0066] FIG. 11 shows a block diagram of a speech translation system 100 according to the second embodiment. As shown in FIG. 11, the components of the second embodiment are basically the same as those of the first embodiment shown in FIG. 2. However, the second embodiment further includes a modified word estimation unit 31d and a natural language processing model 32c in addition to the components of the first embodiment. The function of the modified word estimation unit 31d, like the speech recognition unit 31a, translation unit 31b, and reverse translation unit 31c, is executed by the central control unit 31 of the server device 30 according to a program. The natural language processing model 32c, like the speech recognition model 32a and translation model 32b, is stored in the storage unit 32 of the server device 30.
[0067] When a speaker specifies via the first terminal that a part of a spoken word or a backward-translated word contains an error, the modification word estimation unit 31d estimates a modification word candidate to replace the specified word. The number of post-correction candidates may be one or more. The modification word estimation unit 31d preferably generates a modification word candidate to replace the specified word using a natural language processing model 32c stored in the storage unit. A trained, publicly known model such as BERT or LLM (Large Language Model) can be used as the natural language processing model 32c. BERT and LLM can process and generate sentences while understanding the context. This embodiment assumes a case where a part of a sentence containing spoken words or a sentence containing backward-translated words contains an error due to a speech recognition error or a translation error. Therefore, when a speaker specifies an error in a phrase in a sentence, the modification word estimation unit 31d can mask the specified phrase and estimate a candidate modification word to replace the specified phrase from the context before and after the masked specified phrase using the natural language processing model 32c. In other words, the estimation process of a candidate modification word can be formulated as a mask content estimation problem. Note that, while BERT only takes into consideration language model aspects when estimating a modification word, LLM also takes into comprehensive consideration aspects such as whether the modification word sounds similar to the specified phrase when estimating a modification word.
[0068] FIG. 12 shows a process that can be executed in the second embodiment. Among the processing steps shown in FIG. 12, steps S1 to S9, which are basic functions from speech recognition of an utterance to displaying a reverse-translated word, are the same as the processing steps shown in FIG. 3. Therefore, detailed description of steps S1 to S9 will be omitted. In the example shown in FIG. 13, as shown in FIG. 13(a), Japanese speaker 1 speaks in Japanese to terminal A, "Will tomorrow's weather be rainy?". As a result of speech recognition processing by server device 30, the spoken words converted into text, "Will tomorrow's weather be candy?", are displayed on terminal A. At this time, an error occurs in the speech recognition, and the Japanese word "ame" (rain) is recognized as the homonym "ame" (candy). Furthermore, as shown in FIG. 13(b), the erroneously recognized spoken words in Japanese are translated by server device 30, and the translated words, "Will tomorrow's weather be candy?", are displayed on terminal B of English speaker 2 as a result of the translation. 13(c), the translated English word is subjected to a reverse translation process by the server device 30, and the reverse-translated word "Ashita no tenki wa ame desuka?" (Will the weather be like candy tomorrow?) is displayed on terminal A of the Japanese speaker 1 as a result of the reverse translation. By checking the spoken word and the reverse-translated word displayed on terminal A, the Japanese speaker 1 can recognize that an error occurred during speech recognition.
[0069] Among the processing steps shown in FIG. 12, steps S17 to S26 relate to a function for partially correcting errors in the speech recognition process or translation process, as described above. When a speaker finds a misrecognition or mistranslation, he or she inputs an operation to specify the misrecognized part or the mistranslated part into the first terminal 10 (step S17). In the example shown in FIG. 13(d), as described above, the Japanese utterance "Will the weather be rainy tomorrow?" is misrecognized as "Will the weather be candy tomorrow?" In this case, the Japanese speaker 1 (speaker) operates terminal A (first terminal 10) to specify the misrecognized phrase from the speech-recognized utterance. For example, as shown in FIG. 13(d), when the misrecognized phrase "candy" is specified by dragging, a selection bar including an error specification button G5 labeled "misrecognition" is displayed nearby. The Japanese speaker selects the error specification button G1 from this selection bar. In this way, the misrecognized phrase is specified from the speech-recognized utterance.
[0070] Next, the first terminal 10 transmits an input signal for specifying the misrecognition or mistranslation to the server device 30 (step S18), and the server device 30 receives the signal (step S19). Upon receiving the input signal, the server device 30 performs a display process to notify the second terminal 20 that a misrecognition or mistranslation has been specified (step S20). In this embodiment, when a misrecognition or mistranslation is detected, a process for correcting the erroneous phrase is performed. Therefore, the second display control unit 21a of the second terminal 20, under the control of the server device 30, displays on the display unit of the second terminal 20 a message indicating that the phrase is being corrected (step S21). For example, in the example of FIG. 14(a), the word or image "Correcting" is displayed on the display screen of terminal B (second terminal 20) in the language (English) used by the English speaker 2 (listener) who operates terminal B. This allows the English speaker 2 to immediately recognize that there is an error somewhere in the translation currently displayed on terminal B and that the translation is being corrected.
[0071] Next, the modification word estimation unit 31d of the server device 30 uses the natural language processing model 32c to estimate a modification word candidate to replace the phrase designated by the speaker as being misrecognized or mistranslated (step S22). The natural language processing model 32c can generate one or more modification word candidates to replace the designated phrase from the context before and after the designated phrase. In the example shown in FIG. 14(a), the phrase "ame" (candy) in the utterance "will the weather be like candy tomorrow?" is designated as having been misrecognized. The modification word estimation unit 31d masks the phrase "ame" (candy) or designates that the phrase "ame" (candy) is incorrect, and inputs the utterance to the natural language processing model 32c. The natural language processing model 32c then outputs modification word candidates such as "ame" (rain), "sanreru" (clear), and "do" (how) as a replacement for "ame" (candy).
[0072] Next, the server device 30 transmits the candidate correction words obtained by the correction word estimation unit 31d to the speaker first terminal 10. The first display control unit 11a of the first terminal 10 displays the candidate correction words on the display screen of the first terminal 10 (step S23). In the example shown in FIG. 14(b), candidate correction words such as "rain," "clear," and "how" are displayed on terminal A (first terminal 10) of the Japanese speaker (speaker). In this example, these candidate correction words are displayed in a list format in a drop-down menu G6, which is a screen element.
[0073] Next, the first terminal 10 accepts an operation to select one correct modifying word from among multiple modifying word candidates (step S24). The first terminal 10 transmits information about the one modifying word selected by the speaker to the server device 30. When the server device 30 receives the information about the one modifying word from the first terminal 10, the translation unit 31b regenerates a translated sentence referring to the modified word (step S25). In the example shown in FIG. 14(b), the Japanese speaker 1 (speaker) selected the modifying word "rain" from among the modifying word candidates. As a result, the server device 30, referring to the modifying word "rain," identifies the correct utterance as "Will tomorrow's weather be rainy?". The translation unit 31b of the server device 30 then retranslates this correct utterance to generate the translated word "Will tomorrow's weather be rainy?" The translated term thus regenerated is transmitted from the server device 30 to the second terminal 20 (terminal B), and is displayed on the display screen of the second terminal 20 by the second display control unit 21a (step S26).
[0074] In this way, in this embodiment, if a speaker finds an error in a part of a speech that has been recognized, the speaker can select and correct the part where the error is located, thereby saving the speaker the trouble of having to repeat the entire speech when a recognition error occurs.
[0075] 13 and 14 show the case where an error occurs in speech recognition, but a similar correction can be made when an error occurs in translation. FIG. 15 shows an example of a case where an error occurs in translation. For example, the spoken word obtained by speech recognition of an utterance by Japanese speaker 1 (speaker) is "Ashita no tenki wa ame desuka?" (Will tomorrow's weather be ame desuka?), and the translated word is "Will tomorrow's weather be candy?" In this case, the back-translated word is "Ashita no tenki wa ame desuka?" (Will tomorrow's weather be ame desuka?). In this example, it cannot be said that there is an error in the spoken word resulting from speech recognition, but the translated word is completely different from the intention of the Japanese speaker 1's utterance. For this reason, the Japanese speaker 1 can discover that an error has occurred in the translation process by checking the back-translated word. In this case, the Japanese speaker 1 selects the erroneous word "ame" (candy) from the back-translated words displayed on terminal A (first terminal 10) and selects the error specification button G5. Then, the word "ame" (rain) is displayed in a drop-down menu G6 on the display screen of terminal A as a candidate for a correction word to replace the erroneous word. When Japanese speaker 1 selects the corrected word "ame" (rain) from drop-down menu G6 displayed on terminal A, the selected word is sent to server device 30. Server device 30 performs a retranslation by referring to the corrected word selected by Japanese speaker 1, and displays the resulting translated word on terminal B (second terminal 20) of English speaker 2 (listener). In this way, even if there is an error in the translation process, the speaker can partially select and correct the incorrect part.
[0076] Here, an example has been described in which a speaker (user) can select one correct modifying word from multiple candidate modifying words. However, the present invention is not limited to this. The translation unit 31b of the server device 30 may automatically identify the most likely modifying word from the multiple candidate modifying words obtained by the modifying word estimation unit 31d and regenerate a translated word based on the identified modifying word. For example, such automatic modification word identification may be performed when a user sets automatic modification word selection or when the likelihood of the most likely modifying word is higher than a certain threshold. This allows the user to automate the process from selecting a modifying word to displaying a retranslated word simply by selecting an incorrect word.
[0077] In the above description of the present invention, the embodiments of the present invention have been described with reference to the drawings in order to express the contents of the present invention. However, the present invention is not limited to the above embodiments, and includes modifications and improvements that are obvious to those skilled in the art based on the matters described in the present specification.
[0078] For example, in the examples shown in FIGS. 1 to 3, 11, and 12, a client-server system in which the server device 30 executes the speech recognition process, translation process, and re-translation process has been described. However, the present invention is not limited to a system in which these processes are executed by the server device 30. For example, the speech recognition process may be executed by the first terminal 10, and the translation process and re-translation process may be executed by the server device 30. Alternatively, a peer-to-peer system in which the speech recognition process, translation process, and re-translation process are executed by the terminals alone may be implemented without providing such a server device 30. In this case, for example, the speaker-type first terminal 10 may execute the speech recognition process, the listener-type second terminal 20 may execute the translation process, and the speaker-type first terminal 10 may execute the re-translation process. Alternatively, the speech recognition process, translation process, and re-translation process may all be executed by either the speaker-type first terminal 10 or the listener-type second terminal 20. In this way, whether the speech recognition process, translation process, and retranslation process is executed by the server device 30 or the terminals 10 and 20 can be changed as appropriate depending on the processing capabilities of each device. [Explanation of symbols]
[0079] 10...first terminal 11...control unit 11a...first display control unit 12...storage unit 13...Communication unit 14...Sound collection unit 15...Operation section 16...Display section 17...sound emitting unit 20...second terminal 21a... Second display control unit 30... Server device 31...Central control unit 31a...Speech recognition unit 31b...Translation section 31c...Reverse translation section 31d...Correction word estimation section 32...Storage section 32a...Speech recognition model 32b...Translation model 32c...Natural Language Processing Model 33...Communication Department G1...Respeak button G2...Cancel button G3...Approval button G4...Approval icon G5: Error specification button G6: Drop-down menu 100...Speech translation system
Claims
1. A speech translation system comprising a plurality of terminals connected via a communication line, a speech recognition unit that converts an utterance in a first language input to the first terminal into text and generates a spoken word; a translation unit that generates a translation by translating the utterance into a second language different from the first language; a reverse translation unit that generates a reverse-translated word by reverse-translating the translation word into the first language; a first display control unit that displays the spoken word and / or the backward-translated word on a display screen of the first terminal; a second display control unit that displays the translated words on a display screen of a second terminal; the first display control unit displays, on a display screen of the first terminal, a screen element for specifying that all or part of the spoken word and / or the backward-translated word contains an error, together with the spoken word and / or the backward-translated word; When the screen element is selected via the first terminal, the second display control unit performs a predetermined display together with the translated word on the display screen of the second terminal. Speech translation system.
2. When the screen element is selected via the first terminal, The second display control unit erases the translated word already displayed on the display screen of the second terminal or displays a similar message to erase the translated word. The speech translation system according to claim 1 .
3. the first display control unit includes, as the screen elements, an element for specifying that utterance is to be stopped and an element for specifying that utterance is to be paraphrased; When an element specifying that the speech is to be stopped via the first terminal is selected, The second display control unit erase the translated word already displayed on the display screen of the second terminal or display a similar message indicating that the translated word has been erased; When an element specifying that the paraphrase is to be performed via the first terminal is selected, the speech recognition unit generates a recurrence speech by converting the recurrence speech input to the first terminal into text; the translation unit generates a re-translated word by translating the recurred speech into a language different from the recurred speech; The second display control unit displays the re-translated word on a display screen of the second terminal. The speech translation system according to claim 1 .
4. the first display control unit includes, as the screen element, an element for specifying that a rephrasing be performed; When an element specifying that the paraphrase is to be performed via the first terminal is selected, The second display control unit displays, in the two languages, a message that speech should not be made to the second terminal for a predetermined time.
4. A speech translation system according to claim 1 or 3.
5. When an element specifying that the paraphrase is to be performed via the first terminal is selected, The second terminal disables input of speech for a predetermined time. The speech translation system according to claim 4.
6. the first display control unit includes, as the screen element, an element for specifying that a rephrasing be performed; The speech translation system further includes a storage unit that, when an element specifying the rephrasing is selected via the first terminal, stores the uttered word specified as having an error in association with a re-translated word translated by the translation unit based on the re-translated uttered word subsequently converted into text by the speech recognition unit. The speech translation system according to claim 1 .
7. the screen element is for specifying that a part of the spoken word and / or the back-translated word has an error; the speech translation system further includes a correction word estimation unit that estimates a candidate correction word to replace the part designated as erroneous when the screen element is selected via the first terminal, the translation unit regenerates the translation word by referring to the candidate modification words; The second display control unit displays the regenerated translated word on a display screen of the second terminal. The speech translation system according to claim 1 .
8. the speech translation system further includes a server device connected to the plurality of terminals via a communication line; One or more of the speech recognition unit, the translation unit, and the reverse translation unit are included in the server device. The speech translation system according to claim 1 .
9. A program for causing a mobile terminal to function as a terminal equipped with the first display control unit and / or the second display control unit in the speech translation system according to claim 1.
10. A speech translation method executed by a system including a plurality of terminals connected via a communication line, comprising: a speech recognition unit step of converting an utterance in a first language input to a first terminal into text and generating a spoken word; a translation step of generating a translation of the spoken word into a second language different from the first language; a back-translation step of generating a back-translated word by back-translating the translation word into the first language; a first display step of displaying the spoken word and / or the backward-translated word on a display screen of the first terminal; a second display step of displaying the translated words on a display screen of a second terminal; the first display step displays, on a display screen of the first terminal, a screen element for specifying that all or part of the spoken word and / or the backward-translated word contains an error, together with the spoken word and / or the backward-translated word; When the screen element is selected via the first terminal, the second display step displays a predetermined image together with the translated word on the display screen of the second terminal. Voice translation methods.
Citation Information
Patent Citations
Translation apparatus
JP2009122989A
Word registering apparatus, and computer program for the same
JP2014048506A
Voice translation system, method, and program
JP2015069600A
Call system, call method, and computer program
JP2023124145A
Translation device and translation method
WO2018110096A1