First user-side device, translation system, and translation method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 瀧澤 清美
- Filing Date
- 2024-08-27
- Publication Date
- 2026-08-07
AI Technical Summary
[0015] According to the present invention, even if the parties use different languages, they can communicate smoothly.
Smart Images

Figure CN121844379B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a first user-side device, a translation system, and a translation method. Background Technology
[0002] As existing technology, it is known that in IP (Internet Protocol) telephony that utilizes the Internet, an IP network translation method is used to automatically identify the language used by the caller and the recipient based on the IP number at the time of the IP phone connection and to automatically translate the telephone conversation over the IP network (for example, see Patent Document 1).
[0003] This IP network translation method enables the translation server on the IP network to automatically translate the language used by the caller and the language used by the receiver, and to conduct voice calls while automatically translating the voice between IP phones.
[0004] Existing technical documents
[0005] Patent documents
[0006] Patent Document 1: Japanese Patent Application Publication No. 2004-260430 Summary of the Invention
[0007] The problem the invention aims to solve
[0008] Existing IP network translation methods have the following problems when identifying the language used by the caller and receiver based on IP numbers: between IP phones within the same country, even if the languages used are different, they will be considered to be the same language.
[0009] Therefore, the object of the present invention is to provide a first user-side device, translation system, and translation method that enable smooth communication even when the users use different languages.
[0010] Solution for solving the problem
[0011] One aspect of the present invention provides a first user-side device comprising: a first user-side voice input unit for inputting first user-side voice; a first user-side output unit for outputting second user-side synthesized voice; a first user-side communication unit for communicating with a second user-side device via a network; and a translation unit for generating first user-side synthesized voice that translates the first user-side voice input to the first user-side voice input unit into the second user-side language, generating second user-side synthesized voice that translates the second user-side voice input via the first user-side communication unit into the first user-side language, outputting the second user-side synthesized voice from the first user-side output unit, and outputting the second user-side synthesized voice to the second user-side device.
[0012] Another aspect of the present invention provides a translation system comprising: a first user-side device, including: a first user-side voice input unit for inputting first user-side voice; a first user-side output unit for outputting second user-side synthesized voice; a first user-side communication unit for communicating with the second user-side device via a network; and a translation unit that generates first user-side synthesized voice after translating the first user-side voice input to the first user-side voice input unit into the second user-side language, generates second user-side synthesized voice after translating the second user-side voice input via the first user-side communication unit into the first user-side language, outputs the second user-side synthesized voice from the first user-side output unit, and outputs the second user-side synthesized voice to the second user-side device via the network; and a translation mechanism thereof. The device is connected to a network and performs speech recognition processing to identify and translate sentences from first user-side speech and second user-side speech, text processing of the speech after speech recognition processing, translation processing of the text obtained from text processing into the other party's language, and speech processing of the translated text obtained from translation processing to generate first user-side synthesized speech and second user-side synthesized speech; and a second user-side device, which includes: a second user-side speech input unit that inputs second user-side speech; a second user-side output unit that outputs first user-side synthesized speech; a second user-side communication unit that communicates with the first user-side device via a network; and a second user-side control unit that outputs the first user-side synthesized speech obtained via the second user-side communication unit from the second user-side output unit.
[0013] Another aspect of the present invention provides a translation method, which inputs first user-side speech and second user-side speech, performs speech recognition processing to identify and translate sentence segments from the input first user-side speech and second user-side speech, performs text processing on the speech after speech recognition processing, performs translation processing on the text obtained from text processing to the other party's language, performs speech processing on the translated text obtained from translation processing to generate first user-side synthesized speech and second user-side synthesized speech, outputs second user-side synthesized speech on the first user-side device, and outputs first user-side synthesized speech on the second user-side device.
[0014] Invention Effects
[0015] According to the present invention, even if the parties use different languages, they can communicate smoothly. Attached Figure Description
[0016] Figure 1 This is a diagram illustrating an example of a translation system.
[0017] Figure 2A This is an example of a block diagram of a user-side device.
[0018] Figure 2B This is an example of a block diagram of the opposite-side device.
[0019] Figure 2C This is a diagram illustrating an example of how the translation API processes data.
[0020] Figure 3 This is an example of a language selection screen displayed on the user-side device.
[0021] Figure 4 This is an example of a standby screen displayed on the user's device.
[0022] Figure 5 This is an example of a phonebook screen displayed on a user-side device.
[0023] Figure 6 This is a diagram showing an example of a control screen displayed on the user-side device.
[0024] Figure 7 This is a timing diagram illustrating an example of the actions of a translation system. Detailed Implementation
[0025] [Summary of Implementation Methods]
[0026] As an example, the translation system with AI (Artificial Intelligence) interpretation in this embodiment is roughly constructed by combining a multi-function mobile phone or IP phone terminal with an AI interpretation device. This translation system uses the AI interpretation device to recognize the voice of a conversation taking place on the multi-function mobile phone or IP phone terminal, and outputs the translated voice through stream processing from the multi-function mobile phone or IP phone terminal. Thus, the translation system of this embodiment can conduct telephone conferences or calls with users who speak different languages without the need for formal interpretation. Furthermore, the so-called AI interpretation, as described later, combines translation from multiple Web APIs (Web Application Programming Interfaces) residing in the cloud.
[0027] [First Implementation]
[0028] (Summary of Translation System 1)
[0029] Figure 1 This is a diagram illustrating an example of the translation system of the first embodiment. Figure 2A This is an example of a block diagram of the user-side device according to the first embodiment. Figure 2B This is an example of a block diagram of the counterpart device. Figure 2C This is a diagram illustrating an example of the processing of a translation API. Furthermore, in the various figures of the embodiments described below, the ratios and shapes between the figures sometimes differ from the actual ratios and shapes. Additionally, in Figures 2A to 2C In the image, arrows indicate the flow of the main speech and synthesized speech.
[0030] The translation system 1 of this embodiment is a system in which a first user using a user-side device 2 as a first user-side device can communicate with a second user using a counterpart-side device 3 as a second user-side device via WebRTC (Web RealTime Communication) in a P2P (Peer to Peer) communication manner. Figure 1 The white arrows shown illustrate a scenario where the user-side device 2 and the counterpart-side device 3 are communicating via P2P. Furthermore, the translation system 1 can translate the first language into the second language and vice versa, even when the first user uses the first language and the second user uses a second language, as described later, enabling smooth communication even if the users do not understand each other's languages. Additionally, if at least one of the user-side device 2 and the counterpart-side device 3 has a camera, the audio is separated from the video and output to WebRTC.
[0031] This embodiment describes a scenario where a first user makes a phone call to a second user, and the first user's user-side device 2 has a translation unit described later. In this embodiment, the first user making the call is recorded as the user, and the second user receiving the call is recorded as the other party. Furthermore, as an example, the user-side device 2 and the other party's device 3 are personal computers, multifunction mobile phones, tablet computers, and tablet computers, etc.
[0032] like Figure 1 As shown, the translation system 1 is generally configured to include a user-side device 2, a counterparty-side device 3, and a translation API 6 that functions as a translation agency. Alternatively, the translation agency may be a translation device connected to a network 4. In this embodiment, the translation API 6 uses multiple cloud APIs from a cloud 5 to perform the processing as a translation agency.
[0033] like Figure 2A As shown, the user-side device 2 is generally configured to include: a voice input unit (first user-side voice input unit) 20, which inputs user-side voice (first user-side voice) S1; and an output unit (first user-side output unit) 21, which outputs synthesized voice from the other side (second user-side synthesized voice) S1. 32 The communication unit (first user-side communication unit) 25 communicates with the counterpart device 3 via network 4; and the user-side control unit 26, which acts as a translation unit, generates user-side synthesized speech (first user-side synthesized speech) S1 after translating the user-side speech S1 input to the speech input unit 20 into the counterpart language (second user-side language). 31 Generates the voice input from the other party (the second user's voice) via the communication unit 25. 11 The synthesized speech S on the other side, translated into the user's language (first user's language), is then used. 32 To make the other side synthesize speech S 32 The output is from the output unit 21, so that the user side synthesizes the speech S. 31 Output is sent to device 3 on the other side. Network 4 is, for example, the Internet.
[0034] As an example, such as Figure 2A As shown, the user-side device 2 also includes a display unit 22, an input unit 23, and a storage unit 24.
[0035] like Figure 2B As shown, the counterparty device 3 is generally configured to include: a voice input unit (second user-side voice input unit) 30, which inputs counterparty voice S. 11 Output unit (second user-side output unit) 31, which outputs user-side synthesized speech S 31The communication unit (second user-side communication unit) 35 communicates with the user-side device 2 via the network 4; and the counterparty control unit (second user-side control unit) 36 enables the user-side synthesized speech S obtained via the communication unit 35 to... 31 Output is made from output unit 31.
[0036] As an example, the user-side device 2 also includes a display unit 32, an input unit 33, and a storage unit 34.
[0037] like Figure 2C As shown, the translation API6 is connected to network 4, and is roughly configured to perform voice translation from the user's side S1 and the other party's side S2. 11 The process involves speech recognition processing to identify sentence segments for translation, text processing of the speech after speech recognition processing, translation of the text obtained from text processing into the target language, and speech processing of the translated text obtained from translation processing to generate synthesized speech S for the user. 31 Synthesized speech with the other side S 32 .
[0038] Specifically, as an example, the Translation API 6 is generally configured to include: a speech recognition processing unit 60, which performs speech recognition processing; a text processing unit 61, which performs text processing; a translation processing unit 62, which performs translation processing; and a speech processing unit 63, which performs speech processing. As an example, the speech recognition processing unit 60, the text processing unit 61, the translation processing unit 62, and the speech processing unit 63 are Web APIs constructed using AI technology.
[0039] As an example, such as Figure 1 As shown, the translation API 6 is composed of multiple APIs, including the speech recognition processing unit 60. Therefore, the speech recognition processing unit 60, the text processing unit 61, the translation processing unit 62, and the speech processing unit 63 are configured such that, when multiple APIs capable of performing the same processing exist, other APIs are selected to perform the processing based on factors such as processing speed and congestion level.
[0040] (The configuration of user-side device 2)
[0041] The voice input unit 20 of the user-side device 2 is a microphone. The output unit 21 is a speaker, but it is not limited to this and can also be configured to connect to headphones via wired or wireless means. The display unit 22 is, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display. The input unit 23 is, for example, a keyboard, a mouse, or a touchpad. In this embodiment, the display unit 22 and the input unit 23 of the user-side device 2 are configured as a touch panel with the input unit 23 superimposed on the display unit 22. The storage unit 24 is, for example, a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). This storage unit 24 stores a browser application 240. The communication unit 25 is configured to exchange information or signals with the network 4 or the cloud 5.
[0042] The user-side control unit 26 is, for example, a microcomputer composed of a CPU (Central Processing Unit) that performs calculations and processing on the acquired data according to a stored application program, RAM (Random Access Memory) and ROM (Read Only Memory) as semiconductor memory. The ROM, for example, stores the program used by the user-side control unit 26 to perform operations. The RAM is used, for example, as a storage area for temporarily storing calculation results.
[0043] The user-side control unit 26 functions as a translation unit by reading the browser application 240 from the storage unit 24 and accessing the website 70 of the web server 7 (described later). The website 70 enables two-way communication between the user-side device 2 and the other party's device 3.
[0044] User-side control unit 26 is connected to translation API 6 via network 4, enabling translation API 6 to transmit user-side voice S1 and other-side voice S2. 11 Perform speech recognition processing, text processing, translation processing, and speech conversion processing to generate synthesized speech S for the user. 31 Synthesized speech with the other side S 32 .
[0045] Users of the translation system 1 have pre-registered their identification number, name, phone number, and language of use in the phone book DB (Data Base) 71 (described later). The user-side control unit 26 identifies the language of use based on the phone book DB 71 and causes the translation API 6 to perform the translation. Furthermore, if the user is not pre-registered in the phone book DB 71, the user-side control unit 26 causes the translation API 6 to perform the translation based on the language selected by the user on the website 70.
[0046] like Figure 2C As shown, the speech recognition processing unit 60 receives input speech S1 from the user side and speech S2 from the other party side. 11 When the speech 600 is input, sentence segmentation is recognized for text conversion. Next, the text conversion processing unit 61 performs text conversion processing on the speech 601 after recognition processing output from the speech recognition processing unit 60. Next, the translation processing unit 62 translates the text 602 output from the text conversion processing unit 61 into either the user's language or the other party's language. Next, the speech conversion processing unit 63 performs speech conversion processing on the translated text 603 output from the text conversion processing unit 61, generating synthesized speech 604. This synthesized speech 604 is the user-side synthesized speech S. 31 Synthesized speech with the other side S 32 .
[0047] Furthermore, the translated text 603 is translated into the user's language and output to the user's device 2, and translated into the other party's language and output to the other party's device 3. Thus, the user's device 2 can display the other party's voice S. 11 The text is translated into its own language. Additionally, the other party's device 3 can display the text of the user's speech S1 translated into its own language.
[0048] User-side control unit 26 synthesizes speech S on the user side 31 Synthesized speech with the other side S 32 During the output period, speech recognition processing is temporarily suspended. This temporary suspension is to prevent a loop where the output synthesized speech is input from the respective speech input unit and processing continues. This temporary suspension occurs after the synthesized speech is output, sentence segmentation is recognized, and the sentence is determined, but is not limited to this. Furthermore, as an example, in the case of a temporary suspension, the synthesized speech to be played next can be saved to storage unit 24 after translation processing.
[0049] User-side control unit 26 combines user-side speech S1 with user-side synthesized speech S 31 Superimposed and output to the other side device 3, transmitting the other side's voice S 11 Synthesized speech with the other side S32 The speech is superimposed and output from the output unit 21 on the user's side. Thus, the user and the other party can communicate as if performing simultaneous interpretation. Furthermore, the superposition of the speech and the synthesized speech can be at the same volume, or one of them can be superimposed at a lower volume.
[0050] (The configuration of the counterpart device 3)
[0051] The voice input unit 30 of the counterpart device 3 is a microphone. The output unit 31 is a speaker, but it is not limited to this; it can also be configured to connect to headphones via wired or wireless means. The display unit 32 is, for example, a liquid crystal display or an organic EL display. The input unit 33 is, for example, a keyboard, a mouse, or a touchpad. As an example, the display unit 32 and the input unit 33 of the counterpart device 3 can be configured as a touch panel with the input unit 33 superimposed on the display unit 32. The storage unit 34 is, for example, a storage device such as an HDD or an SSD. The communication unit 35 is configured to exchange information or signals with the network 4 or the cloud 5.
[0052] The control unit 36 on the other side is, for example, a microcomputer composed of a CPU, RAM, and ROM.
[0053] (Web Server 7)
[0054] Web server 7 is connected to network 4. Web server 7 has a website 70 and a phone book DB71. Website 70 can perform peer-to-peer (P2P) communication with registered phone numbers. These phone numbers are registered in the phone book DB71 of web server 7.
[0055] The user operates the user-side device 2 to display the website 70, selects and decides the phone number of the other party to call from the displayed phone numbers, and then makes a call to the other party.
[0056] (The display screen of the user-side device 2)
[0057] Figure 3 This is a diagram showing an example of a language selection screen displayed on the user-side device according to the first embodiment. Figure 4 This is a diagram showing an example of the standby screen displayed on the user-side device according to the first embodiment. Figure 5 This is a diagram showing an example of a phone book screen displayed on the user-side device according to the first embodiment. Figure 6 This is a diagram showing an example of the control screen displayed on the user-side device according to the first embodiment.
[0058] Figures 3-6The language selection screen, standby screen, phonebook screen, and control screen shown are examples of screens displayed when a user launches the browser application 240 on the user-side device 2 and opens the website 70.
[0059] As an example Figure 3 The display screen 220 shown displays multiple language icons 220a and a scroll bar 220b. The language icons 220a are selected by touch operation on the input unit 23. When a user wants to select language D, they can set the language to be used by touching the language icon 220a displayed as "language D".
[0060] As an example, scroll bar 220b is used when you want to select... Figure 3 The system operates on languages other than those displayed as "Language A" to "Language E". Users can move the blacked-out sections up and down in the diagram to display languages other than those listed. As an example, the translation system 1 of this embodiment supports languages from 130 countries and regions, but by using multiple translation APIs 6, it is possible to translate languages from even more countries and regions.
[0061] As an example Figure 4 The display screen 221 shown is the screen displayed when a call is made from the user-side device 2 to the other-side device 3. As an example, the display screen 221 displays the other-side display area 221a, selection area 221b, status display area 221c, translation display 221d, microphone display 221e, translation stop icon 221f, language switching icon 221g, text save icon 221h, end icon 221i, phone number selection icon 221j, phone number display area 221k, call start icon 221l, enter icon 221m, exit icon 221n, and text display area 221o.
[0062] As an example, the other party display area 221a is an area that displays the start time and other information of images, videos or audio recordings of the other party's operation of the other party's device 3.
[0063] As an example, selection area 221b is an area that displays a "mic" radio button that allows selection of whether the voice input unit 20 is turned on or off, and a "voice" radio button that allows selection of whether the voice and synthesized voice are mixed.
[0064] As an example, status display area 221c is an area that displays the current status. As an example, status display area 221c displays "Standby" while waiting for the other party to answer after making a call, and displays "On Call" when the call has started.
[0065] As an example, translation display 221d shows the language to be translated. Figure 4 The text shows how to translate Japanese into English and English into Japanese.
[0066] As an example, the microphone display 221e shows the microphone's on / off status.
[0067] As an example, the translation stop icon 221f is an icon that stops translation when a touch operation is received while translation is in progress and resumes translation when a touch operation is received again.
[0068] As an example, the language switching icon 221g is an icon that allows you to switch the language being translated.
[0069] As an example, the text save icon 221h is an icon that saves the text displayed in the text display area 221o to the storage unit 24.
[0070] As an example, the end icon 221i is the icon that ends the call.
[0071] As an example, the phone number selection icon 221j is the icon operated when selecting a caller from the phone book. This phone book is based on the phone book DB71 for display.
[0072] As an example, the phone number display area 221k is the area that displays the selected phone number.
[0073] As an example, the call start icon 221l is the icon that is operated when making a call to the selected phone number.
[0074] User-side device 2 and counterpart-side device 3 are configured to connect the PSTN (Public Switched Telephone Network) to WebRTC, performing stream relay for their voice communications. The entry icon 221m is activated when entering the WebRTC room. Furthermore, the exit icon 221n is activated when exiting the WebRTC room. Figure 4 In the middle, since it is the state before entering, the exit icon 221n is displayed in gray.
[0075] As an example, text display area 221o displays the user-side voice S1 and the other-side voice S2. 11 The area of the translated text.
[0076] As an example Figure 5The display screen 222 shown is an overlay of the phone book screen 222a displayed when selecting a phone number for a recipient onto the display screen 221. As an example, the phone book screen 222a displays the close icon 222b, add icon 222c, edit icon 222d, delete icon 222e, search window 222f, filter icon 222g, reset icon 222h, phone number display area 222i, selected icon 222j, and deselect icon 222k. Furthermore, in Figure 5 In China, for the sake of Figure 4 The icons shown in the display screen 221 are the same as those in the original screen, and are labeled with the same symbols. Figure 4 Same reference numerals as shown in the attached figures.
[0077] As an example, the close icon 222b is the icon that is operated when closing the phonebook screen 222a.
[0078] As an example, the append icon 222c is the icon that is operated on when appending a phone number.
[0079] As an example, the edit icon 222d is the icon that is used when editing phone numbers, etc.
[0080] As an example, the delete icon 222e is the icon used when deleting a registered phone number, etc.
[0081] As an example, search window 222f is the area where search terms such as registered phone numbers are entered.
[0082] As an example, the filter icon 222g is the icon that is operated when searching in the phone book using search terms entered into the search window 222f.
[0083] As an example, the reset icon 222h is the icon that is operated on when the search term entered into the search window 222f is deleted.
[0084] As an example, the phone number display area 222i displays "checkboxes," "select name," "phone number," and "language." When a user selects the checkbox next to "select name" for the person they want to call, they can make a call to the corresponding phone number. Furthermore, "language" indicates the language spoken by the person whose name has been selected.
[0085] As an example, the selected icon 222j is the icon that is operated when making a call to the phone number selected in the phone number display area 222i.
[0086] As an example, the cancel icon 222k is the icon that is operated on when ending the editing, selection, or closing of the phonebook screen 222a, such as phone numbers.
[0087] As an example Figure 6 The display screen 223 shown illustrates the control screen displayed in a telephone call. In this display screen 223, as an example, the newly displayed control area 223a and the other party display area 221a, etc., have been moved downwards. Furthermore, in... Figure 6 In China, for the sake of Figure 4 The icons shown in the display screen 221 are the same as those in the original screen, and are labeled with the same symbols. Figure 4 Same reference numerals as shown in the attached figures.
[0088] Control area 223a is the area that displays whether text is being saved, or the other party's phone number, etc.
[0089] (action)
[0090] The following is based on Figure 7 The timing diagram illustrates an example of the operation of the translation system 1 in this embodiment. Figure 7 In this example, for ease of explanation, the order of the call is set from the user to the other party.
[0091] Translation System 1 operates based on the following translation method: It inputs user-side speech S1 and recipient-side speech S... 11 The input user-side voice S1 and the other party-side voice S1 are processed. 11 The speech recognition process identifies the sentence segments to be translated. The speech is then converted into text, and the resulting text is translated into the target language. Finally, the translated text is converted back into speech to generate synthesized speech S for the user. 31 Synthesized speech with the other side S 32 Synthesize the other party's voice S 32 The synthesized speech S is output from the user-side device 2. 31 Output is made to the device 3 on the other side.
[0092] Specifically, the user-side device 2 launches the browser application 240 based on the user's operation, so that the website 70 is displayed on the display unit 22 (Step 1).
[0093] The user-side device 2 initiates a call to the phone number selected by the user from the website 70 displayed on the display unit 22 (Step 2).
[0094] When the other party's device 3 receives a call from the user's device 2, it begins a call (Step 3).
[0095] When the user-side device 2 inputs user-side voice S1 into the voice input unit 20 after the call starts, it outputs user-side voice S1 to the translation API 6 via the communication unit 25 and the network 4 (Step 4).
[0096] Translation API6 performs speech recognition, text processing, translation, and speech conversion on the input user-side speech S1 to generate user-side synthesized speech S. 31 (Step 5). Translation API6 will combine the user-side speech S1 with the user-side synthesized speech S. 31 The translated text is then superimposed and output to the user-side device 2 via network 4 (Step 6).
[0097] User-side device 2 will input the superimposed user-side speech S1 and the user-side synthesized speech S via communication unit 25. 31 The translated text is output to the other-side device 3 via the communication unit 25 (Step 7). Here, the user-side device 2 outputs the user-side voice S1 and the user-side synthesized voice S2 to the other-side device 3. 31 During this period, the translation API6 is controlled to temporarily stop new translations (Step 8).
[0098] The other side device 3 will input the user-side voice S1 and the user-side synthesized voice S via network 4 and communication unit 35. 31 The translated text is output from the output unit 31 and displayed on the display unit 32 (Step 9).
[0099] User-side device 2 provides user-side speech S1 and user-side synthesized speech S. 31 After being output, the temporary pause in translation is lifted (Step 10). As an example, the user-side device 2 lifts the temporary pause in translation based on the voice input unit 30 from the other-side device 3.
[0100] The other side device 3 inputs the other side's voice S into the voice input unit 30. 11 At that time, it is output to the user-side device 2 via the communication unit 35 and the network 4 (Step 11).
[0101] User-side device 2 will input the other party's voice S via network 4 and communication unit 25. 11 Output is made via Communication Section 35 and Network 4 to Translation API 6 (Step 12).
[0102] Translation API 6 is designed for the input of the other party's speech S 11 Speech recognition, text processing, translation, and speech conversion are performed to generate synthesized speech S on the other side. 32(Step 13). Translation API6 will translate the other party's speech S 11 Synthesized speech with the other side S 32 The translated text is then superimposed and output to the user-side device 2 (Step 14).
[0103] User-side device 2 receives the superimposed voice S from the other side via network 4 and communication unit 25. 11 Synthesized speech with the other side S 32 The translated text is output from the output unit 21 and displayed on the display unit 22 (Step 15).
[0104] Translation System 1 repeats these actions until the call ends, and then stops the translation process when the call ends.
[0105] Furthermore, the order of translation is not limited to that from user to recipient; it is arbitrary. In addition, the speech of either the user or the recipient is translated consecutively.
[0106] (Effects of the first embodiment)
[0107] Even when using different languages, the translation system 1 of this embodiment can conduct calls smoothly. Specifically, since the user-side device 2 and the other-side device 3 communicate via P2P, and their voices are automatically translated into the user's language and the other-side's language, compared to situations where people speak alternately and are translated, or where only one side is translated, even those who do not speak the other's language can conduct telephone conferences or calls smoothly and easily without the need for an interpreter.
[0108] Translation System 1 continuously recognizes and translates without requiring any operation other than making a phone call. Therefore, compared to situations where button operations are required for translation, users can make calls without any worries.
[0109] Because the speech recognition unit used in translation, such as in translation system 1, utilizes a translation API 6 located in the cloud 5, it can differentiate between systems integrated into devices by using a Web API that excels in terms of update frequency, number of languages it can translate, translation accuracy, or translation speed. For example, if device 3 is installed in a hospital to answer calls from patients, translation system 1 can accurately translate medical terminology, allowing the hospital to accurately understand the patient's symptoms and provide appropriate responses or effective treatment even when the caller speaks a different language. Furthermore, translation system 1 can be flexibly applied to various occasions, including business and tourism.
[0110] Translation system 1 can temporarily stop translation processing when synthesized speech is output, thus suppressing the mixing of speech used for conversation with synthesized speech and translating it, compared to the case where it is not temporarily stopped.
[0111] Translation System 1 outputs the original speech by superimposing it with the synthesized speech, thus making it easier to obtain subtle differences that cannot be grasped by synthesized speech alone, compared to outputting only synthesized speech.
[0112] [Second Implementation]
[0113] The second embodiment differs from the first embodiment in that the first user-side device is the opposite-side device.
[0114] Furthermore, in the embodiments described below, parts having the same functions and structures as in the first embodiment are marked with the same reference numerals as in the first embodiment, and their descriptions are omitted.
[0115] In the translation system 1 of this embodiment, the first user-side device is the counterpart-side device, and the second user-side device is the user-side device.
[0116] In other words, the second user-side device of translation system 1 is configured as an IP telephony terminal. When a user makes a call from the second user-side device (which functions as an IP telephony terminal) to the first user-side device, the first user-side device performs translation processing of the languages used by both parties. Furthermore, if the user's language is unknown, the translation is automatically performed using AI language recognition from the translation API 6.
[0117] (Effects of the second implementation method)
[0118] Even if the dialing device is an IP phone terminal, the translation system 1 of this embodiment can enable smooth communication between people using different languages without the need for an interpreter.
[0119] Alternatively, other embodiments may be provided as a program or a computer-readable recording medium containing the program that performs the following translation method: inputting first user-side speech and second user-side speech, performing speech recognition processing to identify and translate sentence segments from the input first user-side speech and second user-side speech, performing text processing on the speech after speech recognition processing, performing translation processing on the text obtained from the text processing to the other-side language, performing speech processing on the translated text obtained from the translation processing to generate first user-side synthesized speech and second user-side synthesized speech, outputting the second user-side synthesized speech to the first user-side device, and outputting the first user-side synthesized speech to the second user-side device.
[0120] The translation method is further configured to temporarily suspend speech recognition processing during the period when the synthesized speech on the first user side and the synthesized speech on the second user side are output.
[0121] The first user-side device, translation system 1, and translation method described above may, depending on the application, be partly implemented by a computer-executed program, an ASIC (Application Specific Integrated Circuit), or a FPGA (Field Programmable Gate Array).
[0122] The foregoing has described several embodiments of the present invention, but these embodiments are merely examples and are not intended to limit the invention as described in the claims. These new embodiments can be implemented in various other ways, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. Furthermore, not all combinations of features described in these embodiments are necessary for solving the problems of the invention. Moreover, these embodiments are included within the scope and spirit of the invention, and are encompassed by the invention as described in the claims and its equivalents.
[0123] Industrial availability
[0124] According to the present invention, even if the parties use different languages, they can communicate smoothly.
[0125] Explanation of reference numerals in the attached figures
[0126] 1…Translation system, 2…User-side device, 3…Reciprocal device, 4…Network, 5…Cloud, 6…Translation API, 7…Web server, 20…Voice input unit, 21…Output unit, 22…Display unit, 23…Input unit, 24…Storage unit, 25…Communication unit, 26…User-side control unit, 30…Voice input unit, 31…Output unit, 32…Display unit, 33…Input unit, 34…Storage unit, 35…Communication unit, 36…Reciprocal control unit, 60…Speech recognition processing unit, 61…Text processing unit, 62…Translation processing unit, 63…Speech processing unit, 70…Website, S1…User-side voice, S 11 …Opponent's voice, S 31 …user-side synthesized speech, S 32 …synthesized speech from the other side.
Claims
1. A first user-side device, characterized in that, have: The first user-side voice input unit inputs the first user-side voice. The first user-side output unit outputs the synthesized speech from the second user side; The first user-side communication unit communicates with the second user-side device via WebRTC (Web Real Time Communication) through P2P (Peer to Peer). The first user-side display unit displays a web website provided by a web server connected via a network, based on a phone book database (Data Base) that maps the first user-side language and the second user-side language to the second user-side phone number. The first user-side input unit accepts the selection input of the second user-side phone number displayed on the website; as well as The translation unit, after determining the first user-side language and the second user-side language based on the second user-side phone number selected by the first user-side input unit and referring to the phone book DB on the web server, performs the following actions via a cloud-based translation API (Application Programming Interface): generating first user-side synthesized speech input to the first user-side voice input unit into the second user-side language, and generating second user-side synthesized speech obtained from the WebRTC via the first user-side communication unit into the first user-side language. The second user-side speech obtained from the translation API is superimposed with the second user-side synthesized speech and output from the first user-side output unit. The first user-side speech and the first user-side synthesized speech are also superimposed and output to the second user-side device.
2. The first user-side device according to claim 1, wherein, The translation unit causes the translation API to perform speech recognition processing to identify and translate sentence segments from the speech of the first user and the speech of the second user, text processing of the speech after the speech recognition processing, translation processing of the text obtained from the text processing to the other party's language, and speech processing of the translated text obtained from the translation processing, thereby generating the synthesized speech of the first user and the synthesized speech of the second user.
3. The first user-side device according to claim 2, wherein, The translation unit temporarily suspends the speech recognition process during the period when the synthesized speech on the first user side and the synthesized speech on the second user side are output.
4. A translation system, characterized in that, have: A first user-side device includes: a first user-side voice input unit for inputting first user-side voice; a first user-side output unit for outputting second user-side synthesized voice; a first user-side communication unit for P2P communication with the second user-side device via WebRTC; a first user-side display unit for displaying a webpage provided by a web server connected via a network, based on a phonebook database that maps the first user-side language and the second user-side language to the second user-side phone number; a first user-side input unit for accepting selection input of the second user-side phone number from the displayed webpage; and a translation unit for translating the second user-side phone number selected by the first user-side input unit, referring to the phonebook database on the web server. After the call book DB determines the first user-side language and the second user-side language, it performs the following operations via a cloud-based translation API connected to the network: generating first user-side synthesized speech after translating the first user-side speech input to the first user-side speech input unit into the second user-side language; and generating second user-side synthesized speech after translating the second user-side speech obtained via the first user-side communication unit and WebRTC into the first user-side language. The second user-side speech obtained from the translation API is superimposed with the second user-side synthesized speech and output from the first user-side output unit. The first user-side speech is superimposed with the first user-side synthesized speech and output to the second user-side device. The translation API, connected to the network, performs speech recognition processing to identify and translate sentence segments from the speech of the first user and the speech of the second user, text processing of the speech after speech recognition processing, translation processing of the text obtained from the text processing to the other side's language, and speech processing of the translated text obtained from the translation processing to generate the synthesized speech of the first user and the synthesized speech of the second user. as well as The second user-side device includes: a second user-side voice input unit for inputting the second user-side voice; a second user-side output unit for outputting the first user-side synthesized voice; and a second user-side communication unit for P2P communication with the first user-side device via the WebRTC. And a second user-side control unit, which superimposes the synthesized speech of the first user side obtained via the second user-side communication unit with the speech of the first user side and outputs it from the second user-side output unit.
5. A translation method, characterized in that, Input the first user's voice into the first user's device. Input the second user's voice into the second user's device. A website displaying a phone book database provided by a web server connected via the network, which maps the first user's language and the second user's language to the second user's phone number. Accept the selection input of the second user-side phone number displayed on the website. Based on the selected second user-side phone number, the language of the first user-side and the language of the second user-side are determined by referring to the phone book DB on the web server. The generation of first-user-side synthesized speech, which translates the first-user-side speech input to the first-user-side device into the second-user-side language, and the generation of second-user-side synthesized speech, which translates the second-user-side speech obtained from WebRTC via the first-user-side device into the first-user-side language, are performed using a cloud-based translation API connected via the network. The second user-side speech obtained from the translation API is superimposed with the synthesized second user-side speech and output from the first user-side device. The first user-side speech is superimposed with the first user-side synthesized speech and output from the second user-side device.
Citation Information
Patent Citations
System and method for IP network translation
JP2004260430A
Automatic voice translated telephone call system
JP2002300289A
Translation server, terminal, card and translation service method
JP2003022268A