Spoken language identification system, spoken language identification method, and program
The hands-free speech language identification system quickly identifies and translates languages by setting a first language for automatic detection, addressing the inefficiencies of existing translation devices that require user operation and time-consuming registration or cross-matching.
Patent Information
- Application Number
- JP2024099078
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-19
- Publication Date
- 2026-01-07
AI Technical Summary
Existing translation devices require pressing an operation button each time a speaker changes languages, leading to increased time consumption and reduced communication efficiency, and existing language identification technologies are time-consuming due to the need to register users or perform cross-matching with multiple language features.
A hands-free speech language identification system that sets a first language for identification, allowing automatic language detection without user operation, and includes a translation mechanism to convert speech into the identified language.
Enables rapid language identification and translation without the need for user intervention, allowing seamless communication between speakers of different languages without time constraints.
Smart Images

Figure 2026001611000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech language identification system, a speech language identification method, and a program used when translating the content of speech, such as an utterance, into another language. [Background technology]
[0002] With the increase in the number of foreigners due to business, travel, immigration, etc., there is a demand for smoother communication between people who speak different languages, such as Japanese and foreigners, in medical, local government, hotels, and other settings.
[0003] For this reason, various translation tools have been proposed. For example, a translation machine is provided that inputs speech by pressing an operation button and speaking each time a speaker changes, and then translates the speech.
[0004] Patent document 1 also proposes a translation device that can smoothly transmit speech to a destination device when converting input speech in a first language into output speech in one or more target languages other than the first language and transmitting the speech to the destination device.
[0005] Patent Document 2 proposes a translation device that can automatically change language settings depending on the language used by the other party.
[0006] Patent Document 3 proposes a technique for an interpretation system that can translate input speech in a shorter time and reduce the rate of translation errors. [Prior art documents] [Patent documents]
[0007] [Patent Document 1] Patent Publication No. 2021-117676 [Patent Document 2] Japanese Patent Application Publication No. 2018-124695 [Patent Document 3] Japanese Patent Application Laid-Open No. 2002-229591 Summary of the Invention [Problem to be solved by the invention]
[0008] However, with translation devices that require pressing an operation button to speak each time the speaker changes, it is time-consuming to have to move the translation device back and forth between speakers of different languages. This not only increases the time it takes to serve customers, but also reduces the number of times both speakers can speak, posing a problem of not being able to communicate as needed.
[0009] Furthermore, the technology described in Patent Document 1 requires that users and speech languages be registered in association with each other, and it is not possible to identify and translate between users whose languages are not registered.
[0010] In the technology described in Patent Document 2, language identification and translation are performed on the premise that both the first language and the second language are unknown, and therefore language identification takes time.
[0011] The technology described in Patent Document 3 identifies a language by cross-matching with feature patterns of speech data in multiple languages, so it also takes time to identify the language.
[0012] An object of the present invention is to provide a spoken language identification system, a spoken language identification method, and a program that can identify spoken languages in a short time without the need to press an operation button each time a speaker changes or to register users in association with languages. [Means for solving the problem]
[0013] The above object can be achieved by the following means. (1) a setting means for setting a first language to be identified; a selection means for selecting a second language; an input means for inputting a voice in the first language or the second language; a language identification means for determining whether the language of the voice input by the input means is a first language, and for identifying the language of the voice as the first language if it is determined to be the first language, and for identifying the language of the voice as a second language if it is determined not to be the first language; A spoken language identification system comprising: (2) A speech language identification system as described in the preceding paragraph 1, in which voice input by the input means and language identification by the language identification means are performed in a hands-free state without user operation each time a speaker speaks. (3) The speech language identification system according to the preceding paragraph 1 or 2, further comprising a translation means for translating the content of the speech into the language identified by the language identification means. (4) a setting step of setting a first language to be identified; a selection step for selecting a second language; an input step of inputting a speech in the first language or the second language; a language identification step of determining whether the language of the voice input in the input step is a first language, and if it is determined that the language of the voice is the first language, identifying the language of the voice as the first language, and if it is determined that the language of the voice is not the first language, identifying the language of the voice as a second language; A method for identifying spoken language, comprising: (5) The speech language identification method according to the preceding paragraph 4, wherein the speech input in the input step and the language identification in the language identification step are performed in a hands-free state without any user operation, each time a speaker speaks. (6) The speech language identification method according to the preceding paragraph 4 or 5, further comprising a translation step of translating the speech content into the language identified in the language identification step. (7) a setting step of setting a first language to be identified; a second language receiving step of receiving the selected second language; a voice receiving step of receiving voice data in the first language or the second language; a language identification step of determining whether the language of the voice data received in the voice receiving step is a first language, and if it is determined that the language is the first language, identifying the language of the voice as the first language, and if it is determined that the language is not the first language, identifying the language of the voice as a second language; A program that causes a computer to execute the following. (8) The program according to the preceding paragraph 7, in which voice reception in the reception step and language identification in the language identification step are performed every time a speaker speaks in a hands-free state without user operation. (9) The program according to the preceding paragraph 7 or 8, further causing the computer to execute a translation step of translating the content of the voice data into the language identified by the language identification step. [Effects of the Invention]
[0014] According to this invention, a first language to be identified is set, and a second language is selected based on a selection operation by a user or the like. When a speaker inputs speech in the first language or the second language, it is determined whether the language of the input speech is the first language. If it is determined to be the first language, the language of the speech is identified as the first language. If it is determined not to be the first language, the language of the speech is identified as the second language.
[0015] In this way, since the first language to be identified is set and the second language is selected, if the language of the input speech is not the first language, it can be immediately identified as the second language without the need to identify whether it is the second language. In other words, since it is essentially necessary to determine only whether it is the first language, it is sufficient to build a speech language identification system specialized in identifying the first language. This allows language identification to be performed quickly and with a light load. Furthermore, since it is not necessary to identify whether it is the second language, there is no restriction on the number of second languages.
[0016] Moreover, there is no need to press an operation button each time a speaker changes, or to register a user in association with a language. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a block diagram showing a configuration of a spoken language identification system according to an embodiment of the present invention. [Figure 2] 2 is a flowchart illustrating the operation of the spoken language identification system shown in FIG. [Figure 3] 10 is a language selection screen displayed on a tablet device for selecting a second language. [Figure 4] FIG. 10 is a diagram showing a function selection screen displayed on the tablet terminal. [Figure 5] FIG. 5 is a diagram showing the screen of the tablet terminal immediately after the machine interpretation hands-free mode button is pressed on the screen of FIG. 4. [Figure 6] FIG. 10 is a diagram showing the screen of a tablet terminal on which a translation result is displayed. [Figure 7] FIG. 7 is a diagram showing a screen that is displayed when the start button is pressed on the screen of FIG. 6. [Figure 8] FIG. 10 is a diagram showing screen transitions on a tablet terminal based on a multilingual interpretation service application. DETAILED DESCRIPTION OF THE INVENTION
[0018] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0019] 1 is a block diagram showing the configuration of a spoken language identification system according to an embodiment of the present invention. The spoken language identification system includes a tablet terminal 1, an external display device 2, and a speech recognition translation server 3.
[0020] In this embodiment, a situation will be described in which a Japanese person and a foreigner are conversing with each other, such as when a Japanese person at a counter or a customer service center in a local government or a hotel converses with a foreign visitor.
[0021] The tablet terminal 1 includes a built-in display 11, an audio microphone 12, a network communication unit 13, an external display connection unit 14, a storage unit 15, etc. The tablet terminal 1 may be a smartphone or the like.
[0022] The built-in display 11 is a display unit that displays a language selection screen, translation results, etc. The voice microphone 12 is used to input the speaker's voice.
[0023] The network communication unit 13 is an interface for connecting the tablet terminal 1 to the speech recognition translation server 3 via a network. The external display connection unit 14 is an interface for connecting the tablet terminal 1 and the external display device 2.
[0024] The storage unit 15 is a memory for storing various data, and in this embodiment, particularly stores a multilingual interpretation service application 16 and other programs.
[0025] Although not shown, tablet terminal 1 is equipped with a computer system including a CPU, RAM, etc. Each unit of tablet terminal 1, such as the built-in display 11, audio microphone 12, network communication unit 13, external display connection unit 14, and storage unit 15, is controlled by the CPU.
[0026] The external display device 2 is made up of a liquid crystal display device or the like, and displays the translation results and the like.
[0027] The speech recognition translation server 3 is an information processing device such as a personal computer. The speech recognition translation server 3 is an API (Application Programming Interface) server that works in conjunction with the multilingual interpretation service application 16 on the tablet terminal 1 to recognize speech data, identify the language from the speech data, and translate the content of the speech data based on the language identification results.
[0028] The speech recognition translation server 3 functionally comprises a language recognition unit 31, a machine translation unit 32, and a speech synthesis unit 33. Each of these units fulfills its function when the CPU of the speech recognition translation server 3 operates according to a program.
[0029] In the language recognition unit 31, a first language to be identified is set in advance. In this embodiment, Japanese is set as the language to be identified. However, a plurality of identifiable languages may be prepared, and a language selected from among them may be set as the first language to be identified. The language recognition unit 31 also receives from the tablet terminal 1 information about a second language selected by the user on a language selection screen displayed on the tablet terminal 1. Furthermore, the language recognition unit 31 accepts the speaker's voice data input and transmitted by the tablet terminal 1, and determines whether the language of the accepted voice data is Japanese or the second language.
[0030] In this embodiment, it is assumed that the user has selected English as the second language on the language selection screen of the tablet terminal 1. The language recognition unit 31 identifies the language by determining whether the language of the voice data is Japanese. If it is determined to be Japanese, the language recognition unit 31 identifies the language of the voice data as Japanese. If it is determined not to be Japanese, the language recognition unit 31 identifies the language as English without performing the task of determining whether it is English.
[0031] When determining whether a word is Japanese, the language recognition unit 31 refers to a Japanese / non-Japanese classifier 4. This classifier 4 has a large amount of Japanese feature data, and if there is a high degree of match between the voice data and the feature data, the language recognition unit 32 determines that the word is Japanese. If there is not a high degree of match between the voice data and the feature data, the language recognition unit 31 determines that the word is not Japanese.
[0032] The Japanese / non-Japanese classifier 4 may be stored in the speech recognition translation server, or may be stored in an external device, and the language recognition unit 31 may issue a task request to the external device. Alternatively, the language recognition unit 31 may obtain Japanese feature data from an external device.
[0033] The machine translation unit 32 performs machine translation of the content of the voice data based on the language identification result by the language recognition unit 31. If the language of the voice data is identified as Japanese, the machine translation unit 32 translates the content of the voice data into English. If the language of the voice data is identified as English, the machine translation unit 32 translates the content of the voice data into Japanese. The machine translation unit 32 may have a built-in translation engine, or may perform translation using a known translation engine.
[0034] The voice synthesis unit 33 synthesizes and converts the content machine-translated by the machine translation unit 32 into voice data, and transmits this to the tablet terminal 1 together with the machine-translated text data.
[0035] The tablet terminal 1 displays the received text data and outputs the received audio data from a speaker (not shown). Note that it is possible to display the text data or output the audio data only.
[0036] Next, the operation of the spoken language identification system shown in FIG. 1 will be described with reference to the flowchart of FIG.
[0037] In step S01, when a Japanese counter clerk or customer service representative starts the multilingual interpretation service application 16 on the tablet terminal 1, a language selection screen to be used is displayed on the tablet terminal 1. The language selection screen is shown in Figure 3. On this selection screen, each language type is displayed in Japanese paired with its representation in that language.
[0038] In step S02, a second language (foreign language) to be used is selected by a counter clerk, a customer service representative, or another user on the language selection screen shown in Fig. 3. In this embodiment, it is assumed that English is selected.
[0039] After the selection, a function selection screen such as that shown in Fig. 4 is displayed on the tablet terminal 1. On this screen, a machine interpretation mode button 111, a machine interpretation hands-free mode button 112, and a video interpretation mode button 113 are displayed.
[0040] In the machine interpretation button mode, the speaker presses a button to set the language while translation is performed, and the button must be pressed each time the speaker changes, which requires the tablet device 1 to be moved back and forth between speakers.
[0041] The hands-free mode for machine interpretation is a hands-free mode that does not require user operation, and performs voice input, voice reception, language identification and specification, translation, and display of translation results every time a speaker speaks. In this mode, no user operation such as pressing a button is required, and combined with the fact that processing proceeds automatically, high-speed processing is possible.
[0042] The video interpretation mode is a mode in which interpretation is performed while having a direct conversation with an interpreter via video call.
[0043] In this embodiment, it is assumed that the machine interpretation hands-free mode is selected, and in step S03, the user presses the machine interpretation hands-free mode button 112. Fig. 5 shows the screen of the tablet terminal 1 immediately after the machine interpretation hands-free mode button 112 is pressed. On this screen, the "Start button" 114 is displayed, but the machine interpretation hands-free mode is still stopped (step S04 in Fig. 2).
[0044] Although not shown in the figure, the selection of English and the selection of machine interpretation hands-free mode are notified to the speech recognition translation server 3. Upon receiving this notification, the speech recognition translation server 3 sets the second language to English and the machine interpretation hands-free mode to be active. However, the notification from the tablet terminal 1 to the speech recognition translation server 3 that English has been selected and that the machine interpretation hands-free mode has been selected may be made at any time. As mentioned above, the fact that the language to be identified is Japanese is preset in the speech recognition translation server 3.
[0045] When the "Start button" 114 on the screen of the tablet terminal 1 is pressed, the machine interpretation hands-free mode starts in step S05, but the tablet terminal 1 goes into a standby state in step S06 until speech is input into the tablet terminal 1.
[0046] When speech is input, the tablet terminal 1 transmits the speech data to the speech recognition translation server 3. In step S07, the speech recognition translation server 3 determines whether speech has been recognized, in other words, performs speech detection. If no speech is detected (detection determination), the server remains in the standby state of step S06. When the speech data is received from the tablet terminal 1, the speech recognition translation server 3 determines that speech has been detected, and performs language recognition (language identification) in step S08.
[0047] The speech recognition translation server 3 performs language identification by determining whether the language of the speech data is Japanese. If the speech recognition translation server 3 identifies and specifies that the speech data is Japanese, it proceeds to step S09 and performs machine translation of the speech data from Japanese to the language other than Japanese selected in step S02 (English in this embodiment).
[0048] The speech recognition translation server 3 synthesizes and converts the machine-translated data into voice data, and transmits the synthesized voice data and the machine-translated text data to the tablet terminal 1.
[0049] In step S10, the tablet terminal 1 receives the voice data and text data and outputs the voice data in Japanese, which is the translation result. The tablet terminal 1 also displays the Japanese text data, which is the translation result, on its display screen, as shown in FIG. 6. Note that while the voice recognition translation server 3 transmitted the Japanese voice data and text data to the tablet terminal 1, it may transmit only either the voice data or the text data. Also, while the tablet terminal 1 is described as outputting both the Japanese voice data and text data, it may output only one of them.
[0050] If the speech recognition translation server 3 determines in step S08 that the language of the speech data is not Japanese, it determines that the language of the speech data is English. Then, in step S11, the speech recognition translation server 3 performs a machine translation of the speech data into Japanese.
[0051] The speech recognition translation server 3 synthesizes and converts the machine-translated data into voice data, and transmits the synthesized voice data and the machine-translated text data to the tablet terminal 1.
[0052] In step S12, the tablet terminal 1 receives the voice data and text data and outputs the translation result in English voice data. The tablet terminal 1 also displays the translation result in English text data on its display screen, as shown in FIG. 6. Note that while the voice recognition translation server 3 transmitted English voice data and text data to the tablet terminal 1, it may transmit only either the voice data or the text data. Also, while the tablet terminal 1 is described as outputting both English voice data and text data, it may output only one of them.
[0053] Thus, as shown on the screen in FIG. 6, Japanese text data is displayed on the left edge and English text data is displayed on the right edge in chronological order up and down.
[0054] It should be noted that when the "Start button" 114 is pressed on the screen of FIG. 6, the hands-free mode ends (is interrupted), as in step S13 of FIG. 2, and the process returns to step S04, where the hands-free mode is stopped. The screen in this stopped state is shown in FIG. 7. In order to distinguish between the operating state of the hands-free mode in FIG. 6 and the stopped state in FIG. 7, the display manner may be changed. In this example, the display manner of the "Start button" 114, for example, the color, is changed, and the change in color is indicated by hatching the "Start button" 114 in FIG. 7.
[0055] In this embodiment, Japanese is set as the language to be identified, and English is selected as the second language on the selection screen of the tablet device. When a speaker inputs speech in Japanese or English, it is determined whether the language of the input speech is Japanese. If it is determined to be Japanese, the language of the speech is identified as Japanese. If it is determined not to be Japanese, it is identified as English, which is the second language.
[0056] In this way, Japanese is set as the language to be identified, and English is selected as the second language. Therefore, if the language of the input speech is not Japanese, it can be immediately identified as English without the need to identify whether it is English or not. In other words, since it is only necessary to determine whether it is Japanese or not, it is sufficient to build a speech language identification system specialized for identifying Japanese. This allows language identification to be performed at high speed and with a light load. Furthermore, since there is no need to identify whether it is English or not, there is no limit to the number of languages that can be used as the second language.
[0057] Furthermore, since language identification and translation are performed each time a voice is input in hands-free mode, there is no need to press a button to set the language each time a speaker changes, which allows for faster language identification and translation.
[0058] FIG. 8 is a diagram showing screen transitions on the tablet terminal 1 based on the multilingual interpretation service application 16.
[0059] When the multilingual interpretation service application 16 is started, a language selection screen (S21) having a list of languages is displayed on the tablet terminal 1. When a second language to be used is selected, a function selection screen (S22) having buttons for machine interpretation mode 111, machine interpretation hands-free mode 112, and video interpretation mode 113 is displayed on the tablet terminal 1.
[0060] When the machine interpretation hands-free mode is selected on the function selection screen, the screen transitions to a machine interpretation hands-free mode screen (S23).
[0061] This screen is in standby mode (S24), and when the "Start button" 114 is pressed, hands-free mode is initiated (S25). When speech is input in this state, the screen changes to speech recognition in progress (S26), language recognition in progress (S27), or machine translation in progress (S28), depending on the operating state of the language recognition and translation server 3. When language identification and translation by the language recognition and translation server 3 are complete, the translation result is displayed (S29).
[0062] Furthermore, if the "Start button" 114 is pressed during speech recognition, language recognition, or machine translation, the hands-free mode ends (interrupts) (S30) and the device enters a standby state.
[0063] When the multilingual interpretation service application 16 is terminated, the application is terminated.
[0064] Although one embodiment of the present invention has been described above, the present invention is not limited to the above embodiment. For example, while Japanese was set as the language to be identified, a language other than Japanese may also be set as the language to be identified. Furthermore, although English was selected as the second language to be used on the language selection screen of the tablet terminal 1, a language other than English may also be selected. Furthermore, two or more languages may be selected as the second language. In this case, the speech recognition translation server 3 identifies whether the language of the speech data is Japanese, and if it is determined not to be Japanese, it determines that the language is one of multiple other languages (e.g., English and French). The speech recognition translation server 3 then sets one of the other languages (e.g., English) as the language to be identified, identifies whether the language is English, and if it is not English, determines that the language is French. In this case, a classifier for identifying whether the language is English is required, but a classifier for identifying whether the language is French is not necessary.
[0065] Although the example shows a case where processing is performed in a hands-free mode that does not involve user operation, use is not limited to the hands-free mode. For example, it is also possible for processing to be started by the speaker pressing an operation button. [Explanation of symbols]
[0066] 1 tablet device 2 External display devices 3. Speech recognition translation server 4. Japanese / non-Japanese classifier 11 Built-in display 12 Audio microphone 13 Network Communications Department 14 External display connector 15 Storage section 16 Multilingual translation service applications 31 Language Recognition Department 32 Machine Translation Department 33 Speech synthesis unit 111 Machine interpretation mode button 112 Machine interpretation hands-free mode button 113 Video Interpretation Mode Button 114 Start button
Claims
1. a setting means for setting a first language to be identified; a selection means for selecting a second language; an input means for inputting a voice in the first language or the second language; a language identification means for determining whether the language of the voice input by the input means is a first language, and for identifying the language of the voice as the first language if it is determined to be the first language, and for identifying the language of the voice as a second language if it is determined not to be the first language; A spoken language identification system comprising:
2. 2. The spoken language identification system according to claim 1, wherein the voice input by the input means and the language identification by the language identification means are performed in a hands-free state without any user operation, every time a speaker speaks.
3. 3. The speech language identification system according to claim 1, further comprising a translation means for translating the content of the speech into the language identified by the language identification means.
4. a setting step of setting a first language to be identified; a selection step of selecting a second language; an input step of inputting speech in the first language or the second language; a language identification step of determining whether the language of the voice input in the input step is a first language, and if it is determined that the language of the voice is the first language, identifying the language of the voice as the first language, and if it is determined that the language of the voice is not the first language, identifying the language of the voice as a second language; A method for identifying spoken language, comprising:
5. 5. The method of claim 4, wherein the speech input in the input step and the language identification in the language identification step are performed in a hands-free state without any user operation, every time a speaker speaks.
6. 6. The method for identifying a spoken language according to claim 4, further comprising a translation step of translating the content of the speech into the language identified in the language identification step.
7. a setting step of setting a first language to be identified; a second language receiving step of receiving the selected second language; a voice receiving step of receiving voice data in the first language or the second language; a language identification step of determining whether the language of the voice data received in the voice receiving step is a first language, and if it is determined that the language is the first language, identifying the language of the voice as the first language, and if it is determined that the language is not the first language, identifying the language of the voice as a second language; A program that causes a computer to execute the following.
8. 8. The program according to claim 7, wherein the voice reception in the reception step and the language identification in the language identification step are performed every time a speaker speaks in a hands-free state without any user operation.
9. 9. The program according to claim 7, further causing the computer to execute a translation step of translating the content of the voice data into the language identified in the language identification step.
Citation Information
Patent Citations
Speech interpretation system, and speech interpretation method
JP2002229591A
Translator
JP2018124695A
Translator, translation method, and program
JP2021117676A