Interpretation system

By employing misrecognition and mistranslation detection to identify and correct low-likelihood words, the system efficiently improves interpretation accuracy by reducing the need for extensive searches in large text datasets.

JP2026004129APending Publication Date: 2026-01-14BRIDE MULTILINGUAL SOLUTIONS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024102373
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

Existing interpretation systems face inefficiencies in improving speech recognition and translation accuracy due to the inefficiency of searching for unnatural words in large volumes of pre- and post-translation text data.

Method used

The system includes misrecognition and mistranslation detection units that extract low-likelihood words from pre- and post-translation text data, storing them in a readable format for manual review and registration in speech and translation dictionaries, allowing for efficient accuracy improvement through additional learning.

Benefits of technology

This approach reduces the need for exhaustive searches, enabling efficient enhancement of interpretation accuracy by allowing for targeted correction of misrecognized and mistranslated words.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026004129000001_ABST
    Figure 2026004129000001_ABST
Patent Text Reader

Abstract

To provide an interpretation system capable of efficiently improving the accuracy of interpretation.SOLUTION: An interpretation system includes a speech recognition unit (102) that performs speech recognition processing on untranslated speech data acquired by a user terminal (101) and outputs untranslated text data, and a translation unit (103) that performs translation processing on the untranslated text data and outputs translated text data to be transmitted to the user terminal. The interpretation system also includes a false recognition detection unit (111) that extracts a word with low likelihood from the untranslated text data, a storage unit (105) that stores the untranslated text data and the word with low likelihood in a readable format, and a speech recognition dictionary (112) that can be used for additional learning by the speech recognition unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an interpretation system. [Background technology]

[0002] There is known an interpretation system that, between a user and a conversation partner who speaks a different language, performs speech recognition processing on pre-translation speech data acquired by a user terminal, outputs pre-translation text data, performs translation processing on this pre-translation text data, and outputs translated text data that is sent to the user terminal (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2023-022150 Summary of the Invention [Problem to be solved by the invention]

[0004] In such an interpretation system, it is desirable to improve the accuracy of speech recognition processing and translation processing in accordance with the accumulation of data.

[0005] The present invention has been made in view of the above-mentioned problems, and has as its object to provide an interpretation system that can efficiently improve the accuracy of interpretation. [Means for solving the problem]

[0006] An interpretation system according to one embodiment of the present invention includes a speech recognition unit that performs speech recognition processing on pre-translation speech data acquired by a user terminal and outputs pre-translation text data, and a translation unit that performs translation processing on the pre-translation text data and outputs translated text data to be transmitted to the user terminal. The interpretation system also includes a misrecognition detection unit that extracts low-likelihood words from the pre-translation text data, a memory unit that stores the pre-translation text data and low-likelihood words in a readable format, and a speech recognition dictionary that can be used for additional learning by the speech recognition unit. The interpretation system may also include a first terminal that can reference the pre-translation text data and low-likelihood words in the memory unit and register them in the speech recognition dictionary.

[0007] To improve the accuracy of interpretation, it is conceivable to search for unnatural words in pre-translation text data, determine the appropriateness of the word if an unnatural word is found, and if the word is not appropriate, register an appropriate word in a speech recognition dictionary and perform additional training on the speech recognition unit. However, searching for unnatural words in a huge amount of pre-translation text data is inefficient.

[0008] Therefore, in this interpretation system, the misrecognition detection unit extracts words with low likelihood from the pre-translation text data and stores them in a readable format in the memory unit. This method eliminates the need to search for unnatural words from vast amounts of pre-translation text data, making it possible to efficiently improve the accuracy of interpretation.

[0009] An interpretation system according to one embodiment of the present invention includes a speech recognition unit that performs speech recognition processing on pre-translation speech data acquired by a user terminal and outputs pre-translation text data, and a translation unit that performs translation processing on the pre-translation text data and outputs translated text data to be transmitted to the user terminal. The interpretation system also includes a mistranslation detection unit that extracts low-likelihood words from the translated text data, a storage unit that stores the translated text data and low-likelihood words in a readable format, and a translation dictionary that can be used for additional learning by the translation unit. The interpretation system may also include a second terminal that can reference the translated text data and low-likelihood words in the storage unit and register them in the translation dictionary.

[0010] To improve the accuracy of interpretation, it is conceivable to search for unnatural words in the translated text data, determine whether the unnatural words are appropriate if they are found, and if they are not appropriate, register appropriate words in the translation dictionary and perform additional learning in the translation unit. However, searching for unnatural words in a huge amount of translated text data is inefficient.

[0011] Therefore, in this interpretation system, the mistranslation detection unit extracts words with low likelihood from the translated text data and stores them in a readable format in the memory unit. This method eliminates the need to search for unnatural words from a huge amount of translated text data, making it possible to efficiently improve the accuracy of interpretation.

[0012] An interpretation system according to one embodiment of the present invention includes a speech recognition unit that performs speech recognition processing on pre-translation speech data acquired by a user terminal and outputs pre-translation text data, a translation unit that performs translation processing on the pre-translation text data and outputs translated text data to be transmitted to the user terminal, and a storage unit that stores at least one of the pre-translation text data and the translated text data in a readable format. The interpretation system may also include a third terminal that is capable of performing voice calls with the user terminal and that can refer to at least one of the pre-translation text data and the translated text data during the voice calls with the user terminal.

[0013] Here, in an interpretation system, if the content of the interpretation is specialized and sufficient accuracy cannot be ensured, it is possible to switch from using the voice recognition unit and translation unit to voice communication with the interpreter, etc.

[0014] In this interpretation system, at least one of the pre-translation text data and the post-translation text data is stored in a readable format in a storage unit. This method allows the interpreter to understand the content of the conversation in advance, making it possible to smoothly transition to a voice call with the interpreter. Furthermore, based on the content of the conversation, it is also possible to select an interpreter (e.g., a highly specialized translator) who is suitable for the content of the conversation. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a schematic block diagram showing the configuration of an interpretation system according to a first embodiment of the present invention. [Figure 2] 1A and 1B are diagrams illustrating examples of pre-translation text data and post-translation text data. [Figure 3] 1A and 1B are diagrams illustrating examples of pre-translation text data and post-translation text data. [Figure 4] FIG. 10 is a schematic block diagram showing the configuration of an interpretation system according to a second embodiment of the present invention. [Figure 5]FIG. 10 is a schematic block diagram showing the configuration of an interpretation system according to a third embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0016] Next, a translation device according to an embodiment will be described in detail with reference to the drawings. Note that the following embodiment is merely an example and is not intended to limit the present invention. Also, the following drawings are schematic, and for the sake of explanation, some components may be omitted. Furthermore, parts common to the following embodiments are given the same reference numerals, and explanations thereof may be omitted.

[0017] [First embodiment] Fig. 1 is a schematic block diagram showing the configuration of an interpretation system according to a first embodiment of the present invention, Fig. 2 and Fig. 3 are diagrams showing examples of pre-translation text data and post-translation text data, respectively.

[0018] 1 shows a user terminal 101 and an interpretation system according to this embodiment. The interpretation system according to this embodiment is an interpretation system (AI hybrid interpretation system) in which a translator and a translation AI work together to efficiently improve the accuracy of interpretation.

[0019] The user terminal 101 is realized by a smartphone, tablet terminal, PC, etc. The user terminal 101 is equipped with a microphone, speaker, display, communication device, CPU (Central Processing Unit), memory, storage, etc. The user terminal 101 acquires the user's voice via the microphone as pre-translation voice data and transmits it to the voice recognition AI 102. The user terminal 101 also receives translated voice data from the voice conversion unit 104 and outputs it via the speaker. The user terminal 101 also receives translated text data from the translation AI 103 and displays it on the display.

[0020] The interpretation system according to this embodiment includes a speech recognition AI 102, a translation AI 103, and a speech conversion unit 104. The interpretation system according to this embodiment also includes an error recognition detection AI 111 and a speech recognition dictionary 112. The interpretation system according to this embodiment also includes an error translation detection AI 121 and a translation dictionary 122. The speech recognition AI 102, the translation AI 103, the speech conversion unit 104, the error recognition detection AI 111, the speech recognition dictionary 112, the error translation detection AI 121, and the translation dictionary 122 are each realized by a CPU, memory, storage, etc. The interpretation system according to this embodiment also includes a server 105 and a translator terminal 106.

[0021] The speech recognition AI 102 receives pre-translation speech data from the user terminal 101 via a communication device (not shown), performs speech recognition processing on the pre-translation speech data, and outputs pre-translation text data. The speech recognition AI 102 is an example of a speech recognition unit according to the present invention realized by AI (Artificial Intelligence).

[0022] The translation AI 103 receives pre-translation text data from the speech recognition AI 102, performs translation processing on the pre-translation text data, and outputs translated text data. The translation AI 103 is an example of a translation unit according to the present invention realized by AI.

[0023] The speech conversion unit 104 receives the translated text data from the translation AI 103, performs speech conversion processing on the translated text data, and outputs translated speech data.

[0024] The misrecognition detection AI 111 receives the pre-translation text data output from the speech recognition AI 102, extracts low-likelihood words from the pre-translation text data, and outputs them in association with the pre-translation text data. The misrecognition detection AI 111 is an example of an AI implementation of the misrecognition detection unit according to the present invention.

[0025] The mistranslation detection AI 121 receives the translated text data output from the translation AI 103, extracts words with low likelihood from the translated text data, and outputs them in association with the translated text data. The mistranslation detection AI 121 is an example of the mistranslation detection unit of the present invention realized by AI.

[0026] The server 105 stores in a readable format the pre-translation text data (associated with words with low likelihood) output from the misrecognition detection AI 111. The server 105 also stores in a readable format the translated text data (associated with words with low likelihood) output from the mistranslation detection AI 121. The server 105 is an example of a storage unit according to the present invention.

[0027] The translator terminal 106 is realized by a smartphone, tablet terminal, PC, etc. The translator terminal 106 is equipped with a display, a communication device, a CPU, memory, storage, etc. The translator terminal 106 is capable of reading pre-translation text data (associated with words with low likelihood) and post-translation text data (associated with words with low likelihood) in the server 105. The translator terminal 106 is also capable of registering data in the speech recognition dictionary 112 and the translation dictionary 122.

[0028] The translator terminal 106 is used by, for example, a translator.

[0029] The translator, for example, refers to pre-translation text data (associated with words with low likelihood) in the server 105 to determine the appropriateness of words determined to have low likelihood, and if the words are not appropriate, register appropriate words in the voice recognition dictionary 112. Thereafter, the voice recognition AI 102 refers to the voice recognition dictionary 112 to perform additional learning.

[0030] For example, in the example of FIG. 2, the misrecognition detection AI 111 determines that the likelihood of the word "reduced" in the pre-translation text data is low, and this word is highlighted in the pre-translation text data. The translator, for example, determines that the correct word for "reduced" is "low cost" and registers the word "low cost" in the speech recognition dictionary 112. Additional learning is then performed by the speech recognition AI 102. As a result, the next time a similar translation task occurs, it will be possible to correctly recognize the word "low cost," as shown in FIG. 3, for example.

[0031] Furthermore, the translator, for example, refers to the translated text data (associated with words with low likelihood) in the server 105 to determine the appropriateness of words determined to have low likelihood, and if the words are not appropriate, register appropriate words in the translation dictionary 122. Thereafter, the translation AI 103 refers to the translation dictionary 122 to perform additional learning.

[0032] For example, in the example of Figure 2, the mistranslation detection AI 121 determines that the words "reduction" and "reduced" in the translated text data have a low likelihood, and these words are highlighted in the translated text data. The translator, for example, determines that such words are the result of misrecognition by the speech recognition AI 102, and does not register them in the translation dictionary 122. In this example, as a result of additional learning by the speech recognition AI 102, the next time a similar translation task occurs, the word "kogenka" (low cost) will be correctly recognized, as shown in Figure 3, and as a result, the word "low cost" will be correctly output.

[0033] Here, in order to improve the accuracy of interpretation, it is conceivable to, for example, search for unnatural words in the pre-translation text data, and if an unnatural word is found, determine whether the word is appropriate, and if the word is not appropriate, register an appropriate word in the voice recognition dictionary 112 and perform additional learning of the voice recognition AI 102.

[0034] Also, for example, it is conceivable to search for unnatural words in the translated text data, and if an unnatural word is found, determine whether the word is appropriate, and if the word is not appropriate, register an appropriate word in the translation dictionary 122 and perform additional learning of the translation AI 103.

[0035] However, it is inefficient to search for unnatural words from a huge amount of pre-translation text data and post-translation text data.

[0036] Therefore, in the interpretation system according to this embodiment, the misrecognition detection units, misrecognition detection AI 111 and mistranslation detection AI 121, extract words with low likelihood from the pre-translation text data and the translated text data, and store them in a readable format in server 105. This method eliminates the need to search for unnatural words from huge amounts of pre-translation text data and translated text data, thereby enabling efficient improvement in the accuracy of interpretation.

[0037] [Second embodiment] FIG. 4 is a schematic block diagram showing the configuration of an interpretation system according to the second embodiment of the present invention.

[0038] The interpretation system according to the second embodiment is basically configured in the same manner as the interpretation system according to the first embodiment. However, the interpretation system according to the second embodiment further includes an interpreter terminal 107.

[0039] In the interpretation system according to the second embodiment, when the content of the interpretation is specialized and sufficient accuracy cannot be ensured, it is possible to switch from a method using the voice recognition AI 102, translation AI 103, etc. (hereinafter referred to as "automatic interpretation") to a voice call with an interpreter, etc.

[0040] The interpreter terminal 107 is realized by a smartphone, tablet terminal, PC, etc. The interpreter terminal 107 is equipped with a microphone, speaker, display, communication device, CPU, memory, storage, etc. The interpreter terminal 107 is capable of reading pre-translation text data (associated with words with low likelihood) and post-translation text data (associated with words with low likelihood) from the server 105. The interpreter terminal 107 is also capable of registering data in the speech recognition dictionary 112 and the translation dictionary 122.

[0041] The interpreter terminal 107 is used by, for example, an interpreter.

[0042] The interpreter performs a voice call with the user via a microphone and a speaker, and before or after the voice call starts, refers to the pre-translation text data (associated with the low-likelihood words) and the post-translation text data (associated with the low-likelihood words).

[0043] The interpretation system according to this embodiment can achieve the same effects as the interpretation system according to the first embodiment.

[0044] Furthermore, in the interpretation system according to this embodiment, the interpreter or the like can grasp the content of the conversation in advance, and it is possible to smoothly transition from automatic interpretation to a voice call with the interpreter or the like. Furthermore, it is also possible to select an interpreter or the like (for example, a highly specialized translator or the like) that is suitable for the content of the conversation based on the content of the conversation.

[0045] Furthermore, when switching from automatic interpretation to a voice call with an interpreter, etc., it is highly likely that misrecognition by the speech recognition AI 102 or mistranslation by the translation AI 103 occurred at the stage of automatic interpretation. Here, if there was an inappropriate part in the conversation up to that point when switching from automatic interpretation to a voice call with an interpreter, etc., it is conceivable that the interpreter will explain this to the user and begin interpretation after clearing up the misunderstanding. In order to determine whether there was an inappropriate part in the conversation up to that point, for example, it is conceivable to search for unnatural words in the pre-translation text data or the translated text data, and if an unnatural word is found, to determine whether the word is appropriate.

[0046] However, searching for unnatural words from vast amounts of pre-translation text data and post-translation text data is inefficient, and it is considered particularly difficult to perform such a task between the end of automatic interpretation and the transition to a voice call with an interpreter.

[0047] Therefore, in the interpretation system according to this embodiment, the misrecognition detection units, misrecognition detection AI 111 and mistranslation detection AI 121, extract words with low likelihood from the pre-translation text data and the translated text data, and store them in a readable format in server 105. This method eliminates the need to search for unnatural words from huge amounts of pre-translation text data and translated text data, thereby enabling efficient improvement in the accuracy of interpretation.

[0048] [Third embodiment] FIG. 5 is a schematic block diagram showing the configuration of an interpretation system according to the third embodiment of the present invention.

[0049] The interpretation system according to the third embodiment is basically configured in the same way as the interpretation system according to the second embodiment. However, the interpretation system according to the third embodiment does not include the misrecognition detection AI 111, the speech recognition dictionary 112, the mistranslation detection AI 121, the translation dictionary 122, and the translator terminal 106.

[0050] The server 105 according to the third embodiment also stores, in a readable format, the pre-translation text data output from the speech recognition AI 102. The server 105 according to the third embodiment also stores, in a readable format, the translated text data output from the translation AI 103.

[0051] In the interpretation system according to this embodiment, similar to the interpretation system according to the second embodiment, the interpreter or the like can grasp the content of the conversation in advance, and it is possible to smoothly transition from automatic interpretation to a voice call with the interpreter or the like. Furthermore, it is also possible to select an interpreter or the like (for example, a highly specialized translator or the like) suitable for the content of the conversation based on the content of the conversation.

[0052] [others] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims.

[0053] For example, each component of the above embodiments may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. [Explanation of symbols]

[0054] 101...user terminal, 102...speech recognition AI, 103...translation AI, 104...speech conversion unit, 105...server, 106...translator terminal, 111...misrecognition detection AI, 112...speech recognition dictionary, 121...mistranslation detection AI, 122...translation dictionary.

Claims

1. a speech recognition unit that performs speech recognition processing on pre-translation speech data acquired by a user terminal and outputs pre-translation text data; a translation unit that performs a translation process on the pre-translation text data and outputs translated text data to be transmitted to the user terminal; a misrecognition detection unit that extracts words with low likelihood from the pre-translation text data; a storage unit that stores the pre-translation text data and the low likelihood words in a readable format; a voice recognition dictionary that can be used for additional learning by the voice recognition unit; An interpretation system comprising:

2. a first terminal capable of registering the pre-translation text data and the low-likelihood words in the voice recognition dictionary by referring to the pre-translation text data and the low-likelihood words in the storage unit; 2. The interpretation system according to claim 1.

3. a speech recognition unit that performs speech recognition processing on pre-translation speech data acquired by a user terminal and outputs pre-translation text data; a translation unit that performs a translation process on the pre-translation text data and outputs translated text data to be transmitted to the user terminal; a mistranslation detection unit that extracts words with low likelihood from the translated text data; a storage unit that stores the translated text data and the low likelihood words in a readable format; a translation dictionary that can be used for additional learning by the translation unit; An interpretation system comprising:

4. a second terminal capable of registering the translated text data and the low-likelihood words in the translation dictionary by referring to the translated text data and the low-likelihood words in the storage unit; 4. The interpretation system according to claim 3.

5. a speech recognition unit that performs speech recognition processing on pre-translation speech data acquired by a user terminal and outputs pre-translation text data; a translation unit that performs a translation process on the pre-translation text data and outputs translated text data to be transmitted to the user terminal; a storage unit that stores at least one of the pre-translation text data and the translated text data in a readable format; An interpretation system comprising:

6. a third terminal capable of making a voice call with the user terminal and capable of referring to at least one of the pre-translation text data and the translated text data when making a voice call with the user terminal; An interpretation system according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Two-way speech translation system, two-way speech translation method and program

    JP2023022150A