Translation support device and program

The translation support device addresses limitations in conventional methods by automating the generation of translation memories through machine translation and preliminary memory usage, enhancing applicability and yield across diverse documents.

JP2025135499APending Publication Date: 2025-09-18JAPAN LICENSED TRANSLATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024033396
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-05
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Conventional translation methods require matching reference numbers in documents, limiting their applicability, and rely on bilingual dictionaries that are time-consuming and have low yield when selecting sentences with edit distances greater than a threshold.

Method used

A translation support device that automatically generates translation memories by machine-translating segments, reversing language direction, and using a preliminary translation memory to pre-translate and register sentence pairs based on matching criteria, reducing manual intervention.

Benefits of technology

Enables the construction of translation memories across various documents without manual manipulation, improving yield and efficiency by automating the process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025135499000001_ABST
    Figure 2025135499000001_ABST
Patent Text Reader

Abstract

To provide a translation support device and program capable of automatically generating at a high yield a translation memory and a term base that do not require artificial operation, have high versatility, and can be quickly acquired in a wide range of fields.SOLUTION: A translation support device acquires original text data in a first language, which is to be translated, and translated text data in a second language, which is a result of translating the original text data, divides the original text data into predetermined segment units and machine-translates at least some of segments obtained by the division into the second language to generate a preliminary translation memory from data obtained by reversing a language direction of a sentence pair related to the segments of the original text data and the machine translation result, divides the acquired translated data of the second language into predetermined segment units to execute pre-translation by the generated preliminary translation memory, and automatically generates a translation memory on the basis of the original text data of the second language and the sentence pair obtained from the translation result by the preliminary translation memory.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a translation support device and a program for supporting translation. [Background technology]

[0002] Patent document 1 discloses an example in which Japanese patent documents and their corresponding US patent documents are used as a bilingual corpus, corresponding documents selected from this bilingual corpus are divided into sentences using their respective delimiters, reference numbers in the divided sentences are extracted, and bilingual sentence pairs are extracted based on the co-occurrence of the extracted reference numbers. As another example, Patent Document 2 discloses an example in which a bilingual text acquisition unit of a bilingual text extraction device matches sentences in a first language and sentences in a second language that constitute a bilingual document using a bilingual dictionary to acquire one or more bilingual texts in the first language and the second language, a translation model generation unit generates a translation model based on the one or more bilingual texts, a translation unit translates the sentences in the first language that constitute the bilingual text into the second language using the generated translation model for each of the one or more bilingual texts, an edit distance calculation unit calculates the edit distance between the sentence in the first language translated into the second language and the sentence in the second language that corresponds to the sentence in the first language for each of the one or more bilingual texts, and a bilingual text selection unit selects from the one or more bilingual texts those bilingual texts whose calculated edit distance is greater than a threshold. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-348514 [Patent Document 2] Japanese Patent Application Publication No. 2018-32324 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the above conventional example requires documents such as patent documents that have reference numbers, and processing is only possible when the reference numbers match between the original text and the translation, which limits the scope of use. In another example, the use of a bilingual dictionary limited the fields covered and took a long time to generate a language model. Furthermore, there were problems with low yield when selecting bilingual sentences with edit distances greater than a threshold.

[0005] The present invention has been made in consideration of the above-mentioned circumstances, and one of its objectives is to provide a translation support device and program that can automatically generate translation memories and term bases with a high yield rate, which do not require manual operation, are highly versatile, and can be quickly acquired in a wide range of fields. [Means for solving the problem]

[0006] One aspect of the present invention that solves the problems of the above-mentioned conventional examples is a translation support device that acquires original data in a first language to be translated and translated data in a second language that is the result of translating the original data, divides the original data into predetermined segment units, machine-translates at least some of the resulting segments into the second language, generates a preliminary translation memory from data in which the language direction of sentence pairs related to those segments of the original data and the machine-translated results is reversed, divides the acquired translation data in the second language into predetermined segment units, performs pre-translation using the generated preliminary translation memory, and automatically generates a translation memory based on the sentence pairs obtained from the original data in the second language and the translation results using the preliminary translation memory. [Effects of the Invention]

[0007] According to the present invention, a translation memory can be constructed for a wide range of documents, not limited to patent documents, while reducing the need for manual manipulation. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram illustrating an example of the configuration of a translation support device according to an embodiment of the present invention. [Figure 2] 1 is a functional block diagram illustrating an example of a translation support device according to an embodiment of the present invention. [Figure 3] FIG. 2 is a flowchart illustrating an example of the operation of the translation support device according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0009] An embodiment of the present invention will be described with reference to the drawings. A translation support device 1 according to an embodiment of the present invention is realized using a general computer device including a control unit 11, a storage unit 12, an operation unit 13, a display unit 14, and a communication unit 15, as shown in FIG.

[0010] The control unit 11 includes a program control device such as a processor, and operates according to a program stored in the storage unit 12. By processing in accordance with this program, the control unit 11 of this embodiment acquires original data in a first language to be translated and translated data in a second language that is the translation result of the original data, divides the acquired original data into predetermined segment units, and machine-translates at least some of the segments obtained by division into the second language to generate second translated data.

[0011] The control unit 11 also generates text files with the language direction reversed for pairs of original text data and second translated text data, which is the result of machine translation, relating to segments machine-translated into the second language (pairs of original text and its machine translation result for each corresponding segment). Here, the language direction means the direction from original text (before translation) to translated text (after translation), and reversing the language direction means that the second language is the language of the original text and the first language is the language of the translation.

[0012] That is, here, control unit 11 registers the second translated data as new original data (referred to as second original data for distinction) and the original original data as translated data (referred to as corresponding translated data for distinction) in the translation memory. The translation memory generated here will be referred to as the "preliminary translation memory" below.

[0013] Here, the "translation memory" is a database that stores translated sentences etc. in association with original sentences etc., and an existing database that functions as a translation memory may be used here.

[0014] The control unit 11 divides the translation data in the second language acquired together with the original data into segments, performs translation using the preliminary translation memory generated above (called pre-translation because it uses the preliminary translation memory), and generates translation data in the first language (called third translation data).

[0015] The control unit 11 then compares the third translated data obtained by this preliminary translation with the acquired original data on a segment-by-segment basis, and for segments where the comparison results meet predetermined criteria, registers in the translation memory a sentence pair (parallel data) that associates that segment of the acquired original data with that segment of the third translated data. The detailed operation of this control unit 11 will be described later.

[0016] The storage unit 12 includes a memory device and a disk device, and stores programs to be executed by the control unit 11. The operation unit 13 includes a keyboard, a mouse, etc., and receives user operations and outputs information indicating the content of the operations to the control unit 11.

[0017] The display unit 14 is a display or the like, and displays and outputs information in accordance with instructions input from the control unit 11. The communication unit 15 is a network interface or the like, and outputs information received via a network to the control unit 11. The communication unit 15 also sends information via the network in accordance with instructions input from the control unit 11.

[0018] Next, a description will be given of the operation of the control unit 11 of this embodiment. The control unit 11 of this embodiment operates in accordance with a program stored in the storage unit 12, thereby achieving a configuration that functionally includes an acquisition unit 21, a machine translation unit 22, and a data processing unit 23, as shown in FIG.

[0019] The acquisition unit 21 acquires original text data O in a first language to be translated, and translation text data T in a second language different from the first language, which is the translation result of the original text data O.

[0020] The machine translation unit 22 divides the source data into predetermined segment units. Here, the segment units may be, for example, sentence units, and may be divided by periods in English and periods in Japanese. Note that this division process can be performed using a method well known for each language, so a detailed description will be omitted here.

[0021] The machine translation unit 22 machine-translates each of the original data elements Oi (i=1, 2, ...) obtained by dividing the original data O into segments into sentences in the second language to obtain sentence pairs. This machine translation may be performed by a known machine translation system (NMT) that uses a neural network, for example.

[0022] Here, the machine-translated translation of the original data O obtained by the machine translation unit 22 is referred to as second translated data D to distinguish it from the translation data T obtained by the acquisition unit 21. The second translated data D includes second translated data elements Di (i = 1, 2, ...) corresponding to the original data elements Oi (i = 1, 2, ...) for each segment.

[0023] The data processing unit 23 sets up a preliminary translation memory by temporarily registering the segment Di of the second original data and the corresponding segment Oi of the corresponding translated data in a translation memory, with the second translated data as the original (second original data) and the original original data as the translated data (corresponding translated data).

[0024] The data processing unit 23 divides the translation data T acquired by the acquisition unit 21 into segments to obtain translation data elements Ti (i=1, 2, ...). The data processing unit 23 translates (pre-translates) the translation data element Ti for each segment using the pre-translation memory set above, and generates third translation data X including third translation data elements Xi (i=1, 2, ...) corresponding to each translation data element Ti as the pre-translation result.

[0025] The data processing unit 23 then compares this third translation data X with the original data O. Specifically, the data processing unit 23 compares the third translation data element Xi with the corresponding original data element Oi, and determines whether the comparison result meets a predetermined standard.

[0026] As an example, the data processing unit 23 calculates the match rate between the third translated data element Xi and the corresponding original data element Oi (a match rate determined based on word match, word order match, etc., and may be a match rate normally used in so-called translation memory tools).

[0027] The data processing unit 23 then associates pairs of third translated data elements Xi and original data elements Oi whose matching rates satisfy a predetermined standard (for example, the matching rate is within a predetermined range) with each other, and registers them as bilingual data in a translation memory tool (a server that provides translation memory).

[0028] In this embodiment, a translation memory can be constructed for a wide range of documents while reducing the need for manual manipulation.

[0029] [Operation] Next, the operation of the translation support device 1 of this embodiment will be described. The translation support device 1 of this embodiment generates bilingual data to be registered in a translation memory. Specifically, this translation support device 1 receives input of original text data O in a first language to be translated and translated text data T in a second language different from the first language, which is the translation result of this original text data O, and starts the processing illustrated in FIG.

[0030] In the example of this embodiment, it is assumed that this translation data T has been manually confirmed and approved as a translation of the original data O. This translation data will be hereinafter referred to as "first translation data."

[0031] The translation support device 1 divides the original data O into original data elements Oi (i = 1, 2, ...) which are predetermined segment units (S11), and machine-translates at least some (here, all) of the original data elements Oi, which are segment units, into translation elements (second translation data elements) Di in a second language using a known machine translation system such as a neural network translation system (NMT) (S12). The accuracy of this machine translation does not necessarily have to be equivalent to that of a human translation. Hereinafter, the machine-translated translation (second translation data) D obtained in step S12 includes second translation data elements Di (i = 1, 2, ...) which are segment units corresponding to the original data elements Oi.

[0032] The translation support device 1 provisionally registers the second translated text data element Di and the corresponding original text data element Oi as a new original text and a new translated text (i.e., reversing the language direction) in the translation memory, and sets the translation memory (preliminary translation memory) (S13).

[0033] Furthermore, the translation support device 1 divides the input first translated sentence data T into first translated sentence data elements Ti (i=1, 2, . . . ) in segment units (S14).

[0034] The translation support device 1 sequentially translates (pre-translates) the first translated data elements Ti obtained by division using a preliminary translation memory (S15), and generates, as a result of this pre-translation, third translated data X including third translated data elements Xi (i=1, 2...) corresponding to each translated data element Ti.

[0035] The translation support device 1 compares the third translated text data element Xi with the corresponding original text data element Oi for each segment, and determines whether the comparison result satisfies a predetermined standard (S16).

[0036] If, in step S16, the translation support device 1 determines that the matching rate determined from word matching or the like satisfies a predetermined standard (for example, a matching rate of 65% to 85%) (S16: Yes), it associates the third translated text data element Xi with the corresponding original text data element Oi and records them as bilingual data (S17).Then, the translation support device 1 returns to step S16 and continues processing for the next segment (S18).

[0037] If the matching rate does not satisfy the predetermined standard in step S16 (S16: No), the translation support device 1 proceeds to step S18 and continues the process.

[0038] The translation support device 1 registers the bilingual data recorded by this process in a translation memory. In this embodiment, this makes it possible to configure a translation memory while reducing the need for manual operations for various documents, including documents for which bilingual translations are difficult to find.

[0039] [Addition and correction processing] Furthermore, the translation support device 1 of this embodiment may supplement and correct the bilingual data by a neural network using a large-scale language model. Specifically, the translation support device 1 may process either the first language segment or the second language segment recorded as bilingual data.

[0040] The translation support device 1 sends pairs of first language segments and second language segments contained in the bilingual data (hereinafter referred to as "target pairs") to a neural network using a large-scale language model (for example, OpenAI's GPT3.5 service may be used), and requests that at least one of these first language segments and second language segments be "corrected if the bilingual translation is inappropriate."

[0041] The translation support device 1 associates the corrected first language segment and second language segment, which are the response of the neural network using a large-scale language model to this instruction, replaces them with the noteworthy pair in the bilingual data, records them, and updates the bilingual data.

[0042] According to this example, if paragraph numbers (such as paragraph numbers given to patent documents, etc.) or control characters that are inappropriate for inclusion in the translation are present in either the original data input to the translation support device 1 of this embodiment or the corresponding third translation data, it is expected that these will be removed.

[0043] In fact, if a first-language segment extracted from source data includes a paragraph number, such as "

[0007] Fig. 1 is a schematic view of a supporting member according to an embodiment of the present disclosure.", and the corresponding second-language segment has the paragraph number removed, such as "[Figure 1 is a schematic view of a supporting member according to an embodiment of the present disclosure.", when a neural network using a large-scale language model is asked to "correct inappropriate translations" as described above, the neural network using the large-scale language model will delete the paragraph number included in the first-language segment, which is the symbol portion of the first-language segment and the second-language segment that does not correspond to each other, and correct the first-language segment to "Fig. 1 is a schematic view of a supporting member according to an embodiment of the present disclosure."

[0044] The translation support device 1 generates the corrected first language segment: "Fig.1 is a schematic view of a supporting member according to an embodiment of the present disclosure." and a segment in the second language: FIG. 1 is a schematic diagram of a support member according to one embodiment of the present disclosure. and are associated with each other, replaced with the original sentence pair, and recorded as part of the bilingual data. According to this example of the present embodiment, it is possible to reduce the amount of manual correction work.

[0045] [Example 1] Japanese was selected as the first language and English as the second language, and as source data, a Japanese patent specification was prepared as an original data file in Microsoft docx file format, and this original data was manually translated into an English specification in Microsoft docx file format. Using the translation support device of the present invention, a project based on the original data was first automatically generated using the device's API. Each generated segment was then automatically translated using a single language model. The general-purpose NT engine provided by the National Institute of Information and Communications Technology (NICT) was used as the machine translation engine. The language direction of this bilingual text file was reversed, and a translation memory was generated from the machine-translated English translation and the original Japanese translation. As an example of bilingual translation, the following tab-delimited format is used: <tab>An example of a bilingual translation of the Japanese source data in the first language is as follows: [Example 1] REMOTE SUPPORT SYSTEM AND METHOD FOR CONTROLLING REMOTE SUPPORT SYSTEM <tab>Remote support system and control method for remote support system The translation memory tmx (tmx version="1.4") for the above translation was as follows: [Example 2] <tu tuid=""> <tuv xml:lang="en"> <seg> REMOTE SUPPORT SYSTEM AND METHOD FOR CONTROLLING REMOTE SUPPORT SYSTEM< / seg> < / tuv> <tuv xml:lang="ja" creationdate="2024218T11523Z" changedate="2024218T11523Z"> <prop type="filename"> filename.txt< / prop> <seg> Remote support system and control method for remote support system< / seg> < / tuv> < / tu> <tu tuid=""> <tuv xml:lang="en"> <seg> remote support system and control method for remote support system< / seg> < / tuv> <tuv xml:lang="ja" creationdate="2024218T11523Z" changedate="2024218T11523Z"> <prop type="filename"> filename.txt< / prop> <seg> Remote support system and control method for remote support system< / seg> < / tuv>

[0046] A new project using English translation data was automatically created using the translation support device via API. At this time, the translation memory was specified, and after the project was created, pre-translation using the translation memory was automatically performed via API. 50% was specified as the threshold for the translation memory. In this way, sentence pairs of English translation data and original Japanese data were obtained. When the Japanese sentences were judged to be appropriate for the English sentences, it was found that 87% of all segments were correct.

[0047] [Example 2] Japanese was selected as the first language and English as the second language. The source data consisted of a Japanese patent specification in a Microsoft .docx file format, and a translation of the original data into an English specification in a Microsoft .docx file format. A project was then created using the source data using the translation support device of the present invention. Each generated segment was translated using a single language model. Three machine translation engines were used: a general-purpose NT engine provided by the National Institute of Information and Communications Technology (NICT), a patent NT engine, and a science engine. The language direction of this bilingual text file was reversed, and a translation memory was generated from the machine-translated English translation and the original Japanese translation.

[0048] As an example of a parallel translation, the following tab-delimited machine-translated sentence in the second language (English) is <tab>An example of a bilingual translation of the Japanese original data in the first language is as follows: [Example 3] REMOTE SUPPORT SYSTEM AND METHOD FOR CONTROLLING REMOTE SUPPORT SYSTEM <tab>Remote support system and control method for remote support system remote support system and control method for remote support system <tab>Remote support system and control method for remote support system REMOTE SUPPORT SYSTEM AND METHOD FOR CONTROLLING REMOTE SUPPORT SYSTEM <tab>Remote support system and control method for remote support system

[0049] Here, since there are three machine translation engines for the same Japanese sentence, the number of sentence pairs has increased threefold. The translation memory tmx (tmx version="1.4") for the above translation was as follows: [Example 4] <tu tuid=""> <tuv xml:lang="en"> <seg> REMOTE SUPPORT SYSTEM AND METHOD FOR CONTROLLING REMOTE SUPPORT SYSTEM< / seg> < / tuv> <tuv xml:lang="ja" creationdate="2024218T11523Z" changedate="2024218T11523Z"> <prop type="filename"> filename.txt< / prop> <seg> Remote support system and control method for remote support system< / seg> < / tuv> < / tu> <tu tuid=""> <tuv xml:lang="en"> <seg> remote support system and control method for remote support system< / seg> < / tuv> <tuv xml:lang="ja" creationdate="2024218T11523Z" changedate="2024218T11523Z"> <prop type="filename"> filename.txt< / prop> <seg> Remote support system and control method for remote support system< / seg> < / tuv> < / tu> <tu tuid=""> <tuv xml:lang="en"> <seg> REMOTE SUPPORT SYSTEM AND METHOD FOR CONTROLLING REMOTE SUPPORT SYSTEM< / seg> < / tuv> <tuv xml:lang="ja" creationdate="2024218T11523Z" changedate="2024218T11523Z"> <prop type="filename"> filename.txt< / prop> <seg> Remote support system and control method for remote support system< / seg> < / tuv> < / tu>

[0050] A new project using English translation data was automatically created using the translation support device via API. At this time, the translation memory was specified, and after the project was created, pre-translation using the translation memory was performed. 50% was specified as the translation memory threshold. In this way, sentence pairs of English translation data and original Japanese data were obtained. When the Japanese sentences were judged to be appropriate for the English sentences, 98% of all segments were found to be correct, indicating a high yield.

[0051] [Example 3] The bilingual data obtained in Examples 1 and 2 may contain noise such as that shown below. [Translation 1] [[1][2][3][4}0007{5][6]][7]Fig. 1 is a schematic view of a remote support system according to an embodiment of the present disclosure. <tab>FIG. 1 is a schematic diagram of a remote support system according to one embodiment of the present disclosure. [Translation 2] [1]The management apparatus includes: <tab>The management device

[0052] Because it is impossible to predict in advance whether such tag information will be included, expert review and correction are essential. As in the example of Bilingual Translation 2, simply deleting the initial symbol would result in "the management apparatus includes:," but because the initial "the" is at the beginning of the sentence, it must be corrected to "The management apparatus includes:." Because such corrections are difficult to identify as patterns, expert review and correction have been necessary until now.

[0053] Here, the following processing is performed on the parallel translations obtained by matching alone, eliminating the need for symbolic processing by experts. First, the following prompt is created based on the following roughly accurate parallel translations. Example prompt: Incorrect translations where the English and Japanese do not correspond should be revised to create good translations.\n[[1][2][3][4}0007{5][6]][7]Fig. 1 is a schematic view of a remote support system according to an embodiment of the present disclosure. <tab>FIG. 1 is a schematic diagram of a remote support system according to one embodiment of the present disclosure. <tab>The management device

[0054] By processing the above prompt using OpenAI's GPT4 model, we were able to obtain the following results.

[0055] Fig. 1 is a schematic view of a remote support system according to an embodiment of the present disclosure. <tab>FIG. 1 is a schematic diagram of a remote support system according to one embodiment of the present disclosure. <tab>The management device

[0056] The response from OpenAI above revealed that symbols that do not correspond to the translation have been deleted, and that the first letters of English sentences have been converted to uppercase, so it is possible to use the translation as is.

[0057] [Generate terminology data] Furthermore, the translation support device 1 may present the first language segments and second language segments contained in the bilingual data to a neural network using a large-scale language model, requesting it to "extract a term base" and generate a term database.

[0058] The translation support device 1 obtains a term base (a list associating words in the first language with corresponding words in the second language) as a response to the request from a neural network using a large-scale language model.

[0059] In this case, the translation support device 1 may not only "extract a term base," but also "extract a term base and remove unnecessary symbols, etc." According to this example of the present embodiment, a term base can be generated with relatively little manual work.

[0060] [Example 4] As an example of a bilingual sentence pair obtained by the method of Example 2, the following will be taken up and the term base extraction will be described in detail. [Example 5] The image forming apparatus is connected to a support center that provides support for the use of the image forming apparatus. <tab>For example, such an image forming apparatus is connected to a support center that supports use of the image forming apparatus.

[0061] Based on the bilingual data, the following prompts were used for the large-scale language model to extract the term base: Example prompt: Extract the term base to be used in the translation support tool based on the following bilingual translations. Delete unnecessary symbols from the term base as well. The output should start with [Term Base] and be in the order of Japanese term\tEnglish translation.\n\nFor example, such an image forming apparatus is connected to a support center that supports use of the image forming apparatus.

[0062] By processing the above prompts using OpenAI's GPT4 model, we were able to obtain the following term base: [Example 6] Termbase Image forming apparatus support use of Support Center connected to

[0063] Generally, verbs are not needed as a terminology database, so if the above is left as it is, it will be necessary to manually delete unnecessary terminology databases.

[0064] [Example 5] For the same translation as in Example 3, the following was used as a prompt for the large-scale language model for term base extraction: Example prompt: Based on the following translations, extract a term base for nouns containing compound words and phrases to be used in translation support tools. Unnecessary symbols should also be deleted from the term base. The output should start with [Term Base] and be in the order of Japanese term and English translation.\n\nFor example, such an image forming apparatus is connected to a support center that supports use of the image forming apparatus.

[0065] By processing the above prompt using OpenAI's GPT4 model, we were able to obtain the following terminology database: Termbase Image forming apparatus use use Support Center

[0066] In this example, unnecessary verb parts can be deleted from the term database, which eliminates the need for manual deletion of unnecessary term bases.

[0067] [Effects seen in examples] All of Examples 1 to 5 can be executed without human operation by the API, which is an implementation example of the translation support device 1 of this embodiment, and by specifying the source data file and translation data file, it is possible to obtain a translation memory and term base with an accuracy equivalent to that of an expert's confirmation and correction level, without the need for subsequent human work. By automating the work of creating translation memories and term bases, which until now has been done by experts, it has become possible to greatly improve the productivity of creating translation memories and term bases. [Explanation of symbols]

[0068] 1 Translation support device, 11 Control unit, 12 Memory unit, 13 Operation unit, 14 Display unit, 15 Communication unit, 21 Acquisition unit, 22 Machine translation unit, 23 Data processing unit.< / tab> < / tab> < / tab> < / tab> < / tab> < / tab> < / tab> < / tab> < / tab> < / tab> < / tab> < / tu> < / tab> < / tab>

Claims

1. Acquire original data in a first language to be translated and translation data in a second language that is a translation result of the original data; Dividing the source data into predetermined segments, machine-translating at least some of the segments obtained by the division into the second language, and generating a preliminary translation memory from data obtained by reversing the language direction of sentence pairs related to the segments of the source data and the machine-translated results; Dividing the acquired second language translation data into predetermined segments and performing pre-translation using the generated pre-translation memory; A translation support device that automatically generates a translation memory based on sentence pairs obtained from original text data in a second language and translation results obtained using the preliminary translation memory.

2. 2. The translation support device according to claim 1, The translation support device further includes means for acquiring the second translated sentence data by a plurality of translation engines consisting of a plurality of neural network machine translation models.

3. 2. The translation support device according to claim 1, The translation support device further includes means for deleting non-corresponding symbols in the bilingual data of the translation memory using a large-scale language model.

4. 2. The translation support device according to claim 1, A translation support device that generates a term database from the bilingual data using a large-scale language model.

5. 4. A translation support device according to claim 1, The predetermined criterion is that the comparison result is within a predetermined range of matching rate.

6. Computer, an acquisition means for acquiring original data in a first language to be translated and translation data in a second language that is a translation result of the original data; means for dividing the source data into predetermined segments, machine-translating at least some of the segments obtained by dividing into the second language, and generating a preliminary translation memory from a text file in which the language direction of the sentence pair relating to the segments of the source data and the machine-translated results is reversed; means for dividing the acquired second language translation data into predetermined segments and performing pre-translation using the generated pre-translation memory; means for automatically generating a translation memory based on source text data in a second language and sentence pairs obtained from a pre-translation result using the pre-translation memory; A program that functions as a

Citation Information

Patent Citations

  • Parallel translation word extraction method, parallel translation word dictionary construction method, and translation memory construction method

    JP2004348514A

  • Parallel translation extraction apparatus, parallel translation extraction method, and program

    JP2018032324A