Post-editing support system, post-editing support method, post-editing support device, and computer program

The post-editing support system addresses the burden of neural machine translation inconsistencies by generating a unified third text, thereby simplifying the post-editing process.

JP7751906B2Active Publication Date: 2025-10-09NGB 가부시키가이샤
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024112508
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2025-10-09
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

Post-editing of machine-translated texts using neural machine translation models is burdensome due to inconsistent translations, requiring post-editors to address unique inconsistencies not present in human translations.

Method used

A post-editing support system and method that generates a third text where multiple translations for a single original word are unified, displaying it in an editable state to reduce the burden on post-editors.

Benefits of technology

This system reduces the need for post-editors to focus on translation inconsistencies, streamlining the post-editing process by presenting consistent translations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007751906000001
    Figure 0007751906000001
  • Figure 0007751906000002
    Figure 0007751906000002
  • Figure 0007751906000003
    Figure 0007751906000003
Patent Text Reader

Abstract

To reduce a burden of a post-editing work for a text mechanically translated using a neural machine translation model.SOLUTION: A translation device 11 translates a first text T1 written in a first language into a second text T2 written in a second language different from the first language using a neural machine translation model 111. When a plurality of different translated words included in the second text T2 corresponds to one original word included in the first text T1, a post-editing support device 12 generates a third text T3 in which one translated word included in the plurality of different translated words replaces the remaining translated words included in the plurality of translated words, and displays the third text T3 on a display device 13. The replaced translated word is displayed on the display device 13 in a manner distinguishable from the other translated words included in the third text T3.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a system and method for supporting post-editing of machine-translated text, and also to an apparatus configured to support post-editing of machine-translated text, and a computer program executable by a processing unit of the apparatus. [Background technology]

[0002] As disclosed in Patent Document 1, neural machine translation models are becoming increasingly popular. A neural machine translation model is a machine translation method that performs translation modeling in an end-to-end manner by directly using a neural network. Because neural machine translation models learn as a consistent model from receiving an original text to outputting a translation, they are known to have superior translation accuracy and translation fluency compared to conventional statistical machine translation. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2020-140709 Summary of the Invention [Problem to be solved by the invention]

[0004] However, for those who check for mistranslations in machine-translated text and edit it as necessary (so-called post-editors), this can sometimes impose a different burden than when checking text translated by a human.

[0005] For example, texts machine-translated using neural machine translation models often have multiple translations corresponding to a specific word in the source text. Because such inconsistencies are rare in human-translated texts, post-editors must pay attention to details that would not otherwise be necessary. Since neural machine translation models fundamentally have difficulty identifying the causal relationship between input and output seen in rule-based statistical machine translation, the occurrence of inconsistencies in translations is currently accepted as "behavior unique to neural machine translation models."

[0006] The object of the present invention is to reduce the burden of post-editing work on texts that have been machine translated using neural machine translation models. [Means for solving the problem]

[0007] One aspect of the present invention to achieve the above object is a post-editing support system, comprising: a translation device that translates a first text written in a first language into a second text written in a second language different from the first language using a neural machine translation model; a post-editing support device that generates a third text in which, when a plurality of different translated words included in the second text correspond to one original word included in the first text, one translated word included in the plurality of different translated words replaces the remaining translated words included in the plurality of translated words, and displays the third text on a display device in an editable state; It is equipped with:

[0008] One aspect of the present invention to achieve the above object is a post-editing support method, comprising: translating a first text written in a first language into a second text written in a second language different from the first language using a neural machine translation model; determining whether a plurality of different translations included in the second text correspond to a single original word included in the first text; generating a third text in which one of the translations included in the plurality of different translations is replaced with one of the translations included in the plurality of different translations when it is determined that the plurality of translations correspond to the one original language; displaying the third text on a display device in an editable state; It is equipped with:

[0009] One aspect of the present invention to achieve the above object is a post-editing support device, comprising: a first reception unit that receives a first text written in a first language; a second receiving unit that receives second text from a translation device that translates the second text into a second text written in a second language different from the first language using a neural machine translation model; a processing unit that, when a plurality of different translated words included in the second text correspond to one original word included in the first text, generates a third text in which one translated word included in the plurality of different translated words is substituted for the remaining translated words included in the plurality of translated words, and displays the third text on a display device in an editable state; It is equipped with:

[0010] One aspect of the present invention to achieve the above object is a computer program executable by a processing unit of a post-editing support device, By executing this, the post-editing support device Accepting a first text written in a first language; receiving second text from a translation device that translates the second text into a second language different from the first language using a neural machine translation model; determining whether a plurality of different translations contained in the second text correspond to a single original word contained in the first text; If it is determined that the plurality of translated words correspond to the one original word, a third text is generated in which one translated word included in the plurality of different translated words is substituted for the remaining translated words included in the plurality of different translated words; The third text is displayed on a display device in an editable state.

[0011] According to the configurations of the above aspects, the third text with consistent translations is displayed on the display device, regardless of the neural machine translation model used in the translation device. This relieves the post-editor from the need to pay attention to the possibility that the translation may not use consistent translations. This reduces the burden of post-editing work on text machine-translated using a neural machine translation model. [Brief explanation of the drawings]

[0012] [Figure 1] 1 illustrates an example of a functional configuration of a post-editing support system according to an embodiment. [Figure 2] 2 illustrates an example of a flow of processing executed in the post-editing support device of FIG. 1. [Figure 3] 2 illustrates an example of a second text output by the translation device of FIG. 1. [Figure 4] 2 shows an example of a third text displayed on the display device of FIG. 1. [Figure 5] 2 illustrates an example of a flow of processing executed in the post-editing support device of FIG. 1. [Figure 6] 2 illustrates an example of a flow of processing executed in the post-editing support device of FIG. 1. [Figure 7] An example is shown to explain the processing illustrated in FIGS. [Figure 8] 10 shows another example of the third text displayed on the display device of FIG. [Figure 9] 9 illustrates a post-editing process performed on the third text of FIG. 8. [Figure 10] 9 illustrates a post-editing process performed on the third text of FIG. 8. [Figure 11] 9 illustrates a post-editing process performed on the third text of FIG. 8. DETAILED DESCRIPTION OF THE INVENTION

[0013]

[0023] An example embodiment will now be described in detail with reference to the accompanying drawings. Fig. 1 illustrates the functional configuration of a post-editing support system 10 according to one embodiment. The post-editing support system 10 includes a translation device 11, a post-editing support device 12, and a display device 13.

[0014] The translation device 11 is configured to translate a first text T1 written in a first language into a second text T2 written in a second language different from the first language using a neural machine translation model 111. That is, the first text T1 includes an original text, and the second text T2 includes a translation text. The combination of the first language and the second language can be arbitrarily selected from a plurality of languages ​​supported by the neural machine translation model 111. The first language is, for example, Japanese. The second language is, for example, English.

[0015] Examples of the neural machine translation model 111 include a sequence-to-sequence (seq2seq) model, a convolutional sequence-to-sequence (ConvS2S) model, a SliceNet model, a Transformer model, and an (RNMT+) model. The neural machine translation model 111 may be commercially available or not, as long as it is based on a method that performs translation modeling in an end-to-end manner. The neural machine translation model 111 includes models that are available for free or for a fee via a communication network.

[0016] The neural machine translation model 111 may include a trained model that has been adapted to improve the translation accuracy of phrases and expressions specific to a particular field. For example, documents related to intellectual property, investor relations (IR) documents, legal documents, product manuals, etc. tend to contain repetitive formulaic expressions and sentences. The neural machine translation model 111 preferably includes a trained model for translating sentences with such tendencies.

[0017] The second text T2 resulting from machine translation by the translation device 11 may contain mistranslations, omissions, duplicate translations, etc. Therefore, it is common for such errors to be manually discovered and corrected. This process is called post-editing. The person who performs post-editing is sometimes called a post-editor.

[0018] The post-editing support device 12 is a device for supporting post-editing by a post-editor, and includes a first reception unit 121, a second reception unit 122, a processing unit 123, an output unit 124, and an edit reception unit 125.

[0019] The first receiving unit 121 is configured as an interface that receives the first text T1 as input data. The interface may be a physical interface or a logical interface.

[0020] The second receiving unit 122 is configured as an interface that receives, as input data, the second text T2 output from the translation device 11. The interface may be a physical interface or a logical interface.

[0021] When multiple different translated words included in the second text T2 correspond to one original word included in the first text T1, the processing unit 123 is configured to generate a third text T3 in which one translated word included in the multiple different translated words replaces the remaining translated words included in the multiple different translated words. Details of this processing will be described later.

[0022] Additionally, the processing unit 123 is configured to display the third text T3 in an editable state on the display device 13. Specifically, the processing unit 123 outputs data corresponding to the third text T3 from the output unit 124. The display device 13 has a screen for displaying the third text T3 based on the data output from the output unit 124.

[0023] The post-editor checks the content of the third text T3 displayed on the display device 13 and performs post-editing as necessary. The edit receiving unit 125 is configured as an interface that receives input corresponding to post-editing. The input may be made via an appropriate man-machine interface such as a keyboard, a mouse, a touch panel, or a touch pad, or may be made via voice recognition technology or gesture recognition technology.

[0024] The processing unit 123 is configured to execute a process of reflecting the content of the post-editing accepted by the edit accepting unit 125 in the third text T3. The processing unit 123 is configured to output data corresponding to the third text T3 in which the post-processing has been reflected from the output unit 124, and to display the post-processed third text T3 on the display device 13.

[0025] The above processing executed by the processing unit 123 will be specifically described with reference to FIGS.

[0026] The processing unit 123 receives the first text T1 through the first receiving unit 121 (STEP 11). The expression "receiving the first text T1" used in this specification includes the meaning of receiving data corresponding to the first text T1. The timing of receiving the first text T1 may be before or after the translation device 11 translates the first text T1 into the second text T2.

[0027] Next, the processing unit 123 receives the second text T2 from the translation device 11 through the second receiving unit 122 (STEP 12). The expression "receiving the second text T2" used in this specification includes the meaning of receiving data corresponding to the second text T2. The processing of STEP 11 and the processing of STEP 12 may be performed in parallel, or the order may be reversed.

[0028] Next, the processing unit 123 determines whether multiple different translations contained in the second text T2 correspond to one original word contained in the first text T1 (STEP 13). As described above, in translations generated using a neural machine translation model, multiple different translations may be assigned to the same language without any regularity. This process is performed to detect such inconsistencies in translations.

[0029] 3 illustrates second text T2 translated based on first text T1 input to translation device 11. In this example, the original word "light emitting element" included in first text T1 is assigned different translations, "light emitting element," "light emitter," and "photo emitting element" (sentence numbers 1, 2, 5, and 6). Also, the original word "detected" included in first text T1 is assigned different translations, "sensed" and "detected" (sentence numbers 3 and 5). Therefore, processing unit 123 determines that the multiple different translations included in the second text correspond to one original word included in first text T1 (YES in STEP 13).

[0030] In this case, the processing unit 123 performs processing to generate a third text T3 (STEP 14). In the third text T3, one of the multiple different translations contained in the second text T2 is substituted for the remaining translations, thereby unifying the translations. Note that the expression "generating the third text T3" used in this specification also means generating data corresponding to the third text T3. Rules for replacing translations will be described later.

[0031] FIG. 4 illustrates a third text T3 generated based on the second text T2 illustrated in FIG. 3. In the third text T3, only the translation "light emitting element" is assigned to the original word "light emitting element" included in the first text T1. That is, "light emitter" and "photo emitting element" are replaced with "light emitting element." "Light emitting element" is an example of one translation included in the plurality of different translations. "Light emitter" and "photo emitting element" are examples of the remaining translations included in the plurality of different translations.

[0032] Similarly, in the third text T3, only the translation "sensed" is assigned to the original word "detected" contained in the first text T1. That is, "detected" is replaced with "sensed." "Sensed" is an example of one translation included in the multiple different translations. "Detected" is an example of the remaining translation included in the multiple different translations.

[0033] Next, the processing unit 123 outputs data for displaying the third text T3 on the display device 13 from the output unit 124 (STEP 15). The display device 13, which has received the data, displays the third text T3. The display format of the third text T3 on the display device 13 can be determined as appropriate. For example, as illustrated in FIG. 4, the original sentence in the first text T1 and the translation in the third text T3 can be displayed in a table format in which they are associated on a sentence-by-sentence basis. Alternatively, only the third text T3 may be displayed.

[0034] The post-editor performs post-editing, as necessary, on the third text T3 displayed on the display device 13. As described above, the processing unit 123 receives input corresponding to the post-processing through the editing receiving unit 125 and reflects the content of the post-editing in the third text T3. The processing unit 123 outputs data corresponding to the third text T3 in which the post-processing has been reflected from the output unit 124, and displays the third text T3 after the post-processing on the display device 13.

[0035] If it is determined that the multiple different translated words contained in the second text T2 do not correspond to a single original word contained in the first text T1 (NO in STEP 13), that is, if it is determined that the translated words are consistent in the second text T2, the processing unit 123 outputs data for displaying the second text T2 in an editable state on the display device 13 from the output unit 124 (STEP 16). The display device 13 that has received the data displays the second text T2. The display manner of the second text T2 on the display device 13 can be determined appropriately. For example, as illustrated in FIG. 3, the original sentences in the first text T1 and the translated sentences in the second text T2 can be displayed in a table format in which they are associated on a sentence-by-sentence basis. Alternatively, only the second text T2 may be displayed.

[0036] With the above configuration, the third text T3 with consistent translations is displayed on the display device 13, regardless of the neural machine translation model used in the translation device 11. This relieves the post-editor from the need to pay attention to the possibility that the translation may not use consistent translations. This reduces the burden of post-editing work on text machine-translated using a neural machine translation model.

[0037] An example of a specific process executed by the processing unit 123 to determine whether or not it is necessary to generate the third text T3 will be described with reference to FIGS.

[0038] As illustrated in FIG. 5, the processing unit 123 assigns serial numbers to each of the original sentences included in the first text T1 and the translated sentences included in the second text T2 (STEP 21).

[0039] For example, the processing unit 123 detects periods included in the first text T1 based on data corresponding to the first text T1 received by the first receiving unit 121. Each time a period is detected, consecutive numbers are assigned to original sentences that end with the period, thereby assigning consecutive numbers to the multiple original sentences included in the first text T1.

[0040] Similarly, the processing unit 123 detects a period included in the second text T2 based on the data corresponding to the second text T2 received by the second receiving unit 122. Each time a period is detected, consecutive numbers are assigned to translations that end with the period, thereby assigning consecutive numbers to the multiple translations included in the second text T2.

[0041] In principle, periods and full stops are the same, so an original sentence and its translation can be associated with each other using the same sentence number, as shown in the example in Figure 3.

[0042] Next, the processing unit 123 applies morphological analysis to each source sentence included in the first text T1 (STEP 22). As a result, multiple source words that can correspond to morphemes in each source sentence are extracted. At this time, the parts of speech of the extracted source words may be limited. For example, by extracting source words limited to nouns, verbs, adjectives, and adverbs, it is possible to suppress an increase in the processing load and processing time on the processing unit 123.

[0043] Next, the processing unit 123 assigns a serial number N to all source words extracted through the morphological analysis (STEP 23). The processing unit 123 assigns a flag along with the serial number N to each extracted source word. For example, the initial value of the flag is set to 0 (off state). The following description will be given assuming that n source words are extracted throughout the first text T1.

[0044] Next, the processing unit 123 identifies the source language with the smallest serial number Nmin among the source languages ​​whose flags are in the OFF state (STEP 24). At the start of the process, all flags are in the OFF state, so the first source language is identified.

[0045] Next, the processing unit 123 determines whether the source language identified in STEP 24 has the last serial number (N=n) assigned (STEP 25). Since it is usually impossible for the first source language to have the last serial number assigned (NO in STEP 25), the processing unit 123 identifies the source language with the serial number (Nmin+1) (STEP 26). At the start of the process, the second source language is identified.

[0046] Next, the processing unit 123 determines whether the original word identified in STEP 24 and the original word identified in STEP 26 are the same word (STEP 27).

[0047] If the two original languages ​​are different (NO in STEP 27), the processing unit 123 determines whether the original language used in the determination in STEP 27 has the last serial number (N=n) assigned (STEP 28).

[0048] If the source language used in the determination in STEP 27 does not have the final serial number assigned (NO in STEP 28), the processing unit 123 returns the process to STEP 26. That is, the third source language is identified. Thereafter, the processing unit 123 repeats the processes from STEP 26 to STEP 28 until the source language used in the determination in STEP 27 has the final serial number assigned.

[0049] If the original language used in the judgment in STEP 27 has been assigned the last serial number (YES in STEP 28), the processing unit 123 sets the value of the flag assigned to the original language assigned the serial number Nmin to 1 (on state) (STEP 29).

[0050] Next, the processing unit 123 determines whether all flags are in the ON state (STEP 30). If all flags are in the ON state (YES in STEP 30), the processing ends. If all flags are not in the ON state (NO in STEP 30), the processing unit 123 returns the processing to STEP 24.

[0051] If the next source word identified to have the smallest serial number Nmin is the source word to which the last serial number is assigned (YES in STEP 25), there is no need to determine whether the same word exists, so the processing unit 123 sets the value of the flag assigned to that source word to 1 (on state) (STEP 29).

[0052] If it is determined that the source word (N=Nmin) identified in STEP 24 and the source word (N=Nmin+1) identified in STEP 26 are the same word (YES in STEP 27), as illustrated in Figure 6, the processing unit 123 sets the value of the flag assigned to the source word assigned the serial number Nmin+1 to 1 (on state) (STEP 31).

[0053] Next, the processing unit 123 identifies, from the second text T2, a translation that has been assigned the same sentence number as the original sentence containing the original word that has been assigned the serial number Nmin (STEP 32). Similarly, the processing unit 123 identifies, from the second text T2, a translation that has been assigned the same sentence number as the original sentence containing the original word that has been assigned the serial number Nmin+1.

[0054] Next, the processing unit 123 determines whether morphological analysis has been applied to the translation identified in STEP 32 (STEP 33).

[0055] If morphological analysis has not been applied to the translation identified in STEP 32 (NO in STEP 33), the processing unit 123 applies morphological analysis to the translation (STEP 34). This extracts multiple translations that may correspond to the morphemes in the translation. At this time, the parts of speech of the source language to be extracted may be limited. For example, by extracting translations limited to nouns, verbs, adjectives, and adverbs, the processing load and processing time on the processing unit 123 can be reduced.

[0056] If morphological analysis has already been applied to the translation identified in STEP 32 (YES in STEP 33), the processing unit 123 skips STEP .

[0057] Next, the processing unit 123 identifies a translation corresponding to the original word assigned the serial number Nmin from among the multiple translations extracted through the morphological analysis (STEP 35). Similarly, the processing unit 123 identifies a translation corresponding to the original word assigned the serial number Nmin+1 from among the multiple translations extracted through the morphological analysis. As illustrated in FIG. 1, the processing unit 123 is configured to identify a translation by referring to the dictionary database 14. The dictionary database 14 may be available via a communication network free of charge or for a fee, or may be provided as part of a rule-based translation engine.

[0058] Next, the processing unit 123 determines whether the translated word corresponding to the original word assigned the serial number Nmin matches the translated word corresponding to the original word assigned the serial number Nmin+1 (STEP 36).

[0059] If the translation corresponding to the original word assigned the serial number Nmin does not match the translation corresponding to the original word assigned the serial number Nmin+1 (NO in STEP 36), the processing unit 123 creates data corresponding to a list including the multiple translations that differ (STEP 37).

[0060] If the translation corresponding to the original language assigned the serial number Nmin matches the translation corresponding to the original language assigned the serial number Nmin+1 (YES in STEP 36), the processing unit 123 proceeds to STEP 28 in FIG.

[0061] To help understand the above process described with reference to Figures 5 and 6, a simple example is shown in Figure 7. In this example, three source sentences and six source words (n=6) are identified in the first text T1 through the processes in Steps 21 to 23 of Figure 5.

[0062] As mentioned above, at the start of the process, the flags for all serial numbers are in the OFF state, so source language A assigned with serial number N=1 is identified (STEP 24). Since serial number N=1 is not the last serial number (NO in STEP 25), source language B assigned with serial number N=2 is subsequently identified (STEP 26).

[0063] Since original language A assigned serial number N=1 and original language B assigned serial number N=2 are different (NO in STEP 27), and serial number N=2 is not the last serial number (NO in STEP 28), the process returns to STEP 26. That is, original language A assigned serial number N=3 is identified.

[0064] Since the original language A assigned the serial number N=1 and the original language A assigned the serial number N=3 match (YES in STEP 27), the flag assigned to the serial number N=3 is turned on (STEP 31).

[0065] Next, a translation sentence assigned the same sentence number 1 as the original sentence containing the original language A assigned the serial number N=1 is identified from the second text T2 (STEP 32). Similarly, a translation sentence assigned the same sentence number 2 as the original sentence containing the original language A assigned the serial number N=3 is identified from the second text T2.

[0066] Since neither the translated sentence assigned with sentence number 1 nor the translated sentence assigned with sentence number 2 has been subjected to morphological analysis (NO in STEP 33), morphological analysis is applied to both translated sentences (STEP 34).

[0067] Next, by referring to the dictionary database 14, one morpheme included in the translation assigned to sentence number 1 is identified as a translation corresponding to the original language A assigned serial number N=1 (STEP 35). In this example, the translation a1 is identified. Similarly, one morpheme included in the translation assigned to sentence number 2 is identified as a translation corresponding to the original language A assigned serial number N=3. In this example, the translation a1 is identified.

[0068] Since the translated word a1 corresponding to the original language A assigned the serial number N=1 and the translated word a1 corresponding to the original language A assigned the serial number N=3 match (YES in STEP 36), the process proceeds to STEP 28.

[0069] Since the serial number N=3 is not the last serial number (NO in STEP 28), the process returns to STEP 26. That is, the source language C assigned the serial number N=4 is identified.

[0070] Since original language A assigned serial number N=1 and original language C assigned serial number N=4 are different (NO in STEP 27), and since serial number N=4 is not the last serial number (NO in STEP 28), the process returns to STEP 26. That is, original language B assigned serial number N=5 is identified.

[0071] Since original language A assigned serial number N=1 and original language B assigned serial number N=5 are different (NO in STEP 27), and since serial number N=5 is not the last serial number (NO in STEP 28), the process returns to STEP 26. That is, original language D assigned serial number N=6 is identified.

[0072] Since original language A assigned serial number N=1 and original language D assigned serial number N=6 are different (NO in STEP 27), and serial number N=6 is the last serial number (YES in STEP 28), the flag assigned to serial number N=1 is turned on (STEP 29).

[0073] Since not all flags are on yet (NO in STEP 30), the process returns to STEP 24, where the source language with the smallest serial number whose flag is off is identified. In this example, source language B with serial number N=2 is identified.

[0074] Since the serial number N=2 is not the last serial number (NO in STEP 25), the original word A assigned the serial number N=3 is subsequently identified (STEP 26).

[0075] Since original language B assigned with serial number N=2 and original language A assigned with serial number N=3 are different (NO in STEP 27), and since serial number N=3 is not the last serial number (NO in STEP 28), the process returns to STEP 26. That is, original language C assigned with serial number N=4 is identified.

[0076] Since original language B assigned serial number N=2 and original language C assigned serial number N=4 are different (NO in STEP 27), and serial number N=4 is not the last serial number (NO in STEP 28), the process returns to STEP 26. That is, original language B assigned serial number N=5 is identified.

[0077] Since the original language B assigned the serial number N=2 and the original language B assigned the serial number N=5 match (YES in STEP 27), the flag assigned to the serial number N=5 is turned on (STEP 31).

[0078] Next, a translation sentence assigned the same sentence number 1 as the original sentence containing the original language B assigned the serial number N=2 is identified from the second text T2 (STEP 32). Similarly, a translation sentence assigned the same sentence number 2 as the original sentence containing the original language B assigned the serial number N=5 is identified from the second text T2.

[0079] Since morphological analysis has already been applied to both the translation assigned with sentence number 1 and the translation assigned with sentence number 2 (YES in STEP 33), no new morphological analysis is performed.

[0080] Next, by referring to the dictionary database 14, one morpheme included in the translation assigned to sentence number 1 is identified as a translation corresponding to the original language B assigned serial number N=2 (STEP 35). In this example, translation b1 is identified. Similarly, one morpheme included in the translation assigned to sentence number 2 is identified as a translation corresponding to the original language B assigned serial number N=5. In this example, translation b2 is identified.

[0081] Since the translated word b1 corresponding to the original language B assigned the serial number N=2 and the translated word b2 corresponding to the original language B assigned the serial number N=5 are different (NO in STEP 36), data corresponding to a list including the translated words b1 and b2 is generated (STEP 37). Then, the process proceeds to STEP 28.

[0082] Since the serial number N=5 is not the last serial number (NO in STEP 28), the process returns to STEP 26. That is, the original word D assigned the serial number N=6 is identified.

[0083] Since original language B assigned serial number N=2 and original language D assigned serial number N=6 are different (NO in STEP 27), and serial number N=6 is the last serial number (YES in STEP 28), the flag assigned to serial number N=2 is turned on (STEP 29).

[0084] Since not all flags are on yet (NO in STEP 30), the process returns to STEP 24, where the source language with the smallest serial number whose flag is off is identified. In this example, source language C with serial number N=4 is identified.

[0085] Since the serial number N=4 is not the last serial number (NO in STEP 25), source language B, which is assigned the serial number N=5, is subsequently identified (STEP 26).

[0086] Since original language C assigned with serial number N=4 and original language B assigned with serial number N=5 are different (NO in STEP 27), and since serial number N=5 is not the last serial number (NO in STEP 28), the process returns to STEP 26. That is, original language D assigned with serial number N=6 is identified.

[0087] Since the original language C assigned the serial number N=4 and the original language D assigned the serial number N=6 are different (NO in STEP 27), and the serial number N=6 is the last serial number (YES in STEP 28), the flag assigned to the serial number N=4 is turned on (STEP 29).

[0088] Since not all flags are on yet (NO in STEP 30), the process returns to STEP 24, where the source word with the smallest serial number whose flag is off is identified. In this example, source word D with serial number N=6 is identified.

[0089] Since serial number N=6 is the last serial number (YES in STEP 25), the flag assigned to serial number N=6 is turned on (STEP 29). As a result, all flags are turned on (YES in STEP 30), and the process ends.

[0090] 5 to 7, the processing unit 123 of the post-editing support device 12 according to this embodiment may be configured to extract a specific source word by applying morphological analysis to the first text T1. In this case, the processing unit 123 is configured to identify a translation corresponding to a source text containing the specific source word only when the specific source word appears two or more times in the first text T1, and to apply morphological analysis to the identified translation. In other words, for the second text T2, morphological analysis is applied only to translations for which mismatches between translated words need to be verified.

[0091] That is, morphological analysis is applied to all original sentences included in the first text T1, but morphological analysis may not be applied to all translated sentences included in the second text T2. In the example shown in FIG. 7, morphological analysis is not applied to the translated sentence assigned sentence number 3. This is because the original sentence assigned sentence number 3 contains only original words that appear only once in the first text T1, and therefore there is no need to verify mismatches between translated words. By minimizing the frequency at which morphological analysis is applied to the second text T2, it is possible to suppress increases in the processing load and processing time on the processing unit 123.

[0092] The fact that data corresponding to a list including multiple translations has been generated in STEP 37 of Fig. 6 is reflected in the determination as to whether or not third text T3 needs to be generated in STEP 13 of Fig. 2. In other words, if data corresponding to the list has been generated, it is determined that multiple different translations correspond to one source word (YES in STEP 13), and third text T3 is generated (STEP 14).

[0093] As mentioned above, in the third text T3 illustrated in Figure 4, only the translation "light emitting element" is assigned to the original word "light emitter" included in the first text T1. To generate this third text T3, "light emitter" and "photo emitting element" included in the second text T2 illustrated in Figure 3 are replaced with "light emitting element."

[0094] In the third text T3 shown in Figure 4, the translation "light emitting element" is displayed in the same format as the other translations. Because the parts where the replacement process has been performed and the parts where it has not are displayed seamlessly, the post-editor does not know which translation has been replaced. By increasing the stealthiness of the replacement process in this way, the post-editor can concentrate on checking the translation.

[0095] Alternatively, the post-editing support device 12 may be configured to display the replaced translation on the display device 13 in a manner that allows the replacement translation to be distinguished from other translations. For example, at least one of the font size, font type, and font style (italics, bold, etc.) of the replaced translation may be changed to be different from those of the other translations. Additionally or alternatively, the replaced translation may be underlined, or the background color of only the replaced translation may be changed, as illustrated in FIG. 8.

[0096] This configuration meets the needs of post-editors who want to know which translated words have been replaced. Based on the tendency of the translated words to which replacement processing has been applied, it is also possible to estimate the tendencies of the neural machine translation model 111 used by the translation device 11.

[0097] In the above example in which the post-editor can recognize the translation that has been replaced, it may be possible to specify the specific translation that has been replaced. FIG. 9 shows an example in which a specific translation is specified using a cursor displayed on the display device 13 in response to the operation of a pointing device such as a mouse or touchpad. Specifically, the translation "sensed" is specified. If the display device 13 has a touch panel function, the post-editor may specify the specific translation by touching the area where the specific translation is displayed. Alternatively, the post-editor may speak the specific translation, thereby specifying the specific translation through a voice recognition function.

[0098] When a specific translation is specified as described above, the post-editing support device 12 can be configured to display the multiple different translations involved in the replacement process on the display device 13. Specifically, data corresponding to the list generated in STEP 37 of Fig. 6 is read out, and the multiple different translations included in the list are displayed on the display device 13. In the example shown in Fig. 9, in addition to "sensed" used in the third text T3, "detected" included in the second text T2 and replaced by "sensed" is also displayed.

[0099] The display manner of the multiple different translations involved in the replacement process can be determined as appropriate. In the example shown in Fig. 9, the multiple different translations are displayed in a floating manner near the specific translation designated by the cursor. As another example, a dedicated area for displaying the multiple different translations involved in the replacement process may be provided in a position different from the area on the display device 13 where the first text T1 and the third text T3 are displayed.

[0100] With the above configuration, if the post-editor feels uncomfortable with a replaced translation contained in the third text T3, the post-editor can find out the translation contained in the second text T2 before the replacement. In other words, the post-editor can find out what other translation the neural machine translation model 111 of the translation device 11 output for a specific source word before the translation was automatically unified. This can help the post-editor reconsider the translation.

[0101] 10, the post-editing support device 12 can display on the display device 13 one of the multiple different translations involved in the replacement process in a selectable manner. In the example shown in the figure, one of the multiple different translations displayed in a flow chart can be selected with a cursor. Specifically, "detected" is selected.

[0102] When one of the multiple different translations involved in the replacement process is selected, the post-editing support device 12 is configured to replace the original translation displayed in the third text T3 with the selected translation. Figure 11 shows an example in which the translation "detected" originally included in the third text T3 has been converted into the "detected" selected in Figure 10 all at once.

[0103] With this configuration, the post-editor can complete the change to a translation that he or she considers more appropriate simply by converting the translations that have been consistently displayed as the third text T3 into different translations at once. In other words, when the post-editor considers that a translation that was not selected by the post-editing support device 12 is more appropriate than a translation that was automatically selected by the post-editing support device 12, the post-editing work can be performed efficiently.

[0104] As explained above, when generating the third text T3, one of the multiple different translations contained in the second text T2 is substituted for the remaining translations, thereby automatically unifying the translations. The rules for replacing translations will be explained using several examples.

[0105] As an example, the post-editing support device 12 may be configured to replace the remaining translations with the translation that appears most frequently in the second text T2 among the plurality of different translations. For example, when a list of the plurality of different translations is generated in STEP 37 of Fig. 6, the number of times each translation is identified may be included as data, thereby making it possible to identify the translation that appears most frequently in the second text T2.

[0106] In the second text T2 illustrated in Fig. 3, the original word "light emitting element" included in the first text T1 is assigned different translations, such as "light emitting element," "light emitter," and "photo emitting element." Of these, "light emitting element" appears most frequently. Therefore, the post-editing support device 12 replaces "light emitter" and "photo emitting element" with "light emitting element," thereby generating the third text T3 illustrated in Fig. 4.

[0107] According to this configuration, the characteristics of the neural machine translation model 111 used by the translation device 11 can be easily reflected in the third text T3.

[0108] There may be cases where a specific original word contained in the first text T1 is associated with multiple translations, and it is not possible to identify the translation that appears most frequently in the second text T2 among the multiple translations. For example, in the second text T2 illustrated in Figure 3, the original word "kansare" (detected) contained in the first text T1 is assigned different translations, "sensed" and "detected." However, the frequency with which both translations appear in the second text T2 is the same.

[0109] In such a case, the post-editing support device 12 may be configured to replace the remaining translations with the translation that appears first in the second text T2 among the multiple different translations. It does not matter whether the selected translation is appropriate. What is important is that the third text T3 is presented to the post-editor in a state where the inconsistencies in translations have been resolved.

[0110] 1, the post-editing support system 10 may include a storage device 15. As described with reference to FIGS. 10 and 11, when a specific translation in the third text T3 is changed to another translation through post-editing, the post-editing support device 12 receives information related to the change through the edit receiving unit 125. The processing unit 123 is configured to store data corresponding to the changed translation in the storage device 15.

[0111] 10, the change is made to one of the multiple different translations originally included in the second text T2. However, the change may also be made to another translation that the post-editor deems more appropriate. In this case, too, the processing unit 123 stores data corresponding to the changed translation in the storage device 15.

[0112] When it is determined that multiple different translated words included in second text T2 generated based on first text T1 received next or later correspond to a specific original word included in the first text T1, processing unit 123 determines whether the multiple different translated words include a translated word corresponding to the original word stored in storage device 15. When the multiple different translated words include a translated word stored in storage device 15, processing unit 123 replaces the other translated word with the translated word stored in storage device 15.

[0113] For example, if "light emitting element" included in the third text T3 illustrated in Fig. 10 is replaced with "light emitter" in post-editing, "light emitter" is stored in the storage device 15. If the second text T2 generated based on the first text T1 received next time or later includes "light emitting element," "light emitter," and "photo emitting element" as translations of "light emitting element," the other translations will be replaced with "light emitter" regardless of the frequency of appearance of each translation in the second text T2.

[0114] If the original translation is changed in post-editing, the post-editor is likely to find the changed translation more appropriate. With the above configuration, the third text T3 is generated by prioritizing the post-editor's preferences over the characteristics of the neural machine translation model 111, thereby preventing an increase in the workload involved in post-editing.

[0115] The storage device 15 may store dictionary data that allows a user (post-editor) to specify or define the correspondence between original words and translated words. In this case, if it is determined that multiple different translated words included in the second text T2 correspond to a specific original word included in the first text T1, the processing unit 123 determines whether the original word is included in the dictionary data. If the original word is included in the dictionary data, the processing unit 123 replaces the other translated word with the translated word that is associated with the original word in the dictionary data.

[0116] For example, if the dictionary data includes "light emitting element" and "photo emitting element," and the second text T2 includes "light emitting element," "light emitter," and "photo emitting element" as translations of "light emitting element," the other translations will be replaced with "photo emitting element" regardless of the frequency of occurrence of each translation in the second text T2.

[0117] Even with this configuration, the third text T3 is generated by prioritizing the preferences of the post-editor over the characteristics of the neural machine translation model 111, so that an increase in the workload involved in post-editing can be suppressed.

[0118] The processing unit 123 of the post-editing support device 12, which has the various functions described above, may be realized by a general-purpose microprocessor operating in cooperation with a general-purpose memory. At least a portion of the storage device 15 may be realized by the general-purpose memory. Examples of the general-purpose microprocessor include a CPU, an MPU, and a GPU. Examples of the general-purpose memory include a ROM and a RAM. In this case, the ROM may store a computer program that executes the various processes described above. The ROM is an example of a storage medium that stores a computer program. The processor specifies at least a portion of the computer program stored in the ROM, expands it on the RAM, and executes the processes described above in cooperation with the RAM. The computer program may be pre-installed in the general-purpose memory, or may be downloaded from an external server device via a communication network (not shown) and installed in the general-purpose memory. In this case, the external server device is an example of a storage medium that stores a computer program.

[0119] The processing unit 123 may be realized by a dedicated integrated circuit capable of executing the above-mentioned computer program, such as a microcontroller, an ASIC, or an FPGA. In this case, at least a part of the storage device 15 may be realized by a memory element included in the dedicated integrated circuit. The above-mentioned computer program is pre-installed in the memory element. The memory element is an example of a storage medium storing a computer program. The processing unit 123 may also be realized by a combination of a general-purpose microprocessor and a dedicated integrated circuit.

[0120] The above-described embodiments are merely examples for facilitating understanding of the present invention, and the configurations according to the above-described embodiments may be appropriately modified or improved without departing from the spirit and scope of the present invention.

[0121] In the post-editing support system 10, the translation device 11, the post-editing support device 12, the display device 13, and the storage device 15 may each be provided as an independent device, or at least one of the translation device 11, the post-editing support device 12, the display device 13, and the storage device 15 may be provided as different functional units within a single device. [Explanation of symbols]

[0122] 10: Post-editing support system, 11: Translation device, 111: Neural machine translation model, 12: Post-editing support device, 121: First reception unit, 122: Second reception unit, 123: Processing unit, 13: Display device, 15: Storage device, T1: First text, T2: Second text, T3: Third text

Claims

1. a translation device that translates a first text written in a first language into a second text written in a second language different from the first language using a neural machine translation model; a post-editing support device that generates a third text in which the translations are unified by replacing one of the translations included in the plurality of different translations with another translation included in the plurality of different translations when the plurality of different translations included in the second text correspond to one original word included in the first text, and displays the third text on a display device; It is equipped with The post-editing support device extracting a plurality of source words by applying morphological analysis to the first text; If a source word included in the plurality of source words appears more than once in the first text, identifying a translation in the second text that corresponds to a source sentence including the source word; extracting a plurality of translations by applying morphological analysis to the identified translation; by referring to a dictionary database storing correspondences between the one original word and its translations, identifying a plurality of translations corresponding to the one original word from the extracted plurality of translations; determining whether the plurality of corresponding translations match each other; If it is determined that the corresponding translations do not match each other, it is determined that the different translations included in the second text correspond to one original word included in the first text; the post-editing support device displays the replaced translated word on the display device in a manner that allows the replaced translated word to be distinguished from other translated words included in the third text. Post-editing support system.

2. the post-editing support device, when the replaced translation is specified in the third text, causes the display device to display the plurality of different translations. The post-editing support system according to claim 1 .

3. the one translation is a word that appears most frequently in the second text among the plurality of different translations; 3. The post-editing support system according to claim 1.

4. The system is provided with a storage device that stores a dictionary that allows a user to specify the correspondence between source words and translations, When the original word is included in the dictionary, the post-editing support device uses a translation associated with the original word in the dictionary as the translation. The post-editing support system according to any one of claims 1 to 3.

5. the post-editing support device causes the display device to display the original sentence in the first text and the translation sentence in the third text in a format in which the original sentence and the translation sentence are associated with each other on a sentence-by-sentence basis; The post-editing support system according to any one of claims 1 to 4.

6. The neural machine translation model includes a trained model for translating documents related to intellectual property rights. The post-editing support system according to any one of claims 1 to 5.

7. 1. A computer-implemented post-editing support method, comprising: translating a first text written in a first language into a second text written in a second language different from the first language using a neural machine translation model; extracting a plurality of source words by applying morphological analysis to the first text; identifying a translation in the second text that corresponds to a sentence containing one of the plurality of source words, if the one of the source words appears more than once in the first text; extracting a plurality of translations by applying morphological analysis to the identified translation; a step of identifying a plurality of corresponding translated words corresponding to the one original word from the extracted plurality of translated words by referring to a dictionary database storing correspondence relationships between the one original word and translated words; determining whether the plurality of corresponding translations match each other; if it is determined that the corresponding translations do not match each other, determining that the different translations included in the second text correspond to one original word included in the first text; generating a third text in which the translations are unified by replacing the remaining translations included in the plurality of different translations with one translation included in the plurality of different translations when it is determined that the plurality of translations correspond to the one original language; displaying the third text on a display device in a manner that allows the replaced translation to be distinguished from other translations included in the third text; Equipped with Post-editing support methods.

8. a first reception unit that receives a first text written in a first language; a second receiving unit that receives second text from a translation device that translates the second text into a second text written in a second language different from the first language using a neural machine translation model; a processing unit that, when a plurality of different translated words included in the second text correspond to one original word included in the first text, generates a third text in which the translated words are unified by replacing one translated word included in the plurality of different translated words with the remaining translated words included in the plurality of translated words, and displays the third text on a display device; It is equipped with The processing unit extracting a plurality of source words by applying morphological analysis to the first text; If a source word included in the plurality of source words appears more than once in the first text, identifying a translation in the second text that corresponds to a source sentence including the source word; extracting a plurality of translations by applying morphological analysis to the identified translation; by referring to a dictionary database storing correspondences between the one original word and its translations, identifying a plurality of translations corresponding to the one original word from the extracted plurality of translations; determining whether the plurality of corresponding translations match each other; If it is determined that the corresponding translations do not match each other, it is determined that the different translations contained in the second text correspond to one original word contained in the first text; The display device is configured to display the replaced translation in a manner that allows it to be distinguished from other translations included in the third text. Post-editing support device.

9. A computer program executable by a processing unit of a post-editing support device, By executing this, the post-editing support device Accepting a first text written in a first language; receiving second text from a translation device that translates the second text into a second language different from the first language using a neural machine translation model; extracting a plurality of source words by applying morphological analysis to the first text; If one of the plurality of source languages ​​appears more than once in the first text, identifying a translation in the second text that corresponds to the source sentence containing the one of the plurality of source languages; extracting a plurality of translations by applying morphological analysis to the identified translation; by referring to a dictionary database storing correspondences between the one original word and its translations, a plurality of corresponding translations corresponding to the one original word are identified from the plurality of extracted translations; determining whether the plurality of corresponding translations match each other; If it is determined that the plurality of corresponding translations do not match each other, it is determined that the plurality of different translations contained in the second text correspond to one original word contained in the first text; If it is determined that the plurality of translated words correspond to the single original word, one translated word included in the plurality of different translated words is substituted for the remaining translated words included in the plurality of translated words, thereby generating a third text in which the translated words are unified; displaying the third text on a display device in a manner that allows the substituted translation to be distinguished from other translations included in the third text; Computer program.

Citation Information

Patent Citations

  • Machine translation device

    JP1992357566A

  • Method and device for detecting nonuniformity of translated word

    JP1993101095A

  • Machine translation system

    JP2007079825A

  • Machine translation program, and machine translation apparatus

    JP2008027458A

  • Translation apparatus, control program of translation apparatus, and translation method using translation apparatus

    JP2020077134A