Method for generating training data for machine translation, method for creating learnable model for machine translation processing, machine translation processing method, and device for generating training data for machine translation

JP2023183618A5Active Publication Date: 2025-05-22NAT INST OF INFORMATION & COMM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022097221
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-06-16
Publication Date
2025-05-22
Estimated Expiration
2042-06-16

AI Technical Summary

Benefits of technology

【0038】 本発明によれば、タグ付きの対訳文を大量に準備することなく、翻訳対象の原文にマークアップ言語用タグを含んだ原文を、マークアップ言語用タグの情報を保持しつつ、高精度に機械翻訳することを可能にする機械翻訳処理方法、機械翻訳用訓練データ生成方法、機械翻訳処理用の学習可能モデルの作成方法、機械翻訳処理方法、機械翻訳用訓練データ生成装置、および、機械翻訳処理システムを実現することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a machine translation processing system that can accurately translate an original sentence containing a markup language tag for a text to be translated by machine translation, while holding information about the markup language tag without preparing a large number of tagged translations.SOLUTION: In a machine translation processing system 1000, a training data generating device 1 executes training data generation processing to detect a start / end corresponding code in translation data not containing a markup language tag and replace the detected start / end corresponding code with an alternative code, thereby easily generating a large amount of data equivalent to translation data with the inserted markup language tag inserted thereto. A machine translation processing device 2 uses the translation data acquired by the training data generation processing in the training data generating device 1 as training data for learning of a machine translation model, thereby producing the same effect as when learning of the machine translation model is performed using the translation data with the markup language tag as training data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a machine translation processing technology, and more particularly to a machine translation processing technology that corresponds to tags in a markup language. [Background technology]

[0002] In the field of industrial translation, the source text to be translated often contains XML tags (an example of a markup language tag), and there is a high demand for machine translation of source texts containing such tags with high accuracy while retaining the tag information.

[0003] As a method for dealing with cases where the original text to be translated contains XML tags, for example, as disclosed in Non-Patent Document 1, there is a method in which the tags are removed from the original text during machine translation, and then the tags are reinserted into the machine translation results based on word alignment between the original text and the translation.

[0004] Furthermore, Patent Document 1 discloses a technique for training a machine translation engine using bilingual texts into which markup language tags (for example, XML tags) are inserted. With the technique of Patent Document 1, when training the machine translation engine, markup language tags are replaced with placeholders, and the machine translation engine is trained using the bilingual texts in which the markup language tags have been replaced with placeholders. Then, with the technique of Patent Document 1, during machine translation, tags in the original text are replaced with placeholders and translated, and then the placeholders in the translated text are replaced with the original tags. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] U.S. Patent No. 10,963,652 [Non-patent literature]

[0006] [Non-Patent Document 1] Mathias Mueller. Treatment of Markup in Statistical Machine Translation. Proceedings of the Third Workshop on Discourse in Machine Translation, pages 36-46, Copenhagen, Denmark, September 8, 2017. Association for Computational Linguistics. Summary of the Invention [Problem to be solved by the invention]

[0007] However, while the tag reinsertion method disclosed in Non-Patent Document 1 has the advantage of being able to train a machine translation engine even if the bilingual text does not contain tags, it does not take tags into account during machine translation, making it difficult to properly retain tags in the translation.

[0008] On the other hand, the method of training a machine translation engine using tagged bilingual texts disclosed in Patent Document 1 has no problems with translation accuracy or tag retention accuracy, but it has the problem of difficulty in preparing a large number of tagged bilingual texts.

[0009] In view of the above problems, the present invention aims to provide a machine translation processing method, a method for generating training data for machine translation, a method for creating a trainable model for machine translation processing, a machine translation processing method, a training data generation device for machine translation, and a machine translation processing system that enable highly accurate machine translation of an original text to be translated that includes markup language tags while retaining the information of the markup language tags, without having to prepare a large number of tagged bilingual texts. [Means for solving the problem]

[0010] The first invention for solving the above problem is a method for generating training data for training a learnable model for machine translation processing (a method for generating training data for machine translation) in a machine translation processing system for machine translating language data including tags for markup languages, the method comprising a start-end correspondence code detection step and a replacement processing step.

[0011] The start-end correspondence code detection step detects start-end correspondence codes, which are codes whose start and end correspond to each other, in bilingual data that is a combination of first language data and second language data that is data obtained by translating the first language data into a second language and does not include tags for markup languages.

[0012] The replacement processing step performs a replacement process on the bilingual data to replace the start / end corresponding code with a substitute code, thereby obtaining bilingual data after the replacement process.

[0013] This method for generating training data for machine translation detects start-end correspondence symbols (symbols where the left and right correspond, such as () and []) in bilingual texts (bilingual data) that do not contain markup language tags (e.g., XML tags), and replaces the detected start-end correspondence symbols with alternative symbols (placeholders), making it possible to easily generate large quantities of data equivalent to bilingual data with markup language tags (e.g., XML tags) inserted.

[0014] Furthermore, since the bilingual data obtained by this method for generating training data for machine translation contains substitute codes (placeholders) equivalent to tags for markup languages, by using the bilingual data as training data for the learning process of a machine translation model, it is possible to achieve the same effect as when bilingual sentences (bilingual data) with tags for markup languages ​​(e.g., XML tags) are used as training data for the learning process of a machine translation model (it is possible to perform the same learning process).

[0015] A second aspect of the present invention is the first aspect of the present invention, further comprising a replacement ratio setting step of setting a replacement ratio.

[0016] The replacement processing step performs a replacement process of replacing the start / end corresponding code with an alternative code for the parallel translation data at the replacement ratio set in the replacement ratio setting step.

[0017] In this method for generating training data for machine translation, by setting the replacement ratio in the replacement ratio setting step (by setting it to a value less than 1.0), it is guaranteed that not all start / end corresponding codes are replaced with alternative codes (placeholders). As a result, in this method for generating training data for machine translation, it is guaranteed that the start / end corresponding codes are included in the parallel translation data after the replacement process, and appropriate learning processing (training) can be performed for the start / end corresponding codes (it becomes possible to correctly appear (perform machine translation) the start / end corresponding codes of the source language data in the machine translation processing result data (target language data)).

[0018] Note that the replacement ratio may be in units of parallel translation data (parallel sentence units). That is, when there are N1 (N1: natural number) parallel translation data containing start / end corresponding codes among the parallel translation data being processed, and the replacement ratio is r (r: real number, 0 < r < 1), replacement processing may be performed on int(N1 × r) (int(x): a function that obtains the largest integer value not exceeding x) of the parallel translation data containing start / end corresponding codes.

[0019] A third invention is a method for creating a learnable model for machine translation processing in a machine translation processing system that performs machine translation processing on language data including tags for markup language using the training data generated by the method for generating training data for machine translation according to the first or second invention, the method comprising a data input step, an output data acquisition step, a loss evaluation step, and a parameter update step.

[0020] The data input step inputs the first language data included in the parallel translation data after the replacement process into the learnable model for machine translation processing.

[0021] The output data acquisition step acquires output data of a trainable model for machine translation processing for the data input in the data input step.

[0022] The loss evaluation step acquires the output data acquired in the output data acquisition step and the second language data included in the bilingual data after the replacement process as correct data, and evaluates the loss between the output data and the correct data.

[0023] The parameter updating step updates the parameters of the trainable model for machine translation processing so as to reduce the loss obtained by the loss evaluation step.

[0024] In this method of creating a trainable model for machine translation processing, the trainable model for machine translation processing can be trained using the first language data contained in the bilingual data after the replacement process and the second language data contained in the bilingual data after the replacement process as correct answer data, so that a trained model of the trainable model that machine translates the first language data after the replacement process into the second language data after the replacement process can be obtained.

[0025] The fourth invention is a method (machine translation processing method) for performing machine translation processing using a trained model of a trainable model for machine translation processing obtained by training using the method for creating a trainable model for machine translation processing of the third invention, and comprises a forward substitution processing step, a machine translation processing step, and a reverse substitution processing step.

[0026] The forward substitution processing step executes a forward substitution processing for substituting a markup language tag included in the input first language data with an alternative code.

[0027] The machine translation processing step performs machine translation processing on the first language data after the forward substitution processing using a trained model of the trainable model for machine translation processing, thereby obtaining second language data after machine translation processing.

[0028] The reverse substitution step executes a reverse substitution process for replacing substitute codes included in the second language data after the machine translation process obtained in the machine translation process step with the markup language tags replaced in the forward substitution step.

[0029] In this machine translation processing method, for input data containing markup language tags (e.g., XML tags), the markup language tags are replaced with substitution codes (placeholders) similar to those used when generating the training data, and machine translation processing is performed using a trained model of a machine translation model optimized with bilingual data into which the substitution codes have been inserted, making it possible to obtain appropriate machine translation processing result data while appropriately maintaining the substitution codes inserted.In this machine translation processing method, by replacing (restoring) the substitution codes in the machine translation processing result data (machine-translated text) with the substitution codes inserted with XML tags, it is possible to obtain machine translation processing result data (machine-translated text) in which the XML tags have been inserted in an appropriate state.

[0030] In this way, this machine translation processing method makes it possible to perform highly accurate machine translation of an original text to be translated that contains markup language tags while retaining the information of the markup language tags, without having to prepare a large amount of tagged bilingual text.

[0031] The fifth invention is a method for generating training data for training a learnable model for machine translation processing (a method for generating training data for machine translation) in a machine translation processing system for machine translating language data including tags for markup languages, comprising a corresponding element detection step and a replacement processing step.

[0032] The corresponding element detection step detects corresponding elements, which are elements that are determined to correspond between the first language data and the second language data, in bilingual data that is a combination of first language data and second language data that is data obtained by translating the first language data into a second language and does not include tags for markup languages.

[0033] The replacement processing step performs a replacement process on the bilingual data to insert substitute codes before and after the corresponding elements, thereby obtaining bilingual data after the replacement process.

[0034] This method for generating training data for machine translation detects elements that correspond between the original and translated texts in bilingual texts (bilingual data) that do not contain markup language tags (e.g., XML tags), and replaces the elements before and after them with substitute codes (placeholders).This makes it possible to easily generate large quantities of data equivalent to bilingual data with markup language tags (e.g., XML tags) inserted.

[0035] A sixth invention is a device (machine translation training data generation device) that generates training data for training a learnable model for machine translation processing in a machine translation processing system for machine translating language data including tags for markup languages, and is equipped with a replacement processing unit.

[0036] The replacement processing unit detects start-end correspondence codes, which are codes whose start and end correspond to each other, in bilingual data that is a combination of first language data and second language data that is data obtained by translating the first language data into a second language and does not include markup language tags, and A replacement process is performed on the bilingual data to replace the start / end corresponding code with a substitute code, thereby obtaining bilingual data after the replacement process.

[0037] This makes it possible to realize a training data generation device for machine translation that achieves the same effects as the first aspect of the invention. [Effects of the Invention]

[0038] According to the present invention, it is possible to realize a machine translation processing method, a method for generating training data for machine translation, a method for creating a trainable model for machine translation processing, a machine translation processing method, a training data generation device for machine translation, and a machine translation processing system that enable highly accurate machine translation of an original text to be translated that includes markup language tags while retaining the information of the markup language tags, without having to prepare a large amount of tagged bilingual text. [Brief explanation of the drawings]

[0039] [Figure 1] FIG. 1 is a schematic configuration diagram of a machine translation processing system 1000 according to a first embodiment. [Figure 2] 10 is a flowchart of a training data generation process executed in the machine translation processing system 1000. [Figure 3] 1 is a diagram for explaining the replacement process executed by the training data generation device 1 of the machine translation processing system 1000. FIG. [Figure 4] 10 is a flowchart of a prediction process (machine translation execution process) executed by the machine translation processing system 1000. [Figure 5] 1 is a diagram for explaining prediction processing (machine translation execution processing) of the machine translation processing system 1000. FIG. [Figure 6] 10 is a diagram showing the results of machine translation processing of first language data (Japanese data) with XML tags by the machine translation processing system 1000. FIG. [Figure 7] FIG. 10 is a schematic configuration diagram of a machine translation processing system 2000 according to a second embodiment. [Figure 8] 10 is a diagram for explaining the replacement process executed by a training data generation device 1A of a machine translation processing system 2000. FIG. [Figure 9] A diagram showing the CPU bus configuration. DETAILED DESCRIPTION OF THE INVENTION

[0040] [First embodiment] The first embodiment will be described below with reference to the drawings.

[0041] <1.1: Machine translation processing system configuration> FIG. 1 is a schematic diagram of a machine translation processing system 1000 according to the first embodiment.

[0042] 1, the machine translation processing system 1000 includes a training data generation device 1, a data storage unit DB1, and a machine translation processing device 2. Note that, in the following explanation, it is assumed that the target of the machine translation processing is language data that includes markup language tags, but the target of the machine translation processing device 2 does not necessarily have to include markup language tags, and when input data that does not include tags is provided, the machine translation processing is executed without performing replacement processing or the like.

[0043] As shown in FIG. 1, the training data generation device 1 includes a replacement ratio setting unit 11 and a replacement processing unit 12.

[0044] The replacement ratio setting unit 11 sets the ratio at which the start-end corresponding code is replaced with an alternative code (placeholder). Then, the replacement ratio setting unit 11 outputs data indicating the ratio at which the set start-end corresponding code is replaced with an alternative code (placeholder) (this is referred to as "replacement ratio data") to the replacement processing unit 12 as data r_rep.

[0045] The replacement processing unit 12 receives input of bilingual data Din_tr, which is data paired between data in a first language (source language data) and data in a second language (target language data) obtained by translating the data in the first language into the second language, and does not include markup language tags. The replacement processing unit 12 also receives replacement ratio data r_rep output from the replacement ratio setting unit 11. The replacement processing unit 12 performs a process of replacing start-end correspondence codes included in the bilingual data Din_tr with substitute codes (placeholders) at a ratio indicated by the replacement ratio data r_rep. The replacement processing unit 12 then outputs the bilingual data after the replacement process to the data storage unit DB1 as post-replacement bilingual data Do_tr.

[0046] For ease of explanation, the bilingual data Din_tr input to the training data generation device 1 is N sets (N: natural number), and the ith (i: natural number, 1≦i≦N) data in the first language (source language data) of the bilingual data Din_tr is referred to as “src i " and the second language data (target language data) that is the data translated from the first language data into the second language is referred to as "dst i " and the ith bilingual data is written as "{src i ,dst i}".

[0047] In addition, the i-th first language data (first language data of the replacement word) of the post-replacement bilingual data Do_tr is defined as "src_rep i " and the second language data (second language data after replacement processing) that is paired with the first language data (constituting a translation) is represented as "dst_rep i " and the i-th data (parallel data) of the post-replacement bilingual data Do_tr is expressed as "{src_rep i ,dst_rep i}".

[0048] The data storage unit DB1 receives the post-replacement bilingual data Do_tr output from the training data generation device 1 and stores and holds the data. In addition, the data storage unit DB1 reads out the stored data (post-replacement bilingual data Do_tr) in accordance with an instruction from the machine translation processing device 2, and outputs the read-out data to the machine translation processing device 2 as data Din_tr_rep. As shown in FIG. 1, the machine translation processing device 2 includes a training data acquisition unit 21, a forward substitution processing unit 22, a first selector SEL21, a machine translation processing unit 23, a second selector SEL22, a loss evaluation unit 24, and an inverse substitution processing unit 25.

[0049] The training data acquisition unit 21 outputs a data read command to the data storage unit DB1 and reads the post-replacement processing bilingual data stored in the data storage unit DB1 from the data storage unit DB1 as training bilingual data Din_tr_rep. The training data acquisition unit 21 extracts data in a first language (source language data) from the training bilingual data Din_tr_rep and outputs the extracted first language data (source language data) to the first selector SEL21 as training input data Din_tr. The training data acquisition unit 21 also extracts data in a second language (target language data) that is bilingual with the first language data output to the first selector SEL21 from the training bilingual data Din_tr_rep and outputs the extracted second language data (target language data) to the loss evaluation unit 24 as training correct data D_correct.

[0050] For ease of explanation, it is assumed that the training data acquisition unit 21 reads out M sets (M: natural number, M≦N) of post-substitution bilingual data Din_tr from the data storage unit DB1, and the j-th (j: natural number, 1≦j≦M) first language data of the read bilingual data Din_tr is referred to as “src_rep j " and the data in the second language that is paired with the data in the first language (to form a parallel translation) is written as "dst_rep j " and the j-th data (parallel data) of the parallel data Din_tr is written as "{src_rep j ,dst_rep j}".

[0051] The forward substitution processing unit 22 receives as input, as data Din_src, data in a first language to be subjected to machine translation processing (source language data), the data including markup language tags (for example, XML tags). The forward substitution processing unit 22 then performs processing (forward substitution processing) to replace the markup language tags included in the data Din_src with substitute codes (placeholders). The forward substitution processing unit 22 then outputs the first language data after the forward substitution processing to the first selector SEL21 as data Din_rep. In addition, in the forward substitution processing, the forward substitution processing unit 22 generates a list of correspondences between markup language tags and substitute codes (placeholders) that have replaced the markup language tags, and outputs data including the list to the reverse substitution processing unit 25 as data D_list_rep.

[0052] The first selector SEL21 receives as input the data Din_tr output from the training data acquisition unit 21 and the data Din_rep output from the forward substitution processing unit 22. The first selector SEL21 also receives as input a selection signal sel21 output from a control unit (not shown) that controls each functional unit of the machine translation processing device 2. The first selector SEL21 selects either the data Din_tr or the data Din_rep in accordance with the selection signal se21, and outputs the selected data to the machine translation processing unit 23 as data D1.

[0053] It should be noted that (1) when a learning process (training process) is performed in the machine translation processing unit 23 (during the learning process (during training)), the control unit outputs a selection signal sel21 having a signal value of "0" to the first selector SEL21, and the first selector SEL21 selects data Din_tr in accordance with the selection signal and outputs the selected data Din_tr as data D1 to the machine translation processing unit 23. (2) When a prediction process (machine translation process) is performed in the machine translation processing unit 23 (during the prediction process (when machine translation is executed)), the control unit outputs a selection signal sel21 having a signal value of "1" to the first selector SEL21, and the first selector SEL21 selects data Din_rep in accordance with the selection signal and outputs the selected data Din_rep to the machine translation processing unit 23 as data D1.

[0054] The machine translation processing unit 23 includes a machine translation model, and receives as input data D1 output from the first selector SEL21. The machine translation model included in the machine translation processing unit 23 is a trainable model (a model in which a trained model is constructed by optimizing parameters through data-based learning), and is a model for training machine translation (for example, a machine translation model using a neural network).

[0055] (1) During learning processing (training), the machine translation model of the machine translation processing unit 23 receives input of data D1 (=Din_tr) from the first selector SEL21 and outputs the data acquired by the machine translation model as data D2 to the second selector SEL22. Also, during learning processing (training), the machine translation model of the machine translation processing unit 23 receives input of parameter update data update(θ) output from the loss evaluation unit 24 and updates the parameters of the machine translation model based on the parameter update data update(θ) (for example, if the machine translation model of the machine translation processing unit 23 is a model using a neural network, the parameters of the machine translation model of the machine translation processing unit 23 are updated by the backpropagation method).

[0056] (2) During prediction processing (when machine translation processing is executed), the machine translation model of the machine translation processing unit 23 (a machine translation model (trained model) in which optimal parameters obtained by the learning processing are set) inputs data D1 (=Din_rep) from the first selector SEL21 and outputs the data obtained by the machine translation model (trained model) of the machine translation processing unit 23 to the second selector SEL22 as data D2.

[0057] The second selector SEL22 receives as input the data D2 output from the machine translation processing unit 23 and a selection signal sel22 output from a control unit (not shown) that controls each functional unit of the machine translation processing device 2. The second selector SEL22 outputs the data D2 to either the loss evaluation unit 24 or the inverse substitution processing unit 25 in accordance with the selection signal sel22.

[0058] It should be noted that (1) when a learning process (training process) is performed in the machine translation processing unit 23 (during the learning process (during training)), the control unit outputs a selection signal sel22 whose signal value is "0" to the second selector SEL22, and the second selector SEL22 outputs the data D2 as data D21 to the loss evaluation unit 24 in accordance with the selection signal. (2) When a prediction process (machine translation process) is performed in the machine translation processing unit 23 (during the prediction process (when machine translation is executed)), the control unit outputs a selection signal sel22 whose signal value is "1" to the second selector SEL22, and the second selector SEL22 outputs the data D2 as data D22 to the inverse substitution processing unit 25 in accordance with the selection signal.

[0059] The loss evaluation unit 24 receives the training correct data D_correct output from the training data acquisition unit 21 and the data D21 output from the second selector SEL22. The loss evaluation unit 24 evaluates the loss (e.g., error) between the data D21 and the training correct data D_correct using, for example, a loss function, and generates parameter update data update(θ) that is data for updating parameters of the machine translation model of the machine translation processing unit 23 based on the evaluation result. The loss evaluation unit 24 then outputs the generated parameter update data update(θ) to the machine translation processing unit 23. Note that in FIG. 1, the path from the output of the machine translation processing unit 23 to the loss evaluation unit 24 and the path for outputting the parameter update data update(θ) from the loss evaluation unit 24 to the machine translation processing unit 23 are illustrated as separate paths, but this is for convenience (for convenience of illustration) and is not limited to the form of FIG. 1. In the machine translation processing device 2, when updating the parameters of the machine translation model of the machine translation processing unit 23 using the error backpropagation method, the error obtained by the loss evaluation unit 24 (error obtained by an error function (e.g., cross-entropy error)) can be propagated (backpropagated) sequentially along a path that reverses the path (forward propagation path) along which output data was obtained by the machine translation model of the machine translation processing unit 23, thereby updating each parameter of the machine translation model of the machine translation processing unit 23 (parameters of each layer of the machine translation model of the machine translation processing unit 23).

[0060] Furthermore, if the acquired error (loss) (1) falls within a predetermined range, or if (2) the amount of change in the error (loss) falls within a predetermined range, the loss evaluation unit 24 determines that there is no need to continue the learning process and terminates the learning process.

[0061] The inverse substitution processor 25 receives data D22 output from the second selector SEL22 and data D_list_rep output from the forward substitution processor 22. The inverse substitution processor 25 detects, from data D22, substitution codes (placeholders) substituted by the forward substitution processor 22, and performs a process (inverse substitution process) of restoring (substituting) the detected substitution codes into their original markup language tags based on a list included in data D_list_rep (a list of correspondences between markup language tags and substitution codes (placeholders) that substituted the markup language tags in the forward substitution process). The inverse substitution processor 25 then outputs the data D22 after the inverse substitution process has been performed as output data Do_dst.

[0062] <1.2: Operation of machine translation processing system> The operation of the machine translation processing system 1000 configured as above will now be described.

[0063] Below, the operation of the machine translation processing system 1000 will be explained in parts: (1) training data generation processing, (2) machine translation model learning processing (training processing) (creation method), and (3) prediction processing (machine translation execution processing).

[0064] For ease of explanation, it is assumed that the machine translation processing system 1000 is a system for executing a process of machine translating a first language (source language) into a second language (target language).

[0065] (1.2.1: Training data generation process) First, the training data generation process executed in the machine translation processing system 1000 will be described.

[0066] FIG. 2 is a flowchart of the training data generation process executed by the machine translation processing system 1000.

[0067] FIG. 3 is a diagram for explaining the replacement process executed by the training data generation device 1 of the machine translation processing system 1000.

[0068] The training data generation process executed by the machine translation processing system 1000 will be described below with reference to the flowchart of FIG.

[0069] (Step S101): In step S101, a placeholder setting process is executed. Specifically, the process is executed as follows.

[0070] The replacement processing unit 12 of the training data generation device 1 sets start and end correspondence codes to be replaced with alternative codes (placeholders) for bilingual data Din_tr (bilingual data input to the training data generation device 1), which is bilingual data that pairs data in a first language (source language data) with data in a second language (target language data), which is data obtained by translating the data in the first language into the second language, and does not include tags for markup languages.

[0071] A "start-end corresponding code" refers to a code that is paired (combined) with a code (start code) that indicates the start (or starting point) of a word string or character string (including a subword string) and a code (end code) that indicates the end (or ending point) of the word string or character string (including a subword string). For example, the following codes can be given as "start-end corresponding codes". (1) "()" (Left parenthesis (opening symbol) and right parenthesis (closing symbol)) (2) "[]" (Left bracket (opening symbol) and right bracket (closing symbol)) (3) """ (opening double quotation marks and closing double quotation marks) (4) "''" (a left single quotation mark (opening quotation mark) and a right single quotation mark (opening quotation mark) The start / end corresponding code is not limited to the above, and may be any other code as long as the start code and the end code correspond to each other (the code on the left corresponds to the code on the right).

[0072] Furthermore, if the first and second languages ​​use double-byte character codes, the start and end corresponding codes in those languages ​​may be set as double-byte code (character code) codes. For example, if the first language is Japanese and the second language is English and the start and end corresponding codes are "()" (a left parenthesis (start code) and a right parenthesis (end code)), (A) in Japanese (the first language), which is a language using double-byte codes, the start and end corresponding codes may be set as a left parenthesis (start code) and a right parenthesis (end code) in single-byte code (half-width characters) and / or a left parenthesis (start code) and a right parenthesis (end code) in double-byte code (full-width characters), and (B) in the second language (English), the start and end corresponding codes may be set as a left parenthesis (start code) and a right parenthesis (end code) in single-byte code (half-width characters).

[0073] For the sake of convenience, the first language is Japanese, the second language is English, and the start and end correspondence codes are (1) "()" (Left parenthesis (opening symbol) and right parenthesis (closing symbol)) (2) "[]" (Left bracket (opening symbol) and right bracket (closing symbol)) An example will be described in which 1-byte code characters (half-width characters) are set as start and end corresponding codes for both the first and second languages.

[0074] The substitution processing unit 12 of the training data generation device 1 sets the first language as Japanese, the second language as English, and the start / end correspondence code as (1) "()" (Left parenthesis (opening symbol) and right parenthesis (closing symbol)) (2) "[]" (Left bracket (opening symbol) and right bracket (closing symbol)) Set to.

[0075] (Step S102): In step S102, a replacement ratio setting process is executed. Specifically, the process is executed as follows.

[0076] The replacement ratio setting unit 11 sets the ratio at which start / end corresponding codes are replaced with substitute codes (placeholders). Then, the replacement ratio setting unit 11 outputs the set replacement ratio data (data indicating the ratio at which start / end corresponding codes are replaced with substitute codes (placeholders)) as data r_rep to the replacement processing unit 12. In this embodiment, for ease of explanation, the following description will be given assuming that the replacement ratio setting unit 11 has set the ratio at which start / end corresponding codes are replaced with substitute codes (placeholders) to "0.1" (10%).

[0077] The rate set by the replacement rate setting unit 11 (rate indicated by the replacement rate data r_rep) is preferably set so that the probability of occurrence of replacement codes (placeholders) is similar to the probability of occurrence of markup language tags in the first language data (source language data) with markup language tags input to the machine translation processing device 2. In other words, it is preferable that the rate be set so that the probability of occurrence (occurrence probability distribution) of replacement codes (placeholders) in the bilingual data Do_tr after the replacement process is close to the probability of occurrence (occurrence probability distribution) of markup language tags in the first language data (source language data) with markup language tags input to the machine translation processing device 2 (data to be subjected to machine translation). In this way, the probability distribution of occurrence of replacement codes (placeholders) in the training data becomes close to the probability distribution of occurrence of markup language tags in the language data to be actually subjected to machine translation, thereby improving the accuracy of the learning process of the machine translation process using the training data. In addition, research by the inventors has shown that the occurrence probability of "()" and "[]" in a large-scale corpus is about 0.1, and if 10% of them are replaced, 1% will become alternative codes. This ratio is close to the occurrence probability of markup language tags in the language data (including plain text and sentences with markup language tags) that are input to the target machine translation process.

[0078] Furthermore, by setting the replacement ratio using the replacement ratio setting unit 11 (by setting it to a value less than 1.0), it is guaranteed that all start-end correspondence codes will not be replaced with substitute codes (placeholders). This ensures that start-end correspondence codes will be included in the bilingual data after the replacement process, and makes it possible to properly learn (train) the start-end correspondence codes (it becomes possible to make the start-end correspondence codes in the source language data appear correctly (machine translated) in the machine translation processing result data (target language data)).

[0079] (Step S103): In step S103, loop processing (loop 1) is started. When the bilingual data Din_tr input to the training data generation device 1 is N sets (N: natural number), each bilingual data {src_rep i ,dst_rep i} (i: natural number, 1≦i≦N), loop processing (loop 1) is executed N times. That is, from the first bilingual data {src_rep1, dst_rep1} to the Nth bilingual data {src_rep N ,dst_rep N}, the loop processing (loop 1) is executed.

[0080] (Steps S104 and S105): In steps S104 and S105, the first language data (src i ) replacement process and second language data (dst i ) is replaced. Specifically, the following process is performed:

[0081] The replacement processing unit 12 inputs bilingual data Din_tr, which is data pairing data in a first language (source language data) with data in a second language (target language data) that is data obtained by translating the data in the first language into the second language, and does not include tags for markup languages. Note that the bilingual data Din_tr is data (word strings, subword strings, etc.) that has been subjected to morphological analysis processing for both the first and second languages ​​and separated into morphemes.

[0082] Furthermore, the replacement processing unit 12 performs a process of replacing start-end correspondence codes included in the bilingual data Din_tr with alternative codes (placeholders) at a rate indicated by the replacement rate data r_rep output from the replacement rate setting unit 11. In this embodiment, the rate indicated by the replacement rate data r_rep is set to "0.1" (10%), and therefore the replacement processing unit 12 targets 10% of the sentences (bilingual data) containing start-end correspondence codes that have been set to be replaced with alternative codes (placeholders) for replacement processing (process of replacing start-end correspondence codes with alternative codes (placeholders)), and performs the replacement processing on the bilingual data that has been targeted for replacement processing.

[0083] Here, as an example of the replacement process, the case of FIG. 3 will be described.

[0084] As shown in Figure 3, the first language (Japanese) data of the ith bilingual data (src i ), and second language (English) data (dst i ) is as follows: <First language (Japanese) data (src i )> [Generic name] Teriparatide (genetical recombination) <Second language (English) data (dst i )> [Non-proprietary name] Teriparatide (Genetical Recombination) Then, the substitution processing unit 12 converts the start / end correspondence code into (1) "()" (Left parenthesis (opening symbol) and right parenthesis (closing symbol)) (2) "[]" (Left bracket (opening symbol) and right bracket (closing symbol)) Therefore, the symbols (1) and (2) above are replaced with placeholder symbols.

[0085] Specifically, the replacement processing unit 12 replaces the data in the first language (Japanese) (src i), and second language (English) data (dst i ), among the start-end correspondence codes, the start code is replaced with "TAGS_k" (or a string containing "TAGS_k"), and the end code is replaced with "TAGE_k" (or a string containing "TAGE_k"). Note that the subscript k of the alternative start code and the alternative end code is set to the same integer value for the same type of start-end correspondence code within the same sentence (within the same bilingual data), and the subscript k is set to an integer value randomly selected from a specified range.

[0086] The bilingual data in Figure 3 ({src i ,dst i}), the replacement processing unit 12 sets the alternative code (placeholder) for the left parenthesis "(", which is the start code of the start-end correspondence code "()", to "_@@@_TAGS_1", and sets the alternative code (placeholder) for the right parenthesis ")", which is the end code of the start-end correspondence code "()", to "_@@@_TAGE_1".

[0087] In addition, the bilingual data in Figure 3 ({src i ,dst i}), the replacement processing unit 12 sets the alternative code (placeholder) for the left bracket "[", which is the start code of the start-end correspondence code "[]", to "_@@@_TAGS_2", and sets the alternative code (placeholder) for the right round bracket "[]", which is the end code of the start-end correspondence code "[]", to "_@@@_TAGE_2" (setting of the replacement target and alternative code).

[0088] Then, the replacement processing unit 12 replaces the first language (Japanese) data (src i ) and the first language data after replacement is src_rep i That is, the replacement processing unit 12 obtains the following data as the first language data after the replacement process: src_rep i (step S104). <First language (Japanese) data after replacement processing (src i )> _@@@_TAGS_2 Generic name _@@@_TAGE_2 Teriparatide _@@@_TAGS_1 Genetic recombinant _@@@_TAGE_1 Furthermore, the replacement processing unit 12 performs the replacement of the second language (English) data (dst i ) and the second language data after the replacement process is dst_rep i That is, the replacement processing unit 12 replaces the following data with the second language data dst_rep after the replacement processing. i (step S105). <Second language (English) data after replacement processing (dst i )> _@@@_TAGS_2 Non - proprietary name _@@@_TAGE_2 Teriparatide _@@@_TAGS_1 Genetical Recombination _@@@_TAGE_1 (Step S106): In step S106, the replacement processing unit 12 receives the first language data src_rep after the replacement processing acquired in steps S104 and S105. i and the second language data after replacement processing dst_rep i The bilingual data after the replacement process ({src_rep i ,dst_rep i}) and obtain the bilingual data after the replacement process ({src_rep i ,dst_rep i}) is output to the data storage unit DB1 as the post-substitution bilingual data Do_tr, and is stored in the data storage unit DB1.

[0089] (Step S107): In step S107, the substitution processing unit 12 determines whether the termination condition of the loop processing (loop 1) is satisfied (whether the substitution processing has been performed on all of the bilingual data that was the target of the substitution processing), and if it determines that the termination condition of the loop processing is not satisfied, the processing returns to step S103 and executes the processing of steps S104 to S106. On the other hand, if it determines that the termination condition of the loop processing is satisfied, the substitution processing unit 12 ends the processing (ends the training data generation processing).

[0090] As a result of the above, in the training data generation device 1, for example, if N pieces of bilingual data are to be subjected to the replacement process, it is possible to obtain N pieces of bilingual data after the replacement process (the proportion of bilingual data on which the replacement process has been performed is 10% (the proportion set by r_rep) of the bilingual sentences containing the start-end correspondence code set as the replacement target).

[0091] The training data generation device 1, through the above process, can insert substitute codes (placeholders) equivalent to markup language tags (e.g., XML tags) into bilingual sentences (bilingual data) that do not include markup language tags (e.g., XML tags). In other words, the training data generation device 1, through the above process, can acquire bilingual sentences (bilingual data) equivalent to bilingual sentences (bilingual data) with markup language tags (e.g., XML tags). In other words, the bilingual data acquired through the above process by the training data generation device 1 includes substitute codes (placeholders) equivalent to markup language tags. Therefore, by using the bilingual data acquired through the above process as training data for the learning process of a machine translation model, it is possible to achieve the same effect as when the learning process of a machine translation model is performed using bilingual sentences (bilingual data) with markup language tags (e.g., XML tags) as training data (it is possible to perform an equivalent learning process).

[0092] (1.2.2: Machine translation model learning process (training process) (creation method)) Next, the learning process (training process) (creation method) of the machine translation model executed in the machine translation processing system 1000 will be described.

[0093] The training data acquisition unit 21 outputs a data read command to the data storage unit DB1, and reads the post-replacement bilingual data stored in the data storage unit DB1 from the data storage unit DB1 as training bilingual data Din_tr_rep(={src_rep j ,dst_rep j The training data acquisition unit 21 reads out the first language data (source language data) (src_rep) from the training bilingual data Din_tr_rep. j ) and use the extracted first language data (the source language data) as the training input data Din_tr (= src_rep j ) to the first selector SEL21. The training data acquisition unit 21 also selects, from the training bilingual data Din_tr_rep, data in a second language (target language data) (dst_rep j ) and use the extracted second language data (translation target language data) as the training correct data D_correct (=dst_rep j ) and outputs it to the loss evaluation unit 24.

[0094] For ease of explanation, it is assumed that the training data acquisition unit 21 reads out M sets (M: natural number, M≦N) of post-substitution bilingual data Din_tr from the data storage unit DB1, and the j-th (j: natural number, 1≦j≦M) first language data of the read bilingual data Din_tr is referred to as “src_rep j " and the data in the second language that is paired with the data in the first language (to form a parallel translation) is written as "dst_rep j " and the j-th data (parallel data) of the parallel data Din_tr is written as "{src_rep j ,dst_rep j}".

[0095] A control unit (not shown) that controls each functional unit of the machine translation processing device 2 outputs a selection signal sel21 whose signal value is set to "0" to the first selector SEL21. The first selector SEL21 selects the data Din_tr in accordance with the selection signal and outputs the selected data Din_tr (=src_rep j ) is output to the machine translation processing unit 23 as data D1.

[0096] The machine translation model of the machine translation processing unit 23 inputs data D1 (=Din_tr) from the first selector SEL21, performs machine translation processing using the machine translation model, and outputs the data obtained by the machine translation processing as data D2 to the second selector SEL22.

[0097] A control unit (not shown) that controls each functional unit of the machine translation processing device 2 outputs a selection signal sel22 whose signal value is "0" to the second selector SEL22. In accordance with the selection signal, the second selector SEL22 selects a path for outputting the data D2 output from the machine translation processing unit 23 to the loss evaluation unit 24, and outputs the data D2 to the loss evaluation unit 24.

[0098] The loss evaluation unit 24 receives the training correct data D_correct output from the training data acquisition unit 21 and the data D21 output from the second selector SEL22. The loss evaluation unit 24 evaluates the loss (e.g., error) between the data D21 and the training correct data D_correct using, for example, a loss function, and generates parameter update data update(θ) that is data for updating parameters of the machine translation model of the machine translation processing unit 23 based on the evaluation result. The loss evaluation unit 24 then outputs the generated parameter update data update(θ) to the machine translation processing unit 23. Note that in FIG. 1, the path from the output of the machine translation processing unit 23 to the loss evaluation unit 24 and the path for outputting the parameter update data update(θ) from the loss evaluation unit 24 to the machine translation processing unit 23 are illustrated as separate paths, but this is for convenience (for convenience of illustration) and is not limited to the form of FIG. 1. In the machine translation processing device 2, when updating the parameters of the machine translation model of the machine translation processing unit 23 using the error backpropagation method, the error obtained by the loss evaluation unit 24 (error obtained by an error function (e.g., cross-entropy error)) can be propagated (backpropagated) sequentially along a path that reverses the path (forward propagation path) along which output data was obtained by the machine translation model of the machine translation processing unit 23, thereby updating each parameter of the machine translation model of the machine translation processing unit 23 (parameters of each layer of the machine translation model of the machine translation processing unit 23).

[0099] In the machine translation processing device 2, the learning process is performed by using the bilingual data ({src_rep j ,dst_rep j}) is executed repeatedly.

[0100] Then, when the error (loss) acquired by the loss evaluation unit 24 (1) falls within a predetermined range, or when (2) the amount of change in the error (loss) acquired by the loss evaluation unit 24 falls within a predetermined range, the loss evaluation unit 24 determines that there is no need to continue the learning process and terminates the learning process. Then, when the learning process is terminated, the parameters set in the machine translation model of the machine translation processing unit 23 are set (fixed) as optimization parameters in the machine translation model of the machine translation processing unit 23, and a trained model of the machine translation model of the machine translation processing unit 23 is acquired.

[0101] As described above, in the machine translation processing system 1000, the learning process (training process) of the machine translation model is executed, and a learned model of the machine translation model of the machine translation processing unit 23 is acquired.

[0102] (1.2.3: Prediction processing (machine translation execution processing)) Next, the prediction process (machine translation execution process) executed by the machine translation processing system 1000 will be described.

[0103] FIG. 4 is a flowchart of the prediction process (machine translation execution process) executed by the machine translation processing system 1000.

[0104] FIG. 5 is a diagram for explaining the prediction process (machine translation execution process) of the machine translation processing system 1000. As shown in FIG.

[0105] The prediction process (machine translation execution process) executed by the machine translation processing system 1000 will be described below with reference to the flowchart of FIG.

[0106] It is assumed that data in the first language (Japanese) including markup language tags (for example, XML tags) is input to the machine translation processing device 2. The following description will be made on the case where the markup language tags are XML tags.

[0107] (Step S201): In step S201, forward permutation processing is performed. Specifically, the following processing is performed.

[0108] The forward substitution processing unit 22 inputs data in the first language (Japanese) to be machine translated (source language data), which includes markup language tags (XML tags), as data Din_src. The data in the first language (source language data) is assumed to be data that has been subjected to morphological analysis and separated into morphemes (word strings, subword strings, etc.).

[0109] The forward substitution processing unit 22 detects markup language tags (XML tags) included in the data Din_src and performs processing (forward substitution processing) to replace the detected markup language tags (XML tags) with substitution codes (placeholders).The forward substitution processing unit 22 then outputs the first language data after the substitution processing to the first selector SEL21 as data Din_rep.

[0110] The forward substitution processing unit 22 performs the forward substitution process by replacing the start and end tags of XML in the data (sentence) of the first language data Din_src containing the input markup language tags (XML tags) with the same substitution codes (placeholders) used in the training data generation process. That is, the forward substitution processing unit 22 (1) replaces the start tag of XML in the data (sentence) of the first language data Din_src containing the input markup language tags (XML tags) with "TAGS_k" (or a character string containing "TAGS_k"), and (2) replaces the end tag of XML in the data (sentence) of the data Din_src with "TAGE_k" (or a character string containing "TAGE_k").

[0111] Then, as in the training data generation process, the subscript k of the alternative code for the XML start tag ("TAGS_k") and the alternative code for the XML end tag ("TAGE_k") is set to the same integer value for the same type of XML start and end tag within the same sentence (within the same input data (within the data of the processing unit that is the target of the forward substitution process)), and the subscript k is set to an integer value randomly taken from a specified range.

[0112] For example, the input data Din_src (= "Today's weather is sunny When the input data Din_src is input to the machine translation processing device 2, the forward substitution processing unit 22 substitutes the XML start tag " " and closing tag " " and finds the opening XML tag " " is replaced with the substitution code "_@@@_TAGS_1", and the closing XML tag " " is replaced with the substitution code "_@@@_TAGE_1", and the data after the forward substitution process Din_rep (="Today's weather is _@@@_TAGS_1 sunny _@@@_TAGE_1.") shown in Figure 5 is obtained.

[0113] The forward substitution processing unit 22 outputs the first language data after the forward substitution processing to the first selector SEL21 as data Din_rep.

[0114] In addition, in the forward substitution process, the forward substitution processing unit 22 generates a list of correspondences between XML tags (tags for markup languages) and substitute codes (placeholders) that replace the XML tags, and outputs data including the list as data D_list_rep to the inverse substitution processing unit 25. In the case of FIG. 5, the forward substitution processing unit 22 generates a list of correspondences between XML tags (tags for markup languages) and substitute codes (placeholders) that replace the XML tags, and outputs data including the list as data D_list_rep to the inverse substitution processing unit 25. " is replaced with the substitution code " _@@@_TAGS_1" and the XML tag " " is replaced with the substitution code " _@@@_TAGE_1 ", and data including this list is output to the inverse replacement processing unit 25 as data D_list_rep.

[0115] A control unit (not shown) that controls each functional unit of the machine translation processing device 2 outputs a selection signal sel21 having a signal value of "0" to the first selector SEL21. The first selector SEL21 selects the data Din_rep output from the forward substitution processing unit 22 in accordance with the selection signal, and outputs the selected data Din_rep to the machine translation processing unit 23 as data D1.

[0116] (Step S202): In step S202, machine translation processing is executed. Specifically, the following processing is executed.

[0117] The machine translation model of the machine translation processing unit 23 receives the data D1 (=Din_tr) from the first selector SEL21 and executes machine translation processing using the machine translation model.

[0118] For example, in the case of Figure 5, when the data Din_rep (="Today's weather is _@@@_TAGS_1 sunny _@@@_TAGE_1.") after the forward substitution process is input to the machine translation model of the machine translation processing unit 23, the machine translation processing unit 23 performs machine translation processing on the input data using the machine translation model (trained model) and obtains the machine translation processing result data (="The weather is _@@@_TAGS_1 fine _@@@_TAGE_1 today.") shown in Figure 5. The machine translation model of the machine translation processing unit 23 is a model that has been optimized by performing a learning process using bilingual data that includes substitution codes (placeholders). Therefore, when data (first language data) in which XML tags have been replaced with substitution codes (placeholders) is input to the machine translation model (trained model), the machine translation model (trained model) outputs (obtains) an appropriate machine-translated sentence (machine translation processing result data (data in the second language (English))) while maintaining the substitution codes (placeholders) in the appropriate position (position within the sentence).

[0119] In this way, the data (data after machine translation processing) acquired by the machine translation model (trained model) of the machine translation processing unit 23 is output as data D2 from the machine translation processing unit 23 to the second selector SEL22.

[0120] A control unit (not shown) that controls each functional unit of the machine translation processing device 2 outputs a selection signal sel22 having a signal value of "1" to the second selector SEL22. In accordance with the selection signal, the second selector SEL22 selects a path for outputting the data D2 output from the machine translation processing unit 23 to the inverse substitution processing unit 25, and outputs the data D2 to the inverse substitution processing unit 25.

[0121] (Step S203): In step S203, the inverse substitution process is performed. Specifically, the following process is performed.

[0122] The inverse substitution processing unit 25 receives data D22 output from the second selector SEL22 and data D_list_rep output from the forward substitution processing unit 22. The inverse substitution processing unit 25 detects, from data D22, substitution codes (placeholders) substituted by the forward substitution processing unit 22, and performs a process (inverse substitution process) of returning (substituting) the detected substitution codes into the original markup language tags based on a list included in data D_list_rep (a list of correspondence between markup language tags and substitution codes (placeholders) that substituted the markup language tags in the forward substitution process).

[0123] For example, in Figure 5, the data D_list_rep contains the XML tag " " is replaced with the substitution code "_@@@_TAGS_1" and the XML tag "Since the data D2 after the machine translation process contains a list indicating that the substitution code "_@@@_TAGS_1" has been replaced with the substitution code "_@@@_TAGE_1", the reverse substitution processing unit 25 acquires the list and performs a process (reverse substitution process) of replacing (reverting) the substitution code included in the data D2 after the machine translation process with the original XML tag. That is, in the case of FIG. 5, in the data D2 after the machine translation process (="The weather is _@@@_TAGS_1 fine _@@@_TAGE_1 today."), the substitution code "_@@@_TAGS_1" has been replaced with the XML tag " " and replace the substitution code "_@@@_TAGE_1" with the XML tag " As a result, the inverse substitution processing unit 25 performs a process of replacing (returning) the data after the inverse substitution process (= "The weather is fine today.")

[0124] Then, the inverse permutation processing unit 25 outputs the data D22 after the inverse permutation processing as output data Do_dst (= “The weather is fine today." (In the case of Figure 5)

[0125] As described above, the machine translation processing system 1000 replaces XML tags in input data containing XML tags with substitution codes (placeholders) similar to those used when generating the training data, and performs machine translation processing using a trained model of a machine translation model optimized with bilingual data into which substitution codes have been inserted, thereby making it possible to obtain appropriate machine translation processing result data while appropriately maintaining the state in which substitution codes have been inserted.The machine translation processing system 1000 then replaces (returns to the original state) substitution codes with XML tags in the machine translation processing result data (machine-translated text) in which substitution codes have been inserted, making it possible to obtain machine translation processing result data (machine-translated text) in which XML tags have been inserted in an appropriate state.

[0126] Fig. 6 shows the results of machine translation of first language data (Japanese data) with XML tags by the machine translation processing system 1000. The upper part of Fig. 6 displays the XML-tagged data (XML source code) of the input data Din_src and the data Do_dst after the reverse substitution process, and the lower part of Fig. 6 displays the interpreted XML tags of the input data Din_src and the data Do_dst after the reverse substitution process. As can be seen from Fig. 6, the machine translation process (from the first language (Japanese) to the second language (English)) has been performed appropriately, with the XML tags maintained in the appropriate positions.

[0127] <Summary> As described above, in the machine translation processing system 1000, the training data generation device 1 performs a training data generation process to detect start-end correspondence symbols (symbols where the left and right correspond, such as () and []) in bilingual sentences (bilingual data) that do not contain markup language tags (for example, XML tags), and by replacing the detected start-end correspondence symbols with substitute symbols (placeholders), it is possible to easily generate large quantities of data equivalent to bilingual data into which markup language tags (for example, XML tags) have been inserted.

[0128] Furthermore, the bilingual data acquired in the training data generation process by the training data generation device 1 of the machine translation processing system 1000 includes substitution codes (placeholders) equivalent to tags for markup languages. Therefore, by using the bilingual data acquired in the training data generation process by the training data generation device 1 as training data for the learning process of the machine translation model, it is possible to achieve the same effect as when the learning process of the machine translation model is performed using bilingual sentences (bilingual data) with tags for markup languages ​​(for example, XML tags) as training data (it is possible to perform the same learning process).

[0129] Furthermore, in the machine translation processing system 1000, for input data containing markup language tags (e.g., XML tags), the markup language tags are replaced with substitution codes (placeholders) similar to those used when generating the training data, and machine translation processing is performed using a trained model of a machine translation model optimized with bilingual data into which the substitution codes have been inserted, so that appropriate machine translation processing result data can be obtained while appropriately maintaining the state in which the substitution codes have been inserted.The machine translation processing system 1000 can then obtain machine translation processing result data (machine-translated text) in which the XML tags have been inserted in an appropriate state by replacing the substitution codes with XML tags (returning them to their original state) in the machine translation processing result data (machine-translated text) in which the substitution codes have been inserted.

[0130] In this way, the machine translation processing system 1000 makes it possible to perform highly accurate machine translation of an original text to be translated that contains markup language tags while retaining the information of the markup language tags, without having to prepare a large amount of tagged bilingual text.

[0131] [Second embodiment] Next, a second embodiment will be described. Note that the same parts as those in the above embodiment are given the same reference numerals and detailed description will be omitted.

[0132] FIG. 7 is a schematic diagram of a machine translation processing system 2000 according to the second embodiment.

[0133] FIG. 8 is a diagram for explaining the replacement process executed by the training data generation device 1A of the machine translation processing system 2000.

[0134] The machine translation processing system 2000 of the second embodiment has a configuration in which the training data generation device 1 in the machine translation processing system 1000 of the first embodiment is replaced with a training data generation device 1A.

[0135] The training data generation device 1A has a configuration in which the replacement processing unit 12 in the training data generation device 1 of the first embodiment is replaced with a replacement processing unit 12A. Otherwise, the machine translation processing system 2000 of the second embodiment is similar to the machine translation processing system 1000 of the first embodiment.

[0136] The replacement processing unit 12A receives bilingual data Din_tr, which is a pair of data in a first language (source language data) and data in a second language (target language data) obtained by translating the first language data into a second language, and does not include markup language tags. The replacement processing unit 12A inserts placeholders (placeholders) around elements that correspond to each other in the bilingual data Din_tr (in the bilingual text). For example, when there is a clear correspondence between the first language data (original text) and the second language data (translated text), such as between proper nouns and numbers, or when word alignment processing is performed and correspondence between words or phrases is established, the replacement processing unit 12A inserts placeholders (placeholders) before and after the corresponding elements. The replacement processing unit 12A uses the same symbols as in the first embodiment as placeholders.

[0137] Specifically, the replacement processing unit 12A (1) inserts the alternative code "TAGS_k" (or a string containing "TAGS_k") for the start code of the first embodiment before an element (word, sub-word, etc.) that corresponds between the first language data (original text) and the second language data (translation), and (2) inserts the alternative code "TAGE_k" (or a string containing "TAGE_k") for the end code of the first embodiment after an element (word, sub-word, etc.) that corresponds between the first language data (original text) and the second language data (translation).

[0138] Here, the case of FIG. 8 will be described as an example of the replacement process by the replacement processor 12A.

[0139] As shown in Figure 8, the first language (Japanese) data of the ith bilingual data (src i ), and second language (English) data (dsti ) is as follows: <First language (Japanese) data (src i )> I work at the National Institute of Information and Communications Technology. <Second language (English) data (dst i )> I am going to work at the National Institute of Information and Communications Technology. Then, the replacement processing unit 12A detects corresponding elements (proper nouns in the above example) between the first language data and the second language data, and performs a process of inserting substitute codes (placeholders) before and after the detected elements. That is, the replacement processing unit 12A detects the proper noun "National Institute of Information and Communications Technology" in the first language data and "the National Institute of Information and Communications Technology" in the second language, which corresponds to the proper noun in the first language (detects the corresponding proper noun), and inserts substitute codes (placeholders) before and after the detected elements (character strings constituting the proper noun in the above example). As a result, the replacement processing unit 12A generates the following post-replacement processing bilingual data ({src_rep i ,dst_rep i}). <First language (Japanese) data after replacement processing (src i )> I will be working at _@@@_TAGS_1 National Institute of Information and Communications Technology _@@@_TAGE_1. <Second language (English) data after replacement processing (dst i )> I am going to work at _@@@_TAGS_1 the National Institute of Information and Communications Technology _@@@_TAGE_1. As in the first embodiment, the replacement processing unit 12A performs the above replacement process (the process of inserting a substitute code (placeholder) to replace the corresponding element) at the ratio set by the replacement ratio setting unit 11 (the ratio indicated by the replacement ratio data r_rep).

[0140] Furthermore, the rate set by the replacement rate setting unit 11 (the rate indicated by the replacement rate data r_rep, 1% in the second embodiment) is preferably set so that the probability of occurrence of substitute codes (placeholders) is similar to the probability of occurrence of markup language tags in the first language data (source language data) with markup language tags that is input to the machine translation processing device 2. In other words, it is preferable that the rate be set so that the probability of occurrence (occurrence probability distribution) of substitute codes (placeholders) in the bilingual data Do_tr after the replacement process is close to the probability of occurrence (occurrence probability distribution) of markup language tags in the first language data (source language data) (data to be subjected to machine translation) that is input to the machine translation processing device 2. In this way, the probability distribution of occurrence of substitute codes (placeholders) in the training data becomes close to the probability distribution of occurrence of markup language tags in the markup language tagged language data that is actually the target of machine translation processing, thereby improving the accuracy of the learning process of the machine translation process using the training data.

[0141] The data Do_tr acquired by the training data generation device 1A through the above process is stored in the data storage unit DB1, and, as in the first embodiment, is used in the learning process (training process) of the machine translation model in the machine translation processing system 2000. Then, in the machine translation processing system 2000 where the learning process has been completed, a prediction process (machine translation execution process) is executed.

[0142] As described above, in the machine translation processing system 2000, the training data generation device 1A performs a training data generation process to detect elements that correspond between the original text and the translation in bilingual text (bilingual data) that does not contain markup language tags (e.g., XML tags), and by replacing the elements before and after the detected elements with substitute codes (placeholders), it is possible to easily generate large quantities of data equivalent to bilingual data in which markup language tags (e.g., XML tags) have been inserted.

[0143] Furthermore, the bilingual data acquired in the training data generation process by the training data generation device 1A of the machine translation processing system 2000 includes substitution codes (placeholders) equivalent to tags for markup languages. Therefore, by using the bilingual data acquired in the training data generation process by the training data generation device 1A as training data for the learning process of the machine translation model, it is possible to achieve the same effect as when the learning process of the machine translation model is performed using bilingual sentences (bilingual data) with tags for markup languages ​​(for example, XML tags) as training data (it is possible to perform the same learning process).

[0144] Furthermore, in the machine translation processing system 2000, for input data containing markup language tags (e.g., XML tags), the markup language tags are replaced with substitution codes (placeholders) similar to those used when generating the training data, and machine translation processing is performed using a trained model of a machine translation model optimized with bilingual data into which substitution codes have been inserted, making it possible to obtain appropriate machine translation processing result data while appropriately maintaining the state in which substitution codes have been inserted.The machine translation processing system 2000 can then obtain machine translation processing result data (machine-translated text) in which XML tags have been inserted in an appropriate state by replacing the substitution codes with XML tags (returning them to their original state) in the machine translation processing result data (machine-translated text) in which substitution codes have been inserted.

[0145] In this way, the machine translation processing system 2000 makes it possible to perform highly accurate machine translation of an original text to be translated that contains markup language tags while retaining the information of the markup language tags, without having to prepare a large amount of tagged bilingual texts.

[0146] [Other embodiments] Each functional unit of the machine translation processing systems 1000 and 2000 described in the above embodiments may be realized by one device (system) or by multiple devices.

[0147] Furthermore, some or all of the above embodiments may be combined.

[0148] In the above embodiment, the case where bilingual data or first language data that has undergone morphological analysis processing is input to the training data generation device 1, 1A and the machine translation processing device 2 has been described. However, this is not limited thereto, and bilingual data or first language data that has not undergone morphological analysis processing may also be input to the training data generation device 1, 1A and the machine translation processing device 2. In this case, the morphological analysis unit may be provided upstream of the substitution processing unit 12, 12A and the forward substitution processing unit 22. Then, bilingual data of data sequences (word sequences, subword sequences) separated into morphemes by the morphological analysis unit, or data in a language to be machine translated (first language data), may be input to the training data generation device 1, 1A or the machine translation processing device 2.

[0149] In the above embodiment, the first language data is Japanese and the second language data is English, but this is not limiting, and the first language data and / or the second language data may be in other languages. In other words, in the machine translation processing systems 1000 and 2000 of the above embodiment, the source language and the target language may be any language.

[0150] In addition, if a start-end correspondence code that is commonly used in the first language data and the second language data exists, the machine translation processing systems 1000 and 2000 may perform a replacement process to replace the start-end correspondence code with an alternative code (placeholder).

[0151] In the machine translation processing systems 1000 and 2000 described in the above embodiments, each block may be individually implemented as a single chip using a semiconductor device such as an LSI, or some or all of the blocks may be integrated into a single chip.

[0152] Although we refer to it as an LSI here, it may also be called an IC, system LSI, super LSI, or ultra LSI depending on the level of integration.

[0153] Furthermore, the method of integration is not limited to LSI, but may be realized by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow the connections and settings of circuit cells within LSI to be reconfigured, may also be used.

[0154] Furthermore, part or all of the processing of each functional block in each of the above embodiments may be realized by a program. And part or all of the processing of each functional block in each of the above embodiments is performed by a central processing unit (CPU) in a computer. Furthermore, the programs for performing each processing are stored in a storage device such as a hard disk or ROM, and are executed in the ROM or read out to the RAM.

[0155] Each process in the above-described embodiments may be realized by hardware, software (including cases where it is realized together with an OS (operating system), middleware, or a predetermined library), or may be realized by a combination of software and hardware.

[0156] For example, when each functional unit in the above embodiment is realized by software, each functional unit may be realized by software processing using the hardware configuration shown in FIG. 9 (for example, a hardware configuration in which a CPU, GPU, ROM, RAM, input unit, output unit, communication unit, memory unit (for example, a memory unit realized by an HDD, SSD, etc.), an external media drive, etc. are connected via a bus).

[0157] Furthermore, when each functional unit of the above embodiment is realized by software, the software may be realized using a single computer having the hardware configuration shown in Figure 9, or may be realized by distributed processing using multiple computers.

[0158] The execution order of the processing method in the above embodiment is not necessarily limited to that described in the above embodiment, and the execution order can be changed within the scope of the gist of the invention. Furthermore, in the processing method in the above embodiment, some steps may be executed in parallel with other steps within the scope of the gist of the invention.

[0159] The scope of the present invention includes a computer program for causing a computer to execute the above-described method, and a computer-readable recording medium having the program recorded thereon. Examples of computer-readable recording media include flexible disks, hard disks, CD-ROMs, MOs, DVDs, DVD-ROMs, DVD-RAMs, large-capacity DVDs, next-generation DVDs, and semiconductor memories.

[0160] The computer program is not limited to one recorded on the recording medium, but may be one transmitted via a telecommunications line, a wireless or wired communication line, a network such as the Internet, or the like.

[0161] The specific configuration of the present invention is not limited to the above-described embodiment, and various changes and modifications are possible without departing from the gist of the invention. [Explanation of symbols]

[0162] 1000, 2000 Machine Translation Processing System 1. 1A Training data generator 11 Replacement ratio setting unit 11 12, 12A Replacement processing section 2. Machine translation processing device 22 Forward substitution processing section 23 Machine translation processing unit 24 Loss Assessment Department 25 Reverse substitution processing section

Claims

1. 1. A method for generating training data for training a trainable model for machine translation processing in a machine translation processing system for machine translation processing of language data including tags for a markup language, comprising: a start-end correspondence code detection step for detecting a start-end correspondence code, which is a code whose start and end correspond to each other, in bilingual data that is a combination of first language data and second language data that is data obtained by translating the first language data into a second language and does not include the markup language tag; a replacement processing step of performing a replacement process on the bilingual data to replace the start / end corresponding code with an alternative code, thereby obtaining the bilingual data after the replacement process; A method for generating training data for machine translation comprising:

2. a replacement ratio setting step for setting a replacement ratio; The replacement processing step includes: performing a replacement process on the bilingual data to replace the start-end corresponding code with an alternative code at the replacement ratio set in the replacement ratio setting step; 2. The method for generating training data for machine translation according to claim 1.

3. 3. A method for training a trainable model for machine translation processing in a machine translation processing system for machine translating language data including markup language tags, using training data generated by the method for generating training data for machine translation according to claim 1 or 2, comprising: a data input step of inputting the first language data included in the bilingual data after the replacement process into the trainable model for machine translation processing; an output data acquisition step of acquiring output data of the trainable model for machine translation processing for the data input in the data input step; a loss evaluation step of acquiring the output data acquired by the output data acquisition step and the second language data included in the bilingual data after the replacement process as correct data, and evaluating a loss between the output data and the correct data; a parameter updating step of updating parameters of the trainable model for machine translation processing so as to reduce the loss obtained by the loss evaluation step; A method for creating a trainable model for machine translation processing comprising:

4. A method for performing machine translation processing using a trained model of a trainable model for machine translation processing obtained by training using the method for creating a trainable model for machine translation processing according to claim 3, a forward substitution processing step of executing a forward substitution process for replacing the markup language tag included in the input first language data with the substitution code; a machine translation processing step of performing machine translation processing on the first language data after the forward substitution processing using a trained model of the trainable model for machine translation processing to obtain second language data after the machine translation processing; a reverse substitution processing step of executing a reverse substitution processing to replace the substitution code included in the second language data after the machine translation processing acquired by the machine translation processing step with the markup language tag replaced in the forward substitution processing step; A machine translation processing method comprising:

5. 1. A method for generating training data for training a trainable model for machine translation processing in a machine translation processing system for machine translation processing of language data including tags for a markup language, comprising: a corresponding element detection step of detecting corresponding elements, which are elements that are determined to correspond between the first language data and the second language data, in bilingual data that is a combination of first language data and second language data that is data obtained by translating the first language data into a second language and does not include the markup language tag; a replacement processing step of performing a replacement process on the bilingual data by inserting replacement codes before and after the corresponding elements, thereby obtaining bilingual data after the replacement process; A method for generating training data for machine translation comprising:

6. 1. A machine translation processing system for machine-translating language data including tags for a markup language, comprising: an apparatus for generating training data for training a trainable model for machine translation processing, the apparatus comprising: Detecting start-end correspondence codes, which are codes whose start and end correspond to each other, in bilingual data that is a combination of first language data and second language data that is data obtained by translating the first language data into a second language and does not include the markup language tag; a replacement processing unit that performs a replacement process on the bilingual data to replace the start / end corresponding code with an alternative code, thereby obtaining the bilingual data after the replacement process; A training data generation device for machine translation comprising: