Progressive training-based chapter-level machine translation model training method and device and medium
Through progressive training and data augmentation methods, the problem of scarce training resources of chapter-level machine translation model is solved, and more efficient translation results are achieved.
Patent Information
- Application Number
- CN202510191353.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-23
AI Technical Summary
The existing technology is difficult to effectively utilize chapter-level corpus, resulting in scarce and difficult to optimize the training resources of chapter-level machine translation model.
Using a progressive training method, we use segmented data augmentation and sentence-level corpus expansion to convert chapter-level corpus into documents of different widths, and use transformer models for progressive training.
By enhancing corpus and gradual training, the problem of lack of corpus in chapter-level model training is solved, making the model easier to learn and improve the translation quality.
Smart Images

Figure CN120031052A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine translation, and in particular, relates to a training method, device and medium for a paragraph-level machine translation model based on progressive training. Background Art
[0002] In the field of machine translation, current researchers mainly focus on sentence-level models, that is, inputting a sentence of original text and outputting a sentence of translation. However, in real translation business, professional translators often need to process the contextual information of the entire article. It is difficult for sentence-level machine translation models to take into account the contextual information at the paragraph level, making it difficult for the human-machine combined MT-PE model to achieve better performance. In the related research on paragraph-level machine translation, people often only focus on the scarce paragraph-level corpus to train the paragraph-level model. A large amount of annotated sentence-level corpus is difficult to use, which makes paragraph-level machine translation transformed into machine translation in a scarce resource scenario and difficult to optimize. Summary of the invention
[0003] The purpose of the present invention is to provide a training method for a paragraph-level machine translation model based on progressive training to solve the technical problems existing in the prior art.
[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows: A training method for a paragraph-level machine translation model based on progressive training, comprising: Step S1: Segmentation-based chapter-level data enhancement: segment the chapter-level machine translation corpus and convert it into documents of different widths to obtain the initial chapter-level corpus; Step S2: Supplement the sentence pair corpus to obtain the chapter-level corpus: Use the sentence pair corpus to supplement the initial chapter-level corpus; Step S3: Based on the chapter-level corpus obtained in step S2, a progressive learning method is used to complete the training of the chapter-level machine translation model. The progressive learning method is as follows: Step S3.1: Calculate the learning difficulty weights of all chapter-level corpora :
[0005] in, is the number of original sentences in the i-th paragraph-level training corpus, is the sentence length; Step S3.2: Sort the learning difficulty weight arrays of all chapter-level corpora from small to large, and obtain the sorted learning difficulty weight arrays of all chapter-level corpora ; Step S3.3: Learning difficulty weight array Perform segmentation to obtain a number of learning segments; Step S3.4: Use the transformer model as the model base and the passage-level corpus, and gradually complete the training of the passage-level machine translation model according to the learning segments from simple to difficult.
[0006] In one implementation, in the step S1, after the passage-level machine translation corpus is segmented, it at least includes the corpus disassembled to the sentence level.
[0007] In one implementation, the specific method of the step S2 is as follows: Parse the syntactic tree structures of the sentence pair corpus and the passage-level corpus at the same time; after obtaining the syntactic tree structures of the sentence pair corpus and the passage-level corpus, calculate the edit distance between the two sequences, and retain 50% of the sentence pair corpus with the smallest edit distance.
[0008] In one implementation, the method of segmentation in the step S3.3 is as follows: Assume that the number of learning segments is k, and the data volume of each segment is , segment the learning difficulty weight value array, then the learning difficulty array of the first learning segment is , and the learning difficulty array of the second learning segment is , and so on until the training data volume reaches the full volume.
[0009] To achieve the above object, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the training method of the passage-level machine translation model based on progressive training as described above.
[0010] To achieve the above object, the present invention also provides a training device for a passage-level machine translation model based on progressive training, including: a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the device for automatically aligning the endnotes and footnotes in the bilingual scenario executes the training method of the passage-level machine translation model based on progressive training as described above.
[0011] Compared with the prior art, the present invention has the following beneficial effects: (1) According to the present invention, the split-type data augmentation scheme and the sentence-level corpus augmentation are used to augment the passage-level corpus. By the split-type passage-level corpus augmentation and the sentence-level corpus supplement, the problem of lack of corpus in the training process of the passage-level model is solved.
[0012] (2) According to the present invention, through the progressive training of the passage-level model, the difficult passage-level training is converted into a training process from easy to difficult, and the model is easier to learn. Description of the Drawings
[0013] Figure 1 This is a schematic diagram of the process of Example 1 of the present invention.
[0014] Figure 2 It is a schematic diagram of training representation of the present invention. DETAILED DESCRIPTION
[0015] In order to enable those skilled in the art to have a clearer understanding and knowledge of the present invention, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described below are only used to explain the present invention, and are convenient for understanding. The technical solutions provided by the present invention are not limited to the technical solutions provided by the following embodiments, and the technical solutions provided by the embodiments should not limit the protection scope of the present invention.
[0016] It should be noted that the illustrations provided in the following embodiments are only used to illustrate the basic concept of the present invention in a schematic manner. Therefore, the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the form, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0017] Example 1 like Figure 1 , 2 As shown, this embodiment provides a training method for a paragraph-level machine translation model based on progressive training, which solves the problems of scarce training resources and difficulty in optimization of the paragraph-level machine translation model through multiple data enhancement methods and progressive training.
[0018] In this embodiment, the training method of the paragraph-level machine translation model based on progressive training mainly includes the following steps: 1. Step S1: Segmentation-based chapter-level data enhancement The chapter-level machine translation corpus is segmented and converted into documents of different widths to obtain the initial chapter-level corpus. After the chapter-level machine translation corpus is segmented, it at least includes the corpus that has been broken down to the sentence level. For example, for a complete chapter data, we will perform operations such as splitting the chapter into two or four parts to convert it into documents of different widths. For example, if an original chapter has 8 sentences, chapters with lengths of 4, 2, and 1 sentences can be obtained at the same time through the above method.
[0019] 2. Supplement sentence-pair corpus to obtain paragraph-level corpus: Sentence pair corpus, that is, the sentence-level translation training data used in historical sentence-level translation tasks, is used to supplement the initial paragraph-level corpus. Specifically, the grammatical tree structure is parsed for both the sentence pair corpus and the paragraph-level corpus at the same time, and the semantic extraction model used is an open source model.
[0020] After obtaining the grammatical tree structures of the sentence pair corpus and the paragraph-level corpus, the edit distance of the two sequences is calculated, and 50% of the sentence pair corpora with the smallest edit distance are retained. The ratio of the edit distance can be adjusted according to the corpus demand, but a higher ratio will introduce some data with dissimilar semantic structures, and a lower ratio will make the corpus relatively scarce. Therefore, the above-mentioned ratio of the edit distance is obtained by the inventor of this application through creative labor.
[0021] Based on the above, the addition of a large amount of sentence-level corpus can make up for the problem of insufficient paragraph-level corpus.
[0022] Step S3: Use progressive learning method to complete the training of the chapter-level machine translation model: The progressive learning method can transform the more complex paragraph-level translation task into a simpler progressive machine translation. In theory, it is easier for the model to learn with increasing difficulty than to directly learn the most difficult level of paragraph-level translation.
[0023] In this embodiment, the progressive learning method mainly includes the following: (1) Calculate the learning difficulty weights of all chapter-level corpora :
[0024] in, is the number of original sentences in the i-th paragraph-level training corpus, is the sentence length; (2) Sort the learning difficulty weight array of all chapter-level corpora from small to large, and obtain the sorted learning difficulty weight array of all chapter-level corpora ; (3) Learning difficulty weight array Segment to obtain several learning segments; the segmentation method is as follows: Assume that the number of learning segments is k, and the amount of data in each segment is , segment the learning difficulty weight array, then the learning difficulty array of the first learning segment is , the learning difficulty array of the second learning segment is , and so on, until the amount of training data reaches the full amount.
[0025] (4) Using the transformer model as the model base and paragraph-level corpus, the training of the paragraph-level machine translation model is gradually completed according to the learning segmentation from simple to easy.
[0026] Through the above, the technical solution provided in this embodiment uses a split-type data augmentation scheme and sentence-level corpus augmentation to augment the document-level corpus; at the same time, a progressive training method is adopted to train the document-level machine translation model, thereby solving the problems of scarce training resources and difficulty in optimization of the document-level machine translation model.
[0027] Embodiment 2 This embodiment provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the training method of the document-level machine translation model based on progressive training provided in Embodiment 1. Those of ordinary skill in the art can understand that all or part of the steps of implementing the method provided in Embodiment 1 can be completed by hardware related to the computer program. The above computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the method provided in Embodiment 1; and the above storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0028] Embodiment 3 This embodiment provides a training device for a document-level machine translation model based on progressive training, including: a processor and a memory; the memory is used to store a computer program; the processor is connected to the memory and is used to execute the computer program stored in the memory, so that the training device for the document-level machine translation model based on progressive training executes the training method of the document-level machine translation model based on progressive training provided in Embodiment 1.
[0029] Specifically, the memory includes: various media such as ROM, RAM, magnetic disk, USB flash drive, memory card, or optical disc that can store program codes.
[0030] Preferably, the processor can be a general-purpose processor, including a central processing unit, a network processor, etc.; it can also be a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0031] The above embodiments only illustrate the principles and effects of the present invention, rather than limiting the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A training method for a paragraph-level machine translation model based on progressive training, characterized in that: include: Step S1: Segmentation-based chapter-level data enhancement: segment the chapter-level machine translation corpus and convert it into documents of different widths to obtain the initial chapter-level corpus; Step S2: Supplement the sentence pair corpus to obtain the chapter-level corpus: Use the sentence pair corpus to supplement the initial chapter-level corpus; Step S3: Based on the chapter-level corpus obtained in step S2, a progressive learning method is used to complete the training of the chapter-level machine translation model. The progressive learning method is as follows: Step S3.1: Calculate the learning difficulty weights of all chapter-level corpora :
2. Among them, is the number of original sentences in the i-th paragraph-level training corpus, is the sentence length; Step S3.2: Sort the learning difficulty weight arrays of all chapter-level corpora from small to large, and obtain the sorted learning difficulty weight arrays of all chapter-level corpora ; Step S3.3: Learning difficulty weight array Perform segmentation to obtain several learning segments; Step S3.4: Using the transformer model as the model base and the paragraph-level corpus, gradually complete the training of the paragraph-level machine translation model according to the learning segmentation from simple to easy.
3. The training method of the paragraph-level machine translation model based on progressive training according to claim 1 is characterized in that: In the step S1, after the segmentation of the paragraph-level machine translation corpus, at least the corpus is broken down to the sentence level.
4. The training method of the paragraph-level machine translation model based on progressive training according to claim 2 is characterized in that: The specific method of step S2 is as follows: performing grammatical tree structure analysis on the sentence pair corpus and the paragraph-level corpus at the same time; after obtaining the grammatical tree structures of the sentence pair corpus and the paragraph-level corpus, performing edit distance calculation on the two sequences, and retaining 50% of the sentence pair corpora with the smallest edit distance.
5. The training method of the paragraph-level machine translation model based on progressive training according to claim 3 is characterized in that: The segmentation method in step S3.3 is as follows: Assuming the number of learning segments is k, the amount of data for each segment is , segment the learning difficulty weight array, then the learning difficulty array of the first learning segment is , the learning difficulty array of the second learning segment is , and so on, until the amount of training data reaches the full amount.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the training method of a paragraph-level machine translation model based on progressive training as described in any one of claims 1 to 4.
7. A training device for a paragraph-level machine translation model based on progressive training, characterized in that: include: Processor and memory; The memory is used to store computer programs; The processor is connected to the memory and is used to execute the computer program stored in the memory, so that the device for automatically aligning endnote and footnote numbers in the bilingual scenario executes the training method of the paragraph-level machine translation model based on progressive training as described in any one of claims 1 to 4.