A method and apparatus for generating a synthetic corpus

By performing quality evaluation and correlation scores on candidate statements and seed statements, a synthetic corpus is generated, which solves the problem of low training quality of NMT model caused by the difficulty of obtaining parallel corpus, and achieves high-quality translation results.

CN114648033BActive Publication Date: 2025-06-27INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210282394.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-22
Publication Date
2025-06-27
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

It is difficult to obtain large-scale parallel corpus in some fields, resulting in low training quality of neural machine translation (NMT) models and insufficient text quality and vocabulary diversity of pseudo-parallel corpus generated by existing back translation methods.

Method used

By performing quality evaluation and correlation scoring of seed statements in the preset seed dataset and candidate statements in the training dataset, a synthetic corpus is generated to improve the translation quality of the NMT model.

Benefits of technology

A synthetic corpus with high text quality and vocabulary diversity was obtained, thereby greatly improving the translation quality of the NMT model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114648033B_ABST
    Figure CN114648033B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a method and apparatus for generating a synthetic corpus, which can be used in the field of artificial intelligence technology. The method includes: evaluating the quality of candidate sentences in a training dataset according to the seed sentences in a preset seed dataset to generate a quality score for each candidate sentence; generating a comprehensive score for each candidate sentence according to the quality score and the relevance score of each candidate sentence and the seed sentence calculated in advance; and generating a synthetic corpus according to the comprehensive scores of each candidate sentence corresponding to each seed sentence, so as to obtain a synthetic corpus with high text quality and high lexical diversity, thereby greatly improving the translation quality of the NMT model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language technology, particularly to the field of artificial intelligence technology, and more particularly to a method and apparatus for generating a synthetic corpus. Background Art

[0002] Currently, machine translation technology is widely used in many fields such as finance. The neural machine translation (NMT) model can convert the source language into the target language. The NMT model needs to be trained through a large number of datasets, that is, parallel corpora. However, it is difficult to obtain a large-scale parallel corpus in some fields, resulting in the NMT model not meeting the training requirements. In related technologies, the method of back translation is usually used to generate a pseudo-parallel corpus for training the NMT model. However, the text quality and lexical diversity of the corpus obtained by the existing back translation methods are not high, so the improvement of the translation quality of the trained NMT model is relatively small. Summary of the Invention

[0003] An object of the present invention is to provide a method for generating a synthetic corpus, which can obtain a synthetic corpus with high text quality and high lexical diversity, thereby greatly improving the translation quality of the NMT model. Another object of the present invention is to provide an apparatus for generating a synthetic corpus. Still another object of the present invention is to provide a computer-readable medium. Yet another object of the present invention is to provide a computer device.

[0004] To achieve the above objectives, on the one hand, the present invention discloses a method for generating a synthetic corpus, including:

[0005] Evaluating the quality of candidate sentences in the training dataset according to the seed sentences in the preset seed dataset to generate a quality score for each candidate sentence;

[0006] Generating a comprehensive score for each candidate sentence according to the quality score and the pre-calculated relevance score between each candidate sentence and the seed sentence;

[0007] Generating a synthetic corpus according to the comprehensive score of each candidate sentence corresponding to each seed sentence.

[0008] Preferably, evaluating the quality of candidate sentences in the training dataset according to the seed sentences in the preset seed dataset to generate a quality score for each candidate sentence includes:

[0009] Evaluating the quality of candidate sentences in the training dataset according to the seed sentences in the preset seed dataset through a quality evaluation algorithm to generate a quality score for each candidate sentence.

[0010] Preferably, through a quality assessment algorithm, according to the seed statements in the preset seed dataset, the quality of the candidate statements in the training dataset is evaluated to generate a quality score for each candidate statement, including:

[0011] Through the quality assessment algorithm, multiple quality metrics are generated based on the candidate statements and the seed statements;

[0012] Calculate multiple quality metrics to obtain the quality score for each candidate statement.

[0013] Preferably, the quality assessment algorithm includes algorithms for measuring text and lexical diversity, machine translation evaluation metric algorithms, and translation error rate algorithms;

[0014] Through the quality assessment algorithm, multiple quality metrics are generated based on the candidate statements and the seed statements, including:

[0015] Through the algorithm for measuring text and lexical diversity, measure the text and lexical diversity of the candidate statements to obtain the first quality metric;

[0016] Through the machine translation evaluation metric algorithm, calculate the candidate statements and the seed statements to obtain the second quality metric for each candidate statement;

[0017] Through the translation error rate algorithm, calculate the candidate statements and the seed statements to obtain the third quality metric for each candidate statement.

[0018] Preferably, before generating the comprehensive score for each candidate statement based on the quality score and the pre-calculated relevance score between each candidate statement and the seed statement, it further includes:

[0019] Through the feature attenuation algorithm, generate the relevance score between each candidate statement and the seed statement based on the candidate statements and the seed statements.

[0020] Preferably, generating the comprehensive score for each candidate statement based on the quality score and the pre-calculated relevance score between each candidate statement and the seed statement includes:

[0021] Multiply the quality score and the relevance score to obtain the comprehensive score for each candidate statement.

[0022] Preferably, generating a synthetic corpus based on the comprehensive score of each candidate statement corresponding to each seed statement includes:

[0023] Sort the candidate statements corresponding to each seed statement according to the comprehensive score, select the candidate statement corresponding to the highest comprehensive score, and use the selected candidate statement as the training statement;

[0024] Generate a synthetic corpus based on the training statements corresponding to each seed statement.

[0025] Preferably, after generating the synthetic corpus according to the comprehensive scores of each candidate sentence corresponding to each seed sentence, the method further includes:

[0026] Training a pre-trained translation model according to the synthetic corpus to obtain a neural machine translation model.

[0027] The present invention also discloses a synthetic corpus generation device, including:

[0028] A quality evaluation unit configured to evaluate the quality of candidate sentences in a training dataset according to seed sentences in a preset seed dataset, and generate a quality score for each candidate sentence;

[0029] A comprehensive evaluation unit configured to generate a comprehensive score for each candidate sentence according to the quality score and the relevance score of each candidate sentence to the seed sentence calculated in advance;

[0030] A generation unit configured to generate a synthetic corpus according to the comprehensive scores of each candidate sentence corresponding to each seed sentence.

[0031] The present invention also discloses a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method is implemented.

[0032] The present invention also discloses a computer device, including a memory and a processor, where the memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the processor executes the program, the above-mentioned method is implemented.

[0033] The present invention evaluates the quality of candidate sentences in a training dataset according to seed sentences in a preset seed dataset, generates a quality score for each candidate sentence; generates a comprehensive score for each candidate sentence according to the quality score and the relevance score of each candidate sentence to the seed sentence calculated in advance; and generates a synthetic corpus according to the comprehensive scores of each candidate sentence corresponding to each seed sentence, so as to obtain a synthetic corpus with high text quality and high lexical diversity, thereby greatly improving the translation quality of the NMT model. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1Flowchart of a method for generating a synthetic corpus provided by an embodiment of the present invention;

[0036] Figure 2 Flowchart of another method for generating a synthetic corpus provided by an embodiment of the present invention;

[0037] Figure 3 Schematic structural diagram of a device for generating a synthetic corpus provided by an embodiment of the present invention;

[0038] Figure 4 Schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0040] It should be noted that a method and device for generating a synthetic corpus disclosed in the present application can be used in the field of artificial intelligence technology, and can also be used in any field other than the field of artificial intelligence technology. The application fields of the method and device for generating a synthetic corpus disclosed in the present application are not limited.

[0041] To facilitate understanding of the technical solutions provided in the present application, the relevant content of the technical solutions in the present application will be described first. Machine translation is a technology that uses a computer to achieve natural language translation. The neural machine translation (NMT) model is a technology that uses a single neural network to maximize the performance of machine translation. Using the neural machine translation (NMT) model, the source language can be converted into the target language. Currently, it is difficult to obtain a large-scale parallel corpus in some fields. For example, the data volume of the parallel corpus in the financial field is insufficient to meet the training requirements of NMT. The cost of manual translation is high and the translation quality cannot be guaranteed, and the translation efficiency is low. The back-translation method is to use a translation model to translate the monolingual corpus of the source language into the target language, thereby obtaining a pseudo-parallel corpus. The text quality and vocabulary diversity of the pseudo-parallel corpus are low, and the translation quality of the NMT model directly trained using the pseudo-parallel corpus is low. To solve the above technical problems, the embodiments of the present invention perform data selection on the pseudo-parallel corpus obtained by the back-translation technology, optimize the pseudo-parallel corpus, and obtain a synthetic corpus; by training the NMT model with the synthetic corpus, a high-quality NMT model can be obtained, thereby greatly improving the translation quality of the NMT model.

[0042] Taking the synthetic corpus generation device as an example of the execution subject, the implementation process of the synthetic corpus generation method provided by the embodiments of the present invention will be described. It can be understood that the execution subject of the synthetic corpus generation method provided by the embodiments of the present invention includes but is not limited to the synthetic corpus generation device.

[0043] Figure 1 The flowchart of a synthetic corpus generation method provided by the embodiments of the present invention is as Figure 1 shown, and the method includes:

[0044] Step 101, perform quality evaluation on the candidate sentences in the training dataset according to the seed sentences in the preset seed dataset, and generate a quality score for each candidate sentence.

[0045] Step 102, generate a comprehensive score for each candidate sentence according to the quality score and the pre-computed relevance score between each candidate sentence and the seed sentence.

[0046] Step 103, generate a synthetic corpus according to the comprehensive scores of each candidate sentence corresponding to each seed sentence.

[0047] In the technical solution provided by the embodiments of the present invention, the quality evaluation is performed on the candidate sentences in the training dataset according to the seed sentences in the preset seed dataset to generate a quality score for each candidate sentence; a comprehensive score for each candidate sentence is generated according to the quality score and the pre-computed relevance score between each candidate sentence and the seed sentence; a synthetic corpus is generated according to the comprehensive scores of each candidate sentence corresponding to each seed sentence, and a synthetic corpus with high text quality and high lexical diversity can be obtained, thereby greatly improving the translation quality of the NMT model.

[0048] Figure 2 The flowchart of another synthetic corpus generation method provided by the embodiments of the present invention is as Figure 2 shown, and the method includes:

[0049] Step 201, perform quality evaluation on the candidate sentences in the training dataset according to the seed sentences in the preset seed dataset, and generate a quality score for each candidate sentence.

[0050] In the embodiments of the present invention, each step is executed by the synthetic corpus generation device.

[0051] In the embodiments of the present invention, the pseudo-parallel corpus obtained by back translation is divided according to a preset ratio to obtain a training data set, a validation data set, and a test data set. The training data set includes multiple candidate sentences for subsequent retrieval according to the seed sentences in the seed data set; in the embodiments of the present invention, the validation data set is used as the seed data set, and the seed data set includes multiple seed sentences; the test data set includes multiple test sentences for subsequent testing of the NMT model. It should be noted that the preset ratio can be set according to the actual situation, and the specific value of the ratio is not limited in the embodiments of the present invention.

[0052] In the embodiments of the present invention, since the seed data set and the training data set are obtained by dividing the same pseudo-parallel corpus, the seed data set and the training data set are corpora in the same field, laying a good quality foundation for the generation of the synthetic corpus.

[0053] Specifically, through a quality evaluation algorithm, according to the seed sentences in the preset seed data set, the quality of the candidate sentences in the training data set is evaluated to generate a quality score for each candidate sentence.

[0054] In the embodiments of the present invention, step 201 specifically includes:

[0055] Step 2011: Through a quality evaluation algorithm, multiple quality indicators are generated according to the candidate sentences and the seed sentences.

[0056] In the embodiments of the present invention, the quality evaluation algorithm includes the measure of text and lexical diversity (MTLD) algorithm, the machine translation evaluation metric (BLEU) algorithm, and the translation error rate (TER) algorithm.

[0057] In the embodiments of the present invention, through the measure of text and lexical diversity (MTLD) algorithm, the text and lexical diversity of the candidate sentences are measured to obtain a first quality indicator.

[0058] Specifically, the candidate sentences are input into the measure of text and lexical diversity (MTLD) algorithm, and the first quality indicator is output. The first quality indicator is a numerical value. The larger the first quality indicator, the higher the text and lexical diversity; the smaller the first quality indicator, the lower the text and lexical diversity.

[0059] In the embodiments of the present invention, through the machine translation evaluation metric (BLEU) algorithm, the candidate sentences and the seed sentences are calculated to obtain a second quality indicator for each candidate sentence.

[0060] Specifically, the candidate sentence and the seed sentence are input into the Bilingual Evaluation Understudy (BLEU) algorithm, and a second quality metric is output. The second quality metric is a numerical value representing the similarity between the seed sentence and the candidate sentence. The larger the second quality metric, the higher the similarity between the seed sentence and the candidate sentence; the smaller the second quality metric, the lower the similarity between the seed sentence and the candidate sentence.

[0061] In an embodiment of the present invention, the translation error rate (TER) algorithm is used to calculate the candidate sentence and the seed sentence to obtain a third quality metric for each candidate sentence.

[0062] Specifically, the candidate sentence and the seed sentence are input into the translation error rate (TER) algorithm, and a third quality metric is output. The third quality metric is a numerical value. The smaller the third quality metric, the lower the error rate and the higher the translation quality; the larger the third quality metric, the higher the error rate.

[0063] Step 2012: Calculate multiple quality metrics to obtain a quality score for each candidate sentence.

[0064] In an embodiment of the present invention, through the formula calculate the first quality metric, the second quality metric, and the third quality metric to obtain the quality score of the candidate sentence. Wherein, is the quality score of the candidate sentence, Q BLEU is the second quality metric, Q TER is the third quality metric, Q MTLD is the first quality metric.

[0065] In an embodiment of the present invention, considering the text and lexical diversity and the translation error rate comprehensively, the obtained quality score can more significantly reflect the correlation between the seed sentence and the candidate sentence, which is beneficial to optimizing the parallel corpus.

[0066] Step 202: Generate a relevance score between each candidate sentence and the seed sentence according to the candidate sentence and the seed sentence through the Feature Decay Algorithms (FDA) algorithm.

[0067] In an embodiment of the present invention, FDA is a data selection technique used to select high-quality candidate sentences from the training dataset as the synthetic corpus for subsequent training of the NMT model.

[0068] Specifically, through generate a relevance score between each candidate sentence and the seed sentence according to the candidate sentence and the seed sentence. s is the relevance score between the candidate sentence and the seed sentence. is the i-th candidate sentence, S seed is the seed sentence, CL (ngr) is the number of times the n-gram appears in the synthetic corpus L, is the candidate statement is the number of words in, ngr is an n-gram, which is a byte segment of length N.

[0069] Step 203: Generate a comprehensive score for each candidate statement according to the quality score and the pre-computed relevance score between each candidate statement and the seed statement.

[0070] Specifically, multiply the quality score by the relevance score s to obtain the comprehensive score for each candidate statement That is: According to the formula in Step 202, it can be obtained that:

[0071]

[0072] where, is the comprehensive score of the candidate statement, is the quality score of the candidate statement, is the i-th candidate statement, S seed is the seed statement, C L (ngr) is the number of times the n-gram appears in the synthetic corpus L, is the candidate statement is the number of words in, ngr is an n-gram, which is a byte segment of length N.

[0073] In the embodiment of the present invention, to avoid having the same n-gram in the selected statements, a penalty term for the n-gram is added to the formula

[0074] In the embodiment of the present invention, the comprehensive score of each candidate statement takes into account text and lexical diversity, translation error rate, and the similarity between the seed statement and the candidate statement, and can better improve the quality of the training statements in the subsequent generated synthetic corpus.

[0075] Step 204: Generate a synthetic corpus according to the comprehensive score of each candidate statement corresponding to each seed statement.

[0076] In the embodiment of the present invention, Step 204 specifically includes:

[0077] Step 2041: Sort the candidate statements corresponding to each seed statement according to the comprehensive score, select the candidate statement corresponding to the highest comprehensive score, and use the selected candidate statement as the training statement.

[0078] In an embodiment of the present invention, one seed statement corresponds to multiple candidate statements, and the candidate statement with the highest degree of relevance to the seed statement is selected from the multiple candidate statements as the training statement.

[0079] Step 2042: Generate a synthetic corpus according to the training statements corresponding to each seed statement.

[0080] Specifically, the training statements corresponding to each seed statement are combined to form a synthetic corpus for subsequent training of the NMT model, so that the NMT model can achieve a better training effect. As an alternative solution, a pre-trained translation model is trained according to the synthetic corpus to obtain a trained neural machine translation model. The trained neural machine translation model can achieve a better translation effect compared with the pre-trained translation model. In an English-German translation task, the corpus used includes open-source corpora in the medical field and the news field; the NMT model used is the Transformer model in the open-source model (OpenNMT).

[0081] First, the NMT model is trained through an English-German parallel corpus to obtain a pre-trained NMT model, denoted as the "pre-trained" model; the English monolingual corpus is translated into German through the "pre-trained" model to obtain a pseudo-parallel corpus; the "pre-trained" model is trained through the pseudo-parallel corpus to obtain the "pre-trained + back-translation" model.

[0082] Second, data selection is performed on the pseudo-parallel corpus through the synthetic corpus generation method provided in the embodiment of the present invention to obtain a synthetic corpus; the "pre-trained" model is trained through the synthetic corpus to obtain the "pre-trained + data selection" model.

[0083] Finally, the "pre-trained + back-translation" model and the "pre-trained + data selection" model are respectively evaluated through a test data set, and the obtained translation results show that the translation result of the "pre-trained + data selection" model is more accurate than that of the "pre-trained + back-translation" model, that is, the translation performance of the "pre-trained + data selection" model is better than that of the "pre-trained + back-translation" model.

[0084] In an embodiment of the present invention, the pseudo-parallel corpus obtained by optimizing the back-translation method through an improved FDA algorithm, and the generated synthetic corpus has higher text quality and lexical diversity, which essentially improves the quality of the corpus, helps to increase the number of parallel corpora in some fields with small data scales, and can effectively improve the translation performance of the NMT model.

[0085] In the technical solution of the synthetic corpus generation method provided by the embodiments of the present invention, according to the seed sentences in the preset seed dataset, the quality of the candidate sentences in the training dataset is evaluated to generate a quality score for each candidate sentence; according to the quality score and the pre-calculated relevance score of each candidate sentence to the seed sentence, a comprehensive score for each candidate sentence is generated; according to the comprehensive score of each candidate sentence corresponding to each seed sentence, a synthetic corpus is generated, and a synthetic corpus with high text quality and high lexical diversity can be obtained, thereby greatly improving the translation quality of the NMT model.

[0086] Figure 3 FIG. is a structural schematic diagram of a synthetic corpus generation device provided by an embodiment of the present invention. The device is used to execute the above synthetic corpus generation method, as Figure 3 shown, the device includes: a quality evaluation unit 11, a comprehensive evaluation unit 12, and a generation unit 13.

[0087] The quality evaluation unit 11 is used to evaluate the quality of the candidate sentences in the training dataset according to the seed sentences in the preset seed dataset, and generate a quality score for each candidate sentence.

[0088] The comprehensive evaluation unit 12 is used to generate a comprehensive score for each candidate sentence according to the quality score and the pre-calculated relevance score of each candidate sentence to the seed sentence.

[0089] The generation unit 13 is used to generate a synthetic corpus according to the comprehensive score of each candidate sentence corresponding to each seed sentence.

[0090] In the embodiments of the present invention, the quality evaluation unit 11 is specifically used to evaluate the quality of the candidate sentences in the training dataset according to the seed sentences in the preset seed dataset through a quality evaluation algorithm, and generate a quality score for each candidate sentence.

[0091] In the embodiments of the present invention, the quality evaluation unit 11 is specifically used to generate multiple quality indicators according to the candidate sentences and the seed sentences through a quality evaluation algorithm; calculate the multiple quality indicators to obtain the quality score of each candidate sentence.

[0092] In the embodiments of the present invention, the quality evaluation algorithm includes an algorithm for measuring text and lexical diversity, a machine translation evaluation index algorithm, and a translation error rate algorithm; the quality evaluation unit 11 is further specifically used to measure the text and lexical diversity of the candidate sentences through the algorithm for measuring text and lexical diversity to obtain a first quality indicator; calculate the candidate sentences and the seed sentences through the machine translation evaluation index algorithm to obtain a second quality indicator for each candidate sentence; calculate the candidate sentences and the seed sentences through the translation error rate algorithm to obtain a third quality indicator for each candidate sentence.

[0093] In an embodiment of the present invention, the apparatus further includes a relevance evaluation unit 14.

[0094] The relevance evaluation unit 14 is configured to generate a relevance score of each candidate statement and the seed statement according to the candidate statement and the seed statement through a feature attenuation algorithm.

[0095] In an embodiment of the present invention, the comprehensive evaluation unit 12 is specifically configured to multiply the quality score and the relevance score to obtain a comprehensive score of each candidate statement.

[0096] In an embodiment of the present invention, the generation unit 13 is specifically configured to sort the candidate statements corresponding to each seed statement according to the comprehensive score, screen out the candidate statement corresponding to the highest comprehensive score, and use the screened candidate statement as a training statement; generate a synthetic corpus according to the training statement corresponding to each seed statement.

[0097] In an embodiment of the present invention, the apparatus further includes a training unit 15.

[0098] The training unit 15 is configured to train a pre-trained translation model according to the synthetic corpus to obtain a neural machine translation model.

[0099] In the solution of the embodiment of the present invention, quality evaluation is performed on candidate statements in a training dataset according to seed statements in a preset seed dataset to generate a quality score for each candidate statement; a comprehensive score for each candidate statement is generated according to the quality score and the pre-calculated relevance score of each candidate statement and the seed statement; a synthetic corpus is generated according to the comprehensive score of each candidate statement corresponding to each seed statement, and a synthetic corpus with high text quality and high lexical diversity can be obtained, thereby greatly improving the translation quality of the NMT model.

[0100] The system, apparatus, module or unit illustrated in the above embodiments may be specifically implemented by a computer chip or an entity, or by a product with certain functions. A typical implementation device is a computer device. Specifically, the computer device may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0101] An embodiment of the present invention provides a computer device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, the steps of the above embodiment of the synthetic corpus generation method are implemented. For specific descriptions, reference may be made to the above embodiment of the synthetic corpus generation method.

[0102] Refer to the following Figure 4 , which shows a schematic structural diagram of a computer device 600 suitable for implementing the embodiments of the present application.

[0103] As Figure 4 shown, the computer device 600 includes a central processing unit (CPU) 601, which can perform various appropriate operations and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer device 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0104] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required, so that the computer program read from it can be installed in the storage section 608 as required.

[0105] Specifically, according to the embodiments of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments of the present invention include a computer program product, which includes a computer program tangibly contained on a machine-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 609, and / or installed from the removable medium 611.

[0106] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0107] For the convenience of description, when describing the above device, it is divided into various units according to functions and described separately. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0108] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0109] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or steps for implementing the functions specified in multiple blocks.

[0111] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0112] In the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.

[0113] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system or a computer program product. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0114] This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0115] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0116] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A synthetic corpus generation method, characterized in that The method includes: Evaluating the quality of candidate sentences in the training dataset according to the seed sentences in the preset seed dataset, generating a quality score for each candidate sentence, where the candidate sentence and the seed sentence are input into a translation error rate algorithm, and a third quality metric is output, the third quality metric is used to indicate the translation error rate, and the third quality metric is used to determine the quality score; Generating a comprehensive score for each candidate sentence according to the quality score and the pre-calculated relevance score of each candidate sentence to the seed sentence; Generating a synthetic corpus according to the comprehensive scores of the candidate sentences corresponding to each seed sentence, where the candidate sentences corresponding to each seed sentence are sorted according to the comprehensive scores, the candidate sentence corresponding to the highest comprehensive score is selected, and the selected candidate sentence is used as a training sentence, and a synthetic corpus is generated according to the training sentences corresponding to each seed sentence; Among them, the quality score of the candidate sentence is calculated by the formula where is the quality score of the candidate sentence, Q BLEU is the second quality indicator, Q TER is the third quality indicator, Q MTLD is the first quality indicator; the first quality indicator is used to measure text and lexical diversity, and the second quality indicator is used to measure the similarity of sentences; Wherein, before generating a comprehensive score for each candidate sentence according to the quality score and the pre-calculated relevance score of each candidate sentence to the seed sentence, the method further includes: Generating a relevance score of each candidate sentence to the seed sentence according to the candidate sentence and the seed sentence through a feature attenuation algorithm; The feature attenuation algorithm includes a penalty term for n-gram frequency.

2. The method for generating a synthetic corpus according to claim 1, wherein The evaluating the quality of candidate sentences in the training dataset according to the seed sentences in the preset seed dataset, generating a quality score for each candidate sentence, includes: Evaluating the quality of candidate sentences in the training dataset according to the seed sentences in the preset seed dataset through a quality evaluation algorithm, generating a quality score for each candidate sentence.

3. The method for generating a synthetic corpus according to claim 2, wherein The evaluating the quality of candidate sentences in the training dataset according to the seed sentences in the preset seed dataset through a quality evaluation algorithm, generating a quality score for each candidate sentence, includes: Generating a plurality of quality metrics according to the candidate sentence and the seed sentence through the quality evaluation algorithm; Calculating the plurality of quality metrics to obtain a quality score for each candidate sentence.

4. The synthetic corpus generation method according to claim 3, characterized in that, The quality evaluation algorithm includes a measure algorithm for text and lexical diversity, a machine translation evaluation metric algorithm, and a translation error rate algorithm; The generating a plurality of quality metrics according to the candidate sentence and the seed sentence through the quality evaluation algorithm, includes: Measuring the text and lexical diversity of the candidate sentence through the measure algorithm for text and lexical diversity to obtain a first quality metric; Calculating the candidate sentence and the seed sentence through the machine translation evaluation metric algorithm to obtain a second quality metric for each candidate sentence; Calculating the candidate sentence and the seed sentence through the translation error rate algorithm to obtain a third quality metric for each candidate sentence.

5. The method for generating a synthetic corpus according to claim 1, wherein, The generating a comprehensive score for each candidate sentence according to the quality score and the pre-calculated relevance score of each candidate sentence to the seed sentence, includes: Multiplying the quality score and the relevance score to obtain a comprehensive score for each candidate sentence.

6. The method for generating a synthetic corpus according to claim 1, wherein After generating a synthetic corpus based on the comprehensive scores of each candidate sentence corresponding to each seed sentence, it further includes: Training a pre-trained translation model based on the synthetic corpus to obtain a neural machine translation model.

7. A synthetic corpus generation device, characterized in that, The device includes: A quality evaluation unit for evaluating the quality of candidate sentences in a training dataset according to the seed sentences in a preset seed dataset, generating a quality score for each candidate sentence. Among them, the candidate sentence and the seed sentence are input into a translation error rate algorithm, and a third quality metric is output. The third quality metric is used to indicate the translation error rate, and the third quality metric is used to determine the quality score; A comprehensive evaluation unit for generating a comprehensive score for each candidate sentence according to the quality score and the pre-calculated relevance score of each candidate sentence to the seed sentence; A generation unit for generating a synthetic corpus according to the comprehensive scores of each candidate sentence corresponding to each seed sentence. Among them, the candidate sentences corresponding to each seed sentence are sorted according to the comprehensive score, the candidate sentence corresponding to the highest comprehensive score is selected, and the selected candidate sentence is used as a training sentence. A synthetic corpus is generated according to the training sentences corresponding to each seed sentence; Among them, the quality score of the candidate sentence is calculated by the formula where is the quality score of the candidate sentence, Q BLEU is the second quality index, Q TER is the third quality index, Q MTLD is the first quality index; the first quality index is used to measure text and lexical diversity, and the second quality index is used to measure the similarity degree of sentences; Among them, the device is further configured to generate a relevance score of each candidate sentence to the seed sentence according to the candidate sentence and the seed sentence through a feature attenuation algorithm; the feature attenuation algorithm includes a penalty term for n-gram frequency.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the synthetic corpus generation method according to any one of claims 1 to 6.

9. A computer device, comprising a memory and a processor, the memory being used for storing information including program instructions, and the processor being used for controlling the execution of the program instructions, characterized in that, When the program instructions are loaded and executed by a processor, they implement the synthetic corpus generation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Sample statement generation method and device, equipment and storage medium

    CN113705191A

  • Transfer text generation method and device, medium and equipment

    CN114139515A