Parallel corpus data set construction method and device, translation model training method and device and related equipment
By integrating evaluation models to select high-quality translation text pairs to construct parallel corpus datasets, and by pruning and distilling the translation models, the problem of low accuracy of corpus datasets is solved, and the translation performance of translation models in resource-constrained environments is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-07
AI Technical Summary
The current corpus datasets used to train translation models have low accuracy and are not well-matched with text translation models, resulting in poor translation performance.
By integrating evaluation models to comprehensively score translated text pairs, high-quality translated text pairs are selected to construct a parallel corpus dataset. The original translation model is then pruned and distilled to obtain an efficient translation model suitable for devices with limited memory.
It improves the accuracy and adaptability of the corpus dataset, and enhances the translation capabilities and accuracy of the translation model in resource-constrained environments.
Smart Images

Figure CN121809496A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence and machine learning, and in particular to a parallel corpus dataset construction method and device, a translation model training method and device, and related equipment. BACKGROUND
[0002] With the deepening of globalization, the demand for cross-language text translation is increasing, especially in resource-constrained "extreme resource" environments such as mobile devices, embedded systems, and edge computing, which urgently require lightweight and efficient text translation models. In order to train the above lightweight and efficient text translation model, a corpus dataset with high accuracy is crucial. However, the current corpus dataset has a lot of noise, and the domain of the corpus dataset is easily mismatched with the text translation model. Therefore, the accuracy of the current corpus dataset used to train the translation model is low. SUMMARY
[0003] The present application provides a parallel corpus dataset construction method and device, a translation model training method and device, and related equipment to solve the problem of low accuracy of the current corpus dataset used to train the translation model.
[0004] To solve the above problems, the present application is implemented as follows: In a first aspect, the present application provides a parallel corpus dataset construction method for training a translation model, comprising: obtaining N translation text pairs, each translation text pair including a first language text and a corresponding second language text, N being an integer greater than 1; inputting each translation text pair into an integrated evaluation model, the integrated evaluation model scoring according to a preset index to obtain a comprehensive score of each translation text pair; wherein the integrated evaluation model integrates H large language models, H being an integer greater than 1, the preset index including at least one of the following: semantic consistency before and after translation, text readability, and translation diversity, and the comprehensive score representing the translation quality between the first language text and the corresponding second language text; constructing a parallel corpus dataset for training a translation model using translation text pairs with a comprehensive score greater than a preset score.
[0005] In a second aspect, the present application provides a translation model training method, comprising: obtaining first language text information to be translated; pruning the original translation model to obtain a pruned model; distilling the pruned model using the first language text information to be translated to obtain a distilled model; The distillation model is trained by using a parallel corpus data set, wherein the parallel corpus data set is obtained according to the method disclosed in the first aspect.
[0006] In a third aspect, the application further provides a parallel corpus data set construction device for training a translation model, comprising: a translation text pair acquisition module configured to acquire N translation text pairs, each translation text pair comprising a first language text and a corresponding second language text, N being an integer greater than 1; a comprehensive score module configured to input each translation text pair into an integrated evaluation model, the integrated evaluation model being configured to score according to a preset index to obtain a comprehensive score of each translation text pair; wherein the integrated evaluation model integrates H large language models, H being an integer greater than 1, the preset index comprising at least one of the following: semantic consistency before and after translation, text readability, and translation diversity, and the comprehensive score representing the translation quality between the first language text and the corresponding second language text; a construction module configured to construct a parallel corpus data set for training a translation model by using translation text pairs with a comprehensive score greater than a preset score.
[0007] In a fourth aspect, the application further provides a translation model training device, comprising: a language text information acquisition module configured to acquire first language text information to be translated; a pruning processing module configured to prune an original translation model to obtain a pruned model; a distillation processing module configured to distill the pruned model by using the first language text information to be translated to obtain a distilled model; a model training module configured to train the distilled model by using a parallel corpus data set, wherein the parallel corpus data set is obtained according to the method disclosed in the first aspect.
[0008] In a fifth aspect, the application further provides an electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the method disclosed in the first aspect, or to read the program in the memory to implement the steps in the method disclosed in the second aspect.
[0009] In a sixth aspect, the application further provides a readable storage medium for storing a program, wherein the program is executed by a processor to implement the steps in the method disclosed in the first aspect, or to implement the steps in the method disclosed in the second aspect.
[0010] In a seventh aspect, the application further provides a computer program product comprising computer instructions, wherein the computer instructions are executed by a processor to implement the steps in the method disclosed in the first aspect or the second aspect. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is one of the flowcharts illustrating the method for constructing a parallel corpus dataset for training a translation model provided in this application embodiment; Figure 2 This is a schematic diagram of the pruning model provided in the embodiments of this application; Figure 3A This is the second flowchart illustrating the method for constructing a parallel corpus dataset for training a translation model provided in this application embodiment; Figure 3B This is the third flowchart illustrating the method for constructing a parallel corpus dataset for training a translation model provided in this application embodiment; Figure 3C This is the fourth flowchart illustrating the method for constructing a parallel corpus dataset for training a translation model, as provided in the embodiments of this application. Figure 4A This is one of the flowcharts illustrating the translation model training method provided in the embodiments of this application; Figure 4B This is the second flowchart illustrating the translation model training method provided in the embodiments of this application; Figure 4C This is a flowchart illustrating the method for constructing a parallel corpus dataset for training a translation model and the method for training a translation model provided in this application embodiment; Figure 5 This is a schematic diagram of the structure of the apparatus for constructing a parallel corpus dataset for training a translation model provided in an embodiment of this application; Figure 6 This is a schematic diagram of the translation model training device provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] The terms "first," "second," etc., used in the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: including A alone, B alone, C alone, and the presence of both A and B, both B and C, both A and C, and the presence of A, B, and C.
[0015] See Figure 1 , Figure 1 This is a flowchart illustrating the method for constructing a parallel corpus dataset for training a translation model, as provided in an embodiment of this application. Figure 1 The method shown for constructing the parallel corpus dataset used to train the translation model can be performed by an electronic device.
[0016] like Figure 1 As shown, the method for constructing a parallel corpus dataset for training a translation model may include the following steps: Step 200: Obtain N translation text pairs, each of which includes the first language text and the corresponding second language text, where N is an integer greater than 1.
[0017] The method of obtaining the aforementioned translated text pairs is not specifically limited here. Optionally, the aforementioned translated text pairs can be data from an open-source dataset or data from a preset dataset, wherein the preset dataset can be any dataset specified by the user or data from a sample set. Alternatively, the aforementioned translated text pairs can be translated text pairs obtained after preprocessing the dataset. The aforementioned preprocessing can include at least one of the following: deleting at least one of two text pairs with a similarity greater than a preset similarity; deleting text pairs with less than a first number of text contents; deleting text pairs with more than a second number of text contents; deleting text pairs that include text in other languages, wherein the aforementioned other language text can be understood as text in languages other than the first language text and the second language text, wherein the aforementioned second number is greater than the first number, and the aforementioned text pairs with less than a first number of text contents can be referred to as excessively short text pairs, and the aforementioned text pairs with more than a second number of text contents can be referred to as excessively long text pairs.
[0018] Alternatively, the translation text pair can be a text pair obtained by rewriting using a text rewriting model. For example, an original translation text pair can be obtained from an open source data set or a preset data set, where the text in the first language corresponds to the source language text, and the text in the second language corresponds to the translation text. The original translation text pair is input into the rewriting model, and the rewriting model rewrites the translation text according to the source language text to make the translation of the source language text more accurate, thereby obtaining the translation text pair.
[0019] Alternatively, the translation text pair can be obtained by translating data obtained from an open source data set or a preset data set using a trained translation model. For example, a first language text can be obtained from an open source data set or a preset data set, and the first language text is input into the translation model for translation, thereby obtaining a second language text corresponding to the first language text, and pairing the two to obtain a translation text pair. For example, a second language text can be obtained from an open source data set or a preset data set, and the second language text is input into the translation model for translation, thereby obtaining a first language text corresponding to the second language text, and pairing the two to obtain a translation text pair. The translation result can also be manually checked.
[0020] The specific types of the first language text and the second language text are not limited herein. For example, the first language text can be Chinese text, and the second language text can be English text.
[0021] As an optional implementation, referring to Figure 3A Step 200, obtaining N translation text pairs, including: Step 210, obtaining a sample set; The sample set can include information of multiple language texts in multiple fields.
[0022] Step 220, splitting the sample set to obtain multiple language text information; Each language text information in the multiple language text information corresponds to multiple fields of a language, and the multiple language text information corresponds to different languages. In this way, the sample set can be split according to the language types, each language text is split into a language text information, and the multiple language text information is obtained.
[0023] Step 230, translating at least part of the multiple language text information to obtain at least part of the N translation text pairs.
[0024] Based on the user's training needs, the system can select a specific language text from the split multi-language text information required for model training and accurately translate that language text, thereby obtaining at least some of the N translated text pairs. The remaining translated text pairs in the N translated text pairs can come from open-source datasets or pre-set datasets, as mentioned earlier, and will not be elaborated further here. The translated text pairs from each source in the N translated text pairs correspond to the same language.
[0025] Optionally, the domains corresponding to the aforementioned language text information can be different. For example, the domain corresponding to a certain language text information may include finance, communications, consumer goods, etc. This allows for a broader distribution of language text information across different domains, avoiding a concentration on only a few domains, and thus resulting in a more diverse range of data sources in the final parallel corpus dataset.
[0026] In this embodiment, the sample set is split to obtain text information in multiple languages. At least some of the text information in multiple languages is translated to obtain at least some of the N translated text pairs. This can further increase the diversity of the sources of the N translated text pairs and improve the accuracy of the N translated text pairs.
[0027] Optionally, the translated text pairs obtained in this embodiment can also be filtered. For example, a comprehensive score can be calculated for the translated text pairs, and the translated text pairs with a comprehensive score greater than a preset score can be identified as at least a portion of the N translated text pairs. The specific method for calculating the comprehensive score can be found in the relevant description below, and will not be repeated here.
[0028] Step 400: Input each translated text pair into the integrated evaluation model. The integrated evaluation model scores each translated text pair according to preset indicators to obtain a comprehensive score. The integrated evaluation model integrates H large language models, where H is an integer greater than 1. The preset indicators include at least one of the following: semantic consistency before and after translation, text readability, and translation diversity. The comprehensive score represents the translation quality between the first language text and the corresponding second language text.
[0029] The types of H large language models are not limited here. Optionally, the types of H large language models can be the same, or alternatively, the types of H large language models can be different. For example, the H large language models can include models such as GPT5 and DeepSeek. The above-mentioned large language models can be called (Large Language Model, LLM).
[0030] Semantic consistency before and after translation can refer to the degree of semantic matching between the first language text and the corresponding second language text. For example, if the first language text is water and the second language text is also water, then the semantic consistency before and after translation is considered to be high. On the other hand, if the first language text is water and the second language text is car, then the semantic consistency before and after translation is considered to be low.
[0031] Among them, text readability can refer to whether the translated text, that is, the second language text, conforms to the user's reading habits and understanding logic. For example, if the text in the first language is ordered in the first order, and the second language text is translated by simply copying the inherent word order of the first language without following the word order rules of the second language, the second language text will be obscure and difficult to understand, and thus the text readability can be determined to be poor.
[0032] Translation diversity can refer to whether the translated text, i.e., the second language text, has the potential to flexibly adapt to different application needs in terms of semantic expression ability of the source language text, i.e., the first language text, and whether the translated text itself reflects expandable and differentiated expressive features at the levels of vocabulary, sentence structure, etc.
[0033] The overall score represents the translation quality between the first language text and the corresponding second language text. The higher the overall score, the better the translation quality between the first language text and the corresponding second language text.
[0034] It should be noted that because the ensemble evaluation model integrates H large language models, it can be called a judge agent.
[0035] Step 600: Construct a parallel corpus dataset for training the translation model using translated text pairs whose overall score is greater than the preset score.
[0036] The specific method for constructing the parallel corpus dataset for training the translation model using translated text pairs with a comprehensive score greater than the preset score is not limited here. Optionally, translated text pairs with a comprehensive score greater than the preset score can be directly stored in the parallel corpus dataset. Alternatively, translated text pairs with a comprehensive score greater than the preset score can be copied, and the copied translated text pairs can be stored in the parallel corpus dataset.
[0037] In this embodiment, through steps 200 to 600, each translated text pair can be input into the ensemble evaluation model, which then performs a comprehensive score on each pair based on preset metrics. The comprehensive score represents the translation quality between the first language text and the corresponding second language text. Based on the scoring results, translated text pairs with a comprehensive score greater than the preset score are selected, thus obtaining high-quality translated text pairs. Using these high-quality translated text pairs to construct a parallel corpus dataset for training the translation model effectively improves the accuracy of the constructed parallel corpus dataset, thereby enhancing the reliability and adaptability of the corpus from the source.
[0038] It should be noted that the specific method by which the integrated evaluation model comprehensively scores each translated text pair according to preset indicators is not limited here. Optionally, each translated text pair can be input into each large language model for scoring, and each large language model can generate a corresponding score according to preset indicators. Then, the generated scores are summed or weighted to obtain the score of each large language model for the translated text pair. Then, the integrated evaluation model sums or weightedly sums the scores of each large language model for the translated text pair to obtain the comprehensive score of the translated text pair.
[0039] As an optional implementation method, see [link to implementation details]. Figure 3B Step 400: Input each translated text pair into the ensemble evaluation model. The ensemble evaluation model scores each translated text pair according to preset indicators, obtaining a comprehensive score for each translated text pair, including: Step 410: Input each translated text pair into each of the H large language models. Each large language model scores according to preset indicators to obtain at least one indicator score. The indicator scores correspond one-to-one with the preset indicators. Step 420: For each translated text pair, generate H self-scoring values from the H large language models corresponding to each translated text pair based on all the index scores of each large language model. Step 430: Obtain the comprehensive score for each translated text pair based on H self-scoring scores.
[0040] The specific method for obtaining the comprehensive score of each translated text pair based on the H self-scorings is not limited here. Optionally, the sum of the H self-scorings can be calculated and the sum of the H self-scorings can be determined as the comprehensive score of the translated text pair. Alternatively, a weighted sum of the H self-scorings can be calculated and the weighted sum of the H self-scorings can be determined as the comprehensive score of the translated text pair. When calculating the weighted sum of the H self-scorings, the weight coefficient of each self-scoring can be related to the corresponding large language model.
[0041] In this embodiment, each translated text pair is input into each of the H large language models. Each large language model scores according to preset indicators to obtain at least one indicator score. For each translated text pair, H self-scoring scores are generated based on all indicator scores of each large language model. The comprehensive score of each translated text pair is obtained based on the H self-scoring scores. This makes the accuracy of the calculated comprehensive score higher.
[0042] It should be noted that each large language model may have biases when scoring translated text pairs, potentially leading to lower accuracy in the scores. To address this issue, the following implementation method is proposed: As an optional implementation method, see [link to implementation details]. Figure 3C Step 430: Obtain a comprehensive score for each translated text pair based on H self-scoring scores, including: Step 431: Obtain H sub-scores for each of the translated text pairs based on the H self-scores, including: Step 431a: For each of the translated text pairs, take each of the H large language models as the target large language model in turn, input the self-scoring criteria of the translated text pair and the target large language model into the H-1 large language models other than the target large language model for scoring, and obtain the H-1 other scores corresponding to the target large language model. Step 431b: Based on the self-score from the target large language model and the corresponding H-1 other scores, obtain H sub-scores for each translated text pair; Step 432: Obtain the comprehensive score for each translated text pair based on the H sub-scores for each translated text pair.
[0043] In this embodiment, each of the H large language models can be sequentially used as the target large language model. The scoring criteria of the translated text pair and the self-scoring from the target large language model are input into the H-1 large language models other than the target large language model for scoring, resulting in H-1 other scores corresponding to the target large language model. The scoring criteria for the self-scoring includes the self-scoring itself, which refers to the judgment criteria, reference dimensions, or rule set relied upon by the target large language model when obtaining the self-scoring; it is the core technical element supporting the rationality of the self-scoring result. This embodiment can use H-1 large language models to evaluate the rationality of the self-scoring of the translated text from the target large language model, thereby obtaining H-1 other scores. Then, based on the self-scoring from the target large language model and the H-1 other scores, each sub-scoring for each translated text pair is obtained, thus avoiding bias of the target large language model towards the translated text pair, and further improving the accuracy of the comprehensive score of the obtained translated text pair.
[0044] It should be noted that the specific method for obtaining each sub-score of each translated text pair based on the self-score and H-1 other-scores from the target large language model is not limited here. Optionally, the sum of the self-score and H-1 other-scores corresponding to the target large language model can be calculated, and this sum can be determined as the comprehensive score of each translated text pair.
[0045] The specific implementation method for obtaining the comprehensive score of each translated text pair based on the H sub-scores is not limited here. Optionally, the comprehensive score of each translated text pair can be obtained by averaging the H sub-scores, or by weighted summing the H sub-scores, or by weighted average of the H sub-scores.
[0046] As an optional implementation, step 431b, obtaining H sub-scores for each translated text pair based on the self-score from the target large language model and the corresponding H-1 other scores, includes: for each of the H sub-scores, obtaining the average of the other scores based on the H-1 other scores corresponding to the target large language model; and obtaining the sub-score based on the average of the self-score from the target large language model and the corresponding other scores.
[0047] The process of obtaining the average of the H-1 other ratings corresponding to the target large language model can be understood as: calculating the average of the H-1 other ratings to obtain the average of the other ratings; or it can be understood as: calculating the weighted average of the H-1 other ratings to obtain the corresponding average of the other ratings.
[0048] Sub-scores are obtained by averaging the self-scores and corresponding peer scores from the target large language model. This can be achieved by weighted summation of the self-scores and peer scores, or by averaging the self-scores and peer scores.
[0049] In this embodiment, a sub-score is obtained based on the average of the self-score and the corresponding other scores from the target large language model. By evaluating the rationality of the self-scores of the translated text from the target large language model by other large language models, a sub-score of the translated text is obtained, which can further reduce the bias of the target large language model on the translated text pair and thus improve the accuracy of the comprehensive score.
[0050] To more fully illustrate the above embodiments, a specific embodiment will be used as an example below.
[0051] For example, the ensemble evaluation model includes m large language models, where m is a positive integer greater than 1. The score for each indicator calculated by each large language model for each translated text pair can be obtained using... Therefore, the self-scoring of each translated text pair by the large language model can be adopted. The formula is expressed as follows: Where k represents the number of preset indicators, and i represents a specific preset indicator. However, large language models may exhibit bias towards translated text pairs. Therefore, in the self-scoring of each large language model, an evaluation of the reasonableness of its scores by other large language models is introduced. That is, the self-scoring criteria for each translated text pair and the target large language model are provided to other large language models. The self-scoring criteria include the self-scoring score, and other large language models provide a numerical value (i.e., their score) to represent the reasonableness of the target large language model's score, denoted as [the score]. And his average rating For each of the multiple large language models other than the target large language model, there is a corresponding score. The mean of the terms is expressed by the following formula: , where n represents the target large language model, j represents any other large language model besides the target large language model, and m represents the number of target large language models.
[0052] Optionally, a weighted sum of the self-score and the average of the other-scores can be calculated to obtain the sub-score of the target large language model, avoiding bias of a single large language model towards the translated text pair. The calculation formula is expressed as follows: , among which, S n Let H represent the sub-scores of the target large language model, and α represent the weighting coefficient. For ease of statistics, viewing, and sorting, the mean of both self-scores and peer scores can be normalized to the interval [0,1]. This ensures that the scores of the resulting sub-self-scores also fall within the [0,1] interval. Each large language model in the integrated evaluation model is then used as the target large language model, resulting in H sub-self-scores. Then, the scores of these H sub-scores are... The average value is taken to obtain the comprehensive score S of the integrated evaluation model for the current translated text pair. The calculation formula is expressed as follows: .
[0053] See Figure 4A , Figure 4A This is a flowchart illustrating the translation model training method provided in the embodiments of this application. Figure 4A The translation model training method shown can be executed by an electronic device, and the electronic device in the embodiments of this application is similar to the one described above. Figure 1 The electronic devices in the illustrated embodiments can be the same electronic device or different electronic devices.
[0054] like Figure 4A As shown, the translation model training method may include the following steps: Step 2000: Obtain the first language text information to be translated.
[0055] The method of acquiring the first-language text information to be translated is not specifically limited here. Optionally, the first-language text information to be translated can be information input via voice, touch, or pressure. Alternatively, the first-language text information to be translated can also be content included in image information captured by a camera. When the acquired information is not directly text information, it is converted into text information through technologies such as speech recognition or image recognition.
[0056] Step 4000: Prune the original translation model to obtain the pruned model.
[0057] The original translation model can be understood as a model with good translation capabilities. However, because the original translation model usually has many parameters, it requires a large amount of memory, making it difficult to apply in electronic devices with limited memory, thus limiting the application scenarios of the original translation model. In this embodiment, the original translation model can be pruned to obtain a pruned model, which allows the pruned model to be applied to electronic devices with limited memory, thereby expanding the application scenarios of the pruned model.
[0058] For example: see Figure 2 , Figure 2 The diagram shows the structure of the pruned model obtained after pruning the original translation model. Figure 2 As shown, the pruning model has fewer model parameters, making it applicable to electronic devices with limited memory space.
[0059] It should be noted that the specific type of the original translation model mentioned above is not limited here.
[0060] Step 6000: Distill the pruning model using the first language text information to be translated to obtain the distillation model.
[0061] In this process, the pruning model is distilled using the first-language text information to be translated. The specific method for obtaining the distillation model is not limited here. Optionally, the first-language text information to be translated can be input into the original translation model and the pruning model respectively. Then, the loss value is calculated based on the translated text information output from the original translation model and the pruning model. The parameters of the pruning model are updated by updating the loss value. Through multiple rounds of iterative training, if the loss value is less than the preset loss value, it can be determined that the distillation model can obtain the language habits of the original translation model. This means that the distillation model can have the same translation ability as the original translation model, thereby improving the translation ability of the distillation model and enhancing the accuracy of the translation results.
[0062] Optionally, the above loss value can be the mean squared error (MSE) loss value.
[0063] Step 8000: Train the distillation model using a parallel corpus dataset, wherein the parallel corpus dataset is obtained according to the method described in the above embodiments.
[0064] In this embodiment, the original translation model is pruned to obtain a pruned model. The pruned model is then distilled using the first language text information to be translated to obtain a distilled model. The distilled model is trained using a parallel corpus dataset, allowing for training and adjustment. This enables the trained distilled model to learn the semantics, diversity, and translation accuracy of large language models, making it suitable for use in electronic devices with limited memory while maintaining excellent translation capabilities and accuracy.
[0065] As an optional implementation method, see [link to implementation details]. Figure 4B Step 4000: Prune the original translation model to obtain a pruned model, including: Step 4100: Perform at least one of the following removal operations on the original translation model: remove some token layers of the original translation model; remove some converter layers of the original translation model; Step 4200: Based on the different objects to be removed, obtain M candidate pruning models, where M is an integer greater than 1; Step 4300: Input the verification first language text information into M candidate pruning models for translation to obtain the verification translation text pair for each candidate pruning model. The verification translation text pair includes the first language text and the corresponding second language text. Step 4400: Input the verification translation text pair of each candidate pruning model into the ensemble evaluation model for scoring, and obtain the score corresponding to each candidate pruning model; Step 4500: Select the candidate pruning model with the highest score as the pruning model.
[0066] It should be noted that, in order to better apply the original translation model to electronic devices with limited memory, the original translation model can be pruned to obtain multiple candidate pruned models. These candidate pruned models are then scored using H large language models, and the final pruned model is determined based on the scores. This reduces the number of model parameters in the pruned model, making it more suitable for electronic devices with limited memory, while also improving its translation capabilities. Specifically, the scoring of the candidate pruned models using H large language models involves inputting the verification translation text pairs obtained after translating the verification first-language text information into the H large language models in the ensemble evaluation model. The scores of these verification translation text pairs are then used as the scores for the candidate pruned models. There can be one or more verification translation text pairs; when there are multiple pairs, the average or weighted average of the scores can be used. The scoring process for the verification translation text pairs in the ensemble evaluation model is described in the dataset construction section above, and will not be repeated here.
[0067] It should be noted that there can be multiple first-language text information verifications, and there can also be multiple corresponding verification translation text pairs. The first-language text information and the verification translation text pairs have a corresponding relationship.
[0068] If the original translation model has too many redundant token layers, it can easily lead to low inference performance when translating between multiple language text information, and the output language text information may be unstable during the translation process. Therefore, by removing some token layers from the original translation model, the parameters of the original translation model can be reduced, which is equivalent to pruning the original translation model. At the same time, it can improve the translation speed and stability of the pruned model obtained after removing some token layers, thus enhancing the translation performance of the pruned model.
[0069] If the original translation model has too many converter layers, it can easily lead to too many parameters and too much memory space required. Therefore, some converter layers of the original translation model can be removed to reduce the parameters and memory space required.
[0070] In this process, the validation translation text pairs of each candidate pruning model are input into the ensemble evaluation model for scoring, resulting in a score for each candidate pruning model. The candidate pruning model with the highest score is selected as the pruning model. This ensures that the selected pruning model has the highest translation accuracy, meaning that the model with the best translation performance can be selected from M candidate pruning models.
[0071] In this embodiment, some token layers of the original translation model are removed, or some converter layers of the original translation model are removed, thereby increasing the diversity and flexibility of obtaining the pruning model; at the same time, the candidate pruning model with the highest score is selected as the pruning model, thus ensuring that the determined pruning model is the model with the best translation effect among the M candidate pruning models.
[0072] As an optional implementation, some token layers of the original translation model are removed, including: The test text information is input into each token layer of the original translation model for word embedding, and the hash statistics of the target token of each token layer are obtained. In the original translation model, token layers whose target token hash statistics are lower than the preset statistical value are removed.
[0073] For example, word embedding is performed on the test text information through each token layer of the original translation model, the frequency of occurrence of individual tokens (i.e., target tokens) in each token layer is counted, and a word frequency hash table is generated. The top n tokens whose word frequency is less than the preset statistical value are counted, and the token layers corresponding to the top n tokens in the original translation model are removed.
[0074] In this embodiment, token layers in the original translation model whose target token hash statistics are lower than a preset statistical value are removed. This removes token layers that have a weak impact on the model's translation performance while retaining effective token layers that make a key contribution to translation accuracy. This achieves model pruning and reduces the model's computational resource consumption without significantly reducing the model's translation performance.
[0075] As an optional implementation, some converter layers of the original translation model are removed, including: Obtain multiple converter layer combinations of the original translation model. Each converter layer combination includes multiple converter layers, and the layer numbers of the multiple converter layers gradually increase. The target converter layer combination is determined from multiple converter layer combinations. Converter layers in the original translation model other than those in the target converter layer combination are removed. The target converter layer combination is the converter layer combination with the highest translation accuracy among multiple converter layer combinations.
[0076] Optionally, the original translation model includes 12 transformer layers. Multiple transformer layer combinations are selected from the 12 transformer layers, such as combinations of transformer layers 1-11, 2-10, and 5-12. The transformer layer combination with the highest translation accuracy is selected from the above multiple transformer layer combinations (i.e., the target transformer layer combination). Then, the transformer layers in the original translation model other than those in the target transformer layer combination are removed.
[0077] In this embodiment, precise screening and removal of the converter layer can effectively ensure the accuracy of the removal operation and avoid accidentally deleting core converter layers that make key contributions to translation performance. At the same time, the target converter layers retained after screening can maintain the core semantic encoding and decoding capabilities of the original model, ensuring the translation accuracy of the pruned model and achieving a balance between lightweight model architecture and high-precision translation effect.
[0078] To more fully illustrate the above embodiments, a specific embodiment will be used as an example below. See details... Figure 4C , Figure 4C This is a flowchart illustrating the method for constructing a parallel corpus dataset for training a translation model, as provided in the embodiments of this application, and the method for training the translation model. Figure 4C As shown, the entire process can be divided into four parts: parallel corpus dataset creation, model pruning, model distillation, and model fine-tuning.
[0079] In the parallel corpus dataset creation section, open-source datasets and monolingual databases can be understood as used to store first-language text. Then, after translation or rewriting based on a text rewriting model, N translated text pairs are obtained. These N translated text pairs are then input into an ensemble evaluation model (such as...). Figure 4C As shown, a comprehensive scoring can be performed using large language model 1, large language model 2, and large language model 3, and a parallel corpus dataset for training the translation model can be constructed using translated text pairs whose comprehensive scores are greater than the preset scores. In the model pruning part, the token layer pruning can be performed on the original translation model first. Then, it is determined whether the model structure grid search (which can be understood as whether the model meets the requirements of the application scenario after pruning, or whether it meets the pruning requirements) has ended. If the search has ended, the text translated by the pruned model is input into the ensemble evaluation model to generate a comprehensive score. The model with the highest comprehensive score is selected as the pruned model. If the search has not ended, the converter layer of the model is pruned, and the pruned model is verified by translating the verification text information. This verification text information can be understood as verifying the first language text information. In the model distillation part, text information from a monolingual database can be input into the original translation model and the pruning model for translation. The original translation model can output translation information 1, and the pruning model can output translation information 2. Then, the loss values of translation information 1 and translation information 2 are calculated. Based on the calculated loss values, the pruning model is trained iteratively for multiple rounds until the loss value converges to a preset threshold or reaches a preset number of training rounds, thereby obtaining the distilled model. In the model fine-tuning part, the constructed parallel corpus dataset can be obtained, and then the parallel corpus dataset can be input into the distilled model for training, thereby obtaining the fine-tuned model, which has better translation ability.
[0080] See Figure 5 , Figure 5 This is a structural diagram of the apparatus for constructing a parallel corpus dataset for training a translation model, as provided in the embodiments of this application. Figure 5 As shown, the apparatus 50 for constructing the parallel corpus dataset for training the translation model includes: The translation text pair acquisition module 51 is used to acquire N translation text pairs, each translation text pair including the first language text and the corresponding second language text, where N is an integer greater than 1; The comprehensive scoring module 52 is used to input each translated text pair into the integrated evaluation model. The integrated evaluation model scores each translated text pair according to preset indicators to obtain a comprehensive score for each translated text pair. The integrated evaluation model integrates H large language models, where H is an integer greater than 1. The preset indicators include at least one of the following: semantic consistency before and after translation, text readability, and translation diversity. The comprehensive score represents the translation quality between the first language text and the corresponding second language text. Module 53 is used to construct a parallel corpus dataset for training the translation model using translated text pairs with a comprehensive score greater than the preset score.
[0081] The apparatus for constructing a parallel corpus dataset for training a translation model provided in this application embodiment can input each translated text pair into an ensemble evaluation model, which then performs a comprehensive score on each translated text pair according to preset metrics. The comprehensive score represents the translation quality between the first language text and the corresponding second language text. Based on the scoring results, translated text pairs with a comprehensive score greater than the preset score are selected, thus obtaining high-quality translated text pairs. Using these high-quality translated text pairs to construct a parallel corpus dataset for training a translation model can effectively improve the accuracy of the constructed parallel corpus dataset, thereby improving the reliability and adaptability of the corpus from the source.
[0082] As an optional implementation, the comprehensive scoring module 52 includes: The first scoring submodule is used to input each translated text pair into each of the H large language models. Each large language model scores according to preset indicators to obtain at least one indicator score. The indicator scores correspond one-to-one with the preset indicators. The self-score generation submodule is used to generate H self-scores from H large language models for each translated text pair based on all the metrics of each large language model. The scoring calculation submodule is used to obtain a comprehensive score for each translated text pair based on H self-scoring scores.
[0083] As an optional implementation, the scoring calculation submodule includes: a sub-scoring calculation unit and a comprehensive calculation subunit.
[0084] The sub-score calculation unit is used to obtain H sub-scores for each translated text pair based on H self-scores; The sub-score calculation unit includes: a sub-unit for generating other scores and a sub-score calculation sub-unit.
[0085] The other scoring generation subunit is used to, for each translated text pair, sequentially take each of the H large language models as the target large language model, input the self-scoring criteria of the translated text pair and the target large language model into the H-1 large language models other than the target large language model for scoring, and obtain the H-1 other scores corresponding to the target large language model. The sub-score calculation sub-unit is used to obtain H sub-scores for each translated text pair based on the self-score from the target large language model and the corresponding H-1 other scores; The comprehensive calculation subunit is used to obtain the comprehensive score of each translated text pair based on the H sub-scores of each translated text pair.
[0086] As an optional implementation, the sub-score calculation subunit includes: The first calculation subunit is used to obtain the average value of the corresponding other score for each of the H sub-scores, based on the H-1 other scores corresponding to the target large language model. The second calculation subunit is used to obtain a sub-score based on the average of the self-score and the corresponding other-score from the target large language model.
[0087] As an optional implementation, the translated text pair acquisition module 51 includes: The `get` submodule is used to retrieve the sample set; The splitting submodule is used to split the sample set to obtain text information in multiple languages; The first translation submodule is used to translate at least some of the language text information from multiple languages to obtain at least some of the translated text pairs among N translated text pairs.
[0088] See Figure 6 , Figure 6 This is a structural diagram of the translation model training device provided in the embodiments of this application, as shown below. Figure 6 As shown, the translation model training device 60 includes: The language text information acquisition module 61 is used to acquire the first language text information to be translated. The pruning module 62 is used to prune the original translation model to obtain a pruned model. Distillation processing module 63 is used to distill the pruning model using the first language text information to be translated, so as to obtain a distilled model. The model training module 64 is used to train the distillation model using a parallel corpus dataset, wherein the parallel corpus dataset is obtained according to the method described above for constructing the parallel corpus dataset used to train the translation model.
[0089] The translation model training device provided in this application prunes the original translation model to obtain a pruned model, distills the pruned model using the first language text information to be translated, and obtains a distilled model. The distilled model is trained using a parallel corpus dataset, thereby allowing the distilled model to be trained and adjusted. This enables the trained distilled model to learn the semantics, diversity, and translation accuracy of large language model translation, thus making the trained distilled model applicable to electronic devices with limited memory space while possessing excellent translation capabilities and accuracy.
[0090] As an optional implementation, the pruning module 62 includes: The removal operation submodule is used to perform at least one of the following removal operations on the original translation model: remove some token layers of the original translation model; remove some converter layers of the original translation model; The model generation submodule is used to obtain M candidate pruning models based on the different objects to be removed, where M is an integer greater than 1; The second translation submodule is used to input the verification first language text information into M candidate pruning models for translation, and obtain the verification translation text pair of each candidate pruning model. The verification translation text pair includes the first language text and the corresponding second language text. The second scoring submodule is used to input the validation translation text pairs of each candidate pruning model into the ensemble evaluation model for scoring, and obtain the score corresponding to each candidate pruning model. The selection submodule is used to select the candidate pruning model with the highest score as the pruning model.
[0091] As an optional implementation, the operation submodule is removed, including: The word embedding unit is used to input the test text information into each token layer in the original translation model for word embedding, and to obtain the hash statistics of the target token in each token layer; The first elimination unit is used to eliminate token layers in the original translation model whose target token's hash statistics are lower than a preset statistical value.
[0092] As an optional implementation, the operation submodule is removed, including: The second acquisition unit is used to acquire multiple converter layer combinations of the original translation model. The converter layer combination includes multiple converter layers, and the layer numbers of the multiple converter layers gradually increase. The second elimination unit is used to determine the target converter layer combination from multiple converter layer combinations. It eliminates converter layers in the original translation model other than those in the target converter layer combination. The target converter layer combination is the converter layer combination with the highest translation accuracy among multiple converter layer combinations.
[0093] This application also provides an electronic device. Please refer to [link to relevant documentation]. Figure 7 The electronic device may include a processor 701, a memory 702, and a program 7021 stored in the memory 702 and executable on the processor 701. When the program 7021 is executed by the processor 701, it can achieve... Figure 1 or Figure 4A Any steps in the corresponding method embodiments and the achievement of the same beneficial effects will not be repeated here.
[0094] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by hardware related to program instructions, and the program can be stored in a readable medium. This application also provides a readable storage medium storing a computer program, which, when executed by a processor, can implement the above... Figure 1 or Figure 4A Any step in the corresponding method embodiment can achieve the same technical effect, and will not be repeated here to avoid repetition.
[0095] The aforementioned storage media include read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0096] This application also provides a computer program product, including computer instructions, which, when executed by a processor, can achieve the above-mentioned functions. Figure 1 or Figure 4A Any step in the corresponding method embodiment can achieve the same technical effect, and will not be repeated here to avoid repetition.
[0097] The above-disclosed content is a preferred embodiment of the present application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles disclosed in this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for constructing a parallel corpus dataset for training a translation model, characterized in that, include: Obtain N pairs of translated texts, each pair of translated texts including a first language text and a corresponding second language text, where N is an integer greater than 1; Each of the translated text pairs is input into an integrated evaluation model, which scores the translated text pairs according to preset indicators to obtain a comprehensive score for each translated text pair. The integrated evaluation model integrates H large language models, where H is an integer greater than 1. The preset indicators include at least one of the following: semantic consistency before and after translation, text readability, and translation diversity. The comprehensive score represents the translation quality between the first language text and the corresponding second language text. The parallel corpus dataset used to train the translation model is constructed using translated text pairs whose overall score is greater than the preset score.
2. The method according to claim 1, characterized in that, The process involves inputting each translated text pair into an integrated evaluation model, which scores the translated text pair according to preset indicators to obtain a comprehensive score for each translated text pair, including: Each of the translated text pairs is input into each of the H large language models. Each of the large language models is scored according to the preset index to obtain at least one index score. The index score corresponds one-to-one with the preset index. For each of the translated text pairs, H self-scoring values from the H large language models are generated for each of the translated text pairs based on all the index scores of each of the large language models. A comprehensive score is obtained for each of the H self-scoring pairs.
3. The method according to claim 2, characterized in that, The step of obtaining a comprehensive score for each of the translated text pairs based on the H self-scoring scores includes: Based on the H self-scoring values, H sub-scoring values are obtained for each of the translated text pairs, including: For each of the translated text pairs, each of the H large language models is taken as the target large language model in turn. The translated text pair and the scoring criteria from the self-scoring of the target large language model are input into the H-1 large language models other than the target large language model for scoring, so as to obtain the H-1 other scores corresponding to the target large language model. Based on the self-score from the target large language model and the corresponding H-1 other scores, H sub-scores are obtained for each translated text pair; A comprehensive score for each translated text pair is obtained based on the H sub-scores of each translated text pair.
4. The method according to claim 3, characterized in that, The step of obtaining H sub-scores for each translated text pair based on the self-score from the target large language model and the corresponding H-1 other scores includes: For each of the H sub-ratings, the average value of the corresponding other rating is obtained based on the H-1 other ratings corresponding to the target large language model; The sub-score is obtained based on the average of the self-score and the corresponding other-score from the target large language model.
5. The method according to claim 1, characterized in that, The acquisition of N translated text pairs includes: Obtain the sample set; The sample set was split to obtain text information in multiple languages; At least some of the language text information in the multiple languages is translated to obtain at least some of the N translated text pairs.
6. A method for training a translation model, characterized in that, include: Obtain the first language text information to be translated; The original translation model is pruned to obtain a pruned model. The pruning model is distilled using the first language text information to be translated to obtain a distillation model; The distillation model is trained using a parallel corpus dataset, wherein the parallel corpus dataset is obtained by the method according to any one of claims 1-5.
7. The method according to claim 6, characterized in that, The process of pruning the original translation model to obtain a pruned model includes: Perform at least one of the following removal operations on the original translation model: Some token layers of the original translation model are removed; Some of the converter layers of the original translation model are removed; Based on the different objects to be removed, M candidate pruning models are obtained, where M is an integer greater than 1; The verification first language text information is input into M candidate pruning models for translation, and the verification translation text pair of each candidate pruning model is obtained. The verification translation text pair includes the first language text and the corresponding second language text. The verification translated text pairs of each candidate pruning model are input into the ensemble evaluation model for scoring, and the score corresponding to each candidate pruning model is obtained. The candidate pruning model with the highest score was selected as the pruning model.
8. An apparatus for constructing a parallel corpus dataset for training a translation model, characterized in that, include: The translation text pair acquisition module is used to acquire N translation text pairs, each of which includes a first language text and a corresponding second language text, where N is an integer greater than 1; The comprehensive scoring module is used to input each of the translated text pairs into the integrated evaluation model. The integrated evaluation model scores each of the translated text pairs according to preset indicators to obtain a comprehensive score. The integrated evaluation model integrates H large language models, where H is an integer greater than 1. The preset indicators include at least one of the following: semantic consistency before and after translation, text readability, and translation diversity. The comprehensive score represents the translation quality between the first language text and the corresponding second language text. A construction module is used to construct the parallel corpus dataset for training the translation model using translated text pairs whose comprehensive score is greater than the preset score.
9. A translation model training device, characterized in that, include: The language text information acquisition module is used to acquire the first language text information to be translated; The pruning module is used to prune the original translation model to obtain a pruned model. The distillation processing module is used to perform distillation processing on the pruning model using the first language text information to be translated, to obtain a distilled model. The model training module is used to train the distillation model using a parallel corpus dataset, wherein the parallel corpus dataset is obtained by the method according to any one of claims 1-5.
10. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; characterized in that the processor is configured to read the program from the memory to implement the steps in the method for constructing a parallel corpus dataset for training a translation model as described in any one of claims 1 to 5; or to implement the steps in the method for training a translation model as described in any one of claims 6 to 7.