Information processing method, apparatus, electronic device, and storage medium
By distilling language sentences within bilingual pairs using large language models, the method improves translation quality while reducing computational and financial costs, addressing the challenges of deploying large language models for machine translation.
Patent Information
- Application Number
- JP2025034590
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-05
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-05
AI Technical Summary
Current large language models for machine translation require high computational resources and are costly to deploy, posing a challenge for efficient and cost-effective translation solutions.
The method involves acquiring a bilingual sentence pair and distilling one of the language sentences using a large language model to produce a second bilingual sentence pair with improved translation quality, thereby reducing computational demands.
This approach enhances translation quality while reducing the computational and financial burdens associated with using large language models, offering a more efficient and cost-effective translation method.
Smart Images

Figure 2025084983000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the fields of machine translation, deep learning, and large language models, and particularly to an information processing method, apparatus, electronic device, and storage medium.
Background Art
[0002] Machine translation is a discipline that uses computers to perform human language translation and is a core technology for breaking down language barriers. Currently, neural network machine translation is the mainstream technology, and neural network machine translation has significantly improved in terms of the quality of translation compared to conventional machine translation methods.
[0003] Currently, large language models can strongly express the abilities of understanding, generation, memory, and reasoning and also achieve excellent effects in cross-language task machine translation. However, large language models also face the challenge of having a large number of parameters and high computing power requirements. For machine translation technology, the cost of directly using large language models is high.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The present disclosure provides an information processing method, apparatus, electronic device, and storage medium.
Means for Solving the Problems
[0005] According to one aspect of the present disclosure, an information processing method is provided, an acquisition step of acquiring a first bilingual sentence pair including a source language sentence and a target language sentence, a step of distilling a first language sentence in the first bilingual sentence pair based on a large language model to obtain a second bilingual sentence pair after distillation, where the first language sentence is the source language sentence or the target language sentence.
[0006] According to another aspect of the present disclosure, an information processing apparatus is provided, an acquisition module for acquiring a first bilingual sentence pair including a source language sentence and a target language sentence, a distillation module for distilling a first language sentence in the first bilingual sentence pair based on a large language model to obtain a second bilingual sentence pair after distillation, where the first language sentence is the source language sentence or the target language sentence.
[0007] According to a third aspect of the present disclosure, an electronic device is provided, including at least one processor and a memory communicatively connected to the at least one processor, where instructions executable by the at least one processor are stored in the memory, and when the instructions are executed by the at least one processor, the at least one processor executes the method described in the first aspect.
[0008] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is further provided, and the computer instructions are used to cause the computer to execute the method described in the first aspect.
[0009] According to a fifth aspect of the present disclosure, a computer program is provided, and when the computer program is executed by a processor, the steps of the method described in the first aspect are realized.
Advantages of the Invention
[0010] In an embodiment of the present disclosure, a first language sentence in a first bilingual sentence pair is distilled by a large language model to improve the translation effect of the first language sentence, obtain a second bilingual sentence pair with more accurate translation, and have higher translation quality compared to conventional machine translation.
[0011] Note that the content described in this part is not intended to identify the essential or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be easily understood through the following description.
Brief Description of the Drawings
[0012] The drawings are for better understanding of the solution means and do not limit the present disclosure.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0013] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. For ease of understanding, various details of the embodiments of the present disclosure are included therein, and they should be regarded as merely exemplary. Therefore, as will be appreciated by those skilled in the art, various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and brevity, the description of well-known functions and structures will be omitted in the following description.
[0014] Artificial Intelligence, abbreviated as AI in English, is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, exploring, and expanding human intelligence. It can simulate the information processes of human consciousness and thinking. The main purpose of artificial intelligence research is for machines to undertake complex tasks that can only be completed when they possess normal human intelligence.
[0015] Machine Translation, also known as automatic translation, is a process of using a computer to convert one natural language (source language) into another natural language (target language). The development of machine translation technology is always closely related to the development of disciplines such as computer technology, information theory, and linguistics.
[0016] Deep Learning is a research direction of information in the field of machine learning. It learns the inherent rules and representation hierarchies of sample data. The information obtained during the learning process is of great help in interpreting data such as text, images, and voices. The ultimate goal is for machines to have the ability to analyze and learn like humans and be able to recognize data such as text, images, and voices.
[0017] Large Language Model (LLM) is an artificial intelligence model aimed at understanding and generating language. It can execute a wide range of tasks, including text summarization, translation, and sentiment analysis. Its feature is that it is extremely large-scale, containing tens of billions to trillions of parameters, and learning complex patterns in language.
[0018] Figure 1 is a schematic diagram of the information processing method provided by an embodiment of the present disclosure. As shown in Figure 1, the method includes the following.
[0019] In S101, obtain a first bilingual sentence pair.
[0020] The first bilingual sentence pair includes a source language sentence and a target language sentence.
[0021] The source language text is the text to be translated, and the target language text is the translated text. For example, when the translation request is to translate a Chinese text into an English text, the source Chinese text to be translated is the source language text, and the translated English text is the target language text.
[0022] In some implementations, the first bilingual text pair can be read from a database or collected from a network. As can be understood, the source language may be an existing language such as Chinese, English, Korean, and Japanese, and correspondingly, the target language may also be an existing language such as Chinese, English, Korean, and Japanese.
[0023] In some implementations, the source language text and the target language text in the first bilingual text are texts that appear in pairs, that is, the source language text and the target language text are texts with the same meaning expressed in different languages. For example, the source language text is A expressed in Chinese, and the target language text is A expressed in English.
[0024] In S102, distill the first language text in the first bilingual text pair based on a large language model to obtain a second bilingual text pair after distillation.
[0025] The first language text is the source language text or the target language text.
[0026] The purpose of distillation is to extract the features and capabilities in the model and compress them into a form that can be processed with less effort. Optionally, the distillation in this embodiment may perform translation distillation or optimization distillation on the text, and obtain a second bilingual text pair with better translation quality through the two types of distillation methods.
[0027] In some implementations, translation distillation is the translation ability to translate sentences and extract large language models. As can be understood, after translating the first language sentence in the first bilingual sentence pair based on the large language model, the translated sentence can be obtained, and the translated sentence and the sentence input into the large language model are used to form the second bilingual sentence pair.
[0028] Illustratively, if the first bilingual sentence pair includes the source language sentence A and the target language sentence B, the source language sentence A in the first bilingual sentence pair is input into the large language model for distillation to obtain the translated sentence C after translation. At this time, the translated sentence C and the source language sentence A may be included in the second bilingual sentence pair.
[0029] In another implementation, the distillation in this embodiment can further optimize or polish the sentence. That is, the optimization distillation optimizes or polishes the first language sentence in the first bilingual sentence pair to obtain a sentence with more accurate and smoother expression.
[0030] Illustratively, assuming that the first bilingual sentence pair includes the source language sentence A and the target language sentence B, the target language sentence B in the first bilingual sentence is input into the large language model for distillation to obtain the optimized sentence D after optimization and polishing. At this time, the optimized sentence D and the target language sentence B can be included in the second bilingual sentence pair.
[0031] In this embodiment, a first bilingual sentence pair including the source language sentence and the target language sentence is obtained, the first bilingual sentence pair is distilled by the large language model, and a second bilingual sentence pair after the distillation process by the large language model is obtained. The second bilingual sentence pair has better translation effect and expression by distilling the large language model, ensuring the translation quality.
[0032] FIG. 2 is a schematic diagram of another information processing method provided by an embodiment of the present disclosure. As shown in FIG. 2, the method includes the following.
[0033] In S201, obtain a first bilingual sentence pair.
[0034] In an embodiment of the present disclosure, the implementation method of step S201 can be implemented by using any one of the methods in each embodiment of the present disclosure, which is not limited herein and the description is omitted.
[0035] In S202, determine the distillation target of the first bilingual sentence pair.
[0036] The distillation target is either translation distillation or polishing distillation.
[0037] In some implementations, when the distillation target is translation distillation, the main task of the large language model is to accurately translate sentences. When the distillation target is polishing distillation, the main task of the large language model is to polish and optimize sentences to make their expression effect better.
[0038] In S203, when the distillation target is translation distillation, generate a first prompt of the large language model based on the distillation target and the second language sentence in the first bilingual sentence pair.
[0039] Optionally, the second language sentence may be the source language sentence or the target language sentence. When the first language sentence is the source language sentence, the second language sentence is the target language sentence, and correspondingly, when the first language sentence is the target language sentence, the second language sentence is the source language sentence.
[0040] In some implementations, when the second language sentence is the source language sentence, it is determined that the translation distillation is the target language distillation, the second language sentence is used as the sentence to be translated, and a first prompt for translating the sentence to be translated into the target language is generated. That is, when the second language sentence is the source language sentence, the translation distillation is the target language distillation, the source language sentence is used as the sentence to be translated, and a first prompt for translating the source language sentence into the source language is generated. For example, the first prompt is "translate the source language sentence into the target language".
[0041] Optionally, when the second language sentence is the target language sentence, it is determined that the translation distillation is the source language distillation, the second language sentence is used as the sentence to be translated, and a first prompt for translating the sentence to be translated into the source language is generated. That is, when the second language sentence is the target language sentence, the translation distillation is the source language distillation, the target language sentence is used as the sentence to be translated, and a first prompt for translating the target language sentence into the source language is generated. For example, the first prompt is "translate the target language sentence into the source language", and the large language model is guided by the prompt to output a third language sentence with a better translation effect.
[0042] In S204, at least one language sentence input from the first bilingual sentence pair to the large language model is determined, the first prompt and at least one language sentence in the first bilingual sentence pair are input to the large language model for distillation, and a third language sentence corresponding to the first language sentence is obtained.
[0043] Optionally, the second language sentence can be determined as the language sentence input from the first bilingual sentence pair to the large language model.
[0044] As can be understood, when the first language sentence is the source language sentence, the second language sentence is the target language sentence, and the translation distillation is the source language distillation. Inputting the target language sentence (the sentence to be translated) and the first prompt into the large language model to obtain the third language sentence corresponding to the first language sentence means reconstructing the first language sentence (the source language sentence), that is, re-translating the second language sentence (the target language sentence) by the large language model to obtain the source language sentence corresponding to the first language sentence (the original source language sentence). Accordingly, when the first language sentence is the target language sentence, the second language sentence is the source language sentence, and the translation distillation is the target language distillation. Inputting the source language sentence (the sentence to be translated) and the first prompt into the large language model to obtain the third language sentence corresponding to the first language sentence, that is, re-translating the second language sentence (the source language sentence) by the large language model to obtain the target language sentence corresponding to the first language sentence (the original target language sentence).
[0045] In S205, generate a second bilingual sentence pair based on the second language sentence and the third language sentence in the first bilingual sentence pair.
[0046] When the first language sentence is the source language sentence, the second language sentence is the target language sentence; when the first language sentence is the target language sentence, the second language sentence is the source language sentence.
[0047] As can be understood, after inputting the first prompt and the second language sentence into the large language model, the third language sentence corresponding to the first language sentence is obtained. At this time, the first language sentence and the third language sentence are different representations of the same type of language. For example, the first language sentence X and the second language sentence Y are a pair of Chinese-English translation sentences. The large language model obtains the third language sentence X1 in Chinese by re-translating the second language sentence Y in English. At this time, the first language sentence X in Chinese and the third language sentence X1 in Chinese are translations of the second language sentence Y in English, and the third language sentence X1 and the second language sentence Y form the second bilingual sentence pair.
[0048] In this embodiment, when the distillation target is translation distillation, it is determined that the translation distillation is target language distillation or source language distillation according to the second language sentence, multi-language distillation is realized, the corresponding first prompt is determined, the second language sentence and the first prompt are input into the large language model, and the third language sentence corresponding to the first language sentence is obtained, that is, the third language sentence after reconstructing the first language sentence. The third language sentence and the original second language sentence are composed into a new second bilingual sentence pair, and the quality of translation during translation is ensured by re-translating the large language model, avoiding the influence of poor translation quality.
[0049] FIG. 3 is a schematic diagram of another information processing method provided by an embodiment of the present disclosure. As shown in FIG. 3, the method includes the following.
[0050] In S301, a first bilingual sentence pair is obtained.
[0051] In the embodiments of the present disclosure, the implementation method of step S301 can be realized by using any one of the methods in the embodiments of the present disclosure respectively, which is not limited here and the description is omitted.
[0052] In S302, the distillation target of the first bilingual sentence pair is determined.
[0053] In the embodiments of the present disclosure, the implementation method of step S302 can be realized by using any one of the methods in the embodiments of the present disclosure respectively, which is not limited here and the description is omitted.
[0054] In S303, when the distillation target is polishing distillation, a second prompt of the large language model is generated based on the distillation target and the first language sentence.
[0055] Optionally, the first language sentence can be used as the sentence to be polished, and a second prompt for polishing the sentence to be polished can be generated. That is, when the distillation target is polished distillation, it is necessary to polish the sentence. The first language sentence in the first bilingual sentence pair is used as the sentence to be polished, and a second prompt for polishing the sentence to be polished is generated.
[0056] In some implementations, the first language sentence may be the source language sentence or the target language sentence. When the first language sentence is the source language sentence, the second prompt may be "Polish the source language sentence according to the expression habits of the source language to make it smoother". When the first language sentence is the target language sentence, the second prompt may be "Polish the target language sentence according to the expression habits of the target language to make it smoother". Guide the large language model based on the second prompt to perform a polishing process on the first language sentence.
[0057] In S304, at least one language sentence input from the first bilingual sentence pair to the large language model is determined, and the second prompt and at least one language sentence in the first bilingual sentence pair are input into the large language model for distillation to obtain a third language sentence corresponding to the first language sentence.
[0058] Optionally, the first language sentence and its second prompt can be input into the large language model for polished distillation to obtain the polished sentence, that is, the third language sentence corresponding to the first language sentence.
[0059] Optionally, both the two language texts in the first bilingual text pair and the corresponding second prompt are input into the large language model, and the large language model can perform polishing distillation on the first language text and the second language text respectively to obtain the polished texts. That is, both the first language text and the second language text in the first bilingual text pair are used as the input texts for the large language model. Then, the second language text, the first language text, and the corresponding second prompt in the first bilingual text pair are all input into the large language model, and the third language text corresponding to the first language text and the second language text are output, achieving the purpose of polishing the text and making the text more vivid and smooth in expression.
[0060] In S305, based on the second language text and the third language text in the first bilingual text pair, a second bilingual text pair is generated.
[0061] As can be understood, the first language text in the first bilingual text pair is the text for performing distillation processing, and the third language text is the text obtained after being polished and distilled by the large language model. The expression of the third language text is more vivid and smooth. Therefore, the third language text and the second language text can form the second bilingual text pair, and the second bilingual text pair is the text after polishing processing, with a more vivid and smooth expression.
[0062] In the embodiments of the present disclosure, the implementation method of step S305 can be realized by using any one of the methods in each embodiment of the present disclosure respectively. Here, it is not limited thereto, and the description is omitted.
[0063] In this embodiment, when the distillation target is polishing distillation, a third language text with a smoother and more vivid expression is obtained by performing polishing distillation on the first language text. The first language text may be the source language text or the target language text. Through the polishing and distillation of the large language model, the expression and translation quality of the text are improved. Thereby, a second bilingual text pair with good quality is generated based on the third language text and the second language text, ensuring the translation quality.
[0064] Figure 4 is a schematic diagram of another information processing method provided by an embodiment of the present disclosure. As shown in Figure 4, the method includes the following.
[0065] In S401, a first bilingual sentence pair is obtained.
[0066] In the embodiments of the present disclosure, the implementation method of step S401 can be implemented by using any one of the methods in each embodiment of the present disclosure. Here, it is not limited thereto, and the description is omitted.
[0067] In S402, based on a large language model, the first language sentence in the first bilingual sentence pair is distilled to obtain a second bilingual sentence pair after distillation.
[0068] In the embodiments of the present disclosure, the implementation method of step S402 can be implemented by using any one of the methods in each embodiment of the present disclosure. Here, it is not limited thereto, and the description is omitted.
[0069] In S403, the second bilingual sentence pair and the first bilingual sentence pair are integrated to generate a reinforcement corpus library, and based on the reinforcement corpus library, a student model is trained.
[0070] As can be understood, the first bilingual sentence pair is an original sentence pair obtained by a database, the second bilingual sentence pair is a sentence pair with better translation quality after distillation by a large language model, the first bilingual sentence pair and the second bilingual sentence pair are integrated to obtain a reinforcement corpus library, the reinforcement corpus library contains multiple sets of corpora, each corpus contains a source language sentence and a target language sentence, that is, multiple bilingual sentence pairs, the student model is trained by the reinforcement corpus library, the student model learns the translation ability of the large language model, and the translation effect of the student model on sentences is ensured.
[0071] Optionally, the student model may be a language model with a small number of parameters (compared to large language models) or a conventional neural network model, such as a Long Short-Term Memory (LSTM) network, a Transformer, etc. As can be understood, the teacher model in this embodiment is a large language model with a large number of parameters, and by distillation, the translation ability of the large language model is migrated to the student model, enabling the student model to have the translation ability of the large language model while retaining the characteristics of high efficiency and easy deployment of the student model.
[0072] In some implementations, it is further possible to evaluate the quality of the corpora in the augmented corpus library, perform a screening operation on the corpora in the augmented corpus library based on the evaluation information of the corpus quality, obtain a target corpus library, and train the student model based on the target augmented corpus library. That is, evaluate the quality of the corpora in the augmented corpus library, select better-quality corpora to obtain a target corpus library, train the student model with the target corpus library, and obtain a student model that retains the translation ability of the large language model.
[0073] Optionally, the purpose of evaluating the quality of the corpora in the augmented corpus library is to evaluate the translation quality in the corpora. That is, the evaluation of the quality of a set of corpora is to characterize the correspondence between the source language and the target language in the corpus. The better the translation between the source language and the target language, the better the quality of the corpus.
[0074] In some implementations, the evaluation information of the quality of the corpus may be a value obtained for the quality of the corpus, and the higher the value obtained for the quality, the better the translation quality of the corresponding corpus. The screening process for the corpus in the reinforced corpus library based on the quality evaluation information determines a corpus set corresponding to the same source language sentence, and based on the quality evaluation information of each corpus in the corpus set, at least one target corpus corresponding to the same source language sentence may be screened from the corpus set.
[0075] Illustratively, for the source language sentence A, the corpus a and the corpus b corresponding to the source language sentence A are determined as the first bilingual sentence pair and the second bilingual sentence pair respectively, the quality evaluation information of the corpus a and the corpus b is obtained respectively, and the target corpus is determined based on the quality evaluation information.
[0076] Optionally, by comparing the quality evaluation information of each corpus in the corpus set, the corpus with the highest quality can be determined as the target corpus, or, based on the quality evaluation information of each corpus in the corpus set, the corpora in the corpus set can be sorted, and the corpora sorted in the front can be selected as the target corpus, or, by comparing the quality evaluation information of each corpus in the corpus set with a preset quality evaluation threshold, the corpora with a quality evaluation threshold or higher can be selected as the target corpus.
[0077] Illustratively, for a corpus set (including corpus a and corpus b) corresponding to source language sentence A, select the corpus with the highest quality in corpus a and corpus b as the target corpus, or when many corpora are included in the corpus set, sort all the corpus quality evaluation information in descending order, and select the corpora sorted in the front as the target corpus. Further, compare the quality evaluation information of each corpus in the corpus set with a preset quality evaluation threshold, and select the corpora with a quality evaluation threshold or higher as the target corpus. The specific selection method of the target corpus is not limited. As can be understood, the purpose of selecting the target corpus is to select a corpus with better translation quality. Use the target corpus library composed based on the target corpus as the training sample, that is, use the corpus with better translation quality as the training sample to ensure the training effect of the student model and enable the student model to learn the translation ability of the large language model.
[0078] In this embodiment, after obtaining the second bilingual sentence pair, based on the comparison between the second bilingual sentence pair and the first bilingual sentence pair, determine a target corpus with better translation quality, and construct a target corpus library based on the target corpus to train the student model, so that the student model can learn the translation ability of the large language model. Use the target corpus with better translation quality to ensure the translation quality of the student model, improve the resource utilization rate, and reduce the cost of directly deploying the large model.
[0079] FIG. 5 is a schematic diagram of another information processing method provided by an embodiment of the present disclosure. As shown in FIG. 5, the method includes the following.
[0080] In S501, obtain a first bilingual sentence pair.
[0081] In an embodiment of the present disclosure, the implementation method of step S501 can be implemented by using any one of the methods in each embodiment of the present disclosure respectively, which is not limited herein and the description is omitted.
[0082] In S502, determine the distillation target of the first bilingual sentence pair.
[0083] In an embodiment of the present disclosure, the implementation method of step S502 can be implemented by using any one of the methods in each embodiment of the present disclosure respectively, which is not limited herein and the description is omitted.
[0084] In S503, when the distillation target is translation distillation, based on the distillation target and the second language sentence in the first bilingual sentence pair, generate the first prompt of the large language model.
[0085] In an embodiment of the present disclosure, the implementation method of step S503 can be implemented by using any one of the methods in each embodiment of the present disclosure respectively, which is not limited herein and the description is omitted.
[0086] In S504, determine at least one language sentence input from the first bilingual sentence pair to the large language model, input the first prompt and at least one language sentence in the first bilingual sentence pair into the large language model for distillation, and obtain the third language sentence corresponding to the first language sentence.
[0087] In an embodiment of the present disclosure, the implementation method of step S504 can be implemented by using any one of the methods in each embodiment of the present disclosure respectively, which is not limited herein and the description is omitted.
[0088] In S505, when the distillation target is polishing distillation, based on the distillation target and the first language sentence, generate the second prompt of the large language model.
[0089] In an embodiment of the present disclosure, the implementation method of step S505 can be implemented by using any one of the methods in each embodiment of the present disclosure respectively, which is not limited herein and the description is omitted here.
[0090] In S506, at least one language sentence input from the first bilingual sentence pair to the large language model is determined, and the second prompt and at least one language sentence in the first bilingual sentence pair are input into the large language model for distillation to obtain a third language sentence corresponding to the first language sentence.
[0091] In an embodiment of the present disclosure, the implementation method of step S506 can be implemented by using any one of the methods in each embodiment of the present disclosure respectively, which is not limited herein and the description is omitted here.
[0092] In S507, a second bilingual sentence pair is generated based on the second language sentence and the third language sentence in the first bilingual sentence pair.
[0093] In an embodiment of the present disclosure, the implementation method of step S507 can be implemented by using any one of the methods in each embodiment of the present disclosure respectively, which is not limited herein and the description is omitted here.
[0094] In S508, the second bilingual sentence pair and the first bilingual sentence pair are integrated to generate a reinforcement corpus library, and a student model is trained based on the reinforcement corpus library.
[0095] In an embodiment of the present disclosure, the implementation method of step S508 can be implemented by using any one of the methods in each embodiment of the present disclosure respectively, which is not limited herein and the description is omitted here.
[0096] In this embodiment, a first bilingual sentence pair including a source language sentence and a target language sentence is obtained, the first bilingual sentence pair is distilled by a large language model, a second bilingual sentence pair after the distillation process of the large language model is obtained, and the second bilingual sentence pair has better sentence translation effects and expressions through the distillation of the large language model, ensuring the translation quality. A student model is trained based on a corpus library with better translation quality to ensure that the student model learns the translation ability of the large language model and provides a better translation effect.
[0097] FIG. 6 is a schematic diagram of an information processing apparatus provided by an embodiment of the present disclosure. As shown in FIG. 6, the information processing apparatus 600 includes an acquisition module 601 for acquiring a first bilingual sentence pair including a source language sentence and a target language sentence, a distillation module 602 for distilling a first language sentence in the first bilingual sentence pair based on a large language model to obtain a second bilingual sentence pair after distillation, where the first language sentence is a source language sentence or a target language sentence.
[0098] In some implementations, the distillation module 602 determines a distillation target for the first bilingual sentence pair, and the distillation target is translation distillation or polishing distillation, and is used to distill the first language sentence according to the distillation target by the large language model to obtain a second bilingual sentence pair.
[0099] In some implementations, the distillation module 602 generates a prompt for the large language model based on the distillation target and the first bilingual sentence pair, inputs the prompt and at least one language sentence in the first bilingual sentence pair into the large language model for distillation to obtain a third language sentence corresponding to the first language sentence, Used to generate a second bilingual sentence pair based on the second language sentence and the third language sentence in the first bilingual sentence pair. When the first language sentence is the source language sentence, the second language sentence is the target language sentence; when the first language sentence is the target language sentence, the second language sentence is the source language sentence.
[0100] In some implementations, the distillation module 602 Determines at least one language sentence input to the large language model from the first bilingual sentence pair based on the distillation target, And is used to input the prompt and at least one language sentence into the large language model for distillation.
[0101] In some implementations, the distillation module 602 When the distillation target is translation distillation, it is used to generate the first prompt of the large language model based on the distillation target and the second language sentence in the first bilingual sentence pair.
[0102] In some implementations, the distillation module 602 When the second language sentence is the source language sentence, it is determined that the translation distillation is source language distillation, And is used to generate the first prompt for translating the second language sentence as the target language sentence to be translated and the target sentence to the source language.
[0103] In some implementations, the distillation module 602 When the second language sentence is the target language sentence, it is determined that the translation distillation is target language distillation, And is used to generate the first prompt for translating the second language sentence as the target language sentence to be translated and the target sentence to the target language.
[0104] In some implementations, the distillation module 602 Is used to determine the second language sentence as the language sentence input to the large language model from the first bilingual sentence pair.
[0105] In some implementations, the distillation module 602 is used to generate a second prompt for the large language model based on the distillation target and the first language sentence when the distillation target is polished distillation.
[0106] In some implementations, the distillation module 602 uses the first language sentence as the sentence to be polished, and is used to generate a second prompt for polishing the sentence to be polished.
[0107] In some implementations, the distillation module 602 is used to use the first language sentence and the second language sentence in the first bilingual sentence pair as the language sentences input to the large language model.
[0108] In some implementations, the distillation module 602 further integrates the second bilingual sentence pair and the first bilingual sentence pair to generate a reinforcement corpus library, and is used to train the student model based on the reinforcement corpus library, and each corpus includes a source language sentence and a target language sentence.
[0109] In some implementations, the distillation module 602 evaluates the quality of the corpus in the reinforcement corpus library, performs a screening operation on the corpus in the reinforcement corpus library based on the evaluation information of the corpus quality to obtain a target corpus library, and is used to train the student model based on the target reinforcement corpus library.
[0110] In some implementations, the distillation module 602 determines a corpus set corresponding to the same source language sentence, It is used to select at least one target corpus corresponding to the same source language sentence from a corpus set based on the evaluation information of the quality of each corpus in the corpus set.
[0111] In some implementations, the distillation module 602 By comparing the evaluation information of the quality of each corpus in the corpus set, the corpus with the highest quality is used as the target corpus, or Based on the evaluation information of the quality of each corpus in the corpus set, by sorting the corpora in the corpus set, the corpora sorted in the front are selected as the target corpus, or It is used to compare the evaluation information of the quality of each corpus in the corpus set with a preset quality evaluation threshold, and select the corpus with a quality evaluation threshold or higher as the target corpus.
[0112] In this embodiment, a first bilingual sentence pair including a source language sentence and a target language sentence is obtained, the first bilingual sentence pair is distilled by a large language model, a second bilingual sentence pair after the distillation process of the large language model is obtained, and the second bilingual sentence pair has better translation effects and expressions due to the distillation of the large language model, guarantees the translation quality, trains a student model based on a corpus library with better translation quality, ensures that the student model learns the translation ability of the large language model, and provides a better translation effect.
[0113] In the technical solution of the present disclosure, the acquisition, storage, and application of such user personal information all comply with the provisions of relevant laws and do not violate public order and good customs.
[0114] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.
[0115] FIG. 7 shows an exemplary block diagram of an exemplary electronic device 700 capable of implementing an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices such as personal digital processors, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the description herein and / or the implementation of the present disclosure required.
[0116] As shown in FIG. 7, the device 700 includes a computing unit 701 capable of performing various appropriate operations and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data necessary for the operation of the device 700 may be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is further connected to the bus 704.
[0117] A plurality of components including an input unit 706 such as a keyboard and a mouse, an output unit 707 such as various types of displays and speakers, a storage unit 708 such as a magnetic disk and an optical disk, and a communication unit 709 such as a network card, a modem, and a wireless communication transceiver are connected to the I / O interface 705 in the device 700. The communication unit 709 enables the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0118] The computing unit 701 may be various general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, computing units that execute various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes each of the above-described methods and processes, for example, the information processing method. For example, in some embodiments, the information processing method may be implemented as a computer software program tangibly incorporated into a machine-readable medium such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed into the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the above-described information processing method may be executed. Optionally, in other embodiments, the computing unit 701 may be configured to execute the information processing method in any other suitable manner (e.g., by firmware).
[0119] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can be implemented in and / or interpreted in one or more computer programs obtained and executed on a programmable system including at least one programmable processor which is a special-purpose or general-purpose programmable processor, and can include receiving data and instructions from a memory system, at least one input device, and at least one output device, and transmitting data and instructions to the memory system, the at least one input device, and the at least one output device.
[0120] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, and when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package, partially on a remote machine, or entirely on a remote machine or server.
[0121] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes, or can store, a program for use by or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include electrical connections based on one or more wirings, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0122] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with a user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user can be in any form, including voice input, speech input, or tactile input.
[0123] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be connected to each other via any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0124] A computer system can include clients and servers. Clients and servers are generally physically separate and typically interact via a communication network. The relationship between a client and a server is generated by computer programs that run on respective computers and have a client-server relationship with each other. A server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0125] It should be understood that the various forms of flow shown above can be used to rearrange, add, or delete steps. For example, each step described in this application may be executed in parallel, sequentially, or in a different order, but is not limited herein as long as the technical solutions disclosed in this application can achieve the desired results.
[0126] The above specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art can make various modifications, combinations, sub-combinations, and substitutions based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure should all be included within the protection scope of the present disclosure.
Claims
1. obtaining a first bilingual sentence pair including a source language sentence and a target language sentence; distilling a first language sentence in the first bilingual sentence pair based on a large-scale language model to obtain a distilled second bilingual sentence pair, wherein the first language sentence is the source language sentence or the target language sentence. Information processing methods.
2. Distilling first language sentences in the first bilingual sentence pair based on the large-scale language model to obtain a distilled second bilingual sentence pair, determining a distillation target for the first bilingual sentence pair, the distillation target being a translation distillation or an embellishment distillation; distilling the first language sentence according to the distillation target by the large-scale language model to obtain the second bilingual sentence pair; The method of claim 1.
3. The step of distilling the first language sentence according to the distillation target by the large-scale language model to obtain the second bilingual sentence pair includes: generating suggested words for the large-scale language model based on the distillation target and the first bilingual sentence pair; inputting the suggested word and at least one language sentence in the first bilingual sentence pair into the large-scale language model for distillation to obtain a third language sentence corresponding to the first language sentence; generating the second bilingual sentence pair based on the second language sentence and the third language sentence in the first bilingual sentence pair; When the first language sentence is the source language sentence, the second language sentence is the target language sentence, and when the first language sentence is the target language sentence, the second language sentence is a source language sentence. The method of claim 2.
4. The step of inputting the suggested word and at least one language sentence in the first bilingual sentence pair into the large-scale language model for distillation includes: determining at least one language sentence from the first bilingual sentence pair to be input to the large-scale language model based on the distillation target; The suggested words and the at least one language sentence are input to the large-scale language model for distillation. The method according to claim 3.
5. generating suggested words for the large-scale language model based on the distillation target and the first bilingual sentence pair, When the distillation target is a translation distillation, generating a first suggested word of the large-scale language model based on the distillation target and a second language sentence in the first bilingual sentence pair; The method according to claim 3.
6. generating first suggested words of the large-scale language model based on the distillation target and a second language sentence in the first bilingual sentence pair, if the second language sentence is the source language sentence, determining that the translation distillation is a source language distillation; and generating a first suggested word for translating the second language sentence into a source language, the first suggested word being a translation target language sentence, the second language sentence being a translation target language sentence, The method according to claim 5.
7. generating first suggested words of the large-scale language model based on the distillation target and a second language sentence in the first bilingual sentence pair, If the second language sentence is the target language sentence, determining that the translation distillation is a target language distillation; and generating a first suggested word for translating the second language sentence into a target language, the second language sentence being a translation target language sentence. The method according to claim 5.
8. Determining at least one language sentence from the first bilingual sentence pair to be input to the large-scale language model based on the distillation target includes: determining the second language sentence from the first bilingual sentence pair as a language sentence to be input to the large-scale language model; The method according to claim 5.
9. The step of generating a suggested word of the large-scale language model based on the distillation target and the first language sentence includes: When the distillation target is an embellishment distillation, generating a second presentation word of the large-scale language model based on the distillation target and the first language sentence; The method of claim 2.
10. The step of generating a second suggested word of the large-scale language model based on the distillation target and the first language sentence includes: The method includes a step of treating the first language sentence as a language sentence to be retouched, and generating a second suggested word for retouching the language sentence to be retouched.
10. The method of claim 9.
11. Determining at least one language sentence from the first bilingual sentence pair to be input to the large-scale language model based on the distillation target includes: the first language sentence and the second language sentence of the first bilingual sentence pair being language sentences input to the large-scale language model; 10. The method of claim 9.
12. Distilling a first language sentence in the first bilingual sentence pair based on a large-scale language model to obtain a distilled second bilingual sentence pair, further comprising: merging the second bilingual sentence pairs with the first bilingual sentence pairs to generate an augmented corpus library; and training a student model based on the augmented corpus library, each corpus including source language sentences and target language sentences. The method of claim 1.
13. Training a student model based on the augmented corpus library includes: performing a quality assessment on the corpora in the augmented corpus library; performing a selection operation on the corpora in the enrichment corpus library based on the evaluation information of the quality of the corpus to obtain a target corpus library; training a student model based on the target augmented corpus library; The method of claim 12.
14. The step of performing a selection operation on the corpus in the enrichment corpus library based on the evaluation information of the quality of the corpus to obtain a target corpus library includes: determining a set of corpora corresponding to a same source language sentence; and selecting at least one target corpus from the corpus set that corresponds to the same source language sentence based on evaluation information of quality of each corpus in the corpus set. The method of claim 13.
15. The step of selecting at least one target corpus corresponding to the same source language sentence from the corpus set based on evaluation information of quality of each corpus in the corpus set includes: determining the corpus with the highest quality as the target corpus by comparing quality evaluation information of each corpus in the set of corpora; or sorting the corpora in the corpus set based on the evaluation information of the quality of each corpus in the corpus set, thereby selecting a top-sorted corpus as the target corpus; or a step of comparing quality evaluation information of each corpus in the corpus set with a preset quality evaluation threshold, and selecting a corpus having a quality evaluation value equal to or greater than the quality evaluation threshold as the target corpus; The method of claim 14.
16. an acquisition module for acquiring a first bilingual sentence pair including a source language sentence and a target language sentence; a distillation module for distilling a first language sentence in the first bilingual sentence pair based on a large-scale language model to obtain a distilled second bilingual sentence pair, the first language sentence being the source language sentence or the target language sentence; Information processing device.
17. The distillation module comprises: Determine a distillation target for the first bilingual sentence pair, the distillation target being a translation distillation or an embellishment distillation; distilling the first language sentence according to the distillation target by the large-scale language model to obtain the second bilingual sentence pair; 17. The apparatus of claim 16.
18. The distillation module comprises: generating suggested words for the large-scale language model based on the distillation target and the first bilingual sentence pair; inputting the suggested word and at least one language sentence in the first bilingual sentence pair into the large-scale language model and distilling the inputted word and the at least one language sentence in the first bilingual sentence pair to obtain a third language sentence corresponding to the first language sentence; to generate the second bilingual sentence pair based on the second language sentence and the third language sentence in the first bilingual sentence pair; When the first language sentence is the source language sentence, the second language sentence is the target language sentence, and when the first language sentence is the target language sentence, the second language sentence is a source language sentence.
20. The apparatus of claim 17.
19. The distillation module comprises: determining at least one language sentence from the first bilingual sentence pair to be input to the large-scale language model based on the distillation target; The suggested words and the at least one language sentence are input to the large-scale language model for distillation.
20. The apparatus of claim 18.
20. The distillation module comprises: When the distillation target is a translation distillation, the distillation target is used to generate a first suggested word of the large-scale language model based on the distillation target and a second language sentence in the first bilingual sentence pair; 20. Apparatus according to claim 18 or 19.
21. The distillation module comprises: If the second language sentence is the source language sentence, determining that the translation distillation is a source language distillation; The second language sentence is a translation target language sentence, and a first suggested word is generated for translating the translation target sentence into a source language.
21. The apparatus of claim 20.
22. The distillation module comprises: If the second language sentence is the target language sentence, determining that the translation distillation is a target language distillation; The second language sentence is a translation target language sentence, and a first suggested word is generated for translating the translation target sentence into a target language.
21. The apparatus of claim 20.
23. The distillation module comprises: used to determine the second language sentence from the first bilingual sentence pair as a language sentence to be input to the large-scale language model.
21. The apparatus of claim 20.
24. The distillation module comprises: When the distillation target is an embellishment distillation, the second presentation word of the large-scale language model is generated based on the distillation target and the first language sentence.
20. The apparatus of claim 17.
25. The distillation module comprises: The first language sentence is a language sentence to be retouched, and a second suggestion word is generated to retouch the first language sentence.
25. The apparatus of claim 24.
26. The distillation module comprises: the first language sentence and the second language sentence in the first bilingual sentence pair are used as language sentences to be input to the large-scale language model; 26. Apparatus according to claim 24 or 25.
27. The distillation module further comprises: merging the second bilingual sentence pairs with the first bilingual sentence pairs to generate an augmented corpus library, and using the augmented corpus library to train a student model, each corpus including source language sentences and target language sentences; 26. Apparatus according to any one of claims 16 to 20 or 24 or 25.
28. The distillation module comprises: evaluating the quality of the corpora in the augmented corpus library; performing a selection operation on the corpus in the enrichment corpus library based on the evaluation information of the quality of the corpus to obtain a target corpus library; used to train a student model based on the target augmented corpus library; 28. The apparatus of claim 27.
29. The distillation module comprises: determining a set of corpora corresponding to the same source language sentence; used to select at least one target corpus from the corpus set corresponding to the same source language sentence based on evaluation information of the quality of each corpus in the corpus set; 30. The apparatus of claim 28.
30. The distillation module comprises: determining the corpus with the highest quality as the target corpus by comparing quality evaluation information of each corpus in the set of corpora; or Sorting the corpora in the corpus set based on the evaluation information of the quality of each corpus in the corpus set, and selecting a corpus sorted first as the target corpus; or The quality evaluation information of each corpus in the corpus set is compared with a predetermined quality evaluation threshold, and a corpus having a quality evaluation value equal to or greater than the quality evaluation threshold is selected as the target corpus.
30. The apparatus of claim 29.
31. An electronic device, At least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform a method according to any one of claims 1 to 15. electronic equipment.
32. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause the computer to carry out a method according to any one of claims 1 to 15. A non-transitory computer-readable storage medium.
33. A computer program comprising: The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15. Computer program.
Citation Information
Patent Citations
Machine translation quality evaluation method, device, equipment and medium
CN112347795A
A method and apparatus for predicting the quality of unsupervised machine translation based on knowledge distillation
CN114936567A
Method and device for constructing machine translation model based on double knowledge distillation
CN116644763A
Text translation method and related device, electronic equipment and storage medium
CN116976364A
Text Generation System
JP2022174244A