Large language model machine translation strengthening method based on prompt optimization

Through the machine translation enhancement method of large language model based on prompt optimization, the problem that traditional fine-tuning methods are difficult to improve model performance is solved, and more accurate machine translation results and more efficient parameter fine-tuning are achieved.

CN120068892AActive Publication Date: 2025-05-30HARBIN INST OF TECH

Patent Information

Application Number
CN202510107860.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

Traditional fine-tuning methods for large language models are difficult to improve model performance, resulting in inaccurate machine translation results.

Method used

The machine translation enhancement method of large language model based on prompt optimization is adopted, and the prompt decoder is pre-trained and fine-tuned, combined with the SVD-LoRA method for end-to-end training, and the prompts of machine translation are optimized using an external knowledge base.

Benefits of technology

Improve the translation performance of large language models, automatically optimize prompts, shorten the input prompt length, improve the efficient parameter fine-tuning method, and improve the performance of the model on target translation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068892A_ABST
    Figure CN120068892A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model machine translation strengthening method based on prompt optimization, and belongs to the technical field of machine translation strengthening. The problem that in the prior art, due to the fact that a traditional fine adjustment method for a large language model is difficult to improve the model performance, the model translation result is inaccurate is solved. According to the method, the prompt decoder is pre-trained and fine-tuned through the prompt decoder, the pre-trained and fine-tuned prompt decoder is obtained, and a large language model based on the prompt decoder is constructed; introducing an SVD-LoRA method, and performing end-to-end training on the large language model based on the prompt decoder to obtain a trained large language model; and on the basis of an external knowledge base, constructing an optimized machine translation prompt, and inputting the optimized machine translation prompt into the trained large language model to obtain a target end statement. The method improves the translation performance of the large language model, can automatically optimize the prompt and shorten the input prompt length, and can be applied to fine tuning of the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for enhancing machine translation of large language models, and particularly to a method for enhancing machine translation of large language models based on prompt optimization, belonging to the technical field of machine translation enhancement. Background Art

[0002] Large Language Models (LLMs) are one of the current research hotspots in the field of artificial intelligence. With the development of deep neural network structures and computing power, neural networks have been widely applied in various industries. Based on the serial modeling scheme of the transformer (self-attention mechanism neural machine translation system) and the good parallel computing efficiency brought by the model structure, large language models have emerged. The large number of parameters and training corpora of large language models enable large language models to have good generalization ability and can also achieve good results when facing complex downstream tasks that have not been trained.

[0003] However, due to the large number of parameters of large language models, training large language models using traditional fine-tuning methods will cause extremely high computing power consumption and may also bring the problem of pre-trained knowledge forgetting, reducing the comprehensive performance of the model.

[0004] In summary, a method for enhancing machine translation of large language models based on prompt optimization is needed. Summary of the Invention

[0005] A brief overview of the present invention is given below to provide a basic understanding of certain aspects of the present invention. It should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is only to present certain concepts in a simplified form as a prelude to the more detailed description to be discussed later.

[0006] In view of this, to solve the problem that the traditional fine-tuning method for large language models in the prior art is difficult to improve the model performance, resulting in inaccurate model translation results, the present invention provides a method for enhancing machine translation of large language models based on prompt optimization.

[0007] The technical solution is as follows: A method for enhancing machine translation of large language models based on prompt optimization, including the following steps:

[0008] S1. Pre-train and fine-tune the prompt decoder through the prompt decoder to obtain the pre-trained and fine-tuned prompt decoder, and construct a large language model based on the prompt decoder;

[0009] S2. Introduce the SVD-LoRA method to perform end-to-end training on the large language model based on the prompt decoder to obtain the trained large language model;

[0010] S3. Based on the external knowledge base, construct the optimized prompts for machine translation, and input the optimized prompts for machine translation into the trained large language model to obtain the target-side sentences.

[0011] Further, in S1, it specifically includes the following steps:

[0012] S11. Input the text into the prompt encoder, and pre-train the prompt encoder with the help of the prompt decoder to obtain the hyperparameters of the prompt encoder and the pre-trained prompt encoder;

[0013] S12. According to the hyperparameters of the prompt encoder, select the input data and fine-tune the pre-trained prompt encoder to obtain the fine-tuned prompt encoder;

[0014] In S11, an unsupervised monolingual corpus is selected as the text based on the main language of the large language model to pre-train the prompt encoder. According to the text sequence of length M output by the prompt encoder, the text sequence of length N at the input end of the prompt encoder is restored through the prompt decoder. According to the effect of the prompt decoder restoring the text sequence during the pre-training process, determine the value of the hyperparameter M, so that the prompt encoder compresses the text sequence of length N to the text sequence of length M;

[0015] In S12, initialize with the hyperparameter M of the prompt encoder in the pre-training stage, and use the p-tuning method for fine-tuning training. During the fine-tuning process, input the instruction fine-tuning datasets in the general domain and different tasks into the large language model, and align the language with the language used in the pre-training stage of the prompt encoder. Use the instruction text in the fine-tuning dataset as the continuous prompt. The input data and output data of the large language model are the source-side sentences and target-side sentences of the parallel corpus respectively, and the parameters of the large language model remain fixed. The adjustable parameters are the parameters of the prompt encoder and the continuous prompt.

[0016] Further, in S2, it specifically includes the following steps:

[0017] S21. Establish the SVD-LoRA method according to the LoRA method, and construct a large language model based on the prompt decoder combined with the SVD-LoRA method;

[0018] S22. Conduct end-to-end training based on prompt optimization on the large language model based on the prompt decoder combined with the SVD-LoRA method to obtain the trained large language model;

[0019] In S21, the process of the LoRA method is expressed as:

[0020] W step=0 = kW0 +ΔW = W 0

[0021] Wherein, W step=0 is the weight before the large language model starts training, ΔW is the weight added by the LoRA bypass, k is a constant, obtained according to the scaling ratio of ΔW to the current layer W 0 ; W 0 is the current layer;

[0022] When k = 1 and ΔW = 0, the process of the LoRA method is expressed as:

[0023] W step=0 = W 0 + BA = W 0 , ΔW = BA = 0, k = 1

[0024] Wherein, A is the first low-rank matrix and B is the second low-rank matrix;

[0025] The process of the SVD-LoRA method is to initialize the first low-rank matrix A and the second low-rank matrix B using the SVD decomposition result SVD'(W 0 ) of the current layer W 0 ;

[0026] The process of the SVD-LoRA method is expressed as:

[0027] ΔW = SVD'(W 0 ) = U'Σ'V' T

[0028] Wherein, U'Σ'V' T are respectively the SVD decomposition results after retaining the top K singular values and the corresponding singular vectors of the singular values;

[0029] In S22, the process of end-to-end training based on prompt optimization is as follows: Align the prompt encoder with the text of the large language model through adjustable parameters, perform end-to-end training of the framework using parallel corpora in the general domain. After the large language model converges, test the translation performance of the large language model using the continuous prompts obtained by training the prompt decoder and the domain translation test set, so as to solve translation tasks in different specific domains and obtain the trained large language model.

[0030] Furthermore, the construction of the optimized prompts for machine translation includes the following steps:

[0031] S31. Define the form of the prompt instruction;

[0032] S32. Define the selection of examples;

[0033] S33. Define the encoding range of the prompt encoder;

[0034] In S31, the method adopted is to manually write the initialization prompt instructions, and rely on the prompt encoder to convert the initialized text prompt into a continuous prompt, which is dynamically iterated during the training process of each stage of the large language model;

[0035] In S32, the method of retrieval-augmented generation is used. The BGE model is used to vectorize the examples in the example library, and the topK results in the example library are selected as the example for the prompt according to the text to be translated;

[0036] In S33, the effective information brought by each component in the prompt to the translation task is sorted in ascending order, which are: instructions, translation examples in the form of bilingual sentence pairs, and external knowledge introduced by the bilingual dictionary. Based on the above order, the component with the lowest information volume is compressed first, and the specific coding range is instructions and translation examples.

[0037] The beneficial effects of the present invention are as follows: The present invention proposes a method for enhancing machine translation of large language models based on prompt optimization, which can automatically optimize the prompt and shorten the input prompt length of the large language model, thereby improving the translation performance of the large language model. The constructed prompt encoder is used to optimize and compress the input prompt of the large language model, and the parameter-efficient fine-tuning method is improved to enhance the translation performance of the large language model through end-to-end training; The present invention also proposes a retrieval-augmented generation method for enhancing the translation performance of large language models based on an external knowledge base, which provides good initialization for the input prompt of the large language model, thereby achieving better performance in the target translation task. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation of the present invention. In the drawings:

[0039] Figure 1 It is a schematic flowchart of a method for enhancing machine translation of large language models based on prompt optimization;

[0040] Figure 2 It is a schematic flowchart of an embodiment of a method for enhancing machine translation of large language models based on prompt optimization;

[0041] Figure 3 It is a schematic structural diagram of a large language model based on a prompt decoder. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] In order to make the technical solutions and advantages in the embodiments of the present invention clearer and more understandable, the exemplary embodiments of the present invention are further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0043] Reference Figures 1 - 3 A method for enhancing machine translation of a large language model based on prompt optimization is described in detail in this embodiment, which specifically includes the following steps:

[0044] S1. The prompt decoder is pre-trained and fine-tuned through the prompt decoder to obtain a pre-trained and fine-tuned prompt decoder, and a large language model based on the prompt decoder is constructed;

[0045] S2. The SVD-LoRA method is introduced to perform end-to-end training on the large language model based on the prompt decoder to obtain a trained large language model;

[0046] S3. Based on an external knowledge base, an optimized prompt for machine translation is constructed, and the optimized prompt for machine translation is input into the trained large language model to obtain the target end statement.

[0047] Furthermore, in S1, the following steps are specifically included:

[0048] S11. The text is input into the prompt encoder of the constructed large language model, and the prompt encoder is pre-trained with the help of the prompt decoder to obtain the hyperparameters of the prompt encoder and compress the text length;

[0049] S12. According to the hyperparameters of the prompt encoder, input data is selected to fine-tune the prompt encoder;

[0050] In S11, unsupervised monolingual corpus is selected as the text based on the main language of the large language model to pre-train the prompt encoder. According to the text sequence of length M output by the prompt encoder, the text sequence of length N at the input end of the prompt encoder is restored through the prompt decoder. According to the effect of the prompt decoder restoring the text sequence during the pre-training process, the value of the hyperparameter M is determined so that the prompt encoder can compress the text sequence of length N to a text sequence of length M;

[0051] In S12, it is initialized with the hyperparameter M of the prompt encoder in the pre-training stage and fine-tuned using the p-tuning method. During the fine-tuning process, an instruction fine-tuning dataset in the general domain and different tasks is input into the large language model, and the language is aligned with the language used in the pre-training stage of the prompt encoder. The instruction text in the fine-tuning dataset is used as a continuous prompt. The input data and output data of the large language model are the source sentences and target sentences of the parallel corpus respectively. The parameters of the large language model remain fixed, and the adjustable parameters are the parameters of the prompt encoder and the continuous prompt.

[0052] Specifically, referring to Figure 2 , Prompt Encoder is the prompt encoder, LLM is the large language model, PromptDecoder is the prompt decoder, soft prompt is the soft prompt. The main purpose of step S1 is to achieve the length compression of the prompt and the continuous optimization of the prompt. These two goals are achieved in the pre-training and fine-tuning stages of the prompt encoder respectively;

[0053] In the pre-training stage, to enable the large language model to compress and represent the input prompt, the structure of the prompt decoder is used. If the large language model can converge in the pre-training stage and the restoration effect of the prompt decoder is relatively good, it can be considered that the prompt encoder can compress the information of the text sequence of length N to a length sequence of length M, and the information loss during the compression process is acceptable. Since the goal of the pre-training stage is only to enable the prompt encoder to obtain the ability of text length compression and the training method does not require supervised data, unsupervised monolingual corpus can be selected for training based on the main language of the large model. The role of the prompt decoder is only to assist the pre-training of the prompt encoder and will not participate in the subsequent large language model training process;

[0054] In the fine-tuning stage, to enable the prompt encoder to obtain the continuous optimization ability of the input prompt, step S1 designs a targeted encoding module, namely the prompt encoder, for the input prompt part of the large language model, effectively compressing and optimizing the input prompt of the large language model.

[0055] Furthermore, in S2, it specifically includes the following steps:

[0056] S21. Establish the SVD-LoRA method according to the LoRA method and construct a large language model based on the prompt decoder combined with the SVD-LoRA method;

[0057] S22. Conduct end-to-end training based on prompt optimization on the large language model based on the prompt decoder combined with the SVD-LoRA method to obtain the trained large language model;

[0058] In S21, the process of the LoRA method is expressed as:

[0059] W step=0 = kW 0 + ΔW = W 0

[0060] Among them, W step=0 is the weight before the training of the large language model, ΔW is the weight added by the LoRA bypass, k is a constant, which is approximately obtained according to the scaling ratio of ΔW and the current layer W 0 , and W 0 is the current layer;

[0061] The process of the LoRA method when k = 1 and ΔW = 0 is expressed as:

[0062] W step=0 = W 0 + BA = W 0 , ΔW = BA = 0, k = 1

[0063] Among them, A is the first low-rank matrix and B is the second low-rank matrix;

[0064] The process of the SVD-LoRA method is to initialize the first low-rank matrix A and the second low-rank matrix B by using the SVD decomposition result SVD'(W 0 ) of the current layer W 0 ;

[0065] The process of the SVD-LoRA method is expressed as:

[0066] ΔW = SVD'(W 0 ) = U'Σ'V' T

[0067] Among them, U'Σ'V' T are the SVD decomposition results after retaining the top K singular values and the corresponding singular vectors of the singular values, which can be approximated as a proportional reduction of W 0 .

[0068] In the above S22, the process of end-to-end training based on prompt optimization is as follows: Align the prompt encoder with the text of the large language model through adjustable parameters, and perform end-to-end training of the framework using parallel corpora in the general domain. After the large language model converges, test the translation performance of the large language model with the continuous prompts obtained by training the prompt decoder and the domain translation test set to solve translation tasks in different specific domains, and obtain the trained large language model.

[0069] Specifically, the LoRA method ensures that the initial parameters of the large language model at the start of training are consistent with the parameters of the pre-trained model, i.e., the prompt encoder. However, in the initial stage of training using the traditional LoRA method, the LoRA method cannot directly utilize the information in the weight matrix W, resulting in the LoRA method not fully leveraging the information of the pre-trained model parameters;

[0070] The SVD-LoRA method overcomes the above problems. The specified K value after SVD decomposition is equivalent to the rank in the LoRA method, and the SVD-LoRA method does not affect the training process of the large language model. Therefore, it can be considered that the computational cost of the SVD-LoRA method is equal to that of the LoRA method;

[0071] The present invention improves the prompt optimization method. Refer to Figure 2 , in the end-to-end training stage, i.e., step S22, it is initialized with the prompt encoder obtained in step S1. The training process refers to the training method of the multi-modal large model and is divided into three stages. The adjustable parameter in the first stage is the alignment layer between the prompt encoder and the large model, which is used for the alignment of two text spaces; in the second stage, a general parallel corpus is used for end-to-end training of the framework to improve the model's translation ability; in the third stage, fine-tuning is performed for the specific field involved in the present invention to improve the final performance of the model;

[0072] Different from the SVD-LoRA method that focuses on improving the overall translation ability of the model, the training in this stage focuses on solving translation tasks in different specific fields. It is initialized with a pre-designed prompt, which is used as the input of the prompt Prompt in Figure 2 . The open-source parallel corpus for the specific field is used, which are the source-end sentences SourceSentence and the target-end sentences Target Sentence of the parallel corpus in Figure 1 respectively. The prompt encoder participates in the inference stage of the model synchronously. Since the translation tasks in different fields only have different parameters for continuous prompts, and the subsequent forward process of the large language model is the same, it is possible to complete the translation tasks in different fields in one batch using a single model, optimizing the inference efficiency of the model. By optimizing the end-to-end training method of the large language model, the performance of the large language model in the target translation task is improved;

[0073] The main purpose of step S2 is to improve the efficient fine-tuning of the parameters of the large language model. In terms of the efficient fine-tuning of the parameters of the large model, the prior art has proposed various different training methods, and the main difference lies in the different positions of the trainable parameters in the model. The adjustable parameters of the LoRA method cover all steps of the forward process of the large language model, and have a stronger representation ability for training data, but it can only process one task within the same batch in the prediction stage;

[0074] In a series of works optimized based on prompts, the adjustable parameters cannot cover all processes of the large language model's forward pass. However, due to this, when the large language model makes predictions, there are more common forward parts for different tasks, which makes it possible to process multiple tasks within the same batch. The present invention simultaneously adopts the above two training strategies, uses the proposed SVD-LoRA method (a method combining singular value decomposition and low-rank adaptation), and trains the large language model using parallel corpora in the general domain to improve the overall translation performance of the large model. For translation tasks in different sub-domains, the method of prompt optimization is used for training, enabling translation tasks in different domains to be grouped into a batch for inference and improving the inference efficiency of the model.

[0075] Furthermore, the construction of the optimized prompt for machine translation includes the following steps:

[0076] S31. Define the form of the prompt instruction;

[0077] S32. Define the selection of examples;

[0078] S33. Define the encoding range of the prompt encoder;

[0079] In S31, the method adopted is to manually write the initialized prompt instruction, and rely on the prompt encoder to convert the initialized text prompt into a continuous prompt, which is dynamically iterated during each stage of the training process of the large language model;

[0080] In S32, the method similar to retrieval-augmented generation is used, and the BGE model is used to vectorize the examples in the example library, and the topK results in the example library are selected as the example of the prompt according to the text to be translated;

[0081] In S33, considering that there may be information loss in the prompt text compressed by the prompt encoder, the effective information content brought by each component in the prompt to the translation task is sorted in ascending order, which is: instruction, translation example in the form of a bilingual sentence pair, external knowledge introduced by the bilingual dictionary. Based on the above order, the component with the lowest information content is compressed first, and the specific encoding range is the instruction and the translation example.

[0082] Specifically, the main purpose of step S3 is to design a good prompt initialization scheme for the continuous optimization training process, improve the information density of the prompt input to the large language model, compress the redundant information part in the prompt, and optimize the overall prompt part. The input to the large language model involved is divided into four parts: instruction, example, external knowledge, and text to be translated. Step S3 focuses on three sub-steps: the form of the prompt instruction, the selection of examples, and the encoding range of the prompt encoder;

[0083] In this embodiment, first, it is necessary to pre-train the prompt encoder based on the process shown in step S1, and based on the pre-trained prompt encoder, perform end-to-end training on the large language model with reference to the method in step S2. For the translation task containing supervised reference examples, with reference to the method in step S3, use the BGE model to establish an example library, and splice the retrieved examples in the prompt part to further improve the translation quality;

[0084] In the English-Chinese translation task, if the large language model is trained according to the traditional training framework, for the source sentence to be translated "Despite the complex geopolitical climate, international cooperation on climate change remains imperative for sustainable development.", the large language model will output "Despite the complex geopolitical climate, international cooperation on climate change remains imperative for sustainable development."; using the method proposed by the present invention for model training, the trained large language model will output "Despite the complex geopolitical environment, international cooperation on climate change is still crucial for sustainable development.", and for parts such as "geopolitical environment" in the sentence, a more accurate translation effect is presented.

[0085] Although the present invention has been described in terms of a limited number of embodiments, those skilled in the art in this technical field will appreciate, from the above description, that other embodiments can be contemplated within the scope of the invention thus described. In addition, it should be noted that the language used in this specification has been principally selected for readability and teaching purposes rather than for the purpose of explaining or limiting the subject matter of the invention. Therefore, many modifications and variations will be apparent to those of ordinary skill in the art in this technical field without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure made of the invention is illustrative, not restrictive, and the scope of the present invention is defined by the appended claims.

Claims

1. A large language model machine translation enhancement method based on prompt optimization, characterized in that: The following steps are involved: S1. Pre-train and fine-tune the prompt decoder through the prompt decoder to obtain the pre-trained and fine-tuned prompt decoder, and construct a large language model based on the prompt decoder; S2. Introduce the SVD-LoRA method to perform end-to-end training on the large language model based on the hint decoder to obtain the trained large language model; S3. Based on the external knowledge base, construct optimized machine translation prompts, input the optimized machine translation prompts into the trained large language model, and obtain the target end sentences.

2. According to the method for enhancing large language model machine translation based on prompt optimization according to claim 1, it is characterized in that: The S1 specifically includes the following steps: S11. Input the text into the prompt encoder, pre-train the prompt encoder with the help of the prompt decoder, and obtain the hyper parameters of the prompt encoder and the pre-trained prompt encoder; S12. Select input data according to the hyperparameters of the prompt encoder, and fine-tune the pre-trained prompt encoder to obtain a fine-tuned prompt encoder; In the S11, an unsupervised monolingual corpus is selected as text based on the main language of the large language model, and the prompt encoder is pre-trained. According to the text sequence of length M output by the prompt encoder, the text sequence of length N at the input end of the prompt encoder is restored by the prompt decoder. According to the effect of restoring the text sequence by the prompt decoder during the pre-training process, the value of the hyperparameter M is determined, so that the prompt encoder compresses the text sequence of length N to a text sequence of length M; In the S12, the hyperparameter M of the prompt encoder in the pre-training stage is initialized, and fine-tuning training is performed by p-tuning. During the fine-tuning process, a fine-tuning dataset of instructions in a general field and different tasks is input into the large language model, and the language is aligned with the language used in the pre-training stage of the prompt encoder. The instruction text of the data in the fine-tuning dataset is used as a continuous prompt. The input data and output data of the large language model are the source end sentences of the parallel corpus and the target end sentences of the parallel corpus, respectively. The parameters of the large language model are fixed, and the adjustable parameters are the parameters of the prompt encoder and the continuous prompt.

3. According to the method for enhancing large language model machine translation based on prompt optimization according to claim 2, it is characterized in that: The S2 specifically includes the following steps: S21. Establish an SVD-LoRA method based on the LoRA method, and construct a large language model based on a prompt decoder combined with the SVD-LoRA method; S22. Performing end-to-end training based on prompt optimization on the large language model based on the prompt decoder combined with the SVD-LoRA method to obtain a trained large language model; In S21, the process of the LoRA method is expressed as: IN step=0 =kW0+ΔW=W0 Among them, W step=0 is the weight before the large language model starts training, ΔW is the weight added by the LoRA bypass, k is a constant, which is obtained according to the scaling ratio of ΔW and the current layer W0, where W0 is the current layer; When k = 1 and ΔW = 0, the process of the LoRA method is expressed as: IN step=0 =W0+BA=W0,ΔW=BA=0,k=1 Among them, A is the first low-rank matrix, and B is the second low-rank matrix; The process of the SVD-LoRA method is to use the SVD decomposition result SVD'(W0) of the current layer W0 to initialize the first low-rank matrix A and the second low-rank matrix B; The process of the SVD-LoRA method is expressed as: ΔW=SVD'(W0)=U'Σ'V' T Among them, U'Σ'V' T The SVD decomposition results after retaining the topK singular values ​​and the corresponding singular vectors for the singular values ​​respectively; In S22, the process of end-to-end training based on prompt optimization is as follows: aligning the text of the prompt encoder and the large language model through adjustable parameters, using parallel corpus in a general field to perform end-to-end training of the framework, and after the large language model converges, testing the translation performance of the large language model through continuous prompts obtained by prompt decoder training and a domain translation test set, so as to solve translation tasks in different specific fields and obtain a trained large language model.

4. According to the method of enhancing large language model machine translation based on prompt optimization according to claim 3, it is characterized in that: The step of constructing the optimized machine translation prompt comprises the following steps: S31. Limit the form of prompt instructions; S32. Selection of limited samples; S33 limits the coding range of the prompt encoder; In S31, the method adopted is to manually write the initialization prompt instruction, rely on the prompt encoder to convert the initialization text prompt into a continuous prompt, and dynamically iterate in the training process of each stage of the large language model; In S32, the retrieval enhanced generation method is used to vectorize the samples in the sample library using the BGE model, and the topK results in the sample library are selected as the prompted samples according to the text to be translated; In S33, the effective information amount brought by each component in the prompt to the translation task is sorted in ascending order, which is: instructions, translation samples in the form of bilingual sentence pairs, and external knowledge introduced by the bilingual dictionary. Based on the above order, the component with the lowest information amount is compressed first, and the specific encoding range is instructions and translation samples.

Citation Information

Patent Citations

  • Fine tuning of diffusion-based generative neural networks for text-to-image generation using singular value decomposition

    CN119053994A

  • LLM-based network threat flow detection rule automatic generation method and system

    CN119299130A

  • MOOC stop prediction method and system based on large language model auxiliary graph node classification

    CN119323293A

  • Systems and methods for tuning parameters of a machine learning model for federated learning

    US20240362487A1

Cited By

  • Privacy risk measurement method and device for federal large language model fine tuning

    CN121936550A

  • Privacy risk measurement method and device for federated large language model fine-tuning

    CN121936550B