Multi-modal machine translation training method combining knowledge graph, large language model and visual imagination mechanism
Through a multimodal machine translation training method combining knowledge graphs, large language models and visual imagination mechanisms, the problem of being unable to generate high-quality pictures in the existing technology is solved, and the image generation quality and machine translation performance are improved.
Patent Information
- Application Number
- CN202510035739.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Existing multimodal machine translation methods cannot generate high-quality corresponding pictures based on the sentences to be translated, thus unable to improve the translation quality.
A multimodal machine translation training method combining knowledge graphs, large language models and visual imagination mechanisms is adopted to improve image generation quality and machine translation performance by constructing training sets, extracting text and image triplets, calculating triplet distance properties, and optimizing image generation models and large language models.
It significantly improves the quality and efficiency of image generation, optimizes the machine translation performance of large language models, and improves the overall quality of multimodal machine translation.
Smart Images

Figure CN119962547A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a multimodal machine translation training method that combines a knowledge graph, a large language model, and a visual imagination mechanism. Background Art
[0002] Large language models are large-scale, pre-trained statistical language models based on neural networks. They are usually pre-trained on large-scale corpora and contain tens to hundreds of billions of parameters. Compared with traditional language models, they are larger in scale, have stronger language understanding and generation capabilities, and have emergent capabilities that small-scale language models do not have. For this reason, they have also become the foundation models of many natural language processing methods.
[0003] Knowledge graph is a knowledge representation method in natural language processing tasks. It can represent a large amount of knowledge as a graph structure. The knowledge graph is composed of triples. The structure of the triple is (h, r, t), where h, r, and t represent the head entity, relationship, and tail entity respectively. Compared with natural language, the triple structure has the characteristics of simplicity and high knowledge density. In addition, the distance between different triples can also be compared through the embedding technology of triples.
[0004] The visual imagination mechanism refers to a phenomenon in which, during the reasoning process of a large language model, pictures related to the question are generated through a picture generation model, which can improve the efficiency and accuracy of the large language model.
[0005] Multimodal machine translation technology refers to machine translation technology that combines non-text modalities such as pictures and voice.
[0006] Extracting triplets from natural language is a very mature technology that can be processed using commonly used natural language processing toolkits; recently, the technology of extracting triplets from images has also become a research focus.
[0007] However, existing multimodal machine translation methods cannot generate high-quality corresponding images based on the sentences to be translated, thereby improving the translation quality. Therefore, it is very meaningful to construct a model that can generate high-quality images based on the sentences to be translated using existing technologies. Summary of the invention
[0008] The problem to be solved by the present invention is to optimize the image generation quality of the multimodal machine translation method, and propose a multimodal machine translation training method that combines knowledge graphs, large language models and visual imagination mechanisms.
[0009] To achieve the above object, the present invention is implemented through the following technical solutions:
[0010] A multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism, comprising the following steps:
[0011] S1. Manually set sentences to be translated and construct training set 1;
[0012] S2. for the training set 1 obtained in step S1, the sentences to be translated in the training set 1 are processed by the text processing library to obtain text triplets, the sentences to be translated in the training set 1 are input into the image generation model to generate pictures, and then processed by the image triple extraction model to obtain image triplets, and then the distance property of the two groups of triplets is applied to train the image generation model to obtain a trained image generation model;
[0013] S3. Input the sentence to be translated in the training set 1 obtained in step S1 into the image generation model trained in step S2 to obtain a generated image corresponding to the sentence to be translated;
[0014] S4. Encode the generated image corresponding to the sentence to be translated obtained in step S3, and then concatenate it with the corresponding sentence to be translated to construct training set 2, use training set 2 to train the large language model, and optimize the multimodal machine translation performance of the large language model.
[0015] Furthermore, the specific implementation method of step S2 includes the following steps:
[0016] S2.1. Apply the text processing library to process the sentence x to be translated in the training set 1, iterate each data in the training set 1, extract the verbs and nouns in the language sentence to be translated according to the grammatical relationship, process the language sentence to be translated, and extract the text triple set LSG=(h l1 , r l1 , t l1 ), (h l2 , r l2 , t l2 ),…,(h ln , r ln , t ln ), where h ln is the nth head entity in the text triple set, r ln is the nth relation in the set of text triples, t ln is the nth tail entity in the set of text triples;
[0017] S2.2. Input the sentences to be translated in the training set 1 into the image generation model to generate images, wherein the image generation model is a Stable Diffusion model, which automatically generates the corresponding images I of the sentences to be translated;
[0018] S2.3. Input the image obtained in step S2.2 into the image triple extraction model to automatically obtain the image triple set VSG = (h v1 , r v1 , t v1 ), (h v2 , r v2 , t v2 ),…,(h vm , r vm , t vm ), where h vm is the mth head entity in the image triple set, r vm is the mth relation in the image triple set, t vm is the mth tail entity in the image triple set;
[0019] S2.4. Calculate the distance between the members of the text triplet set obtained in step S2.1 and the image triplet set obtained in step S2.4 as the image generation score Sim(x, I) of the data. The specific formula is:
[0020]
[0021] Among them, Score is the score function of the similarity between the members of the text triple set and the image triple set;
[0022] The specific calculation method of Score is:
[0023] Score(LSG i , VSG)=max(D(LSG i , VSG 1 ),…,D(LSG i , VSG m ))
[0024] Where D is the triple similarity function,
[0025] The specific calculation method of D is:
[0026]
[0027] Among them, SIM is a function for calculating the distance between natural language words;
[0028] S2.5. The image generation score obtained in step S2.4 is used as the performance score of the model, and the DDPO training method is applied to optimize the image generation model. The optimization formula is:
[0029]
[0030] Among them, θ is the parameter of the Stable Diffusion model, JDDPO (θ) is the optimization target of the Stable Diffusion model parameters, r(x, I) is a manually set reward function used to quantify the matching degree between the sentence to be translated x and the corresponding image I. It means that when the sentence to be translated x satisfies the given parameter θ and the given text condition c, the probability distribution p θ (x|c), and the image I satisfies the image’s own probability distribution p(I), the expected likelihood estimate of the reward function.
[0031] Furthermore, the text processing library used in step S2.1 is the SpaCY natural language text processing library or the NLTK natural language toolkit.
[0032] Furthermore, the image generation model is a Stable Diffusion model.
[0033] Furthermore, the image triplet extraction model in step S2.3 is a Scene-Graph-Benchmark model or a Relationformer model.
[0034] Furthermore, the calculation method of SIM in step S2.4 is to first apply the large language model for encoding, and then apply the cosine similarity for calculation. The specific formula is:
[0035]
[0036] Here, Encode represents the process of automatically encoding words using a large language model, and A and B are two natural language words to be calculated.
[0037] Furthermore, the specific implementation method of step S4 includes the following steps:
[0038] S4.1. Encode the generated image I′ corresponding to the sentence to be translated obtained in step S3 to obtain the image encoding C(I′). The expression of the encoder processing is:
[0039] C(I′)=W′I′+b′
[0040] Among them, W′ and b′ are the first weight matrix and the second weight matrix respectively, and the encoder selects the CLIP model;
[0041] S4.2. Then, the image code C(I′) is processed by the following formula to obtain the final image code C′(I′):
[0042] C′(I′)=WC(I′)+b
[0043] Wherein, W and b are the third weight matrix and the fourth weight matrix respectively;
[0044] S4.3. The final image code is concatenated to the front of the sentence to be translated to construct training set 2, which is used as the input of the large language model for training to obtain the final machine translation result;
[0045] S4.4. Calculate the loss of the final machine translation result w′ and the sentence w after manual translation to get the translation loss Loss LM The expression is:
[0046]
[0047] Among them, C is the large language model parameter;
[0048] Then calculate the overall loss Loss of the large language model, expressed as:
[0049] Loss=Loss LM +Sim(x,I).
[0050] Furthermore, in step S4.3, the large language model is a Llama2 model or a Llama3 model or a Vicuna model further trained on the basis of the Llama2 model, and the large language model used is modeled as follows:
[0051]
[0052] Among them, t represents the current time step, j t represents the tth word in the paragraph, which is generated based on the large model according to the time step, generating one word in one time step, then p(w′) represents the probability of generating w′, and p(j t |j <t ) represents the probability of generating the tth word after generating the first t-1 words, and T is the total time.
[0053] Beneficial effects of the present invention:
[0054] The multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism described in the present invention focuses on utilizing the triple structure of the knowledge graph in the training of the image generation module. Compared with the existing model that uses the imagination mechanism to perform large language model reasoning, during the training of the image generation module, the present invention can utilize the triple structure to significantly improve the quality and efficiency of image generation; during the performance optimization training of the large language model, through the combination of loss functions, the large language model can adapt to the changes of the image generation module, thereby optimizing the machine translation performance of the large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1A flowchart of a multimodal machine translation training method combining a knowledge graph, a large language model and a visual imagination mechanism according to the present invention;
[0056] Figure 2 It is a training flow chart of the image generation module of the present invention;
[0057] Figure 3 The present invention is a process for optimizing the training of a large language model. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solution and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0059] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected specific embodiments of the present invention. Based on the specific embodiments of the present invention, all other specific embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0060] In order to further understand the content, features and effects of the present invention, the following specific implementation methods are given as examples, and the attached Figure 1 -Attached Figure 3 The detailed instructions are as follows:
[0061] Embodiment 1:
[0062] A multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism, comprising the following steps:
[0063] S1. Manually set sentences to be translated and construct training set 1;
[0064] S2. for the training set 1 obtained in step S1, the sentences to be translated in the training set 1 are processed by the text processing library to obtain text triplets, the sentences to be translated in the training set 1 are input into the image generation model to generate pictures, and then processed by the image triple extraction model to obtain image triplets, and then the distance property of the two groups of triplets is applied to train the image generation model to obtain a trained image generation model;
[0065] Furthermore, the specific implementation method of step S2 includes the following steps:
[0066] S2.1. Apply the text processing library to process the sentence x to be translated in the training set 1, iterate each data in the training set 1, extract the verbs and nouns in the language sentence to be translated according to the grammatical relationship, process the language sentence to be translated, and extract the text triple set LSG=(h l1 , r l1 , t l1 ), (h l2 , r l2 , t l2 ),…,(h ln , r ln , t ln ), where h ln is the nth head entity in the text triple set, r ln is the nth relation in the set of text triples, t ln is the nth tail entity in the set of text triples;
[0067] Furthermore, the text processing library used in step S2.1 is the SpaCY natural language text processing library or the NLTK natural language toolkit;
[0068] S2.2. Input the sentences to be translated in the training set 1 into the image generation model to generate images, wherein the image generation model is a Stable Diffusion model, which automatically generates the corresponding images I of the sentences to be translated;
[0069] Furthermore, the image generation model is a Stable Diffusion model;
[0070] S2.3. Input the image I obtained in step S2.2 into the image triple extraction model, and automatically obtain the image triple set VSG = (h v1 , r v1 , t v1 ), (h v2 , r v2 , t v2 ),…,(h vm , r vm , t vm ), where h vm is the mth head entity in the image triple set, r vm is the mth relation in the image triple set, t vm is the mth tail entity in the image triple set;
[0071] Furthermore, the image triplet extraction model in step S2.3 is a Scene-Graph-Benchmark model or a Relationformer model;
[0072] S2.4. Calculate the distance between the members of the text triplet set obtained in step S2.1 and the image triplet set obtained in step S2.4 as the image generation score Sim(x, I) of the data. The specific formula is:
[0073]
[0074] Among them, Score is the score function of the similarity between the members of the text triple set and the image triple set;
[0075] The specific calculation method of Score is:
[0076] Score(LSG i , VSG)=max(D(LSG i , VSG 1 ),…,D(LSG i , VSG m ))
[0077] Where D is the triple similarity function,
[0078] The specific calculation method of D is:
[0079]
[0080] Among them, SIM is a function for calculating the distance between natural language words;
[0081] Furthermore, the calculation method of SIM in step S2.4 is to first apply the large language model for encoding, and then apply the cosine similarity for calculation. The specific formula is:
[0082]
[0083] Among them, Encode represents the process of automatically encoding words using a large language model, and A and B are two natural language words to be calculated;
[0084] S2.5. The image generation score obtained in step S2.4 is used as the performance score of the model, and the DDPO training method is applied to optimize the image generation model. The optimization formula is:
[0085]
[0086] Among them, θ is the parameter of the Stable Diffusion model, J DDPO (θ) is the optimization target of the Stable Diffusion model parameters, r(x, I) is a manually set reward function used to quantify the matching degree between the sentence to be translated x and the corresponding image I. It means that when the sentence to be translated x satisfies the given parameter θ and the given text condition c, the probability distribution p θ (x|c), and the image I satisfies the image’s own probability distribution p(I), the expected likelihood estimate of the reward function.
[0087] Furthermore, the meaning of the above formula is to optimize the likelihood estimation of the reward function under the condition of given text and generated image. In this formula, the given text and generated image are the above-mentioned sentence to be translated and the image consistent with the description of the sentence to be translated.
[0088] S3. Input the sentence to be translated in the training set 1 obtained in step S1 into the image generation model trained in step S2 to obtain a generated image corresponding to the sentence to be translated;
[0089] S4. Encode the generated image corresponding to the sentence to be translated obtained in step S3, and then concatenate it with the corresponding sentence to be translated to construct training set 2, use training set 2 to train the large language model, and optimize the multimodal machine translation performance of the large language model.
[0090] Furthermore, the specific implementation method of step S4 includes the following steps:
[0091] S4.1. Encode the generated image I′ corresponding to the sentence to be translated obtained in step S3 to obtain the image encoding C(I′). The expression of the encoder processing is:
[0092] C(I′)=W′I′+b′
[0093] Among them, W′ and b′ are the first weight matrix and the second weight matrix respectively, and the encoder selects the CLIP model;
[0094] S4.2. Then, the image code C(I′) is processed by the following formula to obtain the final image code C′(I′):
[0095] C′(I′)=WC(I′)+b
[0096] Wherein, W and b are the third weight matrix and the fourth weight matrix respectively;
[0097] S4.3. The final image code is concatenated to the front of the sentence to be translated to construct training set 2, which is used as the input of the large language model for training to obtain the final machine translation result;
[0098] Furthermore, in step S4.3, the large language model is a Llama2 model or a Llama3 model or a Vicuna model further trained on the basis of the Llama2 model, and the large language model used is modeled as follows:
[0099]
[0100] Where t represents the current time step, jt represents the tth word in the paragraph, and the large model generates words according to the time step, generating one word in one time step. Then p(w′) represents the probability of generating w′, and p(jt) represents the probability of generating word w′. t |j <t ) represents the probability of generating the tth word after generating the first t-1 words, and T is the total time.
[0101] S4.4. Calculate the loss of the final machine translation result w′ and the sentence w after manual translation to get the translation loss Loss LM The expression is:
[0102]
[0103] Among them, C is the large language model parameter;
[0104] Then calculate the overall loss Loss of the large language model, expressed as:
[0105] Loss=Loss LM +Sim(x,I).
[0106] Embodiment 2:
[0107] The actual operation using the method of Example 1 is as follows: the large model used is the Vicuna model, the image encoding module uses the CLIP model, and the image triple extraction method used is Scene-Graph-Benchmark.
[0108] First, improve the training interpretation of the image generation module. Assume that there is only one sentence (three women standing under a white wall) in the given data set. First, SpaCY is used to extract the triples in this sentence, and two text triples (three women standing under a wall) and (the wall is white) are obtained, which are then saved in the hardware for standby. Then, the sentence "three women standing under a white wall" is input into the Stable Diffusion model to generate a picture corresponding to the description. Then, the picture obtained by processing the image triple extraction module Scene-Graph-Benchmark is used to obtain about 100 triples, such as: (woman standing under a wall), (woman running)... Then, the previously extracted text triple is read, and the above-mentioned formula is applied to calculate its score with the picture triple. The score obtained in the embodiment is 0.235. Then, the performance of the Stable Diffusion model is optimized by the DDPO training method.
[0109] Secondly, the explanation of the process of Vicuna model optimization training is improved. Assume that there is only one sentence (There is an apple on the table.) in the given data set. First, SpaCY is used to extract the triples in the sentence "There is an apple on the table" to obtain the text triple (an apple, on the table), which is then saved in the hardware for standby. Subsequently, the sentence "There is an apple on the table" is input into the Stable Diffusion model to generate a picture corresponding to the description. Subsequently, the picture obtained by processing the image triple extraction module Scene-Graph-Benchmark is used to obtain about 100 triples, such as: (table, is, brown), (apple, is, red)... Subsequently, the previously extracted text triples are read, and the above-mentioned formula is applied to calculate the score of the text triples and the picture triples. The score obtained in the embodiment is 0.107. Subsequently, the CLIP model is used to encode the generated picture, and then the code is spliced to the front of the sentence to be translated, and the Vicuna model is input for reasoning to obtain the translation "An apple is on the table.". Relying on Loss LM The formula is used to calculate the difference between it and “There is an apple on the table.” In this case, it is 0.605. Then, the 0.107 calculated above is added to 0.605 as the overall loss to optimize the performance of the Vicuna model.
[0110] Embodiment 3:
[0111] The actual operation using the method of Example 1 is as follows: the large model used is the Llama2 model, the image encoding module uses the CLIP model, and the image triple extraction method used is Relationformer.
[0112] First, improve the training interpretation of the image generation module. Assume that there is only one sentence (three women standing under a white wall) in the given data set. First, rely on the NLTK library to extract the triples in this sentence, and obtain three text triples (women, standing, under the wall), (wall, is, white), (women, are, three), and then save them in the hardware for standby. Then, input the sentence "three women standing under the white wall" into the Stable Diffusion model to generate a picture corresponding to the description. Then, rely on the image triple extraction module Relationformer to process the obtained picture, and obtain about 100 triples, such as: (women, standing, under the wall), (women, are, running)... Then, read the previously extracted text triples, and apply the above-mentioned formula to calculate its score with the picture triples. The score obtained in the embodiment is 0.105. Then, rely on the DDPO training method to optimize the performance of the Stable Diffusion model.
[0113] Secondly, the explanation of the process of optimizing the training of the Llama2 model is improved. Assume that there is only one sentence (There is an apple on the table.) in the given data set. First, rely on NLTK to extract the triples in the sentence "There is an apple on the table", and obtain the two text triples (apple, on, table) and (apple, is, one), which are then saved in the hardware for standby. Subsequently, the sentence "There is an apple on the table" is input into the StableDiffusion model to generate a picture corresponding to the description. Subsequently, the picture obtained by processing the image triple extraction module Scene-Graph-Benchmark is relied on to obtain about 100 triples, such as: (table, is, brown), (apple, is, red)... Subsequently, the text triple extracted before is read, and the above-mentioned formula is applied to calculate the score of the triple with the picture. The score obtained in the embodiment is 0.086. Subsequently, the CLIP model is relied on to encode the generated picture, and then the code is spliced to the front of the sentence to be translated, and the Vicuna model is input for reasoning, and the translation is obtained as "An apple is on the table.". Rely on Loss LM The formula is used to calculate the difference between it and “There is an apple on the table.” In this case, it is 0.605. Then, the 0.086 calculated above is added to 0.605 as the overall loss to optimize the performance of the Vicuna model.
[0114] Some necessary terms of the present invention are as follows:
[0115] 1.Stable Diffusion model: An artificial intelligence image generation model that can generate images with corresponding descriptions based on given text.
[0116] 2. CLIP model: An artificial intelligence coding model that can encode images and convert them into vector representations.
[0117] 3. SpaCY: An open source natural language processing library.
[0118] 4.NLTK: An open source natural language processing library.
[0119] 5. Scene-Graph-Benchmark method: A method for extracting triplets from images.
[0120] 6. Relationformer method: A method for extracting triplets from images.
[0121] 7. Vicuna model: The Vicuna model is a large language model that has been fine-tuned for multi-round conversations and has good content generation capabilities.
[0122] 8.DDPO training method: A reinforcement learning method that can optimize the performance of the diffusion model by manually setting the reward function.
[0123] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0124] Although the present application has been described above with reference to specific embodiments, various modifications may be made thereto and parts thereof may be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application may be used in combination with each other in any manner, and the fact that these combinations are not exhaustively described in this specification is only for the sake of omitting space and saving resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism, characterized in that: The steps include: S1. Manually set sentences to be translated and construct training set 1; S2. for the training set 1 obtained in step S1, the sentences to be translated in the training set 1 are processed by the text processing library to obtain text triplets, the sentences to be translated in the training set 1 are input into the image generation model to generate pictures, and then processed by the image triple extraction model to obtain image triplets, and then the distance property of the two groups of triplets is applied to train the image generation model to obtain a trained image generation model; S3. Input the sentence to be translated in the training set 1 obtained in step S1 into the image generation model trained in step S2 to obtain a generated image corresponding to the sentence to be translated; S4. Encode the generated image corresponding to the sentence to be translated obtained in step S3, and then concatenate it with the corresponding sentence to be translated to construct training set 2, use training set 2 to train the large language model, and optimize the multimodal machine translation performance of the large language model.
2. According to claim 1, a multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism is characterized in that: The specific implementation method of step S2 includes the following steps: S2.
1. Apply the text processing library to process the sentence x to be translated in the training set 1, iterate each data in the training set 1, extract the verbs and nouns in the language sentence to be translated according to the grammatical relationship, process the language sentence to be translated, and extract the text triple set LSG=(h l1 ,r l1 ,t l1 ),(h l2 ,r l2 ,t l2 ),…,(h ln ,r ln ,t ln ), where h ln is the nth head entity in the text triple set, r ln is the nth relation in the set of text triples, t ln is the nth tail entity in the set of text triples; S2.
2. Input the sentences to be translated in the training set 1 into the image generation model to generate images, wherein the image generation model is a Stable Diffusion model, which automatically generates the corresponding images I of the sentences to be translated; S2.
3. Input the image I obtained in step S2.2 into the image triple extraction model, and automatically obtain the image triple set VSG = (h v1 ,r v1 ,t v1 ),(h v2 ,r v2 ,t v2 ),…,(h vm ,r vm ,t vm ), where h vm is the mth head entity in the image triple set, r vm is the mth relation in the image triple set, t vm is the mth tail entity in the image triple set; S2.
4. Calculate the distance between the members of the text triplet set obtained in step S2.1 and the image triplet set obtained in step S2.4 as the image generation score Sim(x,I) of the data. The specific formula is: Among them, Score is the score function of the similarity between the members of the text triple set and the image triple set; The specific calculation method of Score is: Score(LSGi,VSG)=max(D(LSG i ,VSG1),…,D(LSG i ,VSG m )) Where D is the triple similarity function, The specific calculation method of D is: Among them, SIM is a function for calculating the distance between natural language words; S2.
5. The image generation score obtained in step S2.4 is used as the performance score of the model, and the DDPO training method is applied to optimize the image generation model. The optimization formula is: Among them, θ is the parameter of the Stable Diffusion model, J DDPO (θ) is the optimization target of the Stable Diffusion model parameters, r(x,I) is a manually set reward function used to quantify the matching degree between the sentence to be translated x and the corresponding image I. It refers to the probability distribution p when the sentence to be translated x satisfies the given parameter θ and the given text condition c. θ (x|c), and the image I satisfies the image’s own probability distribution p(I), the expected likelihood estimate of the reward function.
3. According to claim 2, a multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism is characterized in that: The text processing library used in step S2.1 is the SpaCY natural language text processing library or the NLTK natural language toolkit.
4. According to claim 2, a multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism is characterized in that: The image generation model is a Stable Diffusion model.
5. According to claim 4, a multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism is characterized in that: The image triplet extraction model in step S2.3 is the Scene-Graph-Benchmark model or the Relationformer model.
6. According to claim 5, a multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism is characterized in that: The calculation method of SIM in step S2.4 is to first apply the large language model for encoding, and then apply the cosine similarity for calculation. The specific formula is: Here, Encode represents the process of automatically encoding words using a large language model, and A and B are two natural language words to be calculated.
7. A multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism according to claim 6, characterized in that: The specific implementation method of step S4 includes the following steps: S4.
1. Encode the generated image I' corresponding to the sentence to be translated obtained in step S3 to obtain the image code C(I'). The expression of the encoder processing is: C(I')=W'I'+b' Among them, W' and b' are the first weight matrix and the second weight matrix respectively, and the encoder selects the CLIP model; S4.
2. Then apply the following formula to process the obtained image code C(I') to obtain the final image code C'(I'): C′(I′)=WC(I′)+b Wherein, W and b are the third weight matrix and the fourth weight matrix respectively; S4.
3. The final image code is concatenated to the front of the sentence to be translated to construct training set 2, which is used as the input of the large language model for training to obtain the final machine translation result; S4.
4. Calculate the loss of the final machine translation result w' and the sentence w after manual translation to get the translation loss Loss LM The expression is: Among them, C is the large language model parameter; Then calculate the overall loss Loss of the large language model, expressed as: Loss=Loss LM +Sim(x,I)。 8. The multimodal machine translation training method combining knowledge graph, large language model and visual imagination mechanism according to claim 7, characterized in that: In step S4.3, the large language model is a Llama2 model or a Llama3 model or a Vicuna model further trained on the basis of the Llama2 model. The large language model used is modeled as follows: Among them, t represents the current time step, j t represents the tth word in the paragraph, which is generated based on the large model according to the time step, generating one word in one time step, then p(w') represents the probability of generating w', and p(j t |j <t ) represents the probability of generating the tth word after generating the first t-1 words, and T is the total time.
Citation Information
Patent Citations
Object information translation method and device and derivative information collection method and device
CN113407743A
Training method, translation method and device of neural machine translation model
CN115345181A