A text translation method and related apparatus
Patent Information
- Application Number
- CN202510379437.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]在大模型篇章翻译的技术中,为了在翻译中充分考虑相邻语句之间的关联性,把待翻译的语句前后的语句抽取出来作为该待翻译的语句的跨句上下文,再将跨句上下文和待翻译的语句进行拼接,对拼接后的跨句上下文和待翻译的语句进行翻译,但是待翻译的语句的翻译结果仍不准确
[0078]第四方面,本申请提供了一种计算机可读存储介质,计算机可读存储介质中保存有指令,当指令在处理器上运行时,实现前述第一方面或者第一方面的任一种可能的实现方式所示的方法。
Smart Images

Figure CN122840071A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a text translation method and related apparatus. Background Technology
[0002] With the continuous development of artificial intelligence technology, its application scenarios are constantly expanding. For example, large language models are widely used in text translation. Long text translation, as a relatively special scenario in text translation, requires a full understanding of the semantic relationships between sentences and context in order to translate text that better conforms to the semantics of the source text.
[0003] In large-scale text translation techniques, in order to fully consider the correlation between adjacent sentences during translation, the sentences before and after the sentence to be translated are extracted as the cross-sentence context of the sentence to be translated. The cross-sentence context and the sentence to be translated are then concatenated, and the concatenated cross-sentence context and the sentence to be translated are translated. However, the translation result of the sentence to be translated is still inaccurate. Summary of the Invention
[0004] This application provides a text translation method and related apparatus for improving the accuracy of text translation results.
[0005] Firstly, this application provides a text translation method, including:
[0006] Get the target text;
[0007] The context statement is encoded based on the first vector to obtain the context feature matrix. The context statement includes at least a sentences from the a sentences before or after the target statement in the target text, where a is a positive integer, and the target statement is any sentence in the target text.
[0008] The target statement is encoded based on the context feature matrix and the second vector to obtain the target feature matrix;
[0009] The target feature matrix and the third vector are concatenated to obtain the feature matrix to be inferred;
[0010] The feature matrix to be reasoned is reasoned to obtain a reference feature vector set, where each feature vector in the reference feature vector set corresponds to a word in the target sentence;
[0011] The translation of the target sentence is obtained by decoding each feature vector in the reference feature vector group.
[0012] In this embodiment, when translating the target text, at least a sentences from the preceding or following a sentences of the target statement are used as context statements. These context statements are then encoded based on a first vector to obtain a context feature matrix. The target statement is then encoded based on the context feature matrix and a second vector to obtain a target feature matrix. This target feature matrix is then concatenated with a third vector to obtain a feature matrix to be inferred. This feature matrix is then used to infer a reference feature vector, which is then used for decoding to obtain the translation of the target statement. By adding the first, second, and third vectors, the target statement and context statements learn to different degrees during the translation process, effectively improving translation accuracy.
[0013] In one possible implementation of the first aspect, reasoning on the feature matrix to be reasoned to obtain a target feature vector set includes:
[0014] Feature extraction is performed on the feature matrix to be inferred to obtain the first feature vector, which is the feature vector corresponding to the first word, and the first word is a word in the target sentence;
[0015] Reasoning is performed on the first eigenvector to obtain the first weight reorganization, in which each weight corresponds to an eigenvector.
[0016] The second feature vector is selected based on the first weighting reorganization. The second feature vector is contained in the reference feature vector group. The second feature vector is the feature vector corresponding to the weight with the largest value in the first weighting reorganization.
[0017] In this embodiment of the application, during the reasoning process of the target sentence, based on the sequence of each word in the target sentence and the feature vector corresponding to that word, and after fully learning based on the relationship between the text context, an accurate translation that conforms to the text context is obtained.
[0018] In one possible implementation of the first aspect, the method further includes:
[0019] Pool the context feature matrix to obtain the context feature vector;
[0020] Pool the target feature matrix to obtain the target feature vector;
[0021] Reasoning is performed based on the context feature vector, the target feature vector, and the first feature vector to obtain the second weighted reorganization;
[0022] Reasoning is performed based on the target feature vector and the first feature vector to obtain the third weighted reorganization;
[0023] The selection of the second feature vector based on the first weighted recombination includes:
[0024] The second feature vector is obtained by voting based on the first weighted reorganization, the second weighted reorganization, and the third weighted reorganization.
[0025] In this embodiment, the use of contextual statements and target statements is enhanced during the reasoning process. The second feature vector (reasoning result) is obtained by voting on the reasoning results of multiple cases, which achieves better cross-sentence context modeling effect and further enhances the accuracy of LLM chapter translation.
[0026] In one possible implementation of the first aspect, the method further includes:
[0027] The target text is preprocessed to obtain the context statement and the target statement.
[0028] In this embodiment, the target text is preprocessed to obtain context statements and target statements, so that preprocessing-related operations and reasoning-related operations are performed separately, reducing the workload of the reasoning model and improving the reasoning efficiency of the reasoning model.
[0029] In one possible implementation of the first aspect, encoding the context statement based on the first vector to obtain the context feature matrix includes:
[0030] Feature extraction is performed on the context statements to obtain the feature matrix corresponding to the context statements;
[0031] The feature matrix corresponding to the context statement is concatenated with the first vector to obtain the context feature matrix;
[0032] The target statement is encoded based on the context feature matrix and the second vector to obtain the target feature matrix, which includes:
[0033] Feature extraction is performed on the target statement to obtain the feature matrix corresponding to the target statement;
[0034] The target feature matrix is obtained by concatenating the feature matrix corresponding to the target statement, the context feature matrix, and the second vector.
[0035] In one possible implementation of the first aspect, a is less than or equal to 3.
[0036] In this embodiment of the application, by limiting the number of 'a's, the length of the cross-sentence context can be effectively limited. Since the context that is further away from the target sentence has a lower correlation with the target sentence, introducing an appropriate amount of cross-sentence context can not only avoid the increase in computation caused by introducing too much redundant context, but also effectively improve the inference accuracy and improve the inference efficiency of this solution.
[0037] In one possible implementation of the first aspect, the method further includes:
[0038] Obtain the training text and its corresponding translation;
[0039] Encode the cross-sentence context of the preset sentence and the first vector to be updated to obtain the first feature matrix. The preset sentence is any sentence in the training text. The cross-sentence context of the preset sentence includes at least a sentences in the training text that are either a sentences before or a sentences after the preset sentence.
[0040] The preset sentence is encoded based on the first feature matrix and the second vector to be updated to obtain the second feature matrix;
[0041] Based on the third vector to be updated and the second feature matrix, reason about the preset sentence to obtain the preset feature vector group;
[0042] The first vector to be updated, the second vector to be updated, and the third vector to be updated are updated based on the target feature vector set and the preset feature vector set to obtain the first vector, the second vector, and the third vector. The target feature vector set is obtained by encoding each word in the translation of the preset sentence in the translation of the training text.
[0043] In this embodiment, the training of the LLM text translation model is divided into three different stages, and different adjustable vectors are added to each of the three stages as cues. During the training stage, the three adjustable vectors are updated using labels (translations of the training text) to improve the performance of the LLM text translation model in the inference process.
[0044] In one possible implementation of the first aspect, the method further includes:
[0045] Preprocess the training text and its translation to obtain the preset sentence, the cross-sentence context of the preset sentence, and the translation of the preset sentence.
[0046] Secondly, this application provides a text translation device, comprising:
[0047] The acquisition unit is used to acquire the target text;
[0048] The encoding unit is used to encode the context statement based on the first vector to obtain the context feature matrix. The context statement includes at least a sentences from the a sentences before or after the target statement in the target text, where a is a positive integer, and the target statement is any sentence in the target text.
[0049] The encoding unit is also used to encode the target statement based on the context feature matrix and the second vector to obtain the target feature matrix;
[0050] The encoding unit is also used to concatenate the target feature matrix and the third vector to obtain the feature matrix to be inferred;
[0051] The reasoning unit is also used to reason about the feature matrix to be reasoned, and obtain a reference feature vector group, in which each feature vector corresponds to a word in the target sentence;
[0052] The decoding unit is used to decode each feature vector in the reference feature vector group to obtain the translation of the target sentence.
[0053] In one possible implementation of the second aspect, the inference unit is specifically used for:
[0054] Feature extraction is performed on the feature matrix to be inferred to obtain the first feature vector, which is the feature vector corresponding to the first word, and the first word is a word in the target sentence;
[0055] Reasoning is performed on the first eigenvector to obtain the first weight reorganization, in which each weight corresponds to an eigenvector.
[0056] The second feature vector is selected based on the first weighting reorganization. The second feature vector is contained in the reference feature vector group. The second feature vector is the feature vector corresponding to the weight with the largest value in the first weighting reorganization.
[0057] In a second aspect, in one possible implementation, the apparatus further includes a pooling unit for:
[0058] Pool the context feature matrix to obtain the context feature vector;
[0059] Pool the target feature matrix to obtain the target feature vector;
[0060] The reasoning unit is also used to perform reasoning based on the context feature vector, the target feature vector, and the first feature vector to obtain the second weighted reassembly;
[0061] The reasoning unit is also used to perform reasoning based on the target feature vector and the first feature vector to obtain the third weighted reorganization;
[0062] The reasoning unit is specifically used to vote based on the first weighted reorganization, the second weighted reorganization, and the third weighted reorganization to obtain the second feature vector.
[0063] In one possible implementation of the second aspect, the apparatus further includes a preprocessing unit for preprocessing the target text to obtain context statements and target statements.
[0064] In one possible implementation of the second aspect, the encoding unit is specifically used for:
[0065] Feature extraction is performed on the context statements to obtain the feature matrix corresponding to the context statements;
[0066] The feature matrix corresponding to the context statement is concatenated with the first vector to obtain the context feature matrix;
[0067] Encoding unit, specifically used for:
[0068] Feature extraction is performed on the target statement to obtain the feature matrix corresponding to the target statement;
[0069] The target feature matrix is obtained by concatenating the feature matrix corresponding to the target statement, the context feature matrix, and the second vector.
[0070] In one possible implementation of the second aspect, a is less than or equal to 3.
[0071] In one possible implementation of the second aspect, the acquisition unit is further configured to acquire the training text and the translation corresponding to the training text;
[0072] The encoding unit is also used to encode the cross-sentence context of the preset sentence and the first vector to be updated to obtain the first feature matrix. The preset sentence is any sentence in the training text, and the cross-sentence context of the preset sentence includes at least a sentences in the training text that are either a sentences before or a sentences after the preset sentence.
[0073] The encoding unit is also used to encode the preset sentence based on the first feature matrix and the second vector to be updated to obtain the second feature matrix;
[0074] The reasoning unit is also used to reason about the preset sentence based on the third vector to be updated and the second feature matrix to obtain a preset feature vector set;
[0075] The device also includes an update unit for updating the first vector to be updated, the second vector to be updated, and the third vector to be updated based on the target feature vector set and the preset feature vector set, to obtain the first vector, the second vector, and the third vector. The target feature vector set is obtained by encoding each word in the translation of the preset sentence in the translation of the training text.
[0076] In one possible implementation of the second aspect, the preprocessing unit is further configured to preprocess the training text and its translation to obtain a preset sentence, the cross-sentence context of the preset sentence, and the translation of the preset sentence.
[0077] Thirdly, this application provides a text translation apparatus, including a processor and a memory, wherein the processor stores instructions, and when the instructions stored in the memory are executed on the processor, the method shown in the first aspect or any possible implementation of the first aspect is implemented.
[0078] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a processor, implement the method shown in the first aspect or any possible implementation of the first aspect.
[0079] Fifthly, this application provides a computer program product that, when executed on a processor, implements the method shown in the first aspect or any possible implementation of the first aspect.
[0080] The beneficial effects shown in any of the second to fifth aspects are similar to those of the first aspect or any possible implementation of the first aspect, and will not be repeated here. Attached Figure Description
[0081] Figure 1 A schematic diagram of the architecture of the text translation method provided in this application;
[0082] Figure 2 A flowchart illustrating the text translation method provided in this application;
[0083] Figure 3a Another flowchart illustrating the text translation method provided in this application;
[0084] Figure 3b Another flowchart illustrating the text translation method provided in this application;
[0085] Figure 4 A data illustration of the comparative experiment provided in this application;
[0086] Figure 5 Another data illustration of the comparative experiment provided in this application;
[0087] Figure 6 Another flowchart illustrating the text translation method provided in this application;
[0088] Figure 7 Another flowchart illustrating the text translation method provided in this application;
[0089] Figure 8 Another flowchart illustrating the text translation method provided in this application;
[0090] Figure 9 Another flowchart illustrating the text translation method provided in this application;
[0091] Figure 10 A schematic diagram of the text translation device provided in this application;
[0092] Figure 11 Another structural schematic diagram of the text translation device provided in this application. Detailed Implementation
[0093] This application provides a text translation method and related apparatus for improving the coherence and accuracy of text translation results.
[0094] To facilitate understanding of the solution provided in this application, the technical terms used in this application will be introduced first:
[0095] (1) Large language model (LLM) is a type of machine learning (ML) model that attempts to accomplish text generation tasks. LLM enables computers to process, interpret, and generate human language, thereby improving the efficiency of human-computer interaction. Given input text, the learning process of LLM enables it to predict the optimal possible subsequent words, thereby generating a meaningful response to the input text.
[0096] (2) Prompt tuning (PT) is a fine-tuning technique for LLMs, designed to improve the model's performance on specific artifacts by adjusting input prompts. Phased prompt tuning (PPT) differs from traditional tuning methods in that it focuses primarily on optimizing the prompts rather than making comprehensive adjustments to the model parameters.
[0097] (3) Document-level machine translation (DMT) refers to a translation process that focuses not only on the translation of individual words or phrases, but also on the context, structure, and meaning of the entire text. This translation approach emphasizes the integrity and coherence between sentences, aiming to ensure that the translated content remains consistent with the source text in terms of logic, style, and pragmatics.
[0098] (4) The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion so that a process, method, system, product, or device that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or devices. Additionally, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can be expressed as: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0099] With the rapid development of natural language processing (NLP) and machine learning, especially the emergence of large-scale pre-trained language models, strong support has been provided for high-quality text translation. Models such as BERT, GPT, and T5 learn rich linguistic knowledge and contextual representations through pre-training on large-scale text data. Pre-training enables models to understand the structure, syntax, and semantics of language, allowing for fine-tuning of pre-trained models for specific translation tasks to adapt them to the requirements of specific language pairs or domains. This approach improves the accuracy and naturalness of translation.
[0100] LLM is based on the transformer architecture, which introduces a self-attention mechanism, enabling the model to effectively capture long-distance dependencies. In passage translation scenarios, it can more effectively understand the logical relationships and thematic development within the context of a passage. In large-scale passage translation techniques, the sentences before and after the sentence to be translated are used as the cross-sentence context, and then the cross-sentence context and the sentence to be translated are concatenated. The translation is then performed on the concatenated cross-sentence context and the sentence to be translated, but the translation result is still inaccurate.
[0101] To address the above issues, this application proposes a method for translating target text. This involves taking any sentence in the target text as the target statement and at least a sentences preceding or following the target statement as context statements. First, the context statements are encoded using a first vector to obtain a context feature matrix. Then, the target statement is encoded using a second vector and the context feature matrix to obtain a target feature matrix. The target feature matrix and a third vector are then concatenated to obtain a feature matrix to be inferred. Inference is then performed on this feature matrix to obtain a reference feature vector group. Finally, each feature vector in this reference feature vector group is decoded to obtain the translation of the target statement. By adding the first, second, and third vectors as adjustable cues, the context statements and the target statement have different weights in the translation of the target statement, effectively improving translation accuracy.
[0102] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0103] First, combine Figure 1 The architecture of the text translation method provided in this application is introduced.
[0104] The text translation method provided in this application includes a framework and a decoding enhancement module. The framework divides the text translation method into three stages: cross-sentence context encoding, encoding of the sentence to be translated, and decoding. A cue word, which is a vector, is introduced in each stage.
[0105] The decoding enhancement module extends the decoding framework of the text translation method. Building upon conventional reasoning, it adds additional reasoning to enhance cross-sentence context and the sentence to be translated, combining the results of these three inferences to determine the final word generation. This enhanced reasoning method further improves the efficiency of utilizing source information and enhances the ability to solve textual problems.
[0106] The following is combined Figure 2 One specific implementation of the text translation method provided in this application is described.
[0107] S210, Obtain the target text;
[0108] S220. Encode the context statement based on the first vector to obtain the context feature matrix;
[0109] The context statement includes at least a sentences from the preceding or following sentences of the target statement in the target text, where the target statement is any sentence in the target text, and a is a positive integer. In other words, the context statement is the cross-sentence context of the target statement in the target text.
[0110] In one possible implementation, it is assumed that the target sentence is the Nth sentence of the target text, and the target text has B sentences, where N is a positive integer and B is a positive integer. Since the solution provided in this application is dedicated to solving the problem of inaccurate translation of long texts, B is much greater than a.
[0111] When N is greater than a and BN(B minus N) is greater than or equal to sentence a, the context statement includes the sentences before and after the target statement in the target text.
[0112] When N is less than or equal to a, the context statement includes the first N-1 (N minus 1) sentences and the last a sentences of the target statement. When BN is less than a, the context statement includes the first a sentences and the last BN sentences of the target text. No restrictions are imposed here.
[0113] It should be understood that the descriptions of the statements contained in the context here are for illustrative purposes only. In actual applications, the settings should be tailored to the specific application scenario, and no restrictions are imposed here.
[0114] In one possible implementation, the first vector and the context statement are input into the LLM, which encodes the first vector and the context statement. For example, the LLM extracts features from the context statement to obtain the feature matrix corresponding to the context statement, and then concatenates the feature matrix corresponding to the context statement with the first vector to obtain the context feature matrix.
[0115] In one possible implementation, 'a' is less than or equal to 3. By limiting the number of 'a's, the length of cross-sentence context can be effectively limited. Since contexts further away from the target statement have lower relevance to the target statement, introducing an appropriate amount of cross-sentence context can avoid increasing the computational load caused by introducing too much redundant context, and can also effectively improve the inference accuracy, thus effectively improving the inference efficiency of this scheme.
[0116] S230. Encode the target statement based on the context feature matrix and the second vector to obtain the target feature matrix;
[0117] In one possible implementation, the context feature matrix, the second vector, and the target statement are input into the LLM. The LLM encodes the context feature matrix, the second vector, and the target statement to obtain the target feature matrix. For example, the LLM extracts features from the target statement to obtain the feature matrix corresponding to the target statement. Then, the feature vector corresponding to the target statement, the context feature matrix, and the second vector are concatenated to obtain the target feature matrix. No limitation is imposed here.
[0118] S240. Concatenate the target feature matrix and the third vector to obtain the feature matrix to be inferred.
[0119] S250. Reason the feature matrix to be reasoned to obtain a set of reference feature vectors;
[0120] Each feature vector in the reference feature vector group corresponds to a word in the target sentence.
[0121] In one possible implementation, the feature matrix to be inferred is input into the LLM to obtain a reference feature vector; this is not limited here.
[0122] S260. Decode each feature vector in the reference feature vector group to obtain the translation of the target sentence.
[0123] In this embodiment, when translating the target text, at least a sentences from the preceding or following a sentences of the target statement are used as context statements. These context statements are then encoded based on a first vector to obtain a context feature matrix. The target statement is then encoded based on the context feature matrix and a second vector to obtain a target feature matrix. This target feature matrix is then concatenated with a third vector to obtain a feature matrix to be inferred. This feature matrix is then used to infer a reference feature vector, which is then used for decoding to obtain the translation of the target statement. By adding the first, second, and third vectors, the target statement and context statements learn to different degrees during the translation process, effectively improving translation accuracy.
[0124] In one possible implementation, Figure 2 The text translation method shown may also include step S270, which is performed after step S210 and before step S220.
[0125] S270. Preprocess the target text to obtain the context statement and the target statement.
[0126] In one possible implementation, the preprocessing may include randomly selecting any statement from the target text as the target statement, and then selecting at least a statements from the first a statements or the last a statements of the target text as context statements.
[0127] In this embodiment, the target text is preprocessed to obtain context statements and target statements, so that preprocessing-related operations and reasoning-related operations are performed separately, reducing the workload of the reasoning model and improving the reasoning efficiency of the reasoning model.
[0128] Based on the above scheme, it is easy to see that the translation of the target text in the scheme provided in this application is carried out on a sentence-by-sentence basis. This means that the number of sentences in the output translation is the same as the number of sentences in the target text. When checking the translation results, this feature can effectively improve the convenience of the checking process.
[0129] In the specific implementation of the above scheme, step S250 may include steps S251 to S253, such as... Figure 3a As shown.
[0130] S251. Extract features from the feature matrix to be inferred to obtain the first feature vector;
[0131] Here, the first feature vector is the feature vector corresponding to the first word, and the first word is a word in the target sentence.
[0132] S252. Reason about the first eigenvector to obtain the first weighted reorganization;
[0133] In the first weighting, each weight corresponds to a feature vector.
[0134] In one possible implementation, the method also involves a candidate corpus. In a scenario where the target text is in a first language and the goal of this method is to translate the target text into a second language, assuming that the feature vector corresponding to any weight in the first weight reassembly is a corpus feature vector, then this corpus feature vector corresponds to a candidate word. This candidate word can be any word in the candidate corpus, which is a corpus of the second language. That is, the weights corresponding to all words in the candidate corpus constitute the first weight reassembly.
[0135] It should be understood that the language types of the candidate corpus here can change as the translation target changes. In practical applications, the settings should be combined with the specific application scenario, and no restrictions are imposed here.
[0136] S253. Select the second feature vector based on the first weighted reorganization.
[0137] The second feature vector is contained in the reference feature vector group, and the second feature vector is the feature vector corresponding to the weight with the largest value in the first weighting.
[0138] In one possible implementation, the eigenvector corresponding to the weight with the largest value in the first weighting is directly used as the second eigenvector.
[0139] In this embodiment of the application, during the reasoning process of the target sentence, based on the sequence of each word in the target sentence and the feature vector corresponding to that word, and after fully learning based on the relationship between the text context, an accurate translation that conforms to the text context is obtained.
[0140] To further improve the accuracy of the translation, this application also proposes that, when reasoning about the feature matrix to be reasoned, reasoning is performed based on different features, and the final reasoning result is determined by voting using the reasoning results. That is, based on steps S251 and S252, step S250 can also include steps S254 to S257. In this case, step S253 is as follows: Figure 3b For step S258, please refer to [link / reference]. Figure 3b .
[0141] S254. Pool the context feature matrix to obtain the context feature vector;
[0142] S255. Pool the target feature matrix to obtain the target feature vector;
[0143] S256. Based on the context feature vector, the target feature vector, and the first feature vector, reason to obtain the second weighted reassembly;
[0144] In one possible implementation, the context feature vector, target feature vector, and first feature vector are input into the LLM for inference to obtain a second weighted reassembly. Similar to the first weighted reassembly, each weight in the second weighted reassembly corresponds to a feature vector. For example, if the feature vector corresponding to any weight in the second weighted reassembly is a corpus feature vector, then that corpus feature vector corresponds to a candidate word. That is, the weights corresponding to all words in the candidate corpus constitute the second weighted reassembly, and the weights of these words are obtained based on the analysis of the context and target statements. This enhances the influence of the context on the inference results.
[0145] It should be understood that there is no explicit order between steps S251 to S252 and steps S254 and S256. The description here is only an example. In actual application, it should be set according to the specific application scenario. There are no restrictions here.
[0146] S257. Based on the target feature vector and the first feature vector, reason to obtain the third weighted reorganization;
[0147] In one possible implementation, the target feature vector and the first feature vector are input into the LLM for inference to obtain a third weighted reassembly. Similar to the first weighted reassembly, each weight in the second weighted reassembly corresponds to a feature vector. That is, the weights corresponding to all words in the candidate corpus constitute the third weighted reassembly, and the weights of these words are all obtained based on the analysis of the target sentence. This enhances the influence of the target sentence on the inference results.
[0148] S258. Based on the first weighted reorganization, the second weighted reorganization, and the third weighted reorganization, a second feature vector is obtained by voting.
[0149] In one possible implementation, since the first, second, and third weighted reorganizations are all based on the same candidate corpus for inference—that is, the first weighted reorganization includes the weight 1 of the candidate words, the second weighted reorganization includes the weight 2 of the candidate words, and the third weighted reorganization includes the weight 3 of the candidate words—during voting, the weights 1, 2, and 3 of the candidate words are added together to obtain the total weight of the candidate words. The feature vector corresponding to the word with the highest total weight in the candidate corpus is selected as the second feature vector; no restrictions are imposed here.
[0150] In this embodiment, the use of contextual statements and target statements is enhanced during the reasoning process. The second feature vector (reasoning result) is obtained by voting on the reasoning results of multiple cases, which achieves better cross-sentence context modeling effect and further enhances the accuracy of LLM chapter translation.
[0151] To more intuitively understand the improvement in translation accuracy and optimization of discourse phenomena (an indicator used to describe the consistency of translation of a word or term) brought about by the text translation method provided in this application, evaluations were conducted on this invention and several different models. The evaluation data are as follows:
[0152] Depend on Figure 4 The experimental data shown demonstrates that the cross-language optimization translation evaluation metric (COMET) and machine translation evaluation metric (BLEU) of different models significantly improve the translation quality of the models (MPT and DeMPT) implemented based on the method provided in this application in Chinese-English, French-English, German-English, Spanish-English, and Russian-English translation.
[0153] Depend on Figure 5 The experimental data shown demonstrates improvements in blond debating across different models in Chinese, French, German, Spanish, and Russian translation environments. This indicates that the solution provided in this application effectively enhances the ability of LLM to address blond debating.
[0154] The reasoning process of the text translation method provided in this application has been introduced above. The following section will combine... Figure 6 The training process in the text translation method provided in this application is described.
[0155] S610. Obtain the training text and its translation;
[0156] S620. Preprocess the training text and its translation to obtain the preset sentence, the cross-sentence context of the preset sentence, and the translation of the preset sentence;
[0157] The preset sentence is any sentence in the training text, and the cross-sentence context of the preset sentence includes at least a sentences in the training samples that are either the preceding or following sentences of the preset sentence.
[0158] In one possible implementation, the training text needs to be preprocessed to obtain m triples (C, S, D), where C is the cross-sentence context of the preset sentence, S is the preset sentence, D is the translation of the preset sentence, which is usually referred to as the target sentence, and m is a positive integer.
[0159] It should be understood that since the application scenario of the method provided in this application is long text of the chapter type, m is usually a positive integer much larger than a, and the selection rule of the cross-sentence context of the preset sentence here is similar to the selection rule of the context statement in the aforementioned step S270. In specific scenarios, it should be set according to the specific application scenario, and no restrictions are made here.
[0160] S630. Encode the cross-sentence context of the preset sentence and the first vector to be updated to obtain the first feature matrix;
[0161] In one possible implementation, the cross-sentence context of the preset sentence and the first vector to be updated are input into the LLM. The LLM encodes the cross-sentence context of the preset sentence and the first vector to be updated. For example, the LLM extracts features from the cross-sentence context of the preset sentence to obtain the feature matrix corresponding to the cross-sentence context of the preset sentence, and concatenates the feature vector corresponding to the cross-sentence context of the preset sentence with the first vector to be updated to obtain the first feature matrix. That is, the first feature matrix is obtained by concatenating the feature vector corresponding to the cross-sentence context of the preset sentence and the first vector to be updated. No limitation is made here.
[0162] S640. Encode the preset sentence based on the first feature matrix and the second vector to be updated to obtain the second feature matrix;
[0163] In one possible implementation, the first feature matrix, the second vector to be updated, and the preset sentence are input into the LLM, which encodes the first feature matrix, the second vector to be updated, and the preset sentence. For example, the LLM extracts features from the preset sentence to obtain the feature matrix corresponding to the preset sentence, and then concatenates the second vector to be updated, the first feature matrix, and the feature matrix corresponding to the preset sentence to obtain the second feature matrix. That is, the second feature matrix is obtained by concatenating the feature matrix corresponding to the preset sentence, the second vector to be updated, and the first feature matrix; this is not limited here.
[0164] S650. Based on the third vector to be updated and the second feature matrix, reason about the preset sentence to obtain the preset feature vector group;
[0165] In one possible implementation, the fifth feature vector and the third vector to be updated are input into the LLM. The LLM concatenates the fifth feature vector and the third vector to be updated and then performs word-by-word reasoning on the concatenation result to obtain the second set of feature vectors.
[0166] In one possible implementation, after concatenating the third vector to be updated and the second feature matrix, the concatenation result of the third vector to be updated and the second feature matrix is input into the LLM to obtain a preset feature vector group. The preset feature vector group consists of multiple feature vectors obtained by reasoning on a preset sentence. Each feature vector in the multiple feature vectors corresponds to a translated word, and the translated word corresponding to each feature vector in the multiple vectors is included in the candidate corpus.
[0167] In one possible implementation, when reasoning about the concatenation result of the third vector to be updated and the second feature matrix, reasoning can also be performed based on different features, and the final reasoning result can be determined by voting using the reasoning results. The specific implementation process can be as follows: Figure 7 As shown.
[0168] In execution Figure 7 Before implementing the proposed solution, it is necessary to first pool the first feature matrix to obtain cross-sentence context feature vectors. Repooling the second feature matrix yields the feature vector of the preset sentence. The concatenation result of the third eigenvector to be updated and the second eigenma matrix is processed to obtain the third eigenvector. The third feature vector can be the feature vector corresponding to the d-th word in the preset sentence.
[0169] Execute again Figure 7 The proposed scheme uses cross-sentence context feature vectors, preset sentence feature vectors, and third feature vectors to perform reasoning and obtain a fourth weighted reorganization. Reasoning is performed based on the pre-defined sentence feature vector and the third feature vector to obtain the fifth weighted reorganization. Reasoning is performed based on the third eigenvector to obtain the sixth weighted recombination p(y). d |·).
[0170] Voting was conducted based on the fourth, fifth, and sixth power reorganizations to obtain the seventh power reorganization p. e (y d |·), select the feature vector corresponding to the weight with the largest value in the seventh weight reorganization as the fourth feature vector, and the fourth feature vector is contained in the preset feature vector group.
[0171] It should be understood that here Figure 7 The description of how to obtain a preset feature vector set is only an example. In practical applications, other methods can also be used to obtain the preset feature vector set, and no restrictions are imposed here.
[0172] S660. Update the first vector to be updated, the second vector to be updated, and the third vector to be updated based on the target feature vector group and the preset feature vector group to obtain the first vector, the second vector, and the third vector.
[0173] The target feature vector group is obtained by encoding each word in the translation of the preset sentence in the training text. The first vector is obtained by updating the first vector to be updated, the second vector is obtained by updating the second vector to be updated, and the third vector is obtained by updating the third vector to be updated.
[0174] In this embodiment, the training of the LLM text translation model is divided into three different stages, and different adjustable vectors are added to each of the three stages as cues. During the training stage, the three adjustable vectors are updated using labels (translations of the training text) to improve the performance of the LLM text translation model in the inference process.
[0175] To facilitate understanding of the above Figure 6 The training process described in the article will be discussed below. Figure 8 The architecture of the text translation method provided in this application is introduced.
[0176] In the first stage, the first training vector (P) is input into the LLM. C The first feature matrix (H) of the LLM output is obtained by combining the cross-sentence context (C) of the preset sentence with the cross-sentence context (C). C In the second stage, the second training vector (P) is input into the LLM. S ), preset sentence (S) and first feature matrix (H) C ), to obtain the second feature matrix (H) of the LLM output. S In the third stage, the second feature matrix (H) is input into the LLM. S ) and the third training vector (P) D Obtain the preset feature vector set (y d Based on the preset feature vector set and the target feature vector set (y). d The first, second, and third training vectors are iterated until the preset feature vector group and the target feature vector group are fitted to obtain the first, second, and third vectors.
[0177] From the above Figure 2 , Figure 3a , Figure 3b and Figure 6 This section introduces the text translation method and describes the implementation process of the text translation method provided in this application. Please refer to [link / reference]. Figure 9 ,like Figure 9 As shown, the text translation method provided in this application can be divided into the following four steps:
[0178] Step 1 involves preprocessing the training data, and the implementation of this step is the same as described above. Figure 6 The steps in step S620 are similar and will not be repeated here.
[0179] Step 2 involves fine-tuning the cue vector, and the steps are the same as described above. Figure 6 Steps S630 to S660 are similar and will not be described again here.
[0180] Step 3 involves fine-tuning the decoding enhancement, and its steps are the same as described above. Figure 6 In step S650 Figure 7 The implementation methods are similar and will not be described in detail here.
[0181] Step 4 involves reasoning about the target text, and its implementation is the same as described above. Figure 2 The steps S210-S260 are similar and will not be repeated here.
[0182] The text translation method provided in this application has been described above. The text translation apparatus provided in this application will now be described in conjunction with the accompanying drawings. Please refer to the drawings for details. Figure 10 , Figure 10 A schematic diagram of the structure of the text translation device provided in this application.
[0183] Text translation device 1000, including:
[0184] Acquisition unit 1010 is used to acquire target text;
[0185] The encoding unit 1020 is used to encode the context statement based on the first vector to obtain the context feature matrix. The context statement includes at least a sentences from the first a sentences or the last a sentences of the target statement in the target text, where a is a positive integer and the target statement is any sentence in the target text.
[0186] The encoding unit 1020 is also used to encode the target statement based on the context feature matrix and the second vector to obtain the target feature matrix;
[0187] The encoding unit 1020 is also used to concatenate the target feature matrix and the third vector to obtain the feature matrix to be inferred;
[0188] The reasoning unit 1030 is also used to reason about the feature matrix to be reasoned, and obtain a reference feature vector group, where each feature vector in the reference feature vector group corresponds to a word in the target sentence.
[0189] Decoding unit 1040 is used to decode each feature vector in the reference feature vector group to obtain the translation of the target sentence.
[0190] Optional, inference unit 1030, specifically used for:
[0191] Feature extraction is performed on the feature matrix to be inferred to obtain the first feature vector, which is the feature vector corresponding to the first word, and the first word is a word in the target sentence;
[0192] Reasoning is performed on the first eigenvector to obtain the first weight reorganization, in which each weight corresponds to an eigenvector.
[0193] The second feature vector is selected based on the first weighting reorganization. The second feature vector is contained in the reference feature vector group. The second feature vector is the feature vector corresponding to the weight with the largest value in the first weighting reorganization.
[0194] Optionally, the device also includes a pooling unit 1050 for:
[0195] Pool the context feature matrix to obtain the context feature vector;
[0196] Pool the target feature matrix to obtain the target feature vector;
[0197] The reasoning unit 1030 is also used to perform reasoning based on the context feature vector, the target feature vector and the first feature vector to obtain the second weighted reassembly.
[0198] The reasoning unit 1030 is also used to perform reasoning based on the target feature vector and the first feature vector to obtain the third weighted recombination.
[0199] The reasoning unit 1030 is specifically used to vote based on the first weighted reorganization, the second weighted reorganization, and the third weighted reorganization to obtain the second feature vector.
[0200] Optionally, the apparatus 1000 further includes a preprocessing unit 1060 for preprocessing the target text to obtain context statements and target statements.
[0201] Optionally, encoding unit 1020 is specifically used for:
[0202] Feature extraction is performed on the context statements to obtain the feature matrix corresponding to the context statements;
[0203] The feature matrix corresponding to the context statement is concatenated with the first vector to obtain the context feature matrix;
[0204] Encoding unit 1020 is specifically used for:
[0205] Feature extraction is performed on the target statement to obtain the feature matrix corresponding to the target statement;
[0206] The target feature matrix is obtained by concatenating the feature matrix corresponding to the target statement, the context feature matrix, and the second vector.
[0207] Optional, a is less than or equal to 3.
[0208] Optionally, the acquisition unit 1010 is also used to acquire the training text and the corresponding translation of the training text;
[0209] The encoding unit 1020 is also used to encode the cross-sentence context of the preset sentence and the first vector to be updated to obtain the first feature matrix. The preset sentence is any sentence in the training text, and the cross-sentence context of the preset sentence includes at least a sentences in the training text, either the preceding a sentences or the following a sentences.
[0210] The encoding unit 1020 is also used to encode the preset sentence based on the first feature matrix and the second vector to be updated, so as to obtain the second feature matrix;
[0211] The reasoning unit 1030 is also used to reason about the preset sentence based on the third vector to be updated and the second feature matrix to obtain a preset feature vector group.
[0212] The device also includes an update unit 1070, which is used to update the first vector to be updated, the second vector to be updated, and the third vector to be updated based on the target feature vector set and the preset feature vector set, to obtain the first vector, the second vector, and the third vector. The target feature vector set is obtained by encoding each word in the translation of the preset sentence in the translation of the training text.
[0213] Optionally, the preprocessing unit 1060 is also used to preprocess the training text and its translation to obtain a preset sentence, the cross-sentence context of the preset sentence, and the translation of the preset sentence.
[0214] The text translation device provided in the embodiments of this application will be described below. Please refer to [link / reference]. Figure 11 , Figure 11 This is a schematic diagram of the structure of a text translation device provided in an embodiment of this application. The computer device 1100 includes a processor 1101, a memory 1102, a communication interface 1103, and a bus 1104. The processor 1101, memory 1102, and communication interface 1103 communicate via the bus 1104, or they can communicate via wireless transmission or other means. The memory 1102 stores program code, and the processor 1101 can call the program code stored in the memory 1102 to execute the aforementioned code. Figure 2 , Figure 3a , Figure 3b or Figure 6 The operations performed in the illustrated embodiment will not be described again here.
[0215] It should be understood that in the embodiments of this application, the processor 1101 may be a CPU, or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0216] The memory 1102 may include read-only memory and random access memory, and provides instructions and data to the processor 1101. The memory 1102 may also include non-volatile random access memory. For example, the memory 1102 may also store device type information.
[0217] The memory 1102 can be volatile memory or non-volatile memory, or it can include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0218] In addition to the data bus, bus 1104 may also include a power bus, control bus, and status signal bus. However, for clarity, all buses are labeled as bus 1104 in the diagram. Bus 1140 can be a Peripheral Component Interconnect Express (PCIe) bus, or an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. Bus 1104 can be divided into address bus, data bus, control bus, etc.
[0219] Computer device 1100 may also include one or more communication interfaces and one or more operating systems, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSD TM wait.
[0220] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0221] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0222] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0223] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0224] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A text translation method, characterized in that, include: Get the target text; The context statement is encoded based on the first vector to obtain a context feature matrix. The context statement includes at least a sentences from the a sentences before or after the target statement in the target text, where a is a positive integer, and the target statement is any sentence in the target text. The target statement is encoded based on the context feature matrix and the second vector to obtain the target feature matrix; The target feature matrix and the third vector are concatenated to obtain the feature matrix to be inferred; The feature matrix to be reasoned is reasoned to obtain a reference feature vector group, where each feature vector in the reference feature vector group corresponds to a word in the target sentence; The translation of the target statement is obtained by decoding each feature vector in the reference feature vector group.
2. The method according to claim 1, characterized in that, The step of reasoning on the feature matrix to be reasoned to obtain the target feature vector set includes: Feature extraction is performed on the feature matrix to be reasoned to obtain a first feature vector, which is the feature vector corresponding to the first word, and the first word is a word in the target sentence; Reasoning is performed on the first feature vector to obtain a first weight reorganization, in which each weight corresponds to a feature vector; A second feature vector is selected based on the first weighting reorganization. The second feature vector is included in the reference feature vector group. The second feature vector is the feature vector corresponding to the weight with the largest value in the first weighting reorganization.
3. The method according to claim 2, characterized in that, The method further includes: Pool the context feature matrix to obtain the context feature vector; Pool the target feature matrix to obtain the target feature vector; Reasoning is performed based on the context feature vector, the target feature vector, and the first feature vector to obtain a second weighted reassembly. Based on the target feature vector and the first feature vector, reasoning is performed to obtain a third weighted reorganization; The step of selecting the second feature vector based on the first weighted recombination includes: The second feature vector is obtained by voting based on the first weighted reorganization, the second weighted reorganization, and the third weighted reorganization.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The target text is preprocessed to obtain the context statement and the target statement.
5. The method according to any one of claims 1 to 4, characterized in that, The process of encoding the context statement based on the first vector to obtain the context feature matrix includes: Feature extraction is performed on the context statement to obtain the feature matrix corresponding to the context statement; The feature matrix corresponding to the context statement is concatenated with the first vector to obtain the context feature matrix; The process of encoding the target statement based on the context feature matrix and the second vector to obtain the target feature matrix includes: Feature extraction is performed on the target statement to obtain the feature matrix corresponding to the target statement; The target feature matrix is obtained by concatenating the feature matrix corresponding to the target statement, the context feature matrix, and the second vector.
6. The method according to any one of claims 1 to 5, characterized in that, The value of a is less than or equal to 3.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain the training text and its corresponding translation; Encode the cross-sentence context of the preset sentence and the first vector to be updated to obtain a first feature matrix. The preset sentence is any sentence in the training text. The cross-sentence context of the preset sentence includes at least a sentences in the training text that are either a sentences before or a sentences after the preset sentence. The preset sentence is encoded based on the first feature matrix and the second vector to be updated to obtain the second feature matrix; Based on the third vector to be updated and the second feature matrix, the preset sentence is inferred to obtain a preset feature vector group; The first vector to be updated, the second vector to be updated, and the third vector to be updated are updated based on the target feature vector set and the preset feature vector set to obtain the first vector, the second vector, and the third vector. The target feature vector set is obtained by encoding each word in the translation of the preset sentence in the translation of the training text.
8. The method according to claim 6, characterized in that, The method further includes: The training text and its translation are preprocessed to obtain a preset sentence, the cross-sentence context of the preset sentence, and the translation of the preset sentence.
9. A text translation device, characterized in that, include: The acquisition unit is used to acquire the target text; The encoding unit is used to encode the context statement based on the first vector to obtain a context feature matrix. The context statement includes at least a sentences from the first a sentences or the last a sentences of the target statement in the target text, where a is a positive integer, and the target statement is any sentence in the target text. The encoding unit is also used to encode the target statement based on the context feature matrix and the second vector to obtain the target feature matrix; The encoding unit is also used to concatenate the target feature matrix and the third vector to obtain the feature matrix to be inferred; The reasoning unit is also used to reason about the feature matrix to be reasoned to obtain a reference feature vector group, wherein each feature vector in the reference feature vector group corresponds to a word in the target sentence; The decoding unit is used to decode each feature vector in the reference feature vector group to obtain the translation of the target statement.
10. A communication device, characterized in that, Includes a processor, which is coupled to a memory; The memory stores instructions that, when executed on the processor, cause the communication device to perform the method of any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a processor, cause the method of any one of claims 1 to 8 to be implemented.
12. A computer program product, characterized in that, When the computer program product is executed on a computer, the method of any one of claims 1 to 8 is implemented.