Multi-language deep translation method and system based on large model and electronic equipment
By obtaining target sentence groups for semantic enhancement and translation repair, the semantic offset and consistency problems of multilingual translation systems in complex language structures and professional fields are solved, and the context-aware chapter-level translation and consistency self-evaluation are realized, which improves the accuracy and professionalism of translation.
Patent Information
- Application Number
- CN202510822214.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-19
AI Technical Summary
When dealing with complex language structures, professional field terms and long documents, the existing multilingual translation system lacks context modeling capabilities, resulting in inconsistencies in terms of terminology, confusion of pronouns, mutation of style, and lacks self-verification and adjustment mechanisms, and cannot identify and correct semantic offsets, mistranslation and other problems.
By obtaining the target sentence group for semantic enhancement, the context-enhanced representation vector is generated, the semantic deviation is calculated using the initial translation model and the back-translation model, the user control prompt vector is obtained, and the repair model is input for translation and repair, so as to realize the deep collaboration mechanism of context-aware chapter-level translation, consistency self-evaluation and user prompt control.
It improves the accuracy of multilingual translation, solves the problems of lack of context, unverifiable results and uncontrollable output, realizes context-aware chapter-level translation and consistency self-evaluation, and enhances the professionalism and coherence of translation.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and in particular, to a multi-language deep translation method, system, and electronic device based on a large model. Background Art
[0002] With the continuous evolution of large-scale pre-trained language models based on the Transformer architecture (such as mBART, mT5, GPT, etc.), the multi-language translation task has entered the stage driven by large models. Existing systems can achieve high-quality sentence-level translations between mainstream language pairs and are widely used in scenarios such as web translation, APP real-time conversations, and preliminary document transcribing. However, when dealing with translation tasks involving complex language structures, professional domain terms, and long documents, the existing sentence-splitting-based translation strategies gradually expose their core defects: on the one hand, the lack of context modeling ability leads to frequent problems such as inconsistent terms before and after, chaotic pronoun references, and sudden changes in writing styles, seriously affecting the professionalism and coherence of the translated text; on the other hand, translation systems are generally "one-way generation" and lack a self-verification and adjustment mechanism, unable to identify and correct problems such as semantic drift and mistranslation. In summary, the current multi-language translation systems still have obvious deficiencies in terms of semantic understanding depth, translation stability, and user controllability. Therefore, how to improve the accuracy of multi-language translation has become an urgent technical problem to be solved. Summary of the Invention
[0003] The main purpose of the embodiments of this application is to propose a multi-language deep translation method, system, and electronic device based on a large model, aiming to improve the accuracy of multi-language translation.
[0004] To achieve the above object, the first aspect of the embodiments of this application proposes a multi-language deep translation method based on a large model, and the method includes:
[0005] Obtain a target sentence group; wherein, the target sentence group includes a current translation target sentence and context window sentences;
[0006] Perform semantic enhancement according to the target sentence group to obtain a context-enhanced representation vector;
[0007] Input the context-enhanced representation vector into a preset initial translation model to obtain an initial translation result;
[0008] Input the initial translation result into a preset back-translation model to obtain a first back-translated sentence;
[0009] Calculate the average residual according to the current translation target sentence and the first back-translated sentence to obtain a first semantic deviation amount;
[0010] Score the initial translation result based on the first semantic deviation amount, the current translation target sentence, and the first back-translated sentence to obtain first consistency scoring data;
[0011] If the first consistency scoring data is lower than a preset scoring threshold, obtain a control prompt vector input by the user;
[0012] Obtain a repair control vector based on the first semantic deviation amount and the control prompt vector;
[0013] Input the repair control vector into a preset repair model to obtain a translation repair result.
[0014] In some embodiments, the initial translation model further outputs a first attention matrix, and the calculating the average residual based on the current translation target sentence and the first back-translated sentence to obtain the first semantic deviation amount includes:
[0015] Encode the current translation target sentence and the first back-translated sentence to obtain a target embedding sequence and a first back-translated embedding sequence;
[0016] Obtain a first alignment token pair of the target embedding sequence and the first back-translated embedding sequence according to the first attention matrix;
[0017] Calculate the average residual according to the first alignment token pair, the target embedding sequence, and the first back-translated embedding sequence to obtain the first semantic deviation amount.
[0018] In some embodiments, the calculating the average residual according to the first alignment token pair, the target embedding sequence, and the first back-translated embedding sequence to obtain the first semantic deviation amount includes:
[0019] The average residual calculation formula is:
[0020]
[0021] where, Δ i represents the first semantic deviation amount, represents the set of first alignment token pairs, u j represents the target embedding sequence, represents the first back-translated embedding sequence.
[0022] In some embodiments, the repair model further outputs a second attention matrix. After inputting the repair control vector into the preset repair model to obtain the translation repair result, it includes:
[0023] Input the translation repair result into the back-translation model to obtain a second back-translated sentence;
[0024] Encode the second back-translated sentence to obtain a second back-translated embedding sequence;
[0025] Obtain a second alignment token pair of the target embedding sequence and the second back-translated embedding sequence according to the second attention matrix;
[0026] Calculate the average residual according to the second alignment token pair, the target embedding sequence and the second back-translated embedding sequence to obtain a second semantic deviation;
[0027] Score the translation repair result according to the second semantic deviation, the current translation target sentence, and the second back-translated sentence to obtain second consistency scoring data.
[0028] In some embodiments, after scoring the translation repair result according to the second semantic deviation, the current translation target sentence, and the second back-translated sentence to obtain second consistency scoring data, the method further includes:
[0029] If the second consistency scoring data is greater than the scoring threshold, obtain the terms of the current translation target sentence to obtain a first term, and obtain the terms of the translation repair result to obtain a second term;
[0030] Match the first term and the second term to obtain a plurality of term key-value pairs;
[0031] Store the term key-value pairs in a preset term status cache;
[0032] Construct and store a pronoun mapping relationship according to the pronouns and corresponding entity phrases of the translation repair result.
[0033] In some embodiments, the semantic enhancement according to the target sentence group to obtain a context-enhanced representation vector includes:
[0034] Obtain a structured representation vector of each current translation target sentence and the context window sentence; wherein, the structured representation vector is the sum of a token embedding vector and an absolute position encoding;
[0035] Concatenate each structured representation vector to obtain an input sequence;
[0036] Encode the input sequence to obtain a context-enhanced representation vector.
[0037] In some embodiments, the scoring of the initial translation result according to the first semantic deviation, the current translation target sentence, and the first back-translated sentence to obtain first consistency scoring data includes:
[0038] The calculation formula of the first consistency scoring data is:
[0039] Score i = 1 - (Δ i + γ·δ i + μ·R i );
[0040]
[0041] Wherein, Score i represents the first consistency scoring data, Δ i represents the first semantic deviation, γ represents the term penalty coefficient, δ i represents the term matching deviation, represents the current translation target sentence, represents the set of terms successfully recovered in the first back-translated sentence, μ represents the weight controlling the impact of structure rearrangement, R i represents the sentence rearrangement penalty term.
[0042] To achieve the above object, a second aspect of the embodiments of the present application proposes a multi-language deep translation system based on a large model, and the system includes:
[0043] An acquisition module, configured to acquire a target sentence group; wherein, the target sentence group includes the current translation target sentence and context window sentences;
[0044] An enhancement module, configured to perform semantic enhancement according to the target sentence group to obtain a context-enhanced representation vector;
[0045] A translation module, configured to input the context-enhanced representation vector into a preset initial translation model to obtain an initial translation result;
[0046] A back-translation module, configured to input the initial translation result into a preset back-translation model to obtain a first back-translated sentence;
[0047] A calculation module, configured to calculate an average residual according to the current translation target sentence and the first back-translated sentence to obtain a first semantic deviation;
[0048] A scoring module, configured to score the initial translation result according to the first semantic deviation, the current translation target sentence, and the first back-translated sentence to obtain a first consistency scoring data;
[0049] A first input module, configured to obtain a control prompt vector input by the user if the first consistency scoring data is lower than a preset scoring threshold;
[0050] A repair module, configured to obtain a repair control vector according to the first semantic deviation and the control prompt vector;
[0051] A second input module, configured to input the repair control vector into a preset repair model to obtain a translation repair result.
[0052] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0053] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0054] The multi-language deep translation method, system and electronic device based on a large model proposed in the present application obtain a target sentence group, where the target sentence group includes a current translation target sentence and context window sentences. Semantic enhancement is performed on the target sentence group to obtain a context-enhanced representation vector. The context-enhanced representation vector is input into a preset initial translation model to obtain an initial translation result. The initial translation result is input into a preset back-translation model to obtain a first back-translated sentence. An average residual is calculated based on the current translation target sentence and the first back-translated sentence to obtain a first semantic deviation amount. The initial translation result is scored according to the first semantic deviation amount, the current translation target sentence, and the first back-translated sentence to obtain first consistency scoring data. If the first consistency scoring data is lower than a preset scoring threshold, a control prompt vector input by the user is obtained. A repair control vector is obtained based on the first semantic deviation amount and the control prompt vector. The repair control vector is input into a preset repair model to obtain a translation repair result. It solves the core defects of existing multi-language translation systems in aspects such as context loss, result unverifiability, and output uncontrollability, realizes a deep cooperation mechanism of context-aware paragraph-level translation, consistency self-evaluation, and user prompt control, and improves the accuracy of multi-language translation. Description of the Drawings
[0055] Figure 1 is a flowchart of the multi-language deep translation method based on a large model provided by the embodiments of the present application;
[0056] Figure 2 is Figure 1 a flowchart of step S102 in
[0057] Figure 3 is Figure 1 a flowchart of step S105 in
[0058] Figure 4 is a flowchart of the multi-language deep translation method based on a large model provided by another embodiment of the present application;
[0059] Figure 5 It is a flowchart of the multi - language deep translation method based on a large model provided by the third embodiment of the present application;
[0060] Figure 6 It is a schematic structural diagram of the multi - language deep translation system based on a large model provided by the embodiment of the present application;
[0061] Figure 7 It is a schematic hardware structure diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0062] In order to make the purpose, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0063] It should be noted that although the functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the system or the order in the flowchart. Terms such as "first", "second", etc. in the specification, claims and the above - mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0065] When dealing with translation tasks involving complex language structures, technical - field terms and long documents, the existing translation strategies based on sentence segmentation gradually expose their core defects: on the one hand, the lack of context - modeling ability leads to frequent problems such as inconsistent terms before and after, chaotic pronoun references, and sudden changes in writing styles, seriously affecting the professionalism and coherence of the translation; on the other hand, translation systems are generally "one - way generation" and lack self - verification and adjustment mechanisms, unable to identify and correct semantic drifts, mistranslations, etc. In summary, the current multi - language translation systems still have obvious deficiencies in terms of semantic understanding depth, translation stability and user controllability.
[0066] Based on this, the embodiments of the present application provide a multi - language deep translation method, system and electronic device based on a large model, aiming to enhance the semantics according to the target sentence group to obtain a context - enhanced representation vector. Translate the current translation target sentence based on the context - enhanced representation vector to obtain an initial translation result. Back - translate the initial translation result to obtain a first back - translated sentence, and calculate the first semantic deviation amount between the current translation target sentence and the first back - translated sentence. Score the initial translation result based on the first semantic deviation amount, the current translation target sentence, and the first back - translated sentence to obtain the first consistency scoring data. If the first consistency scoring data is lower than a preset scoring threshold, obtain the control prompt vector input by the user. Obtain a repair control vector according to the first semantic deviation amount and the control prompt vector. Input the repair control vector into a preset repair model to obtain a translation repair result. It solves the core defects of existing multi - language translation systems in aspects such as context loss, result unverifiability, and output uncontrollability, realizes a deep cooperation mechanism integrating context - aware discourse - level translation, consistency self - evaluation, and user - prompt control, and improves the accuracy of multi - language translation.
[0067] The multi - language deep translation method, system and electronic device based on a large model provided by the embodiments of the present application are specifically described through the following embodiments. First, the multi - language deep translation method in the embodiments of the present application is described.
[0068] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0069] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0070] The multi - language deep translation method based on large models provided by the embodiments of the present application can be applied to terminals, can also be applied to server - sides, or can be software running on terminals or server - sides. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server - side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the multi - language deep translation method based on large models, etc., but is not limited to the above forms.
[0071] The present application can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet - type devices, multi - processor systems, micro - processor - based systems, set - top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer - executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In distributed computing environments, program modules can be located in local and remote computer storage media including storage devices.
[0072] Please refer to Figure 1 , Figure 1 which is a flowchart of the multi - language deep translation method based on large models provided by the embodiments of the present application. Figure 1 The method in
[0073] may include but is not limited to steps S101 to S109.
[0074] Step S101, obtain a target sentence group; where the target sentence group includes a current translation target sentence and context window sentences.
[0075] Step S102, perform semantic enhancement according to the target sentence group to obtain a context - enhanced representation vector.
[0076] Step S103, input the context - enhanced representation vector into a preset initial translation model to obtain an initial translation result.
[0077] Step S105: Calculate the average residual based on the current translation target sentence and the first back-translated sentence to obtain the first semantic deviation amount;
[0078] Step S106: Score the initial translation result according to the first semantic deviation amount, the current translation target sentence, and the first back-translated sentence to obtain the first consistency scoring data;
[0079] Step S107: If the first consistency scoring data is lower than the preset scoring threshold, obtain the control prompt vector input by the user;
[0080] Step S108: Obtain the repair control vector according to the first semantic deviation amount and the control prompt vector;
[0081] Step S109: Input the repair control vector into the preset repair model to obtain the translation repair result.
[0082] In step S101 of some embodiments, the current translation target sentence T i and its context window sentences C i ={T i-k ,…,T i-1 ,T i+1 ,…,T i+k} form a target sentence group for subsequent splicing. Where k is the window width, usually set to 2. The original paragraph text is from the translation task request received by this system, specifically it can be: the uploaded local document (such as.docx,.pdf file, processed by the language detection and paragraph division module into sentences), the text extracted from the web page body (obtained by parsing the DOM structure to get the body block), or the OCR scan result (filtered out natural language sentences by combining character confidence). The text processing module uses a regular syntax divider (based on language-specific separation rules) to cut the paragraph into a sentence sequence to form a sentence list [T1, T2,..., T n . If the current translation target sentence is the second sentence T2, the system automatically extracts the context window sentence C2={T1, T3} to form a target sentence group for subsequent splicing.
[0083] Please refer to Figure 2 , in some embodiments, step S102 may include but is not limited to steps S201 to S203:
[0084] Step S201: Obtain the structured representation vectors of each current translation target sentence and context window sentence; among them, the structured representation vector is the sum of the token embedding vector and the absolute position encoding;
[0085] Step S202: Concatenate each structured representation vector to obtain the input sequence;
[0086] Step S203: Encode the input sequence to obtain a context-enhanced representation vector.
[0087] In step S201 of some embodiments, a token is the smallest unit in text processing and can be a word, a character, or a word-and-character combination. Each sentence will generate several tokens, that is, each sentence is a token sequence. Tokenize the text in each sentence of the target sentence group into a token sequence, and encode it using the same tokenizer as the main model (e.g., mBART uses the SentencePiece tokenizer). For each sentence T j ∈{T i}∪C i , use the shared embedding layer to map the token to a token embedding vector where l j is the number of tokens in this sentence, and d is the embedding dimension, usually 768 or 1024. The embedding layer is implemented using a pre-loaded multi-lingual word vector matrix and is initialized during the model training phase.
[0088] Furthermore, to maintain the boundary information of each sentence in the concatenated input, generate independent absolute position encodings for each sentence which not only reflect the position of the token within the sentence but also incorporate "the relative position of the sentence in the paragraph". For example, set the current translation target sentence T2 as the reference position 0, and the previous and subsequent sentences are -1 and +1 respectively, to prevent the same-position tokens in different sentences from conflicting in encoding. Add the token embedding vector of each sentence to the absolute position encoding to generate a structured representation vector, as shown in the following formula (1):
[0089] H j =E j +P j , (1)
[0090] where, H j represents the structured representation vector of sentence T j , E j represents the token embedding vector of sentence T j , and P j represents the absolute position encoding of sentence T j , constructed by combining the in-sentence position index and the inter-sentence offset.
[0091] In steps S202 to S203 of some embodiments, concatenate the structured representation vectors of all sentences in the target sentence group to form a complete input sequence, and send it to the encoder part of the initial translation model for encoding to obtain a context-enhanced representation vector, as shown in the following formula (2):
[0092] S i =M(Concat(Hi-k , …, H i , …, H i+k ), (2)
[0093] Among them, represents the context-enhanced representation vector, L = ∑ j l j is the total number of tokens after concatenation. Concat(·) means concatenating the structured representation vectors of multiple sentences into a long sequence in token order. M represents the encoder part of the initial translation model, specifically a 12-layer Transformer structure, with each layer containing 8-head attention mechanism and a feed-forward sub-layer. H i represents the structured representation vector. S i preserves the context semantic structure of the current translation target sentence in the paragraph and serves as the key input generated by the initial translation model.
[0094] Through the above steps S201 to S203, a structured concatenation mechanism is introduced, and the input of the initial translation model is structurally encoded using sentence-level position information, enabling the model to still retain the inter-sentence boundaries and semantic flow when processing the concatenation of multilingual sentences. A context-enhanced representation vector with discourse context information is constructed to improve semantic consistency, term unity, and reference continuity in multilingual translation. Especially in language pairs with large differences in word order structures such as Chinese-Arabic and Chinese-German, problems such as subject confusion and term incoherence caused by isolated sentence translation can be avoided. Different from conventional pre-trained models that only input continuous tokens, this step improves the model's semantic parsing and modeling ability for "target sentences in a multi-sentence context" by integrating multi-sentence structures and explicitly retaining sentence boundaries.
[0095] In step S103 of some embodiments, considering the need for professional text translation scenarios with high context requirements, strong term dependencies, and the need for coherent control of language styles (such as multilingual legal documents, technical documents, patent specifications, etc.), this step needs to ensure that the process of generating the initial translation result is not only based on word meanings, but also must explicitly consider complex semantic structures such as term dependencies, syntactic logical positions, and paragraph styles in the context.
[0096] For this reason, a context alignment and reinforcement translation mechanism is proposed, which introduces a structural position bias term and a term attention penalty term into the conventional Transformer decoding structure to ensure that the context-enhanced representation vector constructed in the above steps is fully utilized during the generation process, and to prevent the model from ignoring the known term distribution or deviating from the context center of gravity.
[0097] The input of this step is the context-enhanced representation vector Where L is the number of concatenated context tokens and d is the vector dimension. This context-enhanced representation vector is generated by the encoder model M, which fuses the structure and semantic information of the current translation target sentence T i and its context window sentence C i .
[0098] The current goal is to generate an initial translation result Y i , whose initial decoding input is a special start symbol, and subsequent generation depends on the forward prediction at each step and the context-enhanced representation vector S i . The Transformer decoder structure in the large language model (initial translation model) is adopted, where each layer contains a multi-head self-attention mechanism and a context cross-attention module
[0099] To ensure the targeted utilization of context, a structure alignment position bias term is introduced in the cross-attention layer. The idea is that when the initial translation model generates the t-th token of the current translation target sentence, it should give priority to focusing on the content in the context that is closest to its semantic / syntactic position, especially near the boundaries of the current translation target sentence T i . Therefore, a bias term B t,j is added when calculating the attention, which is used to adjust the degree to which the model focuses on the j-th token in the context. The attention distribution is shown in the following formula (3):
[0100]
[0101] where α t represents the query vector (query vector) at the current generation position t, which is calculated by the feed-forward layer of the decoder. S i represents the context-enhanced representation vector, represents the transpose of the context-enhanced representation vector. q t represents the query vector at the current decoding position. d represents the vector dimension scaling factor. B t,j represents the structure alignment bias term, whose value is calculated based on the relative sentence distance and the relative position within the sentence of the token j in the context and the current translation target sentence T i , as shown in the following formula (4):
[0102] B t,j = -λ1·|p j - p i |- λ2·|r j - r t |, (4)
[0103] where p j represents the paragraph position number of the sentence to which the token j belongs, and p i represents the current translation target sentence Ti The paragraph position number of r j Indicates the position of token j in its sentence, r t Indicates the position estimate of the currently generated token t in the current target translation sentence (estimated by the number of historical tokens). λ1 and λ2 represent hyperparameters that control the alignment strength in two dimensions (e.g., λ1 = 0.3, λ2 = 0.7).
[0104] This structural alignment bias term has a clear scenario background: in legal texts or patent translations, the same term often appears repeatedly in adjacent sentences or has consistent semantics in a paragraph. Therefore, it is extremely crucial to guide the model to prioritize attention to nearby terms. This mechanism breaks through the way of undifferentiated projection of attention in the standard Transformer and strengthens the context relevance.
[0105] In addition, to prevent the model from generating term drift (such as translating "the present invention" into "the device" and other biased expressions), a term attention sparsity penalty term is introduced. This mechanism is based on the pre-annotated term position index set in the input (synchronously output by the term recognition module in the above step). If the attention distribution α of the currently generated token t is not concentrated on the term positions, the decoding loss will be increased.
[0106] The specific form is shown in the following formula (5). The following regularization term is added during training:
[0107]
[0108] Among them, represents the attention sparsity penalty term. α t,j represents the attention value of the t-th generated token to the j-th context token. β represents the penalty coefficient (e.g., β = 1.5), represents the term position index set. m represents the total number of tokens to be generated in the current target translation sentence.
[0109] This regularization term directly adds a structural preference to the training objective of the model, prompting it to preferentially utilize the known term context in the context structure rather than "inventing" term translations each time, thereby greatly improving the consistency and stability of professional text translation.
[0110] During the generation process, beam search (with a size of 4) is used for candidate search, and finally the sequence with the highest log-likelihood is selected as the translation output to obtain the initial translation result. The translation output also retains the attention vector α of each token t , which are combined to form the first attention matrix
[0111] It should be noted that the above steps propose two innovative mechanisms for multi - language and term - intensive scenarios: structural alignment attention bias and term attention sparse penalty term. On the one hand, it enhances the model's utilization of structural information in the context, and on the other hand, it explicitly guides the consistent tracking of terms during the generation process, thus solving problems such as polysemy of terms, subject errors, and style jumps in traditional large - model translation. In step S104 of some embodiments, the initial translation result Y generated in the previous stage i and the first attention matrix A i Based on this, a semantic consistency detection mechanism for multi - language professional scenarios is designed. Compared with general texts, contents such as patents, laws, and technical documents have the characteristics of highly consistent terms, strong context dependence, and complex syntactic structures, making it difficult for traditional methods based on shallow metrics such as BLEU and ROUGE to detect deep - level semantic shifts. The core objective of this step is to judge: whether the initial translation result Y i faithfully expresses the true meaning of the original text T i (the current target translation sentence), and whether it maintains the consistency of proper nouns, subject - predicate logic, and context terms. For this purpose, a structured analysis of the translation quality is formed through a multi - channel residual detection mechanism based on back - translation, attention alignment, and term comparison.
[0112] To achieve semantic consistency judgment, a back - translation mechanism is used to perform reverse translation of the initial translation result Y from the target language to the source language. The back - translation model M i has the same structure as the initial translation model (for example, using the same pre - training system such as mBART / mT5), and its parameters are frozen during deployment. The initial translation result is input into the preset back - translation model to obtain the first back - translation sentence r whose language is the original source language and the length is n'. Its language is the original source language, and the length is n'.
[0113] Please refer to Figure 3 , in some embodiments, the initial translation model also outputs the first attention matrix. Step S105 may include but is not limited to steps S301 to S303:
[0114] Step S301, encoding the current target translation sentence and the first back - translation sentence to obtain a target embedding sequence and a first back - translation embedding sequence;
[0115] Step S302, obtaining the first alignment token pair of the target embedding sequence and the first back - translation embedding sequence according to the first attention matrix;
[0116] Step S303, calculating the average residual according to the first alignment token pair, the target embedding sequence, and the first back - translation embedding sequence to obtain the first semantic deviation amount.
[0117] In step S301 of some embodiments, due to language differences and structural changes, when judging whether the initial translation result Y i faithfully expresses the original text T i (the current translation target sentence), simple position alignment cannot be used. Therefore, the following multi-level alignment and residual detection process is designed:
[0118] Embedding alignment construction: Use a shared cross-lingual encoder (such as the XLM-R or LaBSE model) to encode the current translation target sentence T i to obtain the target embedding sequence Encode the first back-translated sentence to obtain the first back-translated embedding sequence
[0119] In step S302 of some embodiments, to construct the alignment relationship, relying on the first attention matrix A i , extract the maximum attention position of each generated token as the semantic source position of the token, and then find the reverse mapping point of the corresponding token in the back-translation, so as to establish the cross-lingual token alignment table and obtain the first alignment token pair.
[0120] In step S303 of some embodiments, after obtaining the first alignment token pair, calculate the average residual through the following formula (6):
[0121]
[0122] where Δ i represents the first semantic deviation amount, and the larger the value, the more serious the deviation. represents the set of the first alignment token pairs, including multiple first alignment token pairs, u j represents the target embedding sequence, represents the first back-translated embedding sequence.
[0123] Through the above steps S301 to S303, the first semantic deviation amount can be obtained, so as to judge whether the initial translation result Y i faithfully expresses the true meaning of the current translation target sentence T i . And this structural residual is based on the real semantic space embedding rather than the surface text form, and can identify scenarios with different structures but semantic drift, especially applicable to language pairs with significant word order differences such as Chinese-German and Chinese-Arabic.
[0124] In step S106 of some embodiments, the initial translation result is scored according to the first semantic deviation amount, the current translation target sentence, and the first back-translated sentence to obtain the first consistency scoring data. Specifically, the calculation of the first consistency scoring data is shown in the following formulas (7) and (8):
[0125] Score i =1-(Δ i +γ·δ i +μ·R i ), (7)
[0126]
[0127] Wherein, Score i represents the first consistency scoring data, Δ i represents the first semantic deviation amount, γ represents the term penalty coefficient (such as 1.2), which controls the intensity of the term missing penalty. δ i represents the term matching deviation degree, represents the current translation target sentence, represents the set of terms successfully restored in the first back-translated sentence. μ represents the weight controlling the impact of structure rearrangement (such as 0.5), and R i represents the sentence rearrangement penalty term, which is calculated by comparing the dependency tree sequences of the current translation target sentence and the first back-translated sentence (for example, based on syntactic distance or edit distance).
[0128] The lower the first consistency scoring data Score i , the lower the semantic credibility of the translation. Set the scoring threshold τ (such as 0.75). When Score i ≤τ, mark B i =1, representing that the translation is not credible and needs to enter the next step for repair.
[0129] In step S107 of some embodiments, when the first consistency scoring data is lower than the scoring threshold, it indicates that there is a semantic deviation in the initial translation result Y i . To ensure the semantic faithfulness and term consistency of the final translation, and at the same time integrate user control requirements (such as style, use of professional terms, etc.), a repair control mechanism that fuses the control hint vector P i and the first semantic deviation amount Δ i is designed, so that the translation repair not only has the ability to correct errors, but also has the directivity of style and professional expression. Therefore, if the first consistency scoring data is lower than the scoring threshold, the control hint vector input by the user is obtained.
[0130] In steps S108 to S109 of some embodiments, when B i =1, the system determines that the first back-translated sentence Yi There are semantic problems, and the translation repair process is initiated. This step reuses the decoder architecture of the initial translation model M and injects repair control vectors at the input stage and the intermediate layer to achieve the directional reconstruction of the incorrect translation. The entire repair process is modeled in the following three modules:
[0131] Semantic Deviation Embedding Construction Module:
[0132] To convert the first semantic deviation Δ calculated in the previous step i into a vector input, a 1-layer feedforward network is used to encode it into a control signal, as shown in the following formula (9):
[0133] r i = tanh(W r ·Δ i + b r ), (9)
[0134] where represents the vector of the residual signal, with the dimension consistent with the intermediate layer of the model (e.g., d = 768). This vector is used to indicate the current semantic direction that needs to be repaired and serves as the "semantic adjustment anchor point". represents the learnable parameter, represents the bias term. Δ i represents the first semantic deviation.
[0135] Hint Fusion Module:
[0136] The control hint vector P i comes from the explicit input of the user (such as "using a formal style") or is extracted by the system from the original passage text (such as stylistic feature tags, term lists). After being encoded by the embedding layer, it is transformed into a vector p i with the same dimension as r i . The two are fused into the repair control vector c i , as shown in the following formula (10):
[0137] c i = tanh(W c [r i ; p i + b c ), (10)
[0138] where represents the repair control vector, which is subsequently injected into the intermediate layer of the repair model to guide the generation behavior during the repair process. [·;·] represents vector concatenation. represents the learnable weight matrix for fusing the hint vector and the semantic residual vector, represents the bias term during the fusion process.
[0139] Control injection generation module:
[0140] During translation repair, tokens with errors or ambiguities in the initial translation result Y i need to be replaced, while the problem-free parts can be retained. Therefore, a partial masking and rewriting strategy is adopted. First, based on the first attention matrix A i and the context-enhanced representation vector S i calculate which tokens have too low attention concentration (indicating insufficient reference to the context during generation and a risk of mistranslation), mask their positions to form the rewriting input ( to retain the credible content and mask the intermediate sequence at the position to be repaired). Then, the rewriting input the context-enhanced representation vector S i and the repair control vector c i are input to the decoder end of the repair model M, which performs context-aware rewriting on the masked part and finally outputs a complete and semantically more consistent translation repair result This process ensures that the model only precisely repairs the untrustworthy parts, avoiding style drift or term inconsistency problems caused by rewriting the entire sentence
[0141] It should be noted that at each intermediate layer during the repair process, the intermediate representation h t of the current decoding position t is corrected according to the following formula (11):
[0142] h t ' = h t + α · c i , (11)
[0143] where h t ' represents the corrected state with control information, h t represents the hidden state of the t-th token in the standard Transformer. α represents the injection strength of the repair control vector (empirically set to 0.5), and c i represents the repair control vector. Such a fusion realizes the unified effect of the semantic adjustment direction (from r i ) and the style guidance (from p i ), making the repaired translation consistent with the original text style in form and closer to the source text in semantics
[0144] In a patent scenario, for example, the sentence "This invention relates to a multilingual translation method" in the original text was mistranslated as "this invention talks about multilingual usage". By using the hint control to replace "talks about" with "relates to" and retaining the term "multilingual translation method", the professionalism and consistency of the repair can be demonstrated.
[0145] This step fully utilizes the first semantic offset Δ i , the initial translation result Y i , the first attention matrix A i , and combines the control hint vector P i to construct a semantic + style joint control repair mechanism. The innovation lies in embedding the translation residual signal as a modeling vector into the Transformer structure and fusing it with the hint for local correction and rewriting, which is especially suitable for professional translation tasks with high term accuracy and strong context correlation, such as technical specifications, legal terms, or patent claims, etc., to ensure that the translation quality meets the standards in both the dimensions of faithfulness and style guidance.
[0146] Steps S101 to S109 illustrated in the embodiments of this application obtain the target sentence group by obtaining the target sentence group; wherein, the target sentence group includes the current translation target sentence and the context window sentence. Semantic enhancement is performed on the target sentence group to obtain a context-enhanced representation vector. The context-enhanced representation vector is input into a preset initial translation model to obtain an initial translation result. The initial translation result is input into a preset back-translation model to obtain a first back-translated sentence. The average residual is calculated based on the current translation target sentence and the first back-translated sentence to obtain a first semantic deviation. The initial translation result is scored based on the first semantic deviation, the current translation target sentence, and the first back-translated sentence to obtain a first consistency scoring data. If the first consistency scoring data is lower than a preset scoring threshold, the control hint vector input by the user is obtained. A repair control vector is obtained based on the first semantic deviation and the control hint vector. The repair control vector is input into a preset repair model to obtain a translation repair result. It solves the core defects of existing multilingual translation systems in aspects such as context loss, result unverifiability, and output uncontrollability, realizes a deep coordination mechanism of context-aware discourse-level translation, consistency self-evaluation, and user hint control in one, and improves the accuracy of multilingual translation.
[0147] Please refer to Figure 4 , in some embodiments, the repair model also outputs a second attention matrix. After step S109, the multilingual deep translation method based on a large model may further include but is not limited to steps S401 to S405:
[0148] Step S401: Input the translation repair result into the back-translation model to obtain a second back-translated sentence;
[0149] Step S402: Encode the second back-translated sentence to obtain a second back-translated embedding sequence;
[0150] Step S403: Obtain a second alignment token pair of the target embedding sequence and the second back-translated embedding sequence according to the second attention matrix;
[0151] Step S404: Calculate the average residual according to the second alignment token pair, the target embedding sequence, and the second back-translated embedding sequence to obtain a second semantic deviation amount;
[0152] Step S405: Score the translation repair result according to the second semantic deviation amount, the current translation target sentence, and the second back-translated sentence to obtain second consistency scoring data.
[0153] In steps S401 to S402 of some embodiments, after repairing the initial translation result to obtain a more accurate translation repair result, it is also necessary to verify whether the translation repair result passes the consistency evaluation. Therefore, the translation repair result is input into the back-translation model to obtain a second back-translated sentence. The second back-translated sentence is encoded by the method in step S301 above to obtain a second back-translated embedding sequence.
[0154] In steps S403 to S405 of some embodiments, a second alignment token pair of the target embedding sequence and the second back-translated embedding sequence is obtained according to the second attention matrix by the methods in steps S302 to S303 above. Calculate the average residual according to the second alignment token pair, the target embedding sequence, and the second back-translated embedding sequence to obtain a second semantic deviation amount. Score the translation repair result according to the second semantic deviation amount, the current translation target sentence, and the second back-translated sentence to obtain second consistency scoring data. If the second consistency scoring data is greater than the scoring threshold, it means that the translation repair result passes the consistency evaluation. Otherwise, it means that there is a semantic deviation in the translation repair result, and it is necessary to re-obtain the control prompt vector input by the user for repair.
[0155] In steps S401 to S405 illustrated in this embodiment, after obtaining the second consistency scoring data, it is possible to judge whether to re-repair by judging whether the second consistency scoring data is greater than the scoring threshold, so as to ensure that an accurate translation repair result can be obtained.
[0156] Please refer to Figure 5 , in some embodiments, after step S405, the multi-language deep translation method based on the large model may further include but is not limited to steps S501 to S504:
[0157] Step S501, if the second consistency scoring data is greater than the scoring threshold, obtain the terms of the current translation target sentence to get the first terms, and obtain the terms of the translation repair result to get the second terms;
[0158] Step S502, match the first terms and the second terms to obtain a plurality of term key-value pairs;
[0159] Step S503, store the term key-value pairs into a preset term status cache;
[0160] Step S504, construct and store a pronoun mapping relationship according to the pronouns and corresponding entity phrases in the translation repair result.
[0161] In steps S501 and S503 of some embodiments, the translation objects that the present invention focuses on are usually professional texts at the chapter level and multi-paragraph structures, such as patent specifications, international standard articles, contract legal documents, etc. Terms in these texts are highly repetitive, subjects are frequently omitted, and syntactic structures highly depend on the previous text. Without a status cache mechanism, each sentence translation may introduce style mutations, term drifts, or even meaning misinterpretations due to the lack of context, seriously affecting the accuracy of multilingual output. Therefore, this step focuses on two goals: "explicit storage of term status" and "backtracking mapping of references", ensuring that the model behavior is controllable and the output style is stable.
[0162] Therefore, if the second consistency scoring data is greater than the scoring threshold, obtain the terms of the current translation target sentence to get the first terms, and obtain the terms of the translation repair result to get the second terms. Match the first terms and the second terms through a term alignment module to obtain a plurality of term pairs, and store them in the term status cache S c in the form of term key-value pairs. This cache is a structured dictionary-type data structure, and a persistent memory structure (such as Redis or SQLite embedded cache) can be used in actual deployment. The cache update process is shown in the following formula (12):
[0163]
[0164] Among them, represents the term status cache updated after the i-th sentence translation, including the newly recognized term pairs in the current sentence, represents the existing term status cache when the (i - 1)-th sentence translation is completed. represents the terms extracted from the current translation target sentence, represents the terms extracted from the translation repair result. represents the set of terms extracted from the current translation target sentence, usually obtained through dictionary matching and named entity recognition. Represents a set of terms extracted from the translation repair results, obtained through statistical or rule extraction (such as consecutive phrases with the first letter capitalized). Term matching is based on the attention matrix A of model M i For auxiliary decision-making, high-confidence term pairs are screened based on their alignment strength
[0165] In step S504 of some embodiments, to improve referential consistency, the system extracts pronouns (such as "it", "this module", "they", etc.) from the translation repair results and uses the above context cache to track the corresponding entity phrase r k to construct a pronoun mapping relationship as shown in the following formula (13):
[0166]
[0167] where p k represents a pronoun extracted from the translation repair results and r k represents the entity phrase corresponding to p obtained through dependency syntactic analysis and context term hits k represents the translation repair results represents when the j-th translation repair result before. The pronoun mapping relationship is stored in a FIFO queue structure, maintaining a certain context window length (such as 5 sentences) for reference when translating the next translation target sentence T i+1
[0168] In addition, to ensure the actual application effect of terms in future translations, the system will pass and to the context window construction module in step S101, so that the context window sentence C of the next sentence i+1 not only contains the original text context but also contains structured state information. This state "back-injection" mechanism ensures that the system has memory capabilities and does not generate from scratch for each sentence. For example, when translating the following Chinese paragraph:
[0169] "The present invention relates to an efficient training method for graph neural networks. This method is applicable to multilingual scenarios."
[0170] The system will identify the term "graph neural network" in the current translation target sentence T i and the corresponding translation in the target language is "graph neural network", and store it in S c , and maps "this method" to "the proposed method" to form a pronoun mapping. When processing the next sentence "It is particularly suitable for cross-border text analysis tasks", the system will preferentially use "the proposed method" to replace "it" and incorporate term consistency into the candidate generation sequence.
[0171] The core task of steps S501 to S504 shown in this embodiment is to convert the "practical" results completed in the translation process into a sustainable knowledge state of the system, thereby supporting the translation of subsequent sentences and achieving the paragraph-level terminology consistency and semantic coherence of the "multi-language long text translation system". In the entire solution, the previous steps solve the problems of "accurate translation" and "clear correction", while this step solves the problem of "remembering". Therefore, this step no longer involves modeling or abstract optimization, but directly acts on the semantic content level of the translation results, performing operational data structure updates so that there is a consistent context reference when continuing to process.
[0172] See also Figure 6 The present application also provides a large-model-based multilingual deep translation system that can implement the above-mentioned large-model-based multilingual deep translation method. The system includes:
[0173] Acquisition module 601 is used to acquire a target sentence group; wherein the target sentence group includes the current translation target sentence and the context window sentence;
[0174] An enhancement module 602 is configured to perform semantic enhancement based on the target sentence group to obtain a context-enhanced representation vector;
[0175] The translation module 603 is used to input the context-enhanced representation vector into a preset initial translation model to obtain an initial translation result;
[0176] A back-translation module 604 is configured to input the initial translation result into a preset back-translation model to obtain a first back-translated sentence;
[0177] A calculation module 605 is configured to calculate an average residual based on the current translation target sentence and the first back-translated sentence to obtain a first semantic deviation;
[0178] A scoring module 606 is configured to score the initial translation result according to the first semantic deviation, the current translation target sentence, and the first back-translated sentence to obtain first consistency score data;
[0179] A first input module 607 is configured to obtain a control prompt vector input by a user if the first consistency score data is lower than a preset score threshold;
[0180] A repair module 608 is configured to obtain a repair control vector according to the first semantic deviation and the control hint vector;
[0181] A second input module 609, configured to input a repair control vector into a preset repair model to obtain a translation repair result.
[0182] The specific implementation manner of the large model-based multilingual deep translation system is basically the same as the specific embodiments of the above-mentioned large model-based multilingual deep translation method, and will not be elaborated here.
[0183] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned large model-based multilingual deep translation method. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.
[0184] Please refer to Figure 7 , Figure 7 , which shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0185] A processor 701, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0186] A memory 702, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 702 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 702, and the processor 701 is used to call and execute the large model-based multilingual deep translation method of the embodiments of the present application;
[0187] An input / output interface 703, configured to implement information input and output;
[0188] A communication interface 704, configured to implement communication interaction between this device and other devices, and can implement communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0189] A bus 705, which transmits information between various components of the device (such as the processor 701, the memory 702, the input / output interface 703, and the communication interface 704);
[0190] Among them, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704 are communicatively connected to each other inside the device through the bus 705.
[0191] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned multi-language deep translation method based on a large model.
[0192] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0193] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0194] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.
[0195] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0196] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0197] In the description of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0198] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0199] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of systems or units can be in electrical, mechanical, or other forms.
[0200] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0201] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0202] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0203] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A multi - language deep translation method based on large models, characterized in that, The method includes: Obtaining a target sentence group; wherein, the target sentence group includes a current translation target sentence and context window sentences; Performing semantic enhancement on the target sentence group to obtain a context-enhanced representation vector; Inputting the context-enhanced representation vector into a preset initial translation model to obtain an initial translation result; Inputting the initial translation result into a preset back-translation model to obtain a first back-translated sentence; Calculating an average residual based on the current translation target sentence and the first back-translated sentence to obtain a first semantic deviation amount; Scoring the initial translation result based on the first semantic deviation amount, the current translation target sentence, and the first back-translated sentence to obtain first consistency scoring data; If the first consistency scoring data is lower than a preset scoring threshold, obtaining a control prompt vector input by a user; Obtaining a repair control vector based on the first semantic deviation amount and the control prompt vector; Inputting the repair control vector into a preset repair model to obtain a translation repair result.
2. The method according to claim 1, wherein The initial translation model also outputs a first attention matrix. The calculating an average residual based on the current translation target sentence and the first back-translated sentence to obtain a first semantic deviation amount includes: Encoding the current translation target sentence and the first back-translated sentence to obtain a target embedding sequence and a first back-translated embedding sequence; Obtaining a first alignment token pair of the target embedding sequence and the first back-translated embedding sequence according to the first attention matrix; Calculating an average residual according to the first alignment token pair, the target embedding sequence, and the first back-translated embedding sequence to obtain the first semantic deviation amount.
3. The method according to claim 2, wherein The calculating an average residual according to the first alignment token pair, the target embedding sequence, and the first back-translated embedding sequence to obtain the first semantic deviation amount includes: The average residual calculation formula is: Among them, Δ i represents the first semantic deviation amount, represents the first alignment token pair set, u j represents the target embedding sequence, represents the first back-translation embedding sequence.
4. The method according to claim 1, wherein The repair model also outputs a second attention matrix. After inputting the repair control vector into the preset repair model to obtain a translation repair result, it includes: Inputting the translation repair result into the back-translation model to obtain a second back-translated sentence; Encoding the second back-translated sentence to obtain a second back-translated embedding sequence; Obtaining a second alignment token pair of the target embedding sequence and the second back-translated embedding sequence according to the second attention matrix; Calculating an average residual according to the second alignment token pair, the target embedding sequence, and the second back-translated embedding sequence to obtain a second semantic deviation amount; Scoring the translation repair result based on the second semantic deviation amount, the current translation target sentence, and the second back-translated sentence to obtain second consistency scoring data.
5. The method according to claim 4, wherein After scoring the translation repair result based on the second semantic deviation amount, the current translation target sentence, and the second back-translated sentence to obtain second consistency scoring data, the method further includes: If the second consistency scoring data is greater than the scoring threshold, obtaining terms of the current translation target sentence to obtain first terms, and obtaining terms of the translation repair result to obtain second terms; Matching the first terms and the second terms to obtain multiple term key-value pairs; Store the term key-value pair in a preset term status cache; Construct and store a pronoun mapping relationship according to the pronouns and corresponding entity phrases in the translation repair result.
6. The method according to claim 1, wherein The semantic enhancement based on the target sentence group to obtain a context-enhanced representation vector includes: Obtain the structured representation vectors of each current translation target sentence and the context window sentence; wherein, the structured representation vector is the sum of the token embedding vector and the absolute position encoding; Concatenate each of the structured representation vectors to obtain an input sequence; Encode the input sequence to obtain a context-enhanced representation vector.
7. The method according to claim 1, characterized in that, The scoring of the initial translation result according to the first semantic deviation amount, the current translation target sentence, and the first back-translated sentence to obtain the first consistency scoring data includes: The calculation formula of the first consistency scoring data is: Score i = 1 - (Δ i + γ·δ i + μ·R i ); Among them, Score i represents the first consistency scoring data, Δ i represents the first semantic deviation amount, γ represents the term penalty coefficient, δ i represents the term matching deviation degree, represents the current translation target sentence, represents the set of terms successfully recovered in the first back-translated sentence, μ represents the weight controlling the impact of structure rearrangement, R i represents the sentence rearrangement penalty term.
8. A multilingual deep translation system based on a large model, characterized in that, The system includes: An acquisition module for acquiring a target sentence group; wherein, the target sentence group includes a current translation target sentence and a context window sentence; An enhancement module for performing semantic enhancement based on the target sentence group to obtain a context-enhanced representation vector; A translation module for inputting the context-enhanced representation vector into a preset initial translation model to obtain an initial translation result; A back-translation module for inputting the initial translation result into a preset back-translation model to obtain a first back-translated sentence; A calculation module for calculating the average residual according to the current translation target sentence and the first back-translated sentence to obtain a first semantic deviation amount; A scoring module for scoring the initial translation result according to the first semantic deviation amount, the current translation target sentence, and the first back-translated sentence to obtain the first consistency scoring data; A first input module for obtaining a control prompt vector input by the user if the first consistency scoring data is lower than a preset scoring threshold; A repair module for obtaining a repair control vector according to the first semantic deviation amount and the control prompt vector; A second input module for inputting the repair control vector into a preset repair model to obtain a translation repair result.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the large model-based multi-language deep translation method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the large model-based multi-language deep translation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Translation detection method and related equipment
CN117709365A
Machine translation quality estimation method based on multi-semantic space
CN119005214A
Prompt strategy activation-based multi-language large model translation performance optimization method and system
CN119903855A
Multilingual translation conversation system
JP2023058045A