Translation method and device of fusion term dictionary

By identifying and semantic encoding of terms in neural machine translation systems, combined with hard-constrained Beam Search and MGRC sampling technology, the inconsistency and error problems of existing systems in professional term translation are solved, and more accurate and safe translation results are achieved.

CN119940379APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510027133.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing neural machine translation systems have inconsistent or incorrect translation problems in handling texts in the field of expertise, especially term translation, leading to misunderstandings or legal and medical risks.

Method used

By setting identifiers to identify terms in input sentences, map input text to high-dimensional space using semantic encoders, and combine hard-constrained Beam Search strategy and MGRC sampling technology to ensure the accuracy and consistency of term translation.

Benefits of technology

It realizes the accurate identification and processing of professional terms during the translation process, ensures the professionalism and safety of translation results, and reduces the risks caused by inaccurate terms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940379A_ABST
    Figure CN119940379A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision, and particularly provides a translation method and device for a fusion term dictionary, and the method comprises the following steps: S1, setting a specific identifier to identify terms in an input sentence; s2, encoding an input text, and mapping the input text to a high-dimensional space; s3, decoding the semantic vector to generate a target language, and ensuring the accuracy of term translation through a Beam Search strategy combined with hard constraints; and S4, sampling the semantic vectors of the sentences of the source language by adopting an MGRC technology to generate a plurality of translation variants, and selecting an optimal variant as a result by calculating semantic similarity. Compared with the prior art, the method not only improves the professional property of translation, but also greatly reduces the risk caused by inaccurate terms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and specifically provides a translation method and device integrating a term dictionary. Background Art

[0002] As globalization deepens, the demand for cross-language information exchange in various industries has exploded, and the application scope of translation technology has gradually expanded to all aspects of society. From government departments, multinational companies to ordinary individuals, translation tools have become an indispensable bridge to meet the increasingly frequent cross-language communication needs. With the increase in translation needs, translation technology is also constantly innovating, from the initial rule-based translation system (RBMT), to statistical machine translation (SMT), and then to the current mainstream neural machine translation (Neural Machine Translation, NMT). Among them, the introduction of NMT technology, relying on the powerful performance of deep learning, has achieved a qualitative leap in translation quality.

[0003] Neural network-based translation systems are particularly good at capturing contextual information and maintaining the natural fluency of language. Through complex network structures and large amounts of data training, they are able to handle subtle differences between multiple languages ​​at the lexical and syntactic levels. This has enabled NMT to make significant progress in general text translation, greatly facilitating cross-language communication in daily use scenarios. However, with the widespread application of technology, people have gradually come to realize that NMT systems still have many limitations in processing texts in professional fields, especially in terminology translation.

[0004] In professional fields such as medicine, law, and engineering, the accuracy of terminology is crucial, and terminology is often highly specialized and standardized with fixed translations. Therefore, incorrect terminology translation may not only lead to misunderstandings, but may also bring legal liability or medical risks.

[0005] However, due to their strong generalization, existing NMT systems often fail to accurately identify or process complex contexts when faced with professional terminology, resulting in inconsistent or erroneous translation results. These problems are particularly inconvenient for industries with high professional requirements. Summary of the invention

[0006] The present invention aims at solving the above-mentioned deficiencies of the prior art and provides a translation method of a fusion term dictionary with strong practicality.

[0007] A further technical task of the present invention is to provide a translation device that integrates a terminology dictionary and is reasonably designed, safe and applicable.

[0008] The technical solution adopted by the present invention to solve its technical problem is:

[0009] A translation method integrating a term dictionary comprises the following steps:

[0010] S1. Setting identifiers to identify terms in the input sentence;

[0011] S2, encode the input text and map it to a high-dimensional space;

[0012] S3, decode the semantic vector to generate the target language and ensure the accuracy of term translation by combining the Beam Search strategy with hard constraints;

[0013] S4. The MGRC technique is used to sample the semantic vectors of the source language sentences to generate multiple translation variants, and the optimal variant is selected as the result by calculating the semantic similarity.

[0014] Furthermore, in step S1, in the sentence input in the source language, the terms are firstly identified and marked;

[0015] Assume that the source language sentence is represented by x = {x1, x2, ..., x T}, match through the term dictionary D, if the term x t In D src , the term is marked, and the sentence form after marking is:

[0016] x={x1,x2, <term> x t < / term> ,...,x T},ifx t in D src .;

[0017] Among them, D src Represents the set of source language terms in the dictionary. Term tags are used for dictionary constraints in the decoding stage and also provide a basis for subsequent term consistency checks.

[0018] Furthermore, in step S2, a sentence-level semantic encoder is used to encode the source language sentence as a whole. The semantic encoder f enc (x) Map the source language sentence x into a continuous semantic vector r x , expressed as:

[0019] r x =f enc (x)

[0020] The semantic encoder here captures the global context information by encoding the sentence as a whole and maps it to a high-dimensional semantic space, where the r generated by the semantic encoder is x It is the overall semantic representation of the source language sentence, which is not directly affected by the individual term tags, but these tags will be used in the subsequent decoding process for accurate translation of the terms.

[0021] Furthermore, in step S3, during the target language generation phase, the semantic vector r x Decode and combine the term dictionary for hard constraint translation. When the decoder generates the target language vocabulary y t When , the probability is calculated as follows:

[0022] P(y t |y <t ,rx)=softmax(W o h t )

[0023] Among them, h t is the hidden state of the decoder at the current moment, W o is the output layer weight matrix of the decoder. If the decoder detects that the current generation position corresponds to a term tag, it is forced to use the target language term translation in the term dictionary, that is:

[0024] y t =D tgt (x t ),ifx t is a marked term.

[0025] In this way, it is ensured that the translation of the term in the target language is consistent with the translation in the dictionary.

[0026] Furthermore, in order to further ensure the correctness of term translation, the BeamSearch strategy combined with hard constraints is introduced in the decoding stage. In each step of generation, Beam Search finds the optimal output sequence by evaluating multiple candidate translation paths. When encountering a marked term, the decoder is only allowed to select the appropriate term translation in the dictionary. The formula is expressed as:

[0027] B t ={y t ∈D tgt (x t )}

[0028] This way, the term part of the search path can only pick up predefined target translations.

[0029] Furthermore, in step S3, when translating, the semantic vector of the source language sentence is sampled by combining the mixed Gaussian recursive chain (MGRC) sampling technique to generate multiple different translation variants r x ′;

[0030] However, for the labeled term part, the semantic vector is kept unchanged to ensure that the translation of the term will not be affected by the sampling process. The specific sampling formula is:

[0031] r yi =r x +ω·(r y -r x )

[0032] in, is the ith translation variant, ω is the weight factor, r x and r y They are the semantic vectors of the source language and the target language respectively. For the marked term part, it does not participate in the semantic transformation to ensure the accuracy of the term translation.

[0033] Furthermore, a semantic similarity model is used to evaluate the semantic consistency between each translation variant and the source language sentence, and the variant with the closest semantics to the source sentence is selected:

[0034]

[0035] Among them, Sim is a similarity measure, which performs post-processing and consistency checking on term translation.

[0036] Furthermore, after the translation results are generated, the consistency is ensured by comparing the terms in the target language translation with the standard translation in the terminology dictionary;

[0037] If inconsistent terminology translation is found, the system will make corrections based on the dictionary to ensure that the final translation complies with the regulations for terminology translation. The post-processing formula is:

[0038]

[0039] Through this process, it is ensured that all terms in the output target language sentences conform to the specifications of the dictionary translation.

[0040] A translation device integrating a term dictionary, comprising: at least one memory and at least one processor;

[0041] The at least one memory is used to store a machine-readable program;

[0042] The at least one processor is used to call the machine-readable program to execute a translation method of a fusion term dictionary.

[0043] Compared with the prior art, the translation method and device of the fusion term dictionary of the present invention has the following outstanding beneficial effects:

[0044] The present invention ensures that professional terminology always follows predefined standards during the translation process by introducing a dedicated terminology dictionary. The terminology dictionary not only contains standard terminology comparisons between the source language and the target language, but also can intelligently identify and mark terms in the source text and give priority to these terms during the translation process.

[0045] This application can also ensure that the terms in the target language strictly match the definitions in the terminology dictionary during translation, which not only improves the professionalism of the translation, but also greatly reduces the risks caused by inaccurate terminology. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Attached Figure 1 It is a flowchart of a translation method integrating a terminology dictionary. DETAILED DESCRIPTION

[0048] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention is further described in detail below in conjunction with specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0049] A best embodiment is given below:

[0050] like Figure 1 As shown, a translation method integrating a term dictionary has the following steps:

[0051] S1. Setting specific identifiers to identify terms in the input sentence;

[0052] In the source language input sentence, the terms are first identified and labeled according to the predefined term dictionary. Assume that the source language sentence is represented as x = {x1, x2, ..., x T}, match through the term dictionary D, if the term x t In D src , the term is marked, and the sentence form after marking is:

[0053] x={x1,x2, <term> x t < / term> ,...,x T},ifx t in D src .

[0054] Among them, D srcRepresenting the set of source language terms in the dictionary, term tags are not only used for dictionary constraints in the decoding stage, but also provide a basis for subsequent term consistency checks.

[0055] S2, encode the input text and map it to a high-dimensional space;

[0056] The sentence-level semantic encoder is used to encode the source language sentence as a whole. enc (x) Map the source language sentence x into a continuous semantic vector r x , expressed as:

[0057] r x =f enc (x)

[0058] The semantic encoder here captures global context information by encoding the sentence as a whole and maps it to a high-dimensional semantic space. Unlike word-level encoding, sentence-level encoding can better maintain semantic consistency and stability.

[0059] Among them, the r generated by the semantic encoder x It is the overall semantic representation of the source language sentence, which is not directly affected by the individual term tags, but these tags will be used in the subsequent decoding process for accurate translation of the terms.

[0060] S3, decode the semantic vector to generate the target language and ensure the accuracy of term translation by combining the Beam Search strategy with hard constraints;

[0061] In the target language generation stage, the semantic vector r x Decode and combine the term dictionary for hard-constrained translation. When the decoder generates the target language vocabulary y t When , the probability is calculated as follows:

[0062] P(y t |y <t ,r x )=softmax(W o h t )

[0063] Among them, h t is the hidden state of the decoder at the current moment, W o is the output layer weight matrix of the decoder. If the decoder detects that the current generation position corresponds to a term token, it is forced to use the target language term translation in the term dictionary, that is:

[0064] y t =D tgt (x t ),ifx tis a marked term.

[0065] In this way, it is ensured that the translation of the term in the target language is consistent with the translation in the dictionary.

[0066] To further ensure the correctness of term translation, a beam search strategy combined with hard constraints is introduced in the decoding stage. In each step of generation, beam search finds the optimal output sequence by evaluating multiple candidate translation paths. When encountering a marked term, the decoder only allows the selection of appropriate term translations in the dictionary, which is expressed as:

[0067] B t ={y t ∈D tgt (x t )}

[0068] In this way, the term part in the search path can only select predefined target translations, thus ensuring the consistency of professional terminology.

[0069] S4, using MGRC technology to sample the semantic vector of the source language sentence to generate multiple translation variants, and selecting the optimal variant as the result by calculating the semantic similarity;

[0070] In order to enhance the diversity of the generative model during translation, the present invention combines the MGRC (mixed Gaussian recursive chain) sampling technology to sample the semantic vector of the source language sentence and generate multiple different translation variants. x ′. However, for the labeled term part, its semantic vector is kept unchanged to ensure that the translation of the term will not be affected by the sampling process. The specific sampling formula is:

[0071] r yi =r x +ω·(r y -r x )

[0072] in, is the ith translation variant, r x and r y They are the semantic vectors of the source language and the target language respectively. For the labeled term part, it does not participate in the semantic transformation to ensure the accuracy of the term translation.

[0073] We further use a semantic similarity model (such as Sentence-BERT) to evaluate the semantic consistency between each translation variant and the source language sentence, and select the variant that is semantically closest to the source sentence:

[0074]

[0075] Where Sim is a similarity measure (such as cosine similarity). Post-processing and consistency checking of term translations are performed.

[0076] After the translation results are generated, the terminology in the target language translation is compared with the standard translation in the terminology dictionary to ensure its consistency. If inconsistent terminology translation is found, the system will correct it according to the dictionary to ensure that the final translation meets the regulations of terminology translation. The post-processing formula is:

[0077]

[0078] Through this process, it is ensured that all terms in the output target language sentences conform to the specifications of the dictionary translation.

[0079] Based on the above method, a translation device integrating a term dictionary in this embodiment includes: at least one memory and at least one processor;

[0080] The at least one memory is used to store a machine-readable program;

[0081] The at least one processor is used to call the machine-readable program to execute a translation method of a fusion term dictionary.

[0082] The above-mentioned specific implementations are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above-mentioned specific implementations. Any technical solutions that conform to the above-mentioned specific implementations of the present invention and any appropriate changes or substitutions made by ordinary technicians in the relevant technical field shall fall within the patent protection scope of the present invention.

[0083] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A translation method integrating a term dictionary, characterized in that: The steps are as follows: S1. Setting identifiers to identify terms in the input sentence; S2, encode the input text and map it to a high-dimensional space; S3, decode the semantic vector to generate the target language and ensure the accuracy of term translation by combining the Beam Search strategy with hard constraints; S4. The MGRC technique is used to sample the semantic vectors of the source language sentences to generate multiple translation variants, and the optimal variant is selected as the result by calculating the semantic similarity.

2. The method for translating a fusion term dictionary according to claim 1, characterized in that: In step S1, in the sentence input in the source language, the terms are firstly identified and marked; Assume that the source language sentence is represented by x = {x1, x2, ..., x T }, match through the term dictionary D, if the term x t In D src , the term is marked, and the sentence form after marking is: x={x1,x2, <term> x t < / term> ,...,x T },ifx t in D src .; Among them, D src Represents the set of source language terms in the dictionary. Term tags are used for dictionary constraints in the decoding stage and also provide a basis for subsequent term consistency checks.

3. A method for translating a fusion term dictionary according to claim 2, characterized in that: In step S2, a sentence-level semantic encoder is used to encode the source language sentence as a whole. The semantic encoder f enc (x) Map the source language sentence x into a continuous semantic vector r x , expressed as: r x =f enc (x) The semantic encoder here captures the global context information by encoding the sentence as a whole and maps it to a high-dimensional semantic space, where the r generated by the semantic encoder is x It is the overall semantic representation of the source language sentence, which is not directly affected by the individual term tags, but these tags will be used in the subsequent decoding process for accurate translation of the terms.

4. The method for translating a fusion term dictionary according to claim 3, characterized in that: In step S3, during the target language generation phase, the semantic vector r x Decode and combine the term dictionary for hard constraint translation. When the decoder generates the target language vocabulary y t When , the probability is calculated as follows: P(y t |y <t ,r x )=softmax(W o h t ) Among them, h t is the hidden state of the decoder at the current moment, W o is the output layer weight matrix of the decoder. If the decoder detects that the current generation position corresponds to a term tag, it is forced to use the target language term translation in the term dictionary, that is: y t =D tgt (x t ),ifx t is a marked term. In this way, it is ensured that the translation of the term in the target language is consistent with the translation in the dictionary.

5. The method for translating a fusion term dictionary according to claim 4, characterized in that: To further ensure the correctness of term translation, a Beam Search strategy combined with hard constraints is introduced in the decoding stage. In each generation step, BeamSearch evaluates multiple candidate translation paths to find the optimal output sequence. When encountering a marked term, the decoder is only allowed to select a suitable term translation in the dictionary. The formula is expressed as: B t ={y t ∈D tgt (x t )} This way, the term part of the search path can only pick up predefined target translations.

6. A method for translating a fusion term dictionary according to claim 5, characterized in that: In step S3, when translating, the semantic vector of the source language sentence is sampled by combining the mixed Gaussian recursive chain (MGRC) sampling technique to generate multiple different translation variants r′. x ; However, for the labeled term part, the semantic vector is kept unchanged to ensure that the translation of the term will not be affected by the sampling process. The specific sampling formula is: Among them, r yi is the ith translation variant, ω is the weight factor, r x and r y They are the semantic vectors of the source language and the target language respectively. For the marked term part, it does not participate in the semantic transformation to ensure the accuracy of the term translation.

7. A method for translating a fusion term dictionary according to claim 6, characterized in that: Use the semantic similarity model to evaluate the semantic consistency between each translation variant and the source language sentence, and select the variant that is semantically closest to the source sentence: Among them, Sim is a similarity measure, which performs post-processing and consistency checking on term translation.

8. The method for translating a fusion term dictionary according to claim 7, characterized in that: After the translation results are generated, the terms in the target language translation are compared with the standard translation in the terminology dictionary to ensure their consistency; If inconsistent terminology translation is found, the system will make corrections based on the dictionary to ensure that the final translation complies with the regulations for terminology translation. The post-processing formula is: Through this process, it is ensured that all terms in the output target language sentences conform to the specifications of the dictionary translation.

9. A translation device integrating a term dictionary, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Multi-language text adaptive configuration method and electronic equipment

    CN121031624A