A medical machine translation method based on keywords of translated sentences
By combining professional thesaurus and manual annotation of keywords in the neural machine translation model, using the multi-head self-attention mechanism and masking mechanism, the missed translation problem in the professional field is solved, and the translation accuracy and loyalty is improved. It is suitable for English-Chinese medical translation.
Patent Information
- Application Number
- CN202010856737.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-24
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-08-24
AI Technical Summary
The existing neural machine translation systems have mistranslation and mistranslation in professional fields such as the medical field, especially when translating professional vocabulary, the existing methods are complex and increase model calculation overhead.
By combining professional thesaurus and manual methods to label keywords in the machine translation model, the multi-head self-attention mechanism and masking mechanism are used to enhance the loss function of untranslated keywords and improve the missed translation phenomenon of the translation system, which is especially suitable for translation in professional fields.
It improves translation accuracy in professional fields, reduces missed translation situations, enhances the loyalty and accuracy of neural machine translation, and does not increase model calculation overhead.
Smart Images

Figure CN114091481B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine translation, and particularly to a medical machine translation method based on keywords of translated sentences for improving the problem of missing translation in neural machine translation. Background Art
[0002] Machine translation, also known as automatic translation, is a process of using a computer to convert one natural language into another natural language. With the significant improvement of computer software and hardware performance, neural machine translation models based on large deep neural networks have been trained. Coupled with a large number of high-quality training data sets, the performance of neural machine translation (NMT) models far exceeds that of traditional statistic-based translation models. Nowadays, well-known translation platforms on the market, such as Google, Baidu, and Youdao translation platforms, have long switched their translation systems from statistic-based methods to deep neural network-based methods.
[0003] However, current machine translation systems based on deep neural networks still have many problems and need to be further improved to a great extent. On the one hand, the high translation quality of current models is mainly reflected in general fields, such as daily spoken language, daily news, etc. However, in some special fields (such as the medical field), the translation results are often unsatisfactory. Especially when translating some professional terms, there are relatively serious problems of missing translation and wrong translation in the translation system, and the translation results are unacceptable to users. On the other hand, the problem of missing translation has existed in the field of neural machine translation for a long time and is one of the difficult problems that researchers focus on and study. The so-called missing translation means that in the process of translation, the translation system fails to translate some important words (such as medical professional terms) in the source sentence or translates them incorrectly. This phenomenon is particularly significant when the sentence sequence is long.
[0004] The problem of missing translation has always been a key research issue in the field of machine translation. Past research and corresponding published literature have also proposed some targeted improvement schemes and feasible methods, and there has been a certain amount of technical accumulation. For example, models with maximum likelihood estimation (MLE) as the training objective tend to obtain shorter translation results. To overcome this problem, Wu et al. [Wei He, Zhongjun He, Hua Wu, and Haifeng Wang. Improved neural machine translation with smt features. In Proceedings of AAAI 2016, 2016] proposed a length normalization method. This method can make translated sentences of different lengths compete as fairly as possible, and thereby eliminate the preference of the NMT model for short sentences.
[0005] The Chinese invention patent application document with a publication date of April 21, 2020 and a publication number of CN111046649A discloses a text segmentation method and device. By using specific delimiters to gradually segment a large amount of input text, relatively complete small pieces of text with preserved semantics are obtained. This method can prevent problems such as overly large text length or semantic inversion during the machine translation process, and can effectively improve the phenomenon of missing translations. The Chinese invention patent application document with a publication date of November 15, 2019 and a publication number of CN110457713A discloses a translation method, device, equipment and storage medium based on a machine translation model. In this solution, the i-th source word of the source sentence is first embedded and encoded into an intermediate vector; then the intermediate vector is decoded in combination with the i-th source word embedding to obtain a decoded intermediate vector; then the decoded intermediate vector is fused with the (i - 1)-th target word embedding of the target sentence to obtain a fused intermediate vector; finally, the probability of the i-th target word is predicted based on the decoded word vector to obtain the result. The Chinese invention patent application document with a publication date of October 15, 2019 and a publication number of CN110334362A discloses a method for solving untranslated words based on medical neural machine translation. The implementation principle of this solution is that if untranslated words appear in the translation result <unk>, then calculate through the attention mechanism <unk>At the corresponding position in the source language, then use a medical professional dictionary to query the word and replace it in the translation result <unk>And return it to the user. The Chinese invention patent application document with the publication date of April 20, 2018 and the publication number of CN107943795A discloses a method for improving the accuracy of neural machine translation, a translation method, system and device. This solution introduces the pre-reordering method in statistical machine learning into neural machine translation, greatly alleviating the problems of missing translation and repeated translation. In addition, a coverage vector is added to the attention layer of neural machine translation to further alleviate the problems of missing translation and repeated translation. Among them, the coverage vector is mainly used to measure the degree of reference to each word in the source sentence during decoding.
[0006] In the above method, although the length normalization method can eliminate the preference of the NMT model for short sentences, this method itself is actually coverage-unaware of the translation content. At the same time, some methods also try to use the coverage vector in the attention layer of the neural translation model to measure whether a specific word at the source end has been translated, and the obtained coverage vector can affect and adjust the attention model at subsequent decoding moments. However, this type of method will complicate the model structure and increase the model parameters to a certain extent, and it is not as simple as optimizing the model by comparing the differences between the translation results and the target translation sentences. In summary, there are still certain defects in current machine translation, and further improvement is needed to improve the translation accuracy in professional fields and reduce the situation of missing translation. Summary of the Invention
[0007] The purpose of the present invention is to provide a medical machine translation method based on keywords in the translation sentence. When determining the keywords in the translation sentence, the method of combining the professional thesaurus search method with the manual method can improve the effectiveness and accuracy of the entire data preprocessing. During the training process of the model, without increasing the training model parameters and the model calculation overhead, in the case of the translation bilingual sequence plus the keyword mode of the translation sentence, by increasing the loss corresponding to the untranslated keywords, the missing translation phenomenon of the translation system is improved. It is especially suitable for translation in professional fields, generally improving the translation accuracy in professional fields and reducing the situation of missing translation, and can effectively improve the loyalty and accuracy of neural machine translation.
[0008] The technical solution of the present invention is as follows:
[0009] A medical machine translation method based on keywords in the translation sentence, characterized in that the steps are as follows:
[0010] (1) In the machine translation model, after the training corpus undergoes data preprocessing such as truncation, padding, and keyword annotation, a source sentence sequence, a translation sentence sequence, a translation sentence target word sequence, and a keyword annotation vector K are obtained; among them, the keyword annotation vector K is used to identify whether a certain word in the translation sentence sequence is a keyword, with the keyword annotated as 1 and the non-keyword annotated as 0;
[0011] (2) The encoder based on multi-head self-attention encodes the source sentence sequence to obtain the source sentence encoding information;
[0012] (3) Input the translated sentence sequence and the source sentence encoding information transmitted from the encoder into the decoder, and use the masked multi-head self-attention mechanism to obtain the self-attention output vector of the translated sentence with respect to the source sentence information;
[0013] (4) After the self-attention output vector is processed by a fully connected layer and softmax, the model prediction probability distribution P of the target word of the translated sentence is obtained, and the translation result is obtained from the model prediction probability distribution P by sampling with the maximum probability;
[0014] (5) Further, according to the sampled translation result and the translated sentence target word sequence, the translation degree vector C is calculated;
[0015] (6) Finally, combining the translated sentence target word sequence, the keyword annotation vector K, the translation degree vector C, and the model prediction probability distribution P, by increasing the loss of the keywords not appearing in the translation result during the training of the model, the ability of the untranslated keywords to adjust the parameters of the model is enhanced, so that the training model tends to translate out the keyword.
[0016] When the above translation method is used in the field of medical English-Chinese translation, the specific implementation steps are as follows:
[0017] S1. Preprocess the English-Chinese medical bilingual corpus
[0018] 1.1 Segment the Chinese sequence based on the proprietary medical word library and word segmentation technology; segment the English sequence using the BPE (Byte Pair Encoder) word segmentation technology;
[0019] 1.2 Construct a medical professional vocabulary list based on English-Chinese bilingual;
[0020] 1.3 Process the English-Chinese bilingual into sequences of fixed length to obtain English and Chinese sentence pairs that are mutual translated sentences, where the lengths of the Chinese and English sequences are set to appropriate values respectively;
[0021] 1.4 Keyword annotation of the target sentence, where the target sentence is Chinese, and the Chinese sentence sequence is the target sentence sequence.
[0022] In the above steps 1.1 - 1.4, all medical professional terms belong to keywords and are obtained by searching the medical professional vocabulary list; the target sentence sequence screened by querying the medical professional vocabulary list is subjected to keyword extraction again by medical researchers; the union of the keywords extracted twice is the final keyword set of the target sentence sequence, and the position of each keyword in the target sentence sequence is recorded, and a keyword annotation vector K indicating whether it is a keyword is obtained. Among them, the position corresponding to the keyword is marked as 1, and the position corresponding to the non-keyword is marked as 0;
[0023] 1.5 Use the word2Vec word vector technology to process the bilingual corpus respectively to obtain the initial word vector matrix corresponding to each training corpus and the vocabulary list for training the model.
[0024] S2. Encode source sentence information
[0025] 2.1 According to the constructed vocabulary list and initial word vector matrix, obtain the word vector of each word in the source sentence sequence;
[0026] 2.2 Calculate the position vector of each word in the source sentence sequence;
[0027] 2.3 According to steps 2.1 and 2.2, calculate the sum of the word vector and the position vector corresponding to each word in the source sentence sequence to obtain the source sentence input matrix of the model encoder;
[0028] 2.4 The encoder uses the multi-head self-attention mechanism to encode the source sentence input matrix to obtain the encoding result of the source sentence sequence.
[0029] S3. Decoder execution process
[0030] 3.1 According to the vocabulary list and initial word vector matrix for training the model obtained in 1.5, obtain the word vector of each word in the translation sentence sequence;
[0031] 3.2 Obtain the position vector of each word in the translation sentence sequence;
[0032] 3.3 Obtain the sum of the word vector and the position vector corresponding to each word in the translation sentence sequence to obtain the target translation sentence input matrix corresponding to the translation sentence sequence;
[0033] 3.4 Use the masked multi-head self-attention mechanism to encode the translation sentence input matrix to obtain the encoding result of the translation sentence;
[0034] 3.5 Use the encoding result of the translation sentence to perform multi-head self-attention processing on the source sentence encoding result in the encoder to obtain the self-attention output vector of the decoder;
[0035] 3.6 After the self-attention output vector of the decoder is processed by the fully connected neural network, softmax normalization processing is performed to obtain the predicted probability distribution P of the model decoder at each position;
[0036] 3.7 Calculate the translation degree vector C, where the elements in the translation degree vector C represent the translation degree of each target word; the calculation process is as follows: Determine whether the word corresponding to the maximum probability value in the predicted distribution at each position is the target word; if it is the target word, the translation degree of this keyword is 1; if it is not the target word, it is a numerical value or the calculation result of a function that can represent the translation degree of the target word.
[0037] In the above method, the objective function of the machine translation model is as shown in Equation (1):
[0038]
[0039] Among them, ξ is a coefficient reflecting the overall tuning strength of the keyword in the model, which is composed of hyperparameters or functions; oriLoss() represents the traditional loss function; P i and Y i lab respectively represent the predicted distribution and the target word corresponding to the i-th position of the decoder; L represents the length of the translated sentence sequence; K i represents the element value corresponding to the i-th position of the keyword annotation vector K; C i represents the element value corresponding to the i-th position of the keyword translation degree C.
[0040] The beneficial effects of the present invention are as follows:
[0041] (1) When determining the keywords of the translated sentence in the present invention, the professional thesaurus search method is combined with the manual method, and the entire data preprocessing process is efficient and highly accurate.
[0042] (2) The model of the present invention increases the coefficient of the loss corresponding to the untranslated keyword. During the training process of the model, no additional training model parameters are added, and the model calculation overhead is not increased.
[0043] (3) Compared with the traditional method starting from the source sentence, the present invention starts from the translated sentence level and is directly based on the translation result, which is completely different from the traditional method.
[0044] (4) In the case of the translation bilingual sequence plus translated sentence keyword mode, the idea of improving the missing translation phenomenon of the translation system by increasing the loss corresponding to the untranslated keyword.
[0045] (5) The present invention has strong scalability and is applicable to all machine translation tasks based on the translation bilingual sequence plus translated sentence keyword mode, without language category restrictions, and is applicable to translation between all bilingual languages, such as English-German translation, German-Chinese translation, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a flowchart of the translation training model of the present invention. Detailed Implementation Manner
[0047] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the present invention will be described below in conjunction with embodiments.
[0048] Embodiment 1
[0049] Among the existing methods for improving the problem of missing translations in medical machine translation, they mainly focus on iterative translation of splitting long sentences into multiple short sentences and methods based on coverage. Moreover, the existing methods all have problems such as complex preliminary work and large model calculation overhead, and such methods only strive to improve the translation model from the perspective of the source sentence.
[0050] Compared with the traditional method, this embodiment provides a medical machine translation method based on keywords in the translated sentence. The improvement method based on the translated sentence can simply and effectively improve the phenomenon of missing translations in the translation system. As Figure 1 shown, in the model of this method:
[0051] (1) In the machine translation model, after the training corpus undergoes data preprocessing, a source sentence sequence, a translated sentence sequence, a translated sentence target word sequence, and a keyword annotation vector K are obtained;
[0052] (2) The encoder based on multi-head self-attention encodes the source sentence sequence to obtain source sentence encoding information;
[0053] (3) Input the translated sentence sequence and the source sentence encoding information transmitted from the encoder into the decoder, and use the masked-multi-head attention mechanism to obtain the self-attention output vector of the translated sentence with respect to the source sentence information;
[0054] (4) After the self-attention output vector is processed by a fully connected layer and softmax, the model prediction probability distribution P for the translated sentence target word is obtained. According to the model prediction probability distribution P, maximum probability sampling is performed to obtain the translation result;
[0055] (5) Further, according to the sampled translation result and the translated sentence target word sequence, the translated degree vector C is calculated;
[0056] (6) Finally, by combining the translated sentence target word sequence, the keyword annotation vector K, the translated degree vector C, and the model prediction probability distribution P, the loss of the keywords not appearing in the translation result during training the model is increased, enhancing the ability of the untranslated keywords to adjust the model parameters, so that the model tends to translate out the keyword.
[0057] When performing translation in the field of medical English-Chinese translation according to the above translation method, the specific model is as follows:
[0058] S1. Preprocess the English-Chinese medical bilingual corpus
[0059] 1.1 Segment the Chinese sequence based on a proprietary medical vocabulary and word segmentation technology; segment the English sequence using BPE (Byte Pair Encoder) word segmentation technology;
[0060] 1.2 Construct a medical professional vocabulary list based on English-Chinese bilingualism;
[0061] 1.3 Process the English-Chinese bilingual into sequences of fixed length to obtain English and Chinese sentence pairs that are mutual translation sentences, where the lengths of the Chinese and English sequences are set to appropriate values respectively;
[0062] 1.4 Label the keywords of the target sentence, where the target sentence is in Chinese, and the Chinese sentence sequence is the target sentence sequence;
[0063] 1.5 Use the word2Vec word vector technology to process the bilingual corpus respectively to obtain the initial word vector matrix corresponding to each training corpus and the vocabulary list for training the model.
[0064] S2. Encode the source sentence information
[0065] 2.1 Obtain the word vectors of each word in the source sentence sequence according to the constructed vocabulary list and initial word vector matrix;
[0066] 2.2 Calculate the position vectors of each word in the source sentence sequence;
[0067] 2.3 According to steps 2.1 and 2.2, calculate the sum of the word vectors and position vectors corresponding to each word in the source sentence sequence to obtain the source sentence input matrix of the model encoder;
[0068] 2.4 The encoder uses the multi-head self-attention mechanism to encode the source sentence input matrix to obtain the encoding result of the source sentence sequence.
[0069] S3. Decoder execution process
[0070] 3.1 Obtain the word vectors of each word (including all words in the training corpus) in the translation sentence sequence according to the vocabulary list and initial word vector matrix for training the model obtained in 1.5;
[0071] 3.2 Obtain the position vectors of each word in the translation sentence sequence;
[0072] 3.3 Obtain the sum of the word vectors and position vectors corresponding to each word in the translation sentence sequence to obtain the translation sentence input matrix corresponding to the translation sentence sequence;
[0073] 3.4 Use the masked multi-head self-attention mechanism to encode the target translation sentence input matrix to obtain the encoding result of the translation sentence;
[0074] 3.5 Use the encoded result of the translated sentence to perform multi-head self-attention processing on the encoded result of the source sentence in the encoder to obtain the self-attention output vector of the decoder;
[0075] 3.6 After the self-attention output vector of the decoder is processed by a fully connected neural network, softmax normalization processing is performed to obtain the predicted probability distribution P of each position of the decoder of the model;
[0076] 3.7 Calculate the translation degree vector C, and the elements in the translation degree vector C represent the translation degree of each target word; the calculation process is: judge whether the word corresponding to the maximum probability value in each position prediction distribution is the target word; if it is the target word, the translation degree of this keyword is 1; if it is not the target word, it is the numerical value or function calculation result that can represent the translation degree of the target word.
[0077] In the above method, the objective function of the machine translation model is as shown in Equation (1):
[0078]
[0079] Among them, ξ is a coefficient reflecting the overall tuning strength of the keyword in the model, which is composed of hyperparameters or functions; oriLoss() represents the traditional loss function; P i and Y i lab respectively represent the predicted distribution and the target word corresponding to the i-th position of the decoder; L represents the length of the translated sentence sequence; K i represents the element value corresponding to the i-th position of the keyword annotation vector K; C i represents the element value corresponding to the i-th position of the keyword translation degree C.
[0080] In this embodiment, the word2Vec word embedding technology is adopted, but not limited to, to obtain the initial word vector matrix of the model. In the present invention, any form of definition and instantiation of word vectors are within the protection scope of the present invention.
[0081] In this embodiment, the transformer is adopted, but not limited to, to encode and decode the translation bilingual. For example, methods of encoding the bilingual using models such as RNN (including LSTM, GRU), attention-based RNN, R-CNN, Bert, XLNet, and gpt and their variants are all within the protection scope of the present invention.
[0082] In this embodiment, methods based on changing the number of layers of the neural network, the number of multi-heads, and the number of each module inside the model are all within the protection scope of the present invention.
[0083] Embodiment 2
[0084] Based on the medical machine translation method using translated sentence keywords provided in Embodiment 1, experiments were conducted during the process of training an English-Chinese translation model. Among them, the corpora were all from the medical field and were English-Chinese bilingual parallel corpora. Among them, the training set contained approximately 190,000 English-Chinese sentence pairs, the validation set had 2,000 pairs, and each of the five test sets (Test01 - Test05) had 2,000 pairs.
[0085] The mainstream evaluation metric BLEU was used to evaluate the overall translation quality of the model, and the translation results of the present invention were compared with those of existing methods. The experimental results are shown in Table 1. Among the five groups of BLEU evaluation results, the average scores of the method proposed in the present invention were all higher than those of the existing methods, and the average value (AVE) exceeded 1.1, indicating that this method helps to improve the overall translation quality of the model.
[0086] Table 1
[0087] Test01 Test02 Test03 Test04 Test05 AVE Existing method (transformer) 28.20 28.96 27.56 29.06 27.21 28.20 This method 29.01 30.23 28.72 30.51 28.22 29.34
[0088] To explore whether the present invention can effectively improve the problem of missing translations in the model, after combining the five test sets, 1,000 pairs of English-Chinese sentence sequences were randomly selected as the test set for evaluating missing translations. Finally, by comparing the missing translation situations of the present invention and the existing model, the advantages and disadvantages of the model performance were compared, and thus the effectiveness of the present invention was verified. The experimental results are shown in Table 2. Compared with the existing method, the overall number of missing translations in the present invention decreased by (3250 - 2496) / 3250 = 0.232, and the keyword missing translation rate decreased by (1377 - 629) / 1377 ≈ 0.543. This shows that this method can effectively reduce the number of missing translations, especially the reduction in the number of keyword missing translations is more obvious.
[0089] Table 2
[0090] Total number of missed translations Number of missed translations of keywords Existing method (transformer) 3250 1377 This method 2496 629 < / unk> < / unk> < / unk>
Claims
1. A medical machine translation method based on keywords of translated sentences, characterized in that, The steps are as follows: (1) In the machine translation model, after the training corpus is preprocessed, a source sentence sequence, a target sentence sequence, a target word sequence of the target sentence, and a keyword annotation vector K are obtained; among them, the keyword annotation vector K is used to identify whether a word in the target sentence sequence is a keyword, with the keyword annotated as 1 and the non-keyword annotated as 0; (2) The source sentence sequence is encoded by an encoder based on multi-head self-attention to obtain source sentence encoding information; (3) The target sentence sequence and the source sentence encoding information transmitted from the encoder are input into the decoder to obtain a self-attention output vector of the target sentence with respect to the source sentence information; (4) After the self-attention output vector is processed by a fully connected layer and softmax, a model prediction probability distribution P for the target word of the target sentence is obtained, and the translation result is obtained by performing maximum probability sampling according to the model prediction probability distribution P; (5) Further, according to the sampled translation result and the target word sequence of the target sentence, a translation degree vector C is calculated; (6) Finally, by combining the target word sequence of the target sentence, the keyword annotation vector K, the translation degree vector C, and the model prediction probability distribution P, the loss of the keywords that do not appear in the translation result during training the model is increased to enhance the ability of the untranslated keywords to adjust the model parameters, so that the model tends to translate out the keyword; The objective function of the machine translation model is shown in Equation (1): (1) Among them, ξ is a coefficient reflecting the overall tuning strength of keywords in the model, which is composed of hyperparameters or functions; oriLoss () represents the traditional loss function; P i and Y i lab respectively represent the predicted distribution and the target word corresponding to the i-th position of the decoder; L represents the length of the translated sentence sequence; K i represents the corresponding element value of the i-th position in the keyword annotation vector K in; C i represents the translated degree vector C and the corresponding element value of the i-th position.
2. The medical machine translation method based on translation sentence keywords according to claim 1, wherein: The translation degree vector C is calculated through the probability value of the target word in the prediction probability distribution P.
3. The medical machine translation method based on translated sentence keywords according to claim 1, wherein: In step (3), a self-attention output vector of the decoder is obtained by using a masked multi-head self-attention mechanism.
4. The medical machine translation method based on translated sentence keywords according to claim 1, wherein: The medical machine translation method is at least applicable to the fields of medical English-Chinese translation, medical English-German translation, and medical German-Chinese translation.
5. The medical machine translation method based on translated sentence keywords according to claim 4, wherein When it is used in the field of medical English-Chinese translation, the specific implementation steps are as follows: S1. Preprocess the English-Chinese medical bilingual corpus 1.1 Segment the Chinese sequence based on a proprietary medical vocabulary and word segmentation technology; segment the English sequence using the BPE word segmentation technology; 1.2 Construct a medical professional vocabulary list based on English-Chinese bilingual; 1.3 Process the English-Chinese bilingual into sequences of fixed length to obtain English and Chinese sentence pairs that are mutual target sentences, where the lengths of the Chinese and English sequences are set to appropriate values respectively; 1.4 Keyword annotation of the target sentence, where the target sentence is Chinese, and the Chinese sentence sequence is the target sentence sequence; 1.5 Use the word2Vec word vector technology to process the bilingual corpus respectively to obtain the initial word vector matrix corresponding to each training corpus and the vocabulary list for training the model; S2. Encode the source sentence information 2.1 Obtain the word vector of each word in the source sentence sequence according to the constructed vocabulary list and initial word vector matrix; 2.2 Calculate the position vector of each word in the source sentence sequence; 2.3 According to steps 2.1 and 2.2, calculate the sum of the word vector and the position vector corresponding to each word in the source sentence sequence to obtain the source sentence input matrix of the model encoder; 2.4 The encoder encodes the source sentence input matrix using a multi-head self-attention mechanism to obtain the encoding result of the source sentence sequence; S3. Decoder execution process 3.1 Obtain the word vector of each word in the target sentence sequence according to the vocabulary list for training the model and the initial word vector matrix obtained in 1.5; 3.2 Obtain the position vectors of each word in the translated sentence sequence; 3.3 Obtain the sum of the word vectors and position vectors corresponding to each word in the translated sentence sequence to obtain the translated sentence input matrix corresponding to the translated sentence sequence; 3.4 Encode the translated sentence input matrix using a masked multi-head self-attention mechanism to obtain the encoded result of the translated sentence; 3.5 Perform multi-head self-attention processing on the encoded result of the source sentence in the encoder using the encoded result of the translated sentence to obtain the self-attention output vector of the decoder; 3.6 After the self-attention output vector of the decoder is processed by a fully connected neural network, perform softmax normalization to obtain the predicted probability distribution P of the decoder of the model at each position; 3.7 Calculate the translation degree vector C, and the elements in the translation degree vector C represent the translation degree of each target word; The calculation process is as follows: Determine whether the word corresponding to the maximum probability value in each position prediction distribution is the target word; if it is the target word, the translation degree of this keyword is 1; if it is not the target word, it is the numerical value or the calculation result of the function that can represent the translation degree of the target word.
6. The medical machine translation method based on the keywords of the translated sentence according to claim 5, characterized in that In the above steps 1.1 - 1.4, medical professional vocabulary all belongs to keywords, and the corresponding medical professional vocabulary can be obtained by looking up the medical professional vocabulary list; The target sentence sequence after query screening by the medical professional vocabulary list is subjected to keyword extraction again by medical researchers; the union of the keywords extracted twice is the final keyword set of the target sentence sequence, and the position of each keyword in the target sentence sequence is recorded, and a keyword annotation vector is obtained; among them, the keyword annotation is 1, and the non-keyword annotation is 0.
Citation Information
Patent Citations
Neural machine translation accuracy improvement method, translation method and system, and devices
CN107943795A
Method for solving generation of untranslated words based on medical neural machine translation
CN110334362A
Translation method and device based on machine translation model, equipment and storage medium
CN110457713A
Text segmentation method and device
CN111046649A
An ancient Chinese automatic translation method based on multi-feature fusion
CN109684648A