Knowledge Fusion Method, Apparatus, Device, Medium and Product Based on Machine Translation
By adopting subword-level knowledge fusion method in machine translation technology, integrating subwords from source and target languages, and performing component annotation and vector encoding, the problem that word-level annotation in the existing technology cannot effectively integrate phrase knowledge, and improves the translation accuracy and the translation effect of the whole sentence.
Patent Information
- Application Number
- CN202210645372.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-06-09
AI Technical Summary
In existing machine translation technology, the method of one-to-one annotation in units of words cannot effectively integrate the knowledge of specified phrases, resulting in poor semantic modeling. Due to vocabulary limitations, the model is prone to exceeding the vocabulary during translation, resulting in excessive memory usage, sacrificing the model structure and size, and reducing translation accuracy.
The knowledge fusion method based on subwords is adopted, and the subwords of the source language text and the target language corresponding words or phrases are fused, and the fused subword text sequence is obtained, and the components are marked, and finally translated through vector encoding and decoding.
It improves the translation accuracy of user-specified words or phrases, reduces the integration of terminal vocabulary, solves the OOM problem, and improves the translation effect of the whole sentence.
Smart Images

Figure CN115099246B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine translation, and particularly to a knowledge fusion method, apparatus, electronic device, computer-readable storage medium, and computer program product based on machine translation. Background Art
[0002] With the continuous deepening of the globalization process, the demand for cross-language communication is increasing. With the rapid development of artificial intelligence, it has gradually become possible to use artificial intelligence technology to improve the content in the forms of documents, speech, etc., and to translate from one language to another. Machine translation, also known as automatic translation, is the process of using a computer to convert one language into another language, which is the ultimate goal of artificial intelligence.
[0003] In related technologies, in machine translation, a one-to-one knowledge fusion method specified by a user is usually adopted. This fusion method is that during encoding, the source language text is first segmented by word, and after the source language user-specified word, the corresponding specified word in the target language is added. Then, the segmented words are respectively labeled as: source language unspecified component, source language specified component, and target language specified component, etc. That is, the corresponding specified word in the target language is added to the source language text, and annotations are made according to the components. Then, the text and the component annotations are represented by vectors, and the two vectors are merged. After that, encoding and decoding are performed through a sequence-to-sequence model (such as Transformer, LSTM, etc.) to obtain the corresponding target language text.
[0004] However, in this related fusion method, one-to-one annotation is performed by word, and when entering the model encoding, the word is also the smallest unit of encoding. When the knowledge to be fused into the model is a specified phrase, this one-to-one annotation method is completely inapplicable, and the encoding method by word cannot make good use of the common encoding information between words, thus resulting in the modeling effect of word encoding on semantics (or meaning). At the same time, taking the word as the smallest unit of encoding will cause the problem of out of vocabulary (OOM) during encoding. Due to the limitation of hardware memory, in the machine translation model, the vocabulary at the input end only covers the words in the source language, and the model cannot cover all possible words during translation. Therefore, the model needs to fuse as many vocabularies as possible. Since the vocabulary occupies more memory, only by sacrificing the model structure and size can the memory be reduced, thereby reducing the accuracy of model translation. Summary of the Invention
[0005] The present invention provides a knowledge fusion method, apparatus, electronic device, computer-readable storage medium and computer program product based on machine translation, so as to at least solve the technical problem in the related art that due to the one-to-one annotation method with words as the smallest unit, all possible words cannot be covered during model translation. In order to incorporate more vocabulary into the model, the structure and size of the model need to be sacrificed, resulting in a relatively low model translation accuracy. The technical solution of the present invention is as follows:
[0006] According to the first aspect of the embodiments of the present invention, a knowledge fusion method based on machine translation is provided, including:
[0007] Obtain the source language text to be translated;
[0008] Obtain the corresponding target language word or phrase corresponding to the specified word or phrase in the source language text;
[0009] Fuse the sub-words of the source language text and the corresponding target language word or phrase to obtain a fused sub-word text sequence;
[0010] Perform component annotation on each sub-word in the sub-word text sequence to obtain the corresponding component annotation sequence for each sub-word;
[0011] Translate the sub-word text sequence and the corresponding component annotation sequence to obtain the target language text of the source language text.
[0012] Optionally, the obtaining of the corresponding target language word or phrase corresponding to the specified word or phrase in the source language text includes:
[0013] Extract the specified word or phrase in the source language text; and the position of the specified word or phrase in the source and target language text; and
[0014] Obtain the corresponding target language word or phrase corresponding to the specified word or phrase; or detect the corresponding target language word or phrase input by the user corresponding to the specified word or phrase in the source language text.
[0015] Optionally, the fusing of the sub-words of the source language text and the corresponding target language word or phrase to obtain a fused sub-word text sequence includes:
[0016] Perform sub-word tokenization on the source language text and the corresponding target language word or phrase respectively to obtain the corresponding first sub-word tokenization result and second sub-word tokenization result;
[0017] Concatenate the first sub-word tokenization result and the second sub-word tokenization result to obtain a concatenated sub-word text sequence.
[0018] Optionally, the first sub-word tokenization result includes: the sub-word tokenization result of a specified word or phrase in the source language text, and the sub-word tokenization result of the remaining part of the source language text except the sub-word tokenization result of the specified word or phrase; the second sub-word tokenization result includes: the sub-word tokenization result of the target language specified word or phrase;
[0019] The splicing of the first sub-word tokenization result and the second sub-word tokenization result includes:
[0020] According to the position of the specified word or phrase in the source-target language text, inserting the sub-word tokenization result of the target language specified word or phrase after the sub-word tokenization result of the specified word or phrase in the source language text for splicing to obtain a spliced sub-word text sequence.
[0021] Optionally, translating the sub-word text sequence and the corresponding component annotation sequence to obtain the target language text includes:
[0022] Performing vector encoding on the sub-word text sequence and the corresponding component annotation sequence respectively to obtain a semantic encoding vector and a corresponding component encoding vector;
[0023] Performing vector fusion on the semantic encoding vector and the corresponding component encoding vector to obtain a vector fusion result;
[0024] Decoding the vector fusion result to obtain the target language text of the source language text.
[0025] Optionally, the ways of vector fusion include but are not limited to: splicing, adding or weighting by a neural network, etc.
[0026] According to the second aspect of the embodiments of the present invention, there is provided a knowledge fusion device based on machine translation, including:
[0027] A first acquisition module, configured to acquire a source language text to be translated;
[0028] A second acquisition module, configured to acquire a target language corresponding word or phrase corresponding to the specified word or phrase in the source language text;
[0029] A sub-word fusion module, configured to perform sub-word fusion on the sub-words of the source language text and the target language corresponding word or phrase to obtain a sub-word text sequence after sub-word fusion;
[0030] An annotation module, configured to perform component annotation on each sub-word in the sub-word text sequence to obtain a corresponding component annotation sequence;
[0031] A translation module, configured to translate the sub-word text sequence and the corresponding component annotation sequence to obtain the target language text of the source language text.
[0032] Optionally, the second acquisition module includes: an extraction module and a text acquisition module, and / or an extraction module and a detection module, where,
[0033] The extraction module is used to extract a specified word or phrase from the source language text, and the position of the specified word or phrase in the source target language text;
[0034] The text acquisition module is used to acquire the corresponding target language word or phrase corresponding to the specified word or phrase;
[0035] The detection module is used to detect the corresponding target language word or phrase input by the user corresponding to the specified word or phrase in the source language text.
[0036] Optionally, the sub-word fusion module includes:
[0037] The word segmentation module is used to perform sub-word segmentation on the source language text and the corresponding target language word or phrase of the target language respectively, to obtain corresponding first sub-word segmentation results and second sub-word segmentation results;
[0038] The splicing module is used to splice the first sub-word segmentation result and the second sub-word segmentation result in sequence to obtain a spliced sub-word text sequence.
[0039] Optionally, the first sub-word segmentation result includes: the sub-word segmentation result of the specified word or phrase in the source language text, and the sub-word segmentation result of the remaining part of the source language text except the specified word or phrase; the second sub-word segmentation result includes: the sub-word segmentation result of the target language specified word or phrase;
[0040] The splicing module is specifically used to insert the sub-word segmentation result of the target language specified word or phrase after the sub-word segmentation result of the specified word or phrase in the source language text according to the position of the specified word or phrase in the source target language text, to obtain a spliced sub-word text sequence.
[0041] Optionally, the translation module includes:
[0042] The encoding module is used to perform vector encoding on the sub-word text sequence and the corresponding component annotation sequence respectively, to obtain a semantic encoding vector and a corresponding component encoding vector;
[0043] The vector fusion module is used to perform vector fusion on the semantic encoding vector and the corresponding component encoding vector to obtain a vector fusion result;
[0044] The decoding module is used to decode the vector fusion result to obtain the target language text of the source language text.
[0045] Optionally, the vector fusion method of the vector fusion module includes: splicing, adding, or weighting by a neural network.
[0046] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, including:
[0047] A processor;
[0048] A memory for storing executable instructions of the processor;
[0049] Wherein, the processor is configured to execute the instructions to implement the knowledge fusion method based on machine translation as described above.
[0050] According to a fourth aspect of an embodiment of the present invention, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the knowledge fusion method based on machine translation as described above.
[0051] According to a fifth aspect of an embodiment of the present invention, there is provided a computer program product, including a computer program or instructions, and when the computer program or instructions are executed by a processor, implementing the knowledge fusion method based on machine translation as described above.
[0052] The technical solutions provided by the embodiments of the present invention at least bring the following beneficial effects:
[0053] In the embodiments of the present invention, a source language text to be translated is obtained; a target language corresponding word or phrase corresponding to a specified word or phrase in the source language text is obtained; sub-words of the source language text and the target language corresponding word or phrase are fused to obtain a fused sub-word text sequence; each sub-word in the sub-word text sequence is component-labeled to obtain a component-labeled sequence corresponding to each sub-word; and the sub-word text sequence and the corresponding component-labeled sequence are translated to obtain the target language text of the source language text. That is to say, in the embodiments of the present invention, when obtaining the target language corresponding word or phrase corresponding to a specified word or phrase in the source language text, the sub-words of the source language text and the target language corresponding word or phrase are respectively fused, that is, after sub-word tokenization, even if they are not one-to-one corresponding, translation can be performed according to the specified word or phrase. After using sub-word tokenization in the embodiments of the present invention, not only the integration of the terminal vocabulary is reduced, the problem of out-of-vocabulary during translation is solved, but also the accuracy of the translation of the user-specified word or phrase and the translation effect of the whole sentence are improved.
[0054] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Description of the Drawings
[0055] The accompanying drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention, and do not constitute an improper limitation of the present invention.
[0056] Figure 1 It is a flowchart of a knowledge fusion method based on machine translation provided by an embodiment of the present invention.
[0057] Figure 2 It is a flowchart of sub-word segmentation and fusion provided by an embodiment of the present invention.
[0058] Figure 3 It is a schematic diagram of vector fusion provided by an embodiment of the present invention.
[0059] Figure 4 It is a block diagram of a knowledge fusion device based on machine translation provided by an embodiment of the present invention.
[0060] Figure 5 It is a block diagram of a second acquisition module provided by an embodiment of the present invention.
[0061] Figure 6 It is a block diagram of a sub-word fusion module provided by an embodiment of the present invention.
[0062] Figure 7 It is a block diagram of a translation module provided by an embodiment of the present invention.
[0063] Figure 8 It is a block diagram of an electronic device provided by an embodiment of the present invention.
[0064] Figure 9 It is a block diagram of a device for knowledge fusion in machine translation provided by an embodiment of the present invention. Detailed implementation manners
[0065] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0066] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0067] In recent years, important progress has been made in the research of technologies such as computer vision, deep learning, machine learning, machine translation, image processing, and image recognition based on artificial intelligence. Artificial Intelligence (AI) is a new science and technology that studies and develops theories, methods, technologies, and application systems for simulating and extending human intelligence. The discipline of artificial intelligence is a comprehensive discipline, involving many technical categories such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, machine translation, and neural networks. As an important branch of artificial intelligence, computer vision specifically enables machines to recognize the world. Computer vision technologies usually include face recognition, live detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, object detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, robot navigation and positioning, etc. With the research and progress of artificial intelligence technology, this technology has been applied in many fields, such as security, urban management, traffic management, building management, park management, face access, face attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart home, wearable devices, driverless, autonomous driving, intelligent healthcare, face payment, face unlocking, fingerprint unlocking, person-certificate verification, smart screen, smart TV, cameras, mobile Internet, webcasting, beauty, makeup, medical beauty, intelligent temperature measurement, etc.
[0068] Before describing the embodiments of the present invention, the following technical terms are introduced first:
[0069] Subword: A semantic unit with a finer granularity than a word. A word may have multiple or no subwords, and subword units can also be shared between words. Subwords are obtained by statistics on monolingual corpora.
[0070] User-specified word or phrase: The standard translation specified artificially in translation, mainly manifested as proper nouns, terms, or fixed phrase collocations, etc.
[0071] End-to-end neural network: In contrast to traditional machine learning, which first obtains the vector representation of data through various methods and then trains the model on this vector representation, an end-to-end neural network does not need to construct a vector representation but directly builds a model on the data for training. It can learn more comprehensively the relationships existing between data, avoiding both the use of complex methods to obtain vector representations and improving the performance of the model.
[0072] Knowledge fusion: When using an end-to-end neural network model or other machine learning models, known or specified knowledge is integrated into the model so that the model can refer to the given knowledge to complete reasoning, thereby improving the model effect.
[0073] Automatic machine translation system: Dependent on data and models, after the model training is completed, the source language sentences input into the system can be translated into target language sentences without any manual participation, and the accuracy and fluency of the automatic translation results are guaranteed.
[0074] Figure 1 is a flowchart of a knowledge fusion method based on machine translation provided by an embodiment of the present invention. As Figure 1 shown, the knowledge fusion method based on machine translation includes the following steps:
[0075] Step 101: Obtain the source language text to be translated.
[0076] Step 102: Obtain the corresponding target language words or phrases corresponding to the specified words or phrases in the source language text.
[0077] Step 103: Fuse the sub-words of the source language text and the corresponding target language words or phrases to obtain a fused sub-word text sequence.
[0078] Step 104: Perform component annotation on each sub-word in the sub-word text sequence to obtain a component annotation sequence corresponding to each sub-word.
[0079] Step 105: Translate the sub-word text sequence and the corresponding component annotation sequence to obtain the target language text of the source language text.
[0080] The knowledge fusion method based on machine translation described in the present invention can be applied to terminals, servers, etc., which are not limited here. The terminal implementation devices can be electronic devices such as smart phones, laptops, and tablets, which are not limited here.
[0081] Next, in combination with Figure 1 , the specific implementation steps of a knowledge fusion method based on machine translation provided by an embodiment of the present invention will be described in detail.
[0082] In step 101, obtain the source language text to be translated.
[0083] In this step, the source language to be translated is any natural language, such as English, Chinese, etc. The source language text can be a sentence, phrase, fixed collocation, etc. of the source language.
[0084] Among them, the method for the terminal to obtain the source language text to be translated in this step is not limited in this embodiment. For example, it can be the source language text read, or the source language text input by the user to be translated, etc.
[0085] In step 102, obtain the corresponding word or phrase in the target language corresponding to the specified word or phrase in the source language text.
[0086] In this step, after the terminal obtains the source language text to be translated, first, the terminal extracts the specified word or phrase in the source language text, and the position of the specified word or phrase in the source and target language text; among them, the specified word in this embodiment refers to the standard translation specified manually in translation, usually embodied as proper nouns, technical terms or fixed phrase collocations, etc. That is, the terminal extracts the specified word or phrase and its position in the source language text by specifying the corresponding relationship between the source language and the target language, and its position is convenient for subsequent insertion.
[0087] After that, the terminal obtains the corresponding word or phrase in the target language corresponding to the specified word or phrase; or detects the corresponding word or phrase in the target language input by the user corresponding to the specified word or phrase in the source language text.
[0088] In this step, the terminal can obtain the corresponding word or phrase in the target language corresponding to the specified word or phrase in the source language text by setting the corresponding relationship between the specified source language and the target language. Of course, it can also be detecting the corresponding word or phrase in the target language input by the user corresponding to the specified word or phrase in the source language text. Among them, the corresponding relationship between the specified source language and the target language is the standard translation of translating the specified source language into the target language, which is pre-configured. For example, the corresponding relationship of the standard translation between English and Chinese is specified, etc.
[0089] In step 103, fuse the sub-words of the source language text and the corresponding word or phrase in the target language to obtain a fused sub-word text sequence.
[0090] In this step, the terminal performs sub-word tokenization on the source language text and the corresponding word or phrase in the target language respectively to obtain the corresponding first sub-word tokenization result and second sub-word tokenization result; then, splice the first sub-word tokenization result and the second sub-word tokenization result to obtain the spliced sub-word text sequence.
[0091] In this step, the terminal can perform sub-word segmentation on the source language text through the source language sub-word segmentation model to obtain the first sub-word segmentation result (including the sub-word segmentation result of the specified word or phrase in the source language text and the sub-word segmentation result of the remaining part of the source language text except the sub-word segmentation result of the specified word or phrase), and perform sub-word segmentation on the corresponding word or phrase in the target language through the target language sub-word segmentation model to obtain the second sub-word segmentation result, that is, the sub-word segmentation result of the target language specified word or phrase. Among them, for those skilled in the art, whether performing sub-word segmentation on the source language text through the source language sub-word segmentation model or performing sub-word segmentation on the corresponding word or phrase in the target language through the target language sub-word segmentation model is well-known technology and will not be elaborated here. Among them, the sub-word segmentation method in this embodiment can be sub-word segmentation methods such as BPE, Word Piece, ULM, Sentence Piece, etc. based on text symbols or bytes.
[0092] Of course, in this embodiment, the terminal can also perform sub-word segmentation on the source language text and the specified word or phrase in the source language text respectively through the source language sub-word segmentation model. Among them, when performing sub-word segmentation on the source language text, two parts can be obtained. One part is the sub-word segmentation result of the source language specified word or phrase, and the other part is the sub-word segmentation result of the other part of the source language text except the sub-word segmentation result of the source language specified word or phrase, that is, the sub-word segmentation result of the other part of the source language text. And when performing sub-word segmentation on the specified word or phrase in the source language text, the sub-word segmentation result of the source language specified word or phrase is obtained. Specifically, as Figure 2 shown, Figure 2 is a flowchart of sub-word segmentation and fusion provided by an embodiment of the present invention. Figure 2 is illustrated by taking the sub-word segmentation of the source language text as an example.
[0093] Such as Figure 2As shown, first obtain the source language text, then extract the specified words or phrases in the source language text and their positions. Next, according to the corresponding relationship between the specified source language and the target language, obtain the target corresponding words and phrases of the specified words or phrases in the source language text, and perform sub-word tokenization on the source language text through the source language sub-word tokenization model to obtain the sub-word tokenization results of the other parts in the corresponding source language text, the sub-word tokenization results of the specified words or phrases in the source language text, and perform sub-word tokenization on the target language corresponding words or phrases through the target language sub-word tokenization model to obtain the sub-word tokenization results of the specified words or phrases in the target language. Next, the STF method is used for fusion annotation, that is, according to the positions of the specified words or phrases in the source language text, the sub-word tokenization results of the other parts in the source language text, the sub-word tokenization results of the specified words or phrases in the source language text, and the sub-word tokenization results of the specified words or phrases in the target language are spliced to obtain the spliced sub-word text sequence, and each sub-word in the sub-word text sequence is marked with its component to obtain the component annotation sequence corresponding to each sub-word. The specific component annotation process is shown in the following step 104.
[0094] In step 104, each sub-word in the sub-word text sequence is marked with its component to obtain the component annotation sequence corresponding to each sub-word.
[0095] In this step, when marking each sub-word in the sub-word text sequence with its component, the sub-word terminology fusion (STF) annotation method can be used. This method is a method of fusing the sub-word tokenization results proposed on the basis of performing sub-word tokenization on the source language text and the target language text respectively. This method stipulates the order of fusing the sub-words of the two language texts and the annotation of the components. As shown in Table 1 specifically, Table 1 is an application example provided by the embodiments of the present invention. As shown in Table 1:
[0096] Table 1
[0097]
[0098] As shown in Table 1, the subword segmentation method used in this embodiment takes Sentence Piece as an example, which is not limited to this in practical applications, and other subword segmentation methods can also be used instead. As shown in Table 1, the source language is English and the target language is Chinese. The source language English sentence is Is suing point acceptance credit, the specified phrase in the English sentence is acceptance credit, and the corresponding target language Chinese phrase is specified as acceptance credit letter. The English phrase is acceptance credit after subword segmentation, and the corresponding Chinese phrase is acceptance credit letter after subword segmentation. The other parts of the English sentence are also subword segmented using the same method, and the subword segmentation results are specifically shown in Table 1. Afterwards, splicing is performed in sequence according to each subword segmentation result in the English sentence, and the subword segmentation result of the Chinese phrase is inserted after the subword segmentation result of the English phrase, and then, each subword part is respectively marked with the corresponding component, that is, the component of each subword is marked, as shown in the corresponding part of SFT in Table 1. Among them, in Table 1, w is the unspecified component of the source language, s is the specified component of the source language, t is the specified component of the target language, etc.
[0099] It should be noted that in the embodiments of the present invention, two methods are usually used to annotate training data, one is a semi-supervised automatic annotation method, and the other is a manual annotation method. The semi-supervised automatic annotation method is to collect corresponding data of source language and target language words or phrases, and match and annotate corresponding parts in parallel corpus using text rules based on these corresponding data. The manual annotation method is to annotate manually for the corresponding relationships that are not covered by the semi-supervised automatic annotation.
[0100] In step 105, the subword text sequence and the corresponding component annotation sequence are translated to obtain a target language text of the source language text.
[0101] In this step, the subword text sequence and the corresponding component annotation sequence are translated to obtain the target language text, which specifically includes: first, vector encoding the subword text sequence and the corresponding component annotation sequence respectively to obtain a semantic encoding vector (also called a semantic encoding vector, the same below) and a corresponding component encoding vector; second, vector fusion is performed on the semantic encoding vector and the corresponding component encoding vector to obtain a vector fusion result; finally, the vector fusion result is decoded to obtain the target language text. Specifically, Figure 3 As shown, Figure 3 It is a schematic diagram of vector fusion provided by an embodiment of the present invention.
[0102] like Figure 3As shown, vector encoding is performed on the sub-word text sequence (e.g., in Table 1: -Is su ing) and the corresponding component annotation sequence (e.g., w w w) respectively. That is, sub-word vector encoding is performed on the sub-word text sequence to obtain a semantic encoding vector, and component vector encoding is performed on the component annotation sequence to obtain the corresponding component encoding vector. Then, the semantic encoding vector and the corresponding component encoding vector are vector fused to obtain a vector fusion result. After that, the vector fusion result is decoded through a sequence-to-sequence model (i.e., a way in inference) to obtain the target language text of the source language text.
[0103] Among them, the sequence-to-sequence model can select any sequence-to-sequence end-to-end model, such as Transformer, LSTM, etc. After encoding the input text according to the above method provided by the embodiments of the present invention, the encoding vector enters the sequence-to-sequence model and decodes the corresponding target language prediction result. During training, the text after sub-word tokenization of the target language is fitted. It should be noted that the sub-word tokenization of the target language is the same as the sub-word tokenization inserted at the input end. When using the model for inference, given the corresponding words or phrases of the source language and the target language, the source language text is processed according to the STF method (or algorithm), and input into the model for encoding and decoding, and the prediction result is the complete target language text.
[0104] Among them, the way of the vector fusion in this embodiment may include: concatenation, addition, or weighted by a neural network, etc., but in specific applications, it is not limited to this.
[0105] Among them, the decoder is a long short-term memory network (LSTM, Long Short-Term Memory), and its initial state is initialized to the final state of the encoder LSTM, that is, the context vector of the final unit of the encoder is input to the first unit of the decoder network. Using these initial states, the decoder starts to generate the output sequence, and the future outputs also take these outputs into account.
[0106] That is to say, in this step, the terminal respectively performs vector encoding on the text sequence and the corresponding component annotation sequence, that is, the semantic encoding vector or the semantic encoding vector, and then concatenates the two vectors to facilitate learning the correct vector encoding method during model training.
[0107] Among them, the semantic encoding in this embodiment, which can also be called semantic coding, processes information through words, organizes and summarizes it according to meaning, system classification, or in its own language form of the speech material, finds out the basic arguments, evidence, and logical structure of the material, and encodes it according to semantic features. Semantic encoding is one form of meaning encoding and also the main encoding method for long-term memory. It represents information in a systematic way according to the order of language occurrence, including information in both speech audition and speech movement aspects. The characteristic of semantic encoding is serial processing, which is a meaningful connection according to nodes and lines.
[0108] In the embodiment of the present invention, the source language text to be translated is obtained; the corresponding target language corresponding word or phrase for the specified word or phrase in the source language text is obtained; the sub-words of the source language text and the target language corresponding word or phrase are fused to obtain a fused sub-word text sequence; each sub-word in the sub-word text sequence is component-labeled to obtain the corresponding component-labeled sequence for each sub-word; the sub-word text sequence and the corresponding component-labeled sequence are translated to obtain the target language text of the source language text. That is to say, in the embodiment of the present invention, when obtaining the corresponding target language corresponding word or phrase for the specified word or phrase in the source language text, the sub-words of the source language text and the target language corresponding word or phrase are respectively fused. That is, after sub-word tokenization, even if they are not in one-to-one correspondence, they can be translated according to the specified word or phrase. The embodiment of the present invention uses sub-word tokenization to reduce the incorporation of the terminal vocabulary and solves the OOM problem, improving the accuracy of the translation of the user-specified word or phrase and the translation effect of the whole sentence.
[0109] The embodiment of the present invention uses the method of sub-word tokenization and fusion annotation on an end-to-end neural network automatic machine translation system (STFMT, Subword Terminology Fusion Machine Translation). This method extracts the specified word or phrase from the input source language text, obtains the corresponding target language corresponding word or phrase, and fuses them after sub-word tokenization of the source language text and the target language corresponding word or phrase respectively, reducing the incorporation of the terminal vocabulary, solving the OOM problem, and also improving the accuracy of the translation of the user-specified word or phrase and the translation effect of the whole sentence.
[0110] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the present invention.
[0111] Figure 4 This is a block diagram of a knowledge fusion device based on machine translation provided by an embodiment of the present invention. The device includes: a first acquisition module 401, a second acquisition module 402, a sub-word fusion module 403, a tagging module 404, and a translation module 405. Among them,
[0112] The first acquisition module 401 is configured to acquire a source language text to be translated;
[0113] The second acquisition module 402 is configured to acquire a target language corresponding word or phrase corresponding to a specified word or phrase in the source language text;
[0114] The sub-word fusion module 403 is configured to perform sub-word fusion on the sub-words of the source language text and the target language corresponding word or phrase to obtain a sub-word text sequence after sub-word fusion;
[0115] The tagging module 404 is configured to perform component tagging on each sub-word in the sub-word text sequence to obtain a corresponding component tagging sequence;
[0116] The translation module 405 is configured to translate the sub-word text sequence and the corresponding component tagging sequence to obtain the target language text of the source language text.
[0117] Optionally, in another embodiment, based on the above embodiment, the second acquisition module 402 includes: an extraction module 501 and a text acquisition module 502, and / or an extraction module 501 and a detection module 503. Its block diagram is as Figure 5 shown, where
[0118] The extraction module 501 is configured to extract a specified word or phrase in the source language text, and the position of the specified word or phrase in the source and target language text;
[0119] The text acquisition module 502 is configured to acquire a target language corresponding word or phrase corresponding to the specified word or phrase;
[0120] The detection module 503 is configured to detect a target language corresponding word or phrase input by the user corresponding to a specified word or phrase in the source language text.
[0121] Optionally, in another embodiment, based on the above embodiment, the sub-word fusion module 403 includes: a word segmentation module 601 and a splicing module 602. Its block diagram is as Figure 6 shown, where
[0122] The word segmentation module 601 is configured to perform sub-word segmentation on the source language text and the target language corresponding word or phrase respectively to obtain corresponding first sub-word segmentation results and second sub-word segmentation results;
[0123] A splicing module 602, configured to splice the first sub-word tokenization result and the second sub-word tokenization result in sequence to obtain a spliced sub-word text sequence.
[0124] Optionally, in another embodiment, based on the above embodiment, the first sub-word tokenization result includes: the sub-word tokenization result of a specified word or phrase in the source language text, and the sub-word tokenization result of the remaining part of the source language text except the sub-word tokenization result of the specified word or phrase; the second sub-word tokenization result includes: the sub-word tokenization result of a target language specified word or phrase.
[0125] The splicing module is specifically configured to insert the sub-word tokenization result of the target language specified word or phrase after the sub-word tokenization result of the specified word or phrase in the source language text according to the position of the specified word or phrase in the source-target language text, so as to obtain a spliced sub-word text sequence.
[0126] Optionally, in another embodiment, based on the above embodiment, the translation module 405 includes: an encoding module 701, a vector fusion module 702, and a decoding module 703. The structural block diagram is as Figure 7 shown, where
[0127] The encoding module 701 is configured to perform vector encoding on the sub-word text sequence and the corresponding component annotation sequence respectively to obtain a semantic encoding vector and a corresponding component encoding vector.
[0128] The vector fusion module 702 is configured to perform vector fusion on the semantic encoding vector and the corresponding component encoding vector to obtain a vector fusion result.
[0129] The decoding module 703 is configured to decode the vector fusion result to obtain the target language text of the source language text.
[0130] Optionally, in another embodiment, based on the above embodiment, the vector fusion method of the vector fusion module includes: splicing, adding, or weighting by a neural network.
[0131] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.
[0132] Optionally, an embodiment of the present invention further provides an electronic device, including:
[0133] A processor;
[0134] A memory for storing executable instructions of the processor;
[0135] Wherein, the processor is configured to execute the instructions to implement the machine translation-based knowledge fusion method as described above.
[0136] Optionally, an embodiment of the present invention further provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the machine translation-based knowledge fusion method as described above. Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0137] Optionally, an embodiment of the present invention further provides a computer program product, including a computer program or instructions. When the computer program or instructions are executed by a processor, the machine translation-based knowledge fusion method as described above is implemented.
[0138] Figure 8 FIG. 12 is a block diagram of an electronic device 800 provided by an embodiment of the present invention. For example, the electronic device 800 may be a mobile terminal or a server. In this embodiment of the present invention, the electronic device is taken as a mobile terminal as an example for illustration. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0139] Refer to Figure 8 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0140] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0141] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0142] The power supply component 806 provides power to various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0143] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0144] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0145] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0146] The sensor assembly 814 includes one or more sensors for providing status assessments of various aspects for the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0147] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0148] In an embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the machine translation-based knowledge fusion method shown above.
[0149] In an embodiment, a computer-readable storage medium is also provided, such as a memory 804 including instructions that can be executed by a processor 820 of the electronic device 800 to complete the machine translation-based knowledge fusion method shown above. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0150] In an embodiment, a computer program product is also provided. When the instructions in the computer program product are executed by the processor 820 of the electronic device 800, the electronic device 800 is caused to execute the above-described knowledge fusion method based on machine translation.
[0151] Figure 9 FIG. 4 is a block diagram of a device 900 for knowledge fusion in machine translation provided by an embodiment of the present invention. For example, the device 900 may be provided as a server. Referring to Figure 9 FIG. 4, the device 900 includes a processing component 922, which further includes one or more processors, and memory resources represented by a memory 932 for storing instructions executable by the processing component 922, such as application programs. The application programs stored in the memory 932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 922 is configured to execute instructions to perform the above method.
[0152] The device 900 may further include a power component 926 configured to perform power management of the device 900, a wired or wireless network interface 950 configured to connect the device 900 to a network, and an input / output (I / O) interface 958. The device 900 may operate based on an operating system stored in the memory 932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0153] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.
[0154] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes may be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A knowledge fusion method based on machine translation, characterized in that, comprising: Obtaining the source language text to be translated; Obtaining the corresponding target language word or phrase corresponding to the specified word or phrase in the source language text; Fusing the sub-words of the source language text and the corresponding target language word or phrase to obtain a fused sub-word text sequence; performing component annotation on each sub-word in the sub-word text sequence to obtain the corresponding component annotation sequence for each sub-word; Translating the sub-word text sequence and the corresponding component annotation sequence to obtain the target language text of the source language text; Wherein, fusing the sub-words of the source language text and the corresponding target language word or phrase to obtain a fused sub-word text sequence includes: Performing sub-word tokenization on the source language text and the corresponding target language word or phrase respectively to obtain corresponding first sub-word tokenization results and second sub-word tokenization results; Concatenating the first sub-word tokenization result and the second sub-word tokenization result to obtain a concatenated sub-word text sequence; The first sub-word tokenization result includes: the sub-word tokenization result of the specified word or phrase in the source language text, and the sub-word tokenization result of the remaining part of the source language text except the sub-word tokenization result of the specified word or phrase; The second sub-word tokenization result includes: the sub-word tokenization result of the target language specified word or phrase; The concatenating of the first sub-word tokenization result and the second sub-word tokenization result includes: Inserting the sub-word tokenization result of the target language specified word or phrase behind the sub-word tokenization result of the specified word or phrase in the source language text according to the position of the specified word or phrase in the target language text for concatenation to obtain a concatenated sub-word text sequence.
2. The knowledge fusion method based on machine translation according to claim 1, characterized in that, The obtaining of the corresponding target language word or phrase corresponding to the specified word or phrase in the source language text includes: Extracting the specified word or phrase in the source language text, and the position of the specified word or phrase in the target language text; and Obtaining the corresponding target language word or phrase corresponding to the specified word or phrase; or detecting the target language word or phrase input by the user corresponding to the specified word or phrase in the source language text.
3. The knowledge fusion method based on machine translation according to claim 1 or 2, characterized in that, Translating the sub-word text sequence and the corresponding component annotation sequence to obtain the target language text includes: Performing vector encoding on the sub-word text sequence and the corresponding component annotation sequence respectively to obtain a semantic encoding vector and a corresponding component encoding vector; Performing vector fusion on the semantic encoding vector and the corresponding component encoding vector to obtain a vector fusion result; Decoding the vector fusion result to obtain the target language text of the source language text.
4. The knowledge fusion method based on machine translation according to claim 3, characterized in that, The ways of vector fusion include: concatenation, addition or weighting by a neural network.
5. A knowledge fusion device based on machine translation, It is characterized in that including: A first acquisition module for acquiring the source language text to be translated; A second acquisition module for acquiring the corresponding target language words or phrases corresponding to the specified words or phrases in the source language text; A sub-word fusion module for performing sub-word fusion on the sub-words of the source language text and the corresponding target language words or phrases to obtain a sub-word text sequence after sub-word fusion; A word segmentation module for performing sub-word segmentation on the source language text and the corresponding target language words or phrases respectively to obtain corresponding first sub-word segmentation results and second sub-word segmentation results; A splicing module for sequentially splicing the first sub-word segmentation result and the second sub-word segmentation result to obtain a spliced sub-word text sequence; The first sub-word segmentation result includes: the sub-word segmentation result of the specified word or phrase in the source language text, and the sub-word segmentation result of the remaining part of the source language text except the sub-word segmentation result of the specified word or phrase; The second sub-word segmentation result includes: the sub-word segmentation result of the target language specified word or phrase; The splicing module is further configured to insert the sub-word segmentation result of the target language specified word or phrase behind the sub-word segmentation result of the specified word or phrase in the source language text according to the position of the specified word or phrase in the target language text for splicing to obtain a spliced sub-word text sequence; A labeling module for performing component labeling on each sub-word in the sub-word text sequence to obtain a corresponding component labeling sequence; A translation module for translating the sub-word text sequence and the corresponding component labeling sequence to obtain the target language text of the source language text.
6. An electronic device It is characterized in that including: A processor; A memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the knowledge fusion method based on machine translation according to any one of claims 1 to 4.
7. A computer-readable storage medium It is characterized in that When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute the knowledge fusion method based on machine translation according to any one of claims 1 to 4.
8. A computer program product including a computer program or instructions It is characterized in that When the computer program or instructions are executed by the processor, the knowledge fusion method based on machine translation according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Chinese- Vietnamese unsupervised neural machine translation method fusing EMD minimized bilingual dictionary
CN111753557A
Compound Splitting
US20110202330A1