Method, apparatus, and medium for machine translation
By using bilingual phrase prompts for pre-training in machine translation, the problem of inconsistency in prompt learning based on prompt learning is solved, and the improvement of translation performance in specific fields and the effective application of lexical constraints is achieved.
Patent Information
- Application Number
- CN202111325941.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-10
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-11-10
AI Technical Summary
The prior art is difficult to create effective prompts when applying prompt-based learning to machine translation, and there is an inconsistency between pre-training and prediction, resulting in limited machine translation performance.
Bilingual phrase prompts are used to retrieve relevant bilingual phrases from the pre-constructed bilingual phrase database, and concatenate them to the source language sentences, and pre-training and prediction pre-training are performed to relieve sentence-level translation example sparsity and pre-training and prediction inconsistency.
Without additional training, the translation accuracy of machine translation BLEU scores and lexical constraints in specific fields is significantly improved, 6.2 and 11.5% respectively, and effectively combine lexical constraints into the translation process.
Smart Images

Figure CN113887253B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to methods, devices, and media for machine translation. Background Art
[0002] Pre-trained models (PTMs) have significantly advanced natural language processing. In recent years, prompt-based learning has become an attractive method for adapting PTMs to specific tasks. Using either manually created or automatically created prompts, PTMs can achieve good performance in many downstream tasks without fine-tuning. Different from fine-tuning and feature-based adaptation, prompt-based learning does not require additional training for downstream tasks. It formulates downstream tasks as language model fill-in-the-blank tasks with prompts. Generally, in prompt-based learning, predicting a specific task using a pre-trained language model includes three stages: (i) constructing a prompt with some unfilled gaps based on the input; (ii) using the pre-trained model to fill these unfilled gaps; and (iii) deriving the final prediction based on the filled gaps.
[0003] The prompt format depends on the pre-trained model and the downstream task. There are mainly two types of prompts: cloze-style prompts, where the unfilled gaps are predefined blanks; and prefix-style prompts, where filling the gaps is a generation process that continues using the prefix. Cloze-style prompts are usually used for natural language understanding tasks, while prefix-style prompts are mainly used for natural language generation tasks. Summary of the Invention
[0004] This Summary of the Invention section is provided to introduce concepts in a brief form that will be described in detail in the Detailed Description section later. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0005] According to some embodiments of the present disclosure, a method for machine translation is provided, including: obtaining original training data including source language sentences for training and target language sentences for training; concatenating bilingual phrase prompts related to the source language sentences for training to the source language sentences for training to generate new training data with bilingual phrase prompts, where the bilingual phrase prompts include one or more bilingual phrases, and each bilingual phrase includes a source language phrase and a corresponding target language phrase; pre-training a machine translation model using at least the new training data with bilingual phrase prompts.
[0006] According to some embodiments of the present disclosure, there is provided an apparatus for machine translation, including: an original training data acquisition unit configured to acquire original training data including a source language sentence for training and a target language sentence for training; a new training data generation unit configured to concatenate bilingual phrase prompts related to the source language sentence for training to the source language sentence for training to generate new training data with bilingual phrase prompts, where the bilingual phrase prompts include one or more bilingual phrases, and each bilingual phrase includes a source language phrase and a corresponding target language phrase; and a pre-training unit configured to pre-train a machine translation model at least using the new training data with bilingual phrase prompts.
[0007] According to some embodiments of the present disclosure, there is provided an electronic device, including: a memory; and a processor coupled to the memory, where instructions are stored in the memory, and when the instructions are executed by the processor, the processor is caused to execute the method according to the embodiments of the present disclosure.
[0008] According to some embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon a computer program, which when executed by a processor implements the method according to the embodiments of the present disclosure.
[0009] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of the exemplary embodiments of the present disclosure with reference to the accompanying drawings. Description of the Drawings
[0010] The following describes the preferred embodiments of the present disclosure with reference to the accompanying drawings. The drawings described herein are used to provide a further understanding of the present disclosure, and together with the following specific description, are included in this specification and form a part of this specification, for explaining the present disclosure. It should be understood that the drawings in the following description only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the drawings:
[0011] Figure 1 is a schematic diagram showing a comparison between the method for machine translation according to the embodiments of the present disclosure and the existing Vanilla method;
[0012] Figure 2 is a flowchart showing the method for machine translation according to the embodiments of the present disclosure;
[0013] Figure 3 is a flowchart showing the method for retrieving bilingual phrases from a bilingual phrase database according to the embodiments of the present disclosure;
[0014] Figure 4 is a flowchart showing the process of translating a source language phrase into a target language phrase based on a bilingual phrase database;
[0015] Figure 5 Shows an example of English-German translation in the medical field according to an embodiment of the present disclosure;
[0016] Figure 6 Is a block diagram showing an apparatus for machine translation according to an embodiment of the present disclosure;
[0017] Figure 7 Is a block diagram showing an electronic device according to an embodiment of the present disclosure; and
[0018] Figure 8 Is a block diagram showing an example structure of a computer system that can be adopted in an embodiment of the present disclosure.
[0019] It should be understood that, for the sake of convenience of description, the sizes of the various parts shown in the drawings are not necessarily drawn in actual proportional relationships. The same or similar reference numerals are used in the various drawings to represent the same or similar components. Therefore, once an item is defined in one drawing, it may not be further discussed in subsequent drawings. Detailed implementation manners
[0020] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. However, it is obvious that the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The following description of the embodiments is actually only illustrative and in no way restricts the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0021] It should be understood that the various steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard. Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments should be construed as merely exemplary and do not limit the scope of the present disclosure.
[0022] The term "comprising" and its variants used in the present disclosure mean open terms that include at least the subsequent elements / features, but do not exclude other elements / features, that is, "including but not limited to". In addition, the term "containing" and its variants used in the present disclosure mean open terms that contain at least the subsequent elements / features, but do not exclude other elements / features, that is, "containing but not limited to". Therefore, including and containing are synonymous. The term "based on" means "at least partially based on".
[0023] Throughout the specification, the terms "one embodiment", "some embodiments", or "an embodiment" mean that the specific features, structures, or characteristics described in connection with the embodiment are included in at least one embodiment of the present disclosure. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Moreover, the appearances of the phrases "in one embodiment", "in some embodiments", or "in an embodiment" throughout the specification do not necessarily all refer to the same embodiment, but may also refer to the same embodiment.
[0024] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units. Unless otherwise specified, the concepts such as "first", "second", etc. are not intended to imply that the objects so described must be in a given order in terms of time, space, ranking, or any other way.
[0025] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless clearly specified otherwise in the context, it should be understood as "one or more".
[0026] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0027] The embodiments of the present disclosure will be described in detail below in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner that will be apparent to those of ordinary skill in the art from the present disclosure.
[0028] There are two main difficulties in applying prompt-based learning to machine translation. First, it is difficult to create effective prompts for machine translation. Brown et al. proposed in "Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877-1901" to construct prompts by concatenating some in-context translation examples. Liu et al. revealed in "What makes good in-context examples for gpt-3? arXiv preprint arXiv:2101.06804" that the performance of downstream tasks heavily depends on the selection of in-context examples. However, sentence-level translation examples are sparse. It is difficult to find sentence-level translation examples related to the input sentence to construct effective prompts. Second, the pre-training tasks of language models are not designed jointly with prompt-based machine translation prediction. The inconsistency between pre-training and prediction limits the potential of prompt-based learning.
[0029] Embodiments of the present disclosure can effectively solve the difficulties encountered in applying prompt-based learning to machine translation. Embodiments of the present disclosure utilize bilingual phrase prompts and use prompt-aware pre-training (PAPT) for prompt-based machine translation. Experiments show that embodiments of the present disclosure can improve the BLEU score of domain-specific machine translation by 6.2 without additional training and improve the accuracy of morphological constraint machine translation by 11.5%
[0030] First, embodiments of the present disclosure construct bilingual phrase prompts for machine translation. The bilingual phrase prompts are concatenated by phrase-level translation examples (e.g., bilingual phrases) to alleviate the sparsity of sentence-level translation examples. By retrieving relevant bilingual phrases from a pre-constructed bilingual phrase database, embodiments of the present disclosure can construct prompts related to the input, which can provide useful knowledge for translation generation. Then, in order to alleviate the inconsistency between pre-training and prompt-based prediction, the prompt can be made known during pre-training. Therefore, a prompt-aware pre-training task can be designed, which can be a sequence-to-sequence generation task. Embodiments of the present disclosure pre-train a prompt-aware model for machine translation, thereby being able to alleviate the inconsistency between pre-training and prompt-based prediction.
[0031] Figure 1 is a schematic diagram showing the comparison between the method for machine translation according to an embodiment of the present disclosure and the existing Vanilla method. In Figure 1In it, the input is the source language sentence in the training sample, the output is the target language sentence in the training sample, and the prompt is a bilingual phrase prompt. In the existing Vanilla method, only each training sample including the source language sentence and the target language sentence is trained, without using the bilingual phrase prompt. PAPT is a pre-training method with awareness of prompts according to an embodiment of the present disclosure. In PAPT, not only each training sample including the source language sentence and the target language sentence is used for training, but also for each training sample, the bilingual phrase prompt is concatenated after the source language sentence to form a new training sample with the target language sentence for training. To exemplarily illustrate the embodiment of the present disclosure, concatenating the bilingual phrase prompt to the source language sentence means using the bilingual phrase prompt as a prefix of the source language sentence. However, the embodiment of the present disclosure is not limited thereto. For example, the bilingual phrase prompt can be used as a suffix rather than a prefix of the source language sentence.
[0032] Figure 2 is a flowchart showing a method 200 for machine translation according to an embodiment of the present disclosure. In step S210, the original training data including the source language sentence for training and the target language sentence for training is obtained. The original training data can be translation data in the general domain, for example, the WMT14 EN-DE dataset.
[0033] In step S220, the bilingual phrase prompt related to the source language sentence for training is concatenated to the source language sentence for training to generate new training data with the bilingual phrase prompt.
[0034] The bilingual phrase prompt includes one or more bilingual phrases, and each bilingual phrase includes a source language phrase and a corresponding target language phrase. The bilingual phrase prompt can provide useful knowledge for machine translation.
[0035] Between multiple bilingual phrases, there can be a first marker (for example, <r>) separated. The source language phrase and the corresponding target language phrase in the bilingual phrase can be separated by a second marker (e.g., <q>) separated. The bilingual phrase hint related to the source language sentence for training and the source language sentence for training can be separated by a third marker (e.g., )Separation.
[0036] Bilingual phrases can be retrieved from a pre - built bilingual phrase database to construct bilingual phrase hints. For the source - language sentences used for training, bilingual phrases related to the source - language sentences used for training can be retrieved from the pre - built bilingual phrase database to construct bilingual phrase hints related to the source - language sentences used for training. For the source - language sentences to be translated, bilingual phrases related to the source - language sentences to be translated can be retrieved from the pre - built bilingual phrase database to construct bilingual phrase hints related to the source - language sentences to be translated.
[0037] The bilingual phrase database can be pre - built and can be offline. Multilingual BERT proposed by, for example, Devlin et al. in "BERT: Pre - training of deep bidirectional transformers for language understanding. In Proc. of NAACL - HLT, pages 4171 - 4186" can be used to extract bilingual phrases from parallel translation data and calculate the context representations of source - language phrases to construct the bilingual phrase database. The context representations of source - language phrases and the corresponding bilingual phrases are stored in the bilingual phrase database as key - value pairs. The context representations of source - language phrases are used as keys in the key - value pairs, and the corresponding bilingual phrases are used as values in the key - value pairs. The bilingual phrase database is a collection of key - value pairs created from parallel translation data.
[0038] The method for extracting bilingual phrases may include: First, the awesome-align method described by Dou et al. in "Word alignment by fine-tuning embeddings on parallel corpora. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2112–2128” can be used to extract word alignments; then, the algorithm described by Koehn et al. in "Statistical Machine Translation” can be used to extract bilingual phrases from the word alignments. The context representation of a phrase can be calculated by performing pool averaging on the hidden states of the words in the phrase. The joint byte pair encoding method described by Sennrich et al. in "Neural machine translation of rare words with subword units. In Proc. of ACL, pages 1715–1725” can be used to perform subword segmentation through 32k merge operations.
[0039] Figure 3 FIG. is a flowchart showing a method 300 for retrieving bilingual phrases from a bilingual phrase database according to an embodiment of the present disclosure. In step S310, source language phrases in the bilingual phrase database are loaded into a trie tree. In step S320, source language phrases existing in the trie tree in the source language sentence are extracted. In step S330, the context representation of the source language phrases in the source language sentence is calculated. In step S340, bilingual phrases are retrieved from the bilingual phrase database based on the context representation of the source language phrases in the source language sentence.
[0040] Generally, the most similar bilingual phrase can be retrieved from the bilingual phrase database based on the L 2 distance of the context representation of the source language phrase. However, in the case where the bilingual phrase database is constructed based on the original training data, the second most similar bilingual phrase is retrieved from the bilingual phrase database based on the context representation of the source language phrase to avoid retrieving bilingual phrases extracted from the current translation sample and overfitting the retrieved bilingual phrases.
[0041] After the bilingual phrase is retrieved, a bilingual phrase hint can be constructed using it and the constructed bilingual phrase hint can be concatenated to the source language sentence for training to generate new training data with the bilingual phrase hint. The new training data includes the source language sentence concatenated with the bilingual phrase hint and the target language sentence.
[0042] Return to Figure 2 , in step S230, pre-train the machine translation model using at least new training data with bilingual phrase hints. In some embodiments, the machine translation model can be pre-trained using both the original training data and the new training data with bilingual phrase hints to obtain a model with better prediction accuracy.
[0043] The machine translation model can be pre-trained for one or more rounds (e.g., 10 rounds) to obtain a model with better prediction accuracy. Cross-entropy loss can be used during pre-training. The machine translation model can be an encoder-decoder model. The encoder-decoder model can be pre-trained for prompt-based learning in machine translation. As described by Vaswani et al. in "Attention is all you need. In Proc. of NeurIPS, pages 5998–6008", the encoder-decoder model can utilize the Transformer architecture. The machine translation model can be implemented based on the Fairseq method described by Ott et al. in "fairseq: A fast, extensible toolkit for sequence modeling. In Proc. of NAACL-Demonstrations, pages 48–53". To perform efficient bilingual phrase retrieval, refer to the content described by Johnson et al. in "Billion-scale similarity search with gpus. arXiv preprint arXiv:1702.08734" to construct an IVFPQ index with FAISS.
[0044] In some embodiments of the present disclosure, the method 200 for machine translation may further include step S240. In step S240, receive the source language sentence to be translated, and use the pre-trained machine translation model to translate the source language sentence to be translated into the target language sentence to be output. In some embodiments of the present disclosure, the pre-trained machine translation model can translate the source language sentence with bilingual phrase hints. In this case, concatenate the bilingual phrase hint related to the source language sentence to be translated to the source language sentence to be translated to generate the source language sentence to be translated with bilingual phrase hints. Then, input the source language sentence to be translated with bilingual phrase hints into the pre-trained machine translation model. In some embodiments of the present disclosure, the pre-trained machine translation model can translate the source language sentence without bilingual phrase hints. In this case, input the source language sentence to be translated into the pre-trained machine translation model.
[0045] In some embodiments of the present disclosure, bilingual phrase hints are manually created for translation intervention. For example, lexical constraints can be represented as bilingual phrase hints to intervene in the lexical selection during translation. Suppose the lexical constraint is to translate the word x in the input sentence into the target word y in the output sentence. This lexical constraint can be represented as the bilingual phrase hint "x <q>"y", and translate the input sentence with such a hint. Lexical constraints can be specified as soft lexical constraints. Compared with hard lexical constraints, the advantage of soft lexical constraints is that the word forms of phrases do not need to be specified.
[0046] In some embodiments of the present disclosure, bilingual phrase hints related to the source language sentence to be translated are constructed by retrieving bilingual phrases from a bilingual phrase database. Figure 4 FIG. 4 is a flowchart showing a process 400 of translating a source language phrase into a target language phrase based on a bilingual phrase database. First, the source language phrase 402 in the source language sentence 401 to be translated is extracted, and the bilingual phrase 405 is retrieved from the bilingual phrase database 404 to construct the bilingual phrase hint 406. The most similar bilingual phrase 405 can be retrieved from the bilingual phrase database 404 by calculating the context representation 403 of the source language phrase 402 and based on the context representation 403. Then, the constructed bilingual phrase hint 406 is concatenated to the source language sentence 401 to be translated to generate a source language sentence 407 with a bilingual phrase hint. Finally, the source language sentence 407 with the bilingual phrase hint is input into a pre-trained machine translation model 408 to be translated into the target language sentence 409 to be output.
[0047] Table 1 compares the BLEU scores of the PAPT scheme of the embodiments of the present disclosure with the existing Vanilla scheme. In Table 1, the database size represents the number of bilingual phrases in the database. PAPT (without hint) means that the bilingual phrase hint is not concatenated to the source language sentence to be translated during the translation stage. PAPT (with hint) means that the bilingual phrase hint is concatenated to the source language sentence to be translated during the translation stage. Table 2 shows the sizes of the datasets used for training, development, and testing respectively.
[0048] As shown in Table 1, when translating English to German and German to English in a specific domain, the PAPT (without hint) scheme and the Vanilla scheme have comparable performance. The PAPT (with hint) scheme can be 6.7 points and 5.6 points higher than the average BLEU score of the Vanilla scheme respectively without additional training. The results show that bilingual phrase hints are helpful for machine translation in specific domains.
[0049] Table 1
[0050]
[0051] Table 2
[0052]
[0053] Table 3 compares the performance of the PAPT scheme of the embodiments of the present disclosure with the existing Vanilla scheme in English-German translation with lexical constraints. The test set used in Table 3 was extracted by Susanto et al. from the Wiktionary and IATE (Interactive Terminology for Europe) terminology databases in "Lexically constrained neural machine translation with levenshtein transformer. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3536–3543”. As shown in Table 3, compared with Vanilla, the PAPT scheme of the embodiments of the present disclosure achieved accuracy improvements of 10.7% and 12.3% in Wiktionary and IATE respectively without additional training. This accuracy represents the ratio of the target phrase appearing in the translation output. The overall translation performance measured by the BLEU score improved slightly. This indicates that PAPT can effectively incorporate lexical constraints into the translation process. Moreover, by concatenating multiple bilingual phrases as cues, PAPT can conveniently incorporate multiple lexical constraints.
[0054] Table 3
[0055]
[0056] Figure 5 shows an example of English-German translation in the medical field according to an embodiment of the present disclosure. As can be seen from Figure 5 it, PAPT will output different target language sentences under different bilingual phrase cues. It can be seen that the translation process can be intervened by modifying the bilingual phrase cues.
[0057] Figure 6 is a block diagram showing an apparatus 600 for machine translation according to an embodiment of the present disclosure. As Figure 6 As shown, the device 600 includes an original training data acquisition unit 601, a new training data generation unit 602, and a pre-training unit 603. The original training data acquisition unit 601 is configured to acquire original training data including source language sentences for training and target language sentences for training. The new training data generation unit 602 is configured to concatenate bilingual phrase prompts related to the source language sentences for training to the source language sentences for training to generate new training data with bilingual phrase prompts. The bilingual phrase prompts include one or more bilingual phrases, and each bilingual phrase includes a source language phrase and a corresponding target language phrase. The pre-training unit 603 is configured to pre-train a machine translation model at least using the new training data with bilingual phrase prompts.
[0058] In some embodiments of the present disclosure, the device 600 may further include a translation unit 604. It is configured to receive a source language sentence to be translated and translate the source language sentence to be translated into a target language sentence to be output using the pre-trained machine translation model.
[0059] Since Figure 6 The specific implementation manners of the operations performed by each unit above have been described in detail before, so they will not be elaborated here.
[0060] As described above, the present disclosure discovers the difficulties in applying prompt-based learning to machine translation. The present disclosure proposes effective methods to solve this difficulty. Data shows that the technical solutions of the present disclosure can effectively improve machine translation in specific fields and machine translation based on lexical constraints without additional training.
[0061] It should be noted that the above-mentioned each unit is only a logical module divided according to its specific implemented functions, rather than for limiting the specific implementation manners. For example, it can be implemented in a software, hardware, or a combination of software and hardware manner. In actual implementation, the above-mentioned each unit can be implemented as an independent physical entity, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, the above-mentioned each unit is shown by a dotted line in the drawings indicating that these units may not actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.
[0062] In addition, although not shown, the device may also include a memory, which may store various information generated during the operation of the device and each unit included in the device, programs and data for operation, data to be transmitted by the communication unit, and the like. The memory may be a volatile memory and / or a non-volatile memory. For example, the memory may include, but is not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Of course, the memory may also be located outside the device. Optionally, although not shown, the device may also include a communication unit, which may be used to communicate with other devices. In one example, the communication unit may be implemented in a suitable manner known in the art, such as including communication components such as an antenna array and / or a radio frequency link, various types of interfaces, communication units, and the like. This will not be described in detail here. In addition, the device may also include other components not shown, such as a radio frequency link, a baseband processing unit, a network interface, a processor, a controller, and the like. This will not be described in detail here.
[0063] Some embodiments of the present disclosure also provide an electronic device. Figure 7 FIG. is a block diagram showing an electronic device according to some embodiments of the present disclosure. For example, in some embodiments, the electronic device 700 may be various types of devices, such as, for example, including but not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. For example, the electronic device 700 may include a display panel for displaying data and / or execution results utilized in the solution according to the present disclosure. For example, the display panel may be of various shapes, such as a rectangular panel, an oval panel, or a polygonal panel, etc. In addition, the display panel may not only be a flat panel, but also a curved panel, or even a spherical panel.
[0064] As Figure 7 shown, the electronic device 700 of this embodiment includes: a memory 701 and a processor 702 coupled to the memory 701. It should be noted that Figure 7 the components of the electronic device 700 shown are only exemplary and not restrictive. According to actual application requirements, the electronic device 700 may also have other components. The processor 702 may control other components in the electronic device 700 to perform desired functions.
[0065] In some embodiments, the memory 701 is used to store one or more computer-readable instructions. When the processor 702 is used to run the computer-readable instructions, the computer-readable instructions, when run by the processor 702, implement the method according to any of the above embodiments. For the specific implementation of each step of the method and the related explanatory content, reference may be made to the above embodiments, and the repeated parts will not be elaborated here.
[0066] For example, the processor 702 and the memory 701 can communicate with each other directly or indirectly. For example, the processor 702 and the memory 701 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 702 and the memory 701 can also communicate with each other through a system bus, and the present disclosure does not limit this.
[0067] For example, the processor 702 can be embodied as various appropriate processors, processing devices, etc., such as a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) can be of the X86 or ARM architecture, etc. For example, the memory 701 can include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The memory 701 can include, for example, a system memory, and the system memory stores, for example, an operating system, application programs, a boot loader, a database, and other programs. Various application programs and various data can also be stored in the storage medium.
[0068] In addition, according to some embodiments of the present disclosure, when various operations / processes according to the present disclosure are implemented by software and / or firmware, they can be installed from a storage medium or a network to a computer system with a dedicated hardware structure, such as Figure 8 the computer system 800 shown, and when various programs are installed in the computer system, it can perform various functions, including the functions described above, etc. Figure 8 is a block diagram showing an example structure of a computer system that can be adopted in the embodiments of the present disclosure.
[0069] In Figure 8 In this case, the central processing unit (CPU) 801 performs various processes according to a program stored in the read-only memory (ROM) 802 or a program loaded from the storage section 808 into the random access memory (RAM) 803. In the RAM 803, data required when the CPU 801 performs various processes and the like is also stored as needed. The central processing unit is merely exemplary, and it may also be other types of processors, such as the various processors described above. The ROM 802, RAM 803, and storage section 808 may be various forms of computer-readable storage media, as described below. It should be noted that although Figure 8 the ROM 802, RAM 803, and storage device 808 are shown separately in the figure, one or more of them may be combined or located in the same or different memories or storage modules.
[0070] The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The input / output interface 805 is also connected to the bus 804.
[0071] The following components are connected to the input / output interface 805: an input section 806, such as a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output section 807, including a display, such as a cathode ray tube (CRT), liquid crystal display (LCD), speaker, vibrator, etc.; a storage section 808, including a hard disk, magnetic tape, etc.; and a communication section 809, including a network interface card such as a LAN card, modem, etc. The communication section 809 allows communication processing to be performed via a network such as the Internet. It is easily understood that although Figure 8 the various devices or modules in the computer system 800 are shown to communicate via the bus 804 in the figure, they may also communicate via a network or other means, where the network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0072] As needed, a drive 810 is also connected to the input / output interface 805. A removable medium 811, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is mounted on the drive 810 as needed, so that a computer program read therefrom is installed in the storage section 808 as needed.
[0073] In the case where the above series of processes are implemented by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 811.
[0074] According to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a CPU 801, the above-described functions defined in the method of the embodiment of the present disclosure are performed.
[0075] It should be noted that in the context of the present disclosure, a computer-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium may be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0076] The above computer-readable medium may be included in the above electronic device; or it may exist separately and not be assembled into the electronic device.
[0077] In some embodiments, a computer program is also provided, including: instructions that, when executed by a processor, cause the processor to execute the method of any one of the above embodiments. For example, the instructions may be embodied as computer program code.
[0078] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0079] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0080] The modules, components, or units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the names of the modules, components, or units do not, in some cases, constitute a limitation on the modules, components, or units themselves.
[0081] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0082] According to some embodiments of the present disclosure, a method for machine translation is provided, including: obtaining original training data including a source language sentence for training and a target language sentence for training; concatenating a bilingual phrase prompt related to the source language sentence for training to the source language sentence for training to generate new training data with a bilingual phrase prompt, where the bilingual phrase prompt includes one or more bilingual phrases, and each bilingual phrase includes a source language phrase and a corresponding target language phrase; pre-training a machine translation model at least using the new training data with a bilingual phrase prompt.
[0083] According to some embodiments of the present disclosure, the one or more bilingual phrases are separated by a first marker, the source language phrase and the corresponding target language phrase are separated by a second marker, and the bilingual phrase prompt and the source language sentence for training are separated by a third marker.
[0084] According to some embodiments of the present disclosure, bilingual phrases are retrieved from a bilingual phrase database to construct the bilingual phrase prompt.
[0085] According to some embodiments of the present disclosure, the context representation of the source language phrase of each bilingual phrase and the bilingual phrase are stored in the bilingual phrase database as key-value pairs.
[0086] According to some embodiments of the present disclosure, retrieving bilingual phrases from a bilingual phrase database includes: loading the source language phrases in the bilingual phrase database into a trie tree; extracting the source language phrases existing in the source language sentence in the trie tree; calculating the context representation of the source language phrases in the source language sentence; and retrieving bilingual phrases from the bilingual phrase database based on the context representation of the source language phrases in the source language sentence.
[0087] According to some embodiments of the present disclosure, in the case where the bilingual phrase database is not constructed based on the original training data, retrieving the most similar bilingual phrases from the bilingual phrase database based on the context representation of the source language phrase, and in the case where the bilingual phrase database is constructed based on the original training data, retrieving the second most similar bilingual phrases from the bilingual phrase database based on the context representation of the source language phrase.
[0088] According to some embodiments of the present disclosure, the machine translation model is an encoder-decoder model.
[0089] According to some embodiments of the present disclosure, pre-training a machine translation model using at least new training data with bilingual phrase hints includes: pre-training the machine translation model using both original training data and new training data with bilingual phrase hints.
[0090] According to some embodiments of the present disclosure, a source language sentence to be translated is received, and the source language sentence to be translated is translated into a target language sentence to be output using the pre-trained machine translation model.
[0091] According to some embodiments of the present disclosure, a bilingual phrase hint related to the source language sentence to be translated is concatenated to the source language sentence to be translated to generate a source language sentence to be translated with a bilingual phrase hint; and the source language sentence to be translated with a bilingual phrase hint is input into the pre-trained machine translation model.
[0092] According to some embodiments of the present disclosure, the bilingual phrase hint related to the source language sentence to be translated is manually created.
[0093] According to some embodiments of the present disclosure, a source language phrase in the source language sentence to be translated is extracted; and a bilingual phrase is retrieved from a bilingual phrase database to construct a bilingual phrase hint related to the source language sentence to be translated.
[0094] According to some embodiments of the present disclosure, a context representation of the source language phrase in the source language sentence to be translated is calculated, wherein retrieving a bilingual phrase from the bilingual phrase database includes retrieving the most similar bilingual phrase from the bilingual phrase database based on the context representation of the source language phrase in the source language sentence to be translated.
[0095] According to some embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the processor performs the method according to the embodiments of the present disclosure.
[0096] According to some embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method according to the embodiments of the present disclosure is implemented.
[0097] According to some embodiments of the present disclosure, there is provided an apparatus for machine translation, including: an original training data acquisition unit configured to acquire original training data including source language sentences for training and target language sentences for training; a new training data generation unit configured to concatenate bilingual phrase prompts related to the source language sentences for training to the source language sentences for training to generate new training data with bilingual phrase prompts, where the bilingual phrase prompts include one or more bilingual phrases, and each bilingual phrase includes a source language phrase and a corresponding target language phrase; and a pre-training unit configured to pre-train a machine translation model at least using the new training data with bilingual phrase prompts.
[0098] According to some embodiments of the present disclosure, there is provided a computer program, including: instructions that, when executed by a processor, cause the processor to execute the method according to the embodiments of the present disclosure.
[0099] According to some embodiments of the present disclosure, there is provided a computer program product including instructions that, when executed by a processor, implement the method according to the embodiments of the present disclosure.
[0100] The above description is only some embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present disclosure.
[0101] In the description provided herein, many specific details are set forth. However, it is understood that the embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail in order not to obscure the understanding of this description.
[0102] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0103] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments can be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.< / q> < / q> < / r>
Claims
1. A method for machine translation, comprising: Obtaining original training data including source language sentences for training and target language sentences for training; Concatenating bilingual phrase prompts related to the source language sentences for training to the source language sentences for training to generate new training data with bilingual phrase prompts, wherein the bilingual phrase prompts include one or more bilingual phrases, and each bilingual phrase includes a source language phrase and a corresponding target language phrase; Pre-training a machine translation model at least using the new training data with bilingual phrase prompts, wherein the input of the machine translation model is the source language sentence concatenated with the bilingual phrase prompt, and the output of the machine translation model is the target language sentence.
2. The method according to claim 1, wherein, The one or more bilingual phrases are separated by a first marker, the source language phrase and the corresponding target language phrase are separated by a second marker, and the bilingual phrase prompt and the source language sentence for training are separated by a third marker.
3. The method according to claim 1, further comprising: Retrieving bilingual phrases from a bilingual phrase database to construct the bilingual phrase prompt.
4. The method according to claim 3, further comprising: Storing the context representation of the source language phrase of each bilingual phrase and the bilingual phrase as a key-value pair in the bilingual phrase database.
5. The method according to claim 3, wherein, Retrieving bilingual phrases from a bilingual phrase database includes: Loading the source language phrases in the bilingual phrase database into a trie; Extracting the source language phrases existing in the trie in the source language sentence; Calculating the context representation of the source language phrases in the source language sentence; and Retrieving bilingual phrases from the bilingual phrase database based on the context representation of the source language phrases in the source language sentence.
6. The method according to claim 3, wherein in the case where the bilingual phrase database is not constructed based on the original training data, retrieving the most similar bilingual phrase from the bilingual phrase database based on the context representation of the source language phrase, and in the case where the bilingual phrase database is constructed based on the original training data, retrieving the second most similar bilingual phrase from the bilingual phrase database based on the context representation of the source language phrase.
7. The method according to claim 1, wherein, The machine translation model is an encoder-decoder model.
8. The method according to claim 1, wherein Pre-training the machine translation model at least using the new training data with bilingual phrase prompts includes: Pre-training the machine translation model using both the original training data and the new training data with bilingual phrase prompts.
9. The method according to claim 1, further comprising: Receiving a source language sentence to be translated, and translating the source language sentence to be translated into a target language sentence to be output using the pre-trained machine translation model.
10. The method according to claim 9, further comprising: Concatenating a bilingual phrase prompt related to the source language sentence to be translated to the source language sentence to be translated to generate a source language sentence to be translated with a bilingual phrase prompt; and Inputting the source language sentence to be translated with a bilingual phrase prompt into the pre-trained machine translation model.
11. The method according to claim 10, wherein, The bilingual phrase prompt related to the source language sentence to be translated is manually created.
12. The method according to claim 10, wherein the method further comprises: extracting source language phrases in the source language sentence to be translated; and retrieving bilingual phrases from a bilingual phrase database to construct a bilingual phrase hint related to the source language sentence to be translated.
13. The method according to claim 12, further comprising: calculating a context representation of the source language phrases in the source language sentence to be translated, wherein retrieving bilingual phrases from the bilingual phrase database includes retrieving the most similar bilingual phrases from the bilingual phrase database based on the context representation of the source language phrases in the source language sentence to be translated.
14. An electronic device, comprising: a memory; and a processor coupled to the memory, wherein instructions are stored in the memory, and when executed by the processor, the instructions cause the processor to execute the method according to any one of claims 1-13.
15. A non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method according to any one of claims 1-13 is implemented.
16. A device for machine translation, comprising: an original training data acquisition unit configured to acquire original training data including a source language sentence for training and a target language sentence for training; a new training data generation unit configured to concatenate a bilingual phrase hint related to the source language sentence for training to the source language sentence for training to generate new training data with a bilingual phrase hint, wherein the bilingual phrase hint includes one or more bilingual phrases, and each bilingual phrase includes a source language phrase and a corresponding target language phrase; and a pre-training unit configured to pre-train a machine translation model at least using the new training data with a bilingual phrase hint, wherein the input of the machine translation model is the source language sentence concatenated with the bilingual phrase hint, and the output of the machine translation model is the target language sentence.
Citation Information
Patent Citations
Text translation method and device based on artificial intelligence, medium and electronic equipment
CN112163434A