Training method of translation model, its translation method, device and electronic device

By introducing two-way global context information and confidence thresholds into the translation model, the problem of insufficient accuracy of the translation model is solved, and higher translation accuracy and performance improvements are achieved.

CN113705256BActive Publication Date: 2025-07-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110341084.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-30
Publication Date
2025-07-22
Estimated Expiration
2041-03-30

AI Technical Summary

Technical Problem

The translation accuracy of existing translation models is low, making it difficult to meet users' increasing translation quality needs.

Method used

By introducing two-way global context information into the translation model, the context word at the position to be predicted is determined using the confidence threshold, and forward propagation is performed in the second translation model, and the model parameters are updated in combination with the confidence of the first translation model and the confidence of the second translation model, targeted introduction of two-way global context information is achieved.

Benefits of technology

It improves the accuracy of the translation model, so that it can not only utilize local context information during translation, but also effectively utilize global context information, thereby improving translation performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113705256B_ABST
    Figure CN113705256B_ABST
Patent Text Reader

Abstract

The present application provides a training method for a translation model, as well as a corpus translation method, apparatus, electronic device, and computer-readable storage medium thereof; the method includes: performing forward propagation on a corpus sample in a first translation model to obtain a first confidence level of a corresponding first pre-labeled target word; determining a second to-be-predicted position for a first to-be-predicted position with a first confidence level lower than a confidence threshold, and determining a first pre-labeled target word with a first confidence level not lower than the confidence threshold as a context word for a corresponding second translation model; performing forward propagation on the context word and the corpus sample in the second translation model to obtain a second confidence level of a corresponding second pre-labeled target word; updating parameters of the first translation model and the second translation model based on the first confidence level of the corresponding first pre-labeled target word and the second confidence level of the corresponding second pre-labeled target word. Through the present application, the accuracy of corpus translation by the translation model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to artificial intelligence technology, and in particular, to a training method for a translation model, a translation method, a device, an electronic device, and a computer-readable storage medium thereof. Background Art

[0002] Artificial Intelligence (AI) is a comprehensive technology in computer science. By studying the design principles and implementation methods of various intelligent machines, machines are enabled to have functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, such as natural language processing technology and machine learning / deep learning. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0003] In related technologies, natural language processing technology is applied to provide users with a corpus translation function in various application products. However, at present, the translation accuracy of the translation model trained in natural language processing technology is relatively low, and it is difficult to meet the increasing translation quality requirements of users.

[0004] Content of the Application

[0005] Embodiments of the present application provide a training method for a translation model, a translation method, a device, an electronic device, and a computer-readable storage medium thereof, which can improve the accuracy of corpus translation by the translation model.

[0006] The technical solution of the embodiments of the present application is implemented as follows:

[0007] Embodiments of the present application provide a training method for a translation model, including:

[0008] Performing forward propagation on a corpus sample in a first translation model to obtain a first confidence level of a first pre-marked target word corresponding to each first position to be predicted;

[0009] Determining the first positions to be predicted with the first confidence level lower than a confidence level threshold as second positions to be predicted of a second translation model, and determining the first pre-marked target words corresponding to the first positions to be predicted with the first confidence level not lower than the confidence level threshold as context words corresponding to the second translation model;

[0010] Performing forward propagation on the context words and the corpus sample in the second translation model to obtain a second confidence level of a second pre-marked target word corresponding to each second position to be predicted;

[0011] Update the parameters of the first translation model and the second translation model based on the first confidence of the first pre-labeled target word corresponding to each first position to be predicted and the second confidence of the second pre-labeled target word corresponding to each second position to be predicted.

[0012] An embodiment of the present application provides a training device for a translation model, including:

[0013] A first task module, configured to perform forward propagation of a corpus sample in a first translation model to obtain a first confidence of a first pre-labeled target word corresponding to each first position to be predicted;

[0014] A selection module, configured to determine, as a second position to be predicted of a second translation model, a first position to be predicted with a first confidence lower than a confidence threshold, and determine, as a context word corresponding to the second translation model, a first pre-labeled target word corresponding to a first position to be predicted with a first confidence not lower than the confidence threshold;

[0015] A second task module, configured to perform forward propagation of the context word and the corpus sample in the second translation model to obtain a second confidence of a second pre-labeled target word corresponding to each second position to be predicted;

[0016] An update module, configured to update the parameters of the first translation model and the second translation model based on the first confidence of the first pre-labeled target word corresponding to each first position to be predicted and the second confidence of the second pre-labeled target word corresponding to each second position to be predicted.

[0017] In the above solution, the update module is further configured to: determine a first loss corresponding to the first translation model based on the first confidence; determine a second loss corresponding to the second translation model based on the first confidence lower than the confidence threshold and the second confidence of the second pre-labeled target word corresponding to each second position to be predicted, where the second loss is used to represent the teaching loss of the second translation model to the first translation model; perform an aggregation process on the first loss and the second loss based on aggregation parameters corresponding to the first loss and the second loss respectively to obtain a joint loss; and update the parameters of the first translation model and the second translation model according to the joint loss.

[0018] In the above solution, there are multiple first positions to be predicted with corresponding multiple first pre-labeled target words; the update module is further configured to: perform a fusion process on the first confidences obtained for each first position to be predicted to obtain a first loss corresponding to the first translation model.

[0019] In the above solution, there are multiple second target words to be predicted corresponding one-to-one to multiple second pre-labeled target words; the updating module is further configured to: fuse the first confidence level lower than the confidence threshold and the second confidence level of the second pre-labeled target word corresponding to each of the second positions to be predicted, so as to obtain a second loss corresponding to the second translation model.

[0020] In the above solution, the first translation model includes a first encoding network and a pre-order decoding network; the first task module is further configured to: determine each original word of the corpus sample and the original word vector corresponding to each of the original words, combine the original word vectors corresponding to each of the original words to obtain an original word vector sequence of the corpus sample; perform semantic encoding processing on the original word vector sequence of the corpus sample through the first encoding network to obtain a first source sentence representation corresponding to the corpus sample; perform corpus decoding processing on the first source sentence representation through the pre-order decoding network to obtain a first confidence level of the first pre-labeled target word corresponding to each of the first positions to be predicted; wherein, the first confidence level is generated based on the previous words corresponding to each of the first positions to be predicted.

[0021] In the above solution, the encoding network includes N cascaded sub-encoding networks, where N is an integer greater than or equal to 2; the first task module is further configured to: perform semantic encoding processing on the original word vector sequence of the corpus sample in the following manner through the N cascaded first sub-encoding networks included in the first encoding network: perform self-attention processing on the input of the first sub-encoding network to obtain a self-attention processing result corresponding to the first sub-encoding network, perform hidden state mapping processing on the self-attention processing result to obtain a hidden state vector sequence corresponding to the first sub-encoding network, and use the hidden state vector sequence as the semantic encoding processing result of the first sub-encoding network; wherein, in the N cascaded first sub-encoding networks, the input of the first first sub-encoding network includes the original word vector sequence of the corpus sample, and the semantic encoding processing result of the Nth first sub-encoding network includes a first source sentence representation corresponding to the corpus sample.

[0022] In the above solution, the first task module is further configured to perform the following processing for each of the original words in the corpus sample: perform a linear transformation on the first intermediate vector corresponding to the original word in the input of the first sub-encoding network to obtain a query vector, a key vector, and a value vector corresponding to the original word; perform a dot product on the query vector of the original word and the key vector of each original word, and perform a normalization process based on the maximum likelihood function on the dot product result to obtain the weight of the value vector of the original word; perform a weighted process on the value vector of the original word based on the weight of the value vector of the original word to obtain the self-attention processing result of the sub-encoding network corresponding to each original word.

[0023] In the above solution, the first task module is further configured to perform the following processing for each first to-be-predicted position output by the previous decoding network: obtain a first pre-labeled target word sequence corresponding to the corpus sample from the corpus sample set; extract the first pre-labeled target word before the first to-be-predicted position from the first pre-labeled target word sequence, and use the extracted first pre-labeled target word as the previous word corresponding to the first to-be-predicted position; perform semantic decoding on the previous word corresponding to the first to-be-predicted position and the first source sentence representation through the previous decoding network to obtain a first confidence level that is decoded as the corresponding first pre-labeled target word at the first to-be-predicted position.

[0024] In the above solution, the previous decoding network includes M cascaded sub-previous decoding networks, where M is an integer greater than or equal to 2; the first task module is further configured to perform semantic decoding in the following manner through each sub-previous decoding network: perform masked self-attention processing on the input of the sub-previous decoding network to obtain a masked self-attention processing result corresponding to the sub-previous decoding network, perform cross-attention processing on the masked self-attention processing result to obtain a cross-attention processing result corresponding to the sub-previous decoding network, and perform a hidden state mapping process on the cross-attention processing result; where, in the M cascaded sub-previous decoding networks, the input of the first sub-previous decoding network includes: the previous word corresponding to the first to-be-predicted position and the first source sentence representation; the hidden state mapping processing result of the Mth sub-previous decoding network includes: a first confidence level that is decoded as the corresponding first pre-labeled target word at the first to-be-predicted position.

[0025] In the above solution, the first task module is further configured to: perform a linear transformation on the masked self-attention processing result to obtain a query vector of the masked self-attention processing result; for each of the original words, perform the following processing: perform a linear transformation on the first source sentence representation of the original word to obtain a key vector and a value vector of the first source sentence representation; perform a dot product on the query vector of the masked self-attention processing result and the key vector of the first source sentence representation, and perform normalization processing based on the maximum likelihood function on the dot product result to obtain the weight of the value vector of the first source sentence representation; perform a weighted processing on the value vector of the first source sentence representation based on the weight of the value vector of the first source sentence representation to obtain the cross-attention processing result corresponding to the sub-previous decoding network.

[0026] In the above solution, the second translation model includes a second encoding network and a context decoding network; the second task module is further configured to: obtain each original word of the corpus sample and the original word vector corresponding to each original word, combine the original word vectors corresponding to each original word to obtain the original word vector sequence of the corpus sample; perform semantic encoding on the original word vector sequence of the corpus sample through the second encoding network to obtain a second source sentence representation corresponding to the corpus sample; perform corpus decoding on the second source sentence representation through the context decoding network to obtain the second confidence of the second pre-marked target word corresponding to each second to-be-predicted position; wherein, the second confidence is generated based on context words corresponding to a plurality of the second to-be-predicted positions.

[0027] In the above solution, the second task module is further configured to: for each second to-be-predicted position output by the context decoding network, perform the following processing: perform semantic decoding on the context word and the second source sentence representation through the context decoding network to obtain the second confidence that the second to-be-predicted position is decoded as the corresponding second pre-marked target word.

[0028] In the above solution, the context decoding network includes P cascaded sub-context decoding networks, where P is an integer greater than or equal to 2; the second task module is further configured to: through the P cascaded sub-context decoding networks, perform semantic decoding processing on the context word set corresponding to the second to-be-predicted position and the source sentence representation in the following manner: perform context self-attention processing on the input of the sub-context decoding network to obtain a context self-attention processing result corresponding to the sub-context decoding network; perform cross-attention processing on the context self-attention processing result to obtain a cross-attention processing result corresponding to the sub-context decoding network; perform hidden state mapping processing on the cross-attention processing result; wherein, in the P cascaded sub-context decoding networks, the input of the first sub-context decoding network includes: the context word set corresponding to the second to-be-predicted position and the second source sentence representation; the hidden state mapping processing result of the P-th sub-context decoding network includes: the second confidence level that the second to-be-predicted position is decoded into the corresponding second pre-marked target word.

[0029] An embodiment of the present application provides a corpus translation method, including:

[0030] In response to a translation request for a target corpus, call a first translation model or a second translation model to perform translation processing on the target corpus to obtain a translation result for the target corpus;

[0031] wherein, the first translation model and the second translation model are trained according to the translation model training method provided by the embodiments of the present application.

[0032] An embodiment of the present application provides a corpus translation device, including:

[0033] An application module, configured to, in response to a translation request for a target corpus, call a first translation model or a second translation model to perform translation processing on the target corpus to obtain a translation result for the target corpus; wherein, the first translation model and the second translation model are trained according to the translation model training method provided by the embodiments of the present application.

[0034] An embodiment of the present application provides an electronic device, including:

[0035] A memory, configured to store executable instructions;

[0036] A processor, configured to, when executing the executable instructions stored in the memory, implement the translation model training method or the corpus translation method provided by the embodiments of the present application.

[0037] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the training method of the translation model or the corpus translation method provided by the embodiment of the present application.

[0038] The embodiment of the present application has the following beneficial effects:

[0039] By using the characteristics and confidence of the translation task of a neural network model, another neural network model is assisted in training in a targeted manner. Since the second translation model uses the context word set for translation, bidirectional global context information is effectively introduced through the context word set, and through the confidence threshold, the second translation model introduces bidirectional global context information based on context words for the first translation model at positions with relatively low confidence at the target end, so that the first translation model after joint training can not only use the local context information of the previous word corresponding to each position to be predicted during translation, but also use global context information in a targeted manner, thereby effectively improving the accuracy of translation by the first translation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic structural diagram of the corpus translation system provided by the embodiment of the present application;

[0041] Figure 2 It is a schematic structural diagram of the server 200 provided by the embodiment of the present application;

[0042] Figures 3A - 3D It is a schematic flowchart of the training method of the translation model provided by the embodiment of the present application;

[0043] Figure 4 It is a schematic structural diagram of the joint training model of the training method of the translation model provided by the embodiment of the present application;

[0044] Figure 5 It is a schematic structural diagram of the first translation model provided by the embodiment of the present application;

[0045] Figure 6 It is a schematic structural diagram of the sub-encoding network provided by the embodiment of the present application;

[0046] Figure 7 It is a schematic structural diagram of the sub-previous decoding network provided by the embodiment of the present application;

[0047] Figure 8 It is a schematic structural diagram of the sub-context decoding network provided by the embodiment of the present application;

[0048] Figure 9 It is a confidence distribution diagram provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] To make the objectives, technical solutions and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0050] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0051] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0053] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are described. The nouns and terms involved in the embodiments of this application are subject to the following explanations.

[0054] 1) Neural machine translation: Neural Machine Translation (NMT) is a machine translation method proposed in recent years. Compared with traditional statistical machine translation, NMT can train a neural network that can map from one sequence to another sequence, and the output can be a variable-length sequence, which can achieve very good performance in translation, dialogue and text summarization.

[0055] 2) Word embedding vector: It is an important concept in natural language processing technology. A word can be converted into a fixed-length vector representation using a word embedding vector, which is convenient for mathematical processing.

[0056] In the related art, prediction is performed in a left-to-right manner solely relying on the local historical context of previous words, and there are still deficiencies in how to more effectively utilize bidirectional global context information. In the related art, a reverse decoder for reverse decoding is introduced at the target end of the translation model. The reverse decoder first generates a sequence of hidden state vectors from right to left, and then the forward decoder uses the sequence of hidden state vectors from right to left for left-to-right decoding processing, so that the forward decoding process fully considers the subsequent information at the target end to improve the translation quality. Since only additional reverse global context information is introduced for the forward decoder by the reverse decoder at each decoding moment in the related art, these reverse global context information and the local context information of previous words are actually independent of each other. The translation model cannot effectively comprehensively consider the reverse global context information and the local context information of previous words, thus unable to effectively improve the translation quality. Moreover, the related art does not consider the confidence of the first translation model itself in predicting the first pre-marked target word, so additional bidirectional global context information is introduced at the decoding moment where bidirectional global context information is not necessarily introduced during the joint training process.

[0057] The embodiments of the present application provide a training method, device, electronic device, and computer-readable storage medium for a translation model, which can, through confidence-based knowledge distillation, specifically introduce bidirectional global context information (based on context words) for the positions where the confidence of the neural machine translation model in predicting the target answer at the target end is relatively low, so that the translation model can not only utilize the local context information of the corresponding previous word but also the global context information of the corresponding context word when making predictions at each position, thereby improving the translation performance. Next, an exemplary application of the electronic device provided by the embodiments of the present application will be described. The electronic device provided by the embodiments of the present application can be implemented as a server. Next, the exemplary application when the electronic device is implemented as a server will be described.

[0058] See Figure 1 , Figure 1 is a schematic structural diagram of the corpus translation system provided by the embodiments of the present application. The corpus translation system can be used in social scenarios. In the corpus translation system, the terminal 400 is connected to the server 200 through the network 300. The network can be a wide area network, a local area network, or a combination of the two.

[0059] In some embodiments, the function of the corpus translation system is implemented based on various modules in the server 200. When the user uses the terminal 400, the terminal 400 collects corpus samples and sends them to the server 200. The server 200 performs joint training on the translation model (the first translation model or the second translation model) based on multiple tasks and confidence levels, and integrates the trained first translation model or the second translation model into the server 200. In response to the terminal 400 receiving a translation operation for a corpus signal in the social client, the terminal 400 sends the corpus signal to the server 200. The server 200 determines the corpus translation result of the corpus signal through the translation model and sends it to the terminal 400, so that the terminal 400 directly presents the corpus translation result.

[0060] In some embodiments, when the corpus translation system is applied to a social scenario, terminal 400 receives a corpus signal sent by other terminals. In response to terminal 400 receiving a translation operation for the corpus signal, terminal 400 sends the corpus signal to server 200. Server 200 determines a corpus translation result of the corpus signal through a translation model, and sends it to terminal 400, so that terminal 400 directly presents the corpus translation result. For example, terminal 400 receives a corpus signal "where are you" sent by other terminals, and terminal 400 sends the corpus signal to server 200. Server 200 determines a corpus translation result of the corpus signal "where are you" through a translation model, and sends it to terminal 400, so that terminal 400 directly presents the corpus translation result "where are you".

[0061] In some embodiments, when the corpus translation system is applied to a web browsing scenario, the terminal 400 presents an English web page. In response to the terminal 400 receiving a translation operation for a corpus signal in the English web page, the terminal 400 sends the corpus signal to the server 200. The server 200 determines the corpus translation result of the corpus signal through the translation model and sends it to the terminal 400, so that the terminal 400 directly presents the corpus translation result.

[0062] In other embodiments, after completing the training process of the translation model, the server 200 sends the translation model to the terminal 400, so that the terminal 400 runs the jointly trained translation model to determine the corpus translation result of the corpus signal and present the corpus translation result.

[0063] In some embodiments, the server 200 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected by wired or wireless communication means, which is not limited in the embodiments of the present application.

[0064] Next, the structure of the electronic device for implementing the training method of the translation model provided in the embodiments of the present application will be described. As before, the electronic device provided in the embodiments of the present application may be Figure 1 the server 200 in Figure 2 , Figure 2 is a schematic structural diagram of the server 200 provided in the embodiments of the present application. Figure 2 The server 200 shown in Figure 2 includes: at least one processor 210, a memory 250, and at least one network interface 220. Each component in the server 200 is coupled together through a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in

[0065] all kinds of buses are labeled as the bus system 240.

[0065] The processor 210 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.

[0066] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 250 optionally includes one or more storage devices that are physically located far from the processor 210.

[0067] The memory 250 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0068] In some embodiments, the memory 250 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which will be described exemplarily below.

[0069] The operating system 251 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks; the network communication module 252 is used to reach other computing devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.

[0070] In some embodiments, the training device of the translation model provided in the embodiments of the present application may be implemented in software. Figure 2 Shown is the training device 255 of the translation model stored in the memory 250, which may be software in the form of programs and plugins, etc., and includes the following software modules: a first task module 2551, a selection module 2552, a second task module 2553, and an update module 2554. Figure 2 Shown is the corpus translation device 256 of the translation model stored in the memory 250, which may be software in the form of programs and plugins, etc., and includes the following software module: an application module 2555. The application module 2555 may also be installed in the terminal 400. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.

[0071] The training method of the translation model provided in the embodiments of the present application will be described in combination with the exemplary applications and implementations of the server 200 provided in the embodiments of the present application.

[0072] See Figure 4 , Figure 4 is a schematic structural diagram of the joint training model of the corpus translation method provided in the embodiments of the present application. First, the first translation model is used to make predictions on the basis of given completely correct previous words corresponding to each first position to be predicted, and a first probability distribution of each first pre-labeled target word at the corresponding first position to be predicted is obtained. The input of the encoding network of the first translation model is the original word vector x, and the input of the previous decoding network of the first translation model is the sequence of previous words corresponding to each first pre-labeled target word. For example, for the first pre-labeled target word y6, the sequence of previous words is the BOS symbol and the first pre-labeled target words y1 - y5. Given a confidence threshold, the first positions to be predicted in the first pre-labeled target words corresponding to the first probability (first confidence) less than this confidence threshold are used as the subset y that will be occluded and input into the second translation model later. m , for example, the first pre-labeled target words y2, y3, and y5, the occluded subset y m When used as the input of the second translation model, it appears in the form of an invisible mask M, and the remaining first pre-labeled target words y1, y4, and y6 are used as the partially visible sequence y that will be input into the second translation model. o , the partially visible sequence y of the second translation model is determined by the first translation model. o , given the source sentence x and the partially visible sequence y o , the input of the encoding network of the second translation model is the original word vector x. The pre-trained second translation model predicts each word y m in the occluded target word subset y t to obtain the corresponding predicted probability distribution. As the second confidence (for example, q2, q3, and q5), next, for the second positions to be predicted at the target end of the second translation model, where the second positions to be predicted are the first positions to be predicted with the first confidence lower than the confidence threshold, knowledge distillation is used to introduce bidirectional global context information (based on context words) for the first translation model in a targeted manner. For other first pre-labeled target words that do not belong to y m , the first loss function of the first translation model is still used for training.

[0073] See Figure 5 , Figure 5 is a schematic structural diagram of the first translation model provided by the embodiment of the present application. The first translation model includes an encoding network and a previous decoding network. The encoding network includes multiple sub-encoding networks (encoders), and the previous decoding network also includes the same number (corresponding to the encoder) of sub-previous decoding networks (decoders). The input of the encoding network is the original word vector of the corpus sample, and the encoding network outputs the source sentence representation of the corpus sample. The input of the previous decoding network is the source sentence representation of the corpus sample and the previous word. Among them, the input of each sub-previous decoding network is the source sentence representation of the corpus sample and the previous word, and the output of the previous decoding network is the first pre-labeled target word of the corpus sample. For example, if the corpus sample is "where are you", the first pre-labeled target word is "where are you".

[0074] See Figure 6 , Figure 6 which is a schematic structural diagram of the first sub-encoding network provided by an embodiment of the present application. The first sub-encoding network includes a self-attention processing layer and a feed-forward processing layer. Among them, the input of this layer (for example, x1, x2, and x3) is subjected to self-attention processing through the self-attention processing layer to obtain corresponding self-attention processing results (for example, z1, z2, and z3), and the self-attention processing results are subjected to hidden state mapping processing through the feed-forward processing layer to obtain a corresponding sequence of hidden state vectors.

[0075] See Figure 7 , Figure 7 which is a schematic structural diagram of the sub-previous decoding network provided by an embodiment of the present application. The sub-previous decoding network includes a masked self-attention processing layer, a cross-attention processing layer, and a feed-forward processing layer. The input of this layer is subjected to masked self-attention processing through the masked self-attention processing layer to obtain a corresponding masked self-attention processing result, and the masked self-attention processing result is subjected to cross-attention processing through the cross-attention processing layer to obtain a corresponding cross-attention processing result. The cross-attention processing result is subjected to hidden state mapping processing through the feed-forward processing layer. The input of the previous decoding network is at least one previous word at each first position to be predicted, and the first pre-labeled target words at different first positions to be predicted are output simultaneously in a parallel processing manner.

[0076] See Figure 8 , Figure 8 which is a schematic structural diagram of the sub-context decoding network provided by an embodiment of the present application. The sub-context decoding network includes a context self-attention processing layer, a cross-attention processing layer, and a feed-forward processing layer. Figure 8 The structure of the context decoding network shown is similar to that of the previous decoding network, except that the context self-attention processing layer included in the sub-context decoding network is different from Figure 7 the masked self-attention processing layer in. The input of the context decoding network is a random set of context words, and the second pre-labeled target words at the second positions to be predicted are output simultaneously.

[0077] See Figure 3A , Figure 3A which is a schematic flowchart of the training method of the translation model provided by an embodiment of the present application, and will be described in conjunction with Figure 3A the steps 101-104 shown.

[0078] In step 101, the corpus sample is propagated forward in the first translation model to obtain the first confidence of the first pre-labeled target word corresponding to each first position to be predicted.

[0079] As an example, when the corpus sample is propagated forward in the first translation model, it needs to go through the encoding network and the pre-decoding network successively. Among them, when decoding through the pre-decoding network, the confidence of the first pre-labeled target word needs to be predicted for each first position to be predicted in a parallel manner.

[0080] In some embodiments, before the corpus sample is propagated forward in the first translation model, the corpus sample can be propagated forward alone in the first translation model to obtain the first forward propagation result of the corpus sample; the first forward propagation result of the corpus sample is propagated backward in the first translation model to update the parameters of the first translation model, and the updated first translation model is used as the first translation model for the corpus sample in processing step 101; before the corpus sample is propagated forward in the second translation model, the corpus sample is propagated forward in the second translation model to obtain the second forward propagation result of the corpus sample; the second forward propagation result of the corpus sample is propagated backward in the second translation model to update the parameters of the second translation model, and the updated second translation model is used as the second translation model for the corpus sample in processing step 101.

[0081] See Figure 3B , Figure 3B is a schematic flowchart of the training method of the translation model provided by the embodiments of the present application. The first translation model includes a first encoding network and a pre-decoding network; in step 101, the corpus sample is propagated forward in the first translation model to obtain the first confidence of the first pre-labeled target word corresponding to each first position to be predicted, which can be realized through Figure 3B the steps 1011-1013 shown.

[0082] In step 1011, each original word of the corpus sample and the original word vector corresponding to each original word are determined, and the original word vectors corresponding to each original word are combined to obtain the original word vector sequence of the corpus sample.

[0083] As an example, in natural language processing tasks, it is necessary to consider how words are represented in a computer. There are usually two representation methods, such as one-hot encoding and distributed encoding. The original word vectors of each original word are obtained through one-hot encoding and distributed encoding. For example, for the source sentence "Where are you", there are four original words "you", "in", "which", and "where", and four original word vectors.

[0084] In step 1012, the first encoding network performs semantic encoding processing on the original word vector sequence of the corpus sample to obtain the first source sentence representation corresponding to the corpus sample.

[0085] In some embodiments, the first encoding network includes N cascaded first sub-encoding networks, where N is an integer greater than or equal to 2. The above semantic encoding process of the original word vector sequence of the corpus sample by the first encoding network to obtain the first source sentence representation corresponding to the corpus sample can be achieved through the following technical solution: Through the N cascaded first sub-encoding networks included in the first encoding network, the following semantic encoding process is performed on the original word vector sequence of the corpus sample: perform self-attention processing on the input of the first sub-encoding network to obtain the self-attention processing result corresponding to the first sub-encoding network, perform hidden state mapping processing on the self-attention processing result to obtain the hidden state vector sequence corresponding to the first sub-encoding network, and use the hidden state vector sequence as the semantic encoding processing result of the first sub-encoding network; among them, in the N cascaded first sub-encoding networks, the input of the first first sub-encoding network includes the original word vector sequence of the corpus sample, and the semantic encoding processing result of the Nth first sub-encoding network includes the first source sentence representation corresponding to the corpus sample. As an alternative implementation, there may be only one first sub-encoding network to perform semantic encoding processing on the original word vector sequence.

[0086] As an example, through the nth first sub-encoding network in the N cascaded first sub-encoding networks, perform semantic encoding processing on the input of the nth first sub-encoding network, and transmit the nth semantic encoding processing result output by the nth first sub-encoding network to the (n + 1)th first sub-encoding network to continue semantic encoding processing to obtain the (n + 1)th semantic encoding processing result; where n is an integer variable starting from 1 and increasing, and the value range of n is 1 ≤ n < N. When n takes the value of 1, the input of the nth first sub-encoding network is the original word vector sequence of the corpus sample. When n takes the value of 2 ≤ n < N, the input of the nth first sub-encoding network is the (n - 1)th semantic encoding processing result output by the (n - 1)th first sub-encoding network. When n takes the value of N - 1, the output of the (n + 1)th first sub-encoding network is the source sentence representation of the corpus sample.

[0087] As an example, each of the first sub-encoding networks includes a self-attention processing layer and a feed-forward processing layer; the semantic encoding process of the input of the nth first sub-encoding network among the N cascaded first sub-encoding networks can be implemented through the following technical solution: through the self-attention processing layer of the nth first sub-encoding network among the N cascaded first sub-encoding networks, perform self-attention processing on the input of the nth first sub-encoding network to obtain the self-attention processing result corresponding to the nth first sub-encoding network; transmit the self-attention processing result corresponding to the nth first sub-encoding network to the feed-forward processing layer of the nth first sub-encoding network for hidden state mapping processing to obtain the hidden state vector sequence corresponding to the nth first sub-encoding network, as the nth semantic encoding processing result output by the nth first sub-encoding network. See Figure 6 , the first sub-encoding network includes a self-attention processing layer and a feed-forward processing layer, wherein, through the self-attention processing layer, self-attention processing is performed on the input of this layer (for example, x1, x2, and x3) to obtain the corresponding self-attention processing results (for example, z1, z2, and z3), and through the feed-forward processing layer, hidden state mapping processing is performed on the self-attention processing results to obtain the corresponding hidden state vector sequence.

[0088] In some embodiments, the above-mentioned self-attention processing of the input of the first sub-encoding network to obtain the self-attention processing result corresponding to the first sub-encoding network can be implemented through the following technical solution: perform the following processing for each original word in the corpus sample: perform linear transformation processing on the first intermediate vector corresponding to the original word in the input of the first sub-encoding network to obtain the query vector, key vector, and value vector corresponding to the original word; perform dot product processing on the query vector of the original word and the key vector of each original word, and perform normalization processing on the dot product processing result based on the maximum likelihood function to obtain the weight of the value vector of the original word; perform weighted processing on the value vector of the original word based on the weight of the value vector of the original word to obtain the self-attention processing result of the sub-encoding network corresponding to each original word.

[0089] As an example, when the first sub-coding network is the first first sub-coding network among multiple first sub-coding networks, the first intermediate vector is the vector of each original word. When the first sub-coding network is not the first first sub-coding network among multiple first sub-coding networks, the first intermediate vector is the hidden state vector output by the previous first sub-coding network. The first intermediate vector corresponds to the original word one by one. The linear transformation processing is actually to multiply the first intermediate vector with the three parameter matrices respectively to obtain the query vector Q, key vector K and value vector V corresponding to the first intermediate vector. Here, the query vector is dot-multiplied with the key vector of each original word, including dot-multiplication with its own key vector and dot-multiplication with the key vectors of other original words. The dot-multiplication result is used. It is used to characterize the relevance, so that the obtained relevance includes autocorrelation and relevance with other original words. Before the relevance is normalized based on the maximum likelihood function, it can also be divided by the square root of the key vector length. The normalization based on the maximum likelihood function is to substitute the result obtained after the relevance or the square root of the key vector length into the Softmax function, so as to obtain the contribution weight of each original word to a certain original word. The value vector of a certain original word is weighted by the contribution weight of each original word to a certain original word, so as to obtain the context relevance corresponding to a certain original word, which is beneficial to the translation accuracy of the subsequent translation model. The above three parameter matrices all need to be obtained through training.

[0090] In step 1013, a corpus decoding process is performed on the first source sentence representation through a pre-order decoding network to obtain a first confidence of a first pre-marked target word corresponding to each first position to be predicted.

[0091] As an example, the first confidence is generated based on the preceding word corresponding to each first position to be predicted, and the generation process of the first confidence is a parallel process. Taking the 5th first position to be predicted as an example, based on the start symbol B and the first pre-marked target words corresponding to the 1st first position to be predicted to the 4th first position to be predicted, the first confidence that the 5th first position to be predicted is translated into the corresponding first pre-marked target word is predicted, for example, the first probability or the entropy of the distribution of probabilities. For example, "where are you" should be translated into "where are you", then there are three first positions to be predicted. For the second first position to be predicted, it should be translated into the corresponding first pre-marked target word "are", then based on the preceding words "B" and "where", the first probability that the second first position to be predicted is predicted to be the corresponding first pre-marked target word "are" is obtained. If the first translation model performs well, the first probability should exceed the probability threshold, indicating that the first translation model has a greater probability of translating correctly.

[0092] In some embodiments, the above-mentioned corpus decoding process of the first source statement representation by the pre-order decoding network to obtain the first confidence of the first pre-marked target word corresponding to each first position to be predicted can be achieved through the following technical solutions: the following process is performed for each first position to be predicted output by the pre-order decoding network: obtain the first pre-marked target word sequence corresponding to the corpus sample from the corpus sample set; extract the first pre-marked target word located before the first position to be predicted from the first pre-marked target word sequence, and use the extracted first pre-marked target word as the pre-order word corresponding to the first position to be predicted; perform semantic decoding processing on the pre-order word corresponding to the first position to be predicted and the first source statement representation through the pre-order decoding network to obtain the first confidence that the first pre-marked target word corresponding to the first position to be predicted is decoded at the first position to be predicted.

[0093] As an example, the corpus sample is "Where are you", and the corresponding first pre-marked target word sequence is obtained. The first pre-marked target word sequence is "where are you". For any one of the first positions to be predicted, the first pre-marked target word located before the first position to be predicted is extracted from the first pre-marked target word sequence. "Where are you" should be translated as "B whereare you E", where B is the start symbol and E is the end symbol. There are three first positions to be predicted. For the second first position to be predicted, the first pre-marked target words before the second first position to be predicted are "B" and "where". The extracted first pre-marked target word is used as the pre-order word corresponding to the second first position to be predicted; semantic decoding processing is performed with the pre-order word corresponding to the second first position to be predicted and the first source statement representation as the input of the pre-order decoding network to obtain the first confidence that the first pre-marked target word "are" corresponding to the second first position to be predicted is decoded at the second first position to be predicted.

[0094] As an example, before performing self-attention processing on the input of the nth first sub-encoding network in N cascaded first sub-encoding networks through the self-attention layer of the nth first sub-encoding network, when n satisfies 2 ≤ n < N, the regularization result of the (n - 1)th semantic encoding processing result is concatenated with the input of the (n - 1)th sub-encoding network; the concatenated result is used as the input of the self-attention layer of the nth sub-encoding network to replace using the (n - 1)th semantic encoding processing result as the input of the nth sub-encoding network. Before transmitting the self-attention processing result corresponding to the nth first sub-encoding network to the first feed-forward processing layer of the nth first sub-encoding network for hidden state mapping processing, the regularization result of the self-attention processing result corresponding to the nth first sub-encoding network is concatenated with the input of the nth first sub-encoding network; the concatenated result is used as the input of the feed-forward processing layer of the nth first sub-encoding network to replace using the self-attention processing result corresponding to the nth first sub-encoding network as the input of the nth first sub-encoding network.

[0095] In some embodiments, the preamble decoding network includes M cascaded sub-preamble decoding networks, where M is an integer greater than or equal to 2; the above semantic decoding processing of the preamble word corresponding to the first to-be-predicted position and the first source sentence representation through the preamble decoding network can be implemented by the following technical solution: each sub-preamble decoding network performs semantic decoding processing in the following manner: performing masked self-attention processing on the input of the sub-preamble decoding network to obtain the masked self-attention processing result of the corresponding sub-preamble decoding network, performing cross-attention processing on the masked self-attention processing result to obtain the cross-attention processing result of the corresponding sub-preamble decoding network, and performing hidden state mapping processing on the cross-attention processing result; wherein, in the M cascaded sub-preamble decoding networks, the input of the first sub-preamble decoding network includes: the preamble word corresponding to the first to-be-predicted position and the first source sentence representation; the hidden state mapping processing result of the Mth sub-preamble decoding network includes: the first confidence that the first to-be-predicted position is decoded into the corresponding first pre-token target word.

[0096] As an example, the preamble decoding network consists of M cascaded sub-preamble decoding networks, where M is an integer greater than or equal to 2; through the Mth sub-preamble decoding network in the M cascaded sub-preamble decoding networks, semantic decoding processing is performed on the input of the Mth sub-preamble decoding network, and the Mth semantic decoding processing result output by the Mth sub-preamble decoding network is transmitted to the (M + 1)th sub-preamble decoding network to continue semantic decoding processing to obtain the (M + 1)th semantic decoding processing result; where M is an integer variable that starts increasing from 1, and the value range of M is 1 ≤ M < M. When M takes the value of 1, the input of the Mth sub-preamble decoding network is the source sentence representation of the encoding network and the preamble word sequence. When M takes the value of 2 ≤ M < M, the input of the Mth sub-preamble decoding network is the (M - 1)th semantic decoding processing result output by the (M - 1)th sub-preamble decoding network. When M takes the value of M - 1, the output of the (M + 1)th sub-preamble decoding network is the first confidence level. As an alternative implementation, there can be only one sub-preamble decoding network for semantic decoding processing.

[0097] As an example, the following processing is performed on the second intermediate vector of each original word in the corpus sample. Where, when the sub-preamble decoding network is the first decoding network, the second intermediate vector is the word vector of each first pre-tokenized target word that serves as a preamble word. When the sub-preamble decoding network is not the first decoding network, the second intermediate vector is the output of the previous sub-preamble decoding network. The second intermediate vector corresponds one-to-one with each first pre-tokenized target word that serves as a preamble word. Linear transformation processing is performed on the second intermediate vector of each preamble word to obtain the key vector and value vector of each original word; linear transformation processing is performed on the second intermediate vector of the last preamble word in the preamble words to obtain the query vector of the last preamble word. The query vector is dot-multiplied with the key vector of each preamble word, and the dot-multiplication processing result is normalized based on the maximum likelihood function to obtain the weight of the value vector of each preamble word; the value vector of each preamble word is weighted based on the weight of the value vector of each preamble word to obtain the masked attention processing result corresponding to the last preamble word of the sub-preamble decoding network.

[0098] In some embodiments, the above-mentioned cross-attention processing of the masked self-attention processing result to obtain the cross-attention processing result corresponding to the sub-previous decoding network can be implemented through the following technical solutions: perform a linear transformation on the masked self-attention processing result to obtain a query vector of the masked self-attention processing result; for each original word, perform the following processing: perform a linear transformation on the first source sentence representation of the original word to obtain a key vector and a value vector of the first source sentence representation; perform a dot product between the query vector of the masked self-attention processing result and the key vector of the first source sentence representation, and perform normalization processing based on the maximum likelihood function on the dot product result to obtain the weight of the value vector of the first source sentence representation; perform a weighted processing on the value vector of the first source sentence representation based on the weight of the value vector of the first source sentence representation to obtain the cross-attention processing result corresponding to the sub-previous decoding network.

[0099] As an example, the masked self-attention processing result is a vector for the last previous word. For example, when decoding and predicting the second first position to be predicted, a linear transformation is performed based on the masked attention processing result corresponding to "where" to obtain a query vector of the masked self-attention processing result; for each original word, perform the following processing, that is, input the first source sentence representation into the decoding network, perform a linear transformation on the first source sentence representation of the original word to obtain a key vector and a value vector of the first source sentence representation. The first source sentence representation is a vector for each original word in "where are you" respectively. Perform a dot product between the query vector of the masked self-attention processing result and the key vector corresponding to each original word in the first source sentence representation, and perform normalization processing based on the maximum likelihood function on the dot product result to obtain the weight of the value vector corresponding to each original word in the first source sentence representation; perform a weighted processing on the value vector of the first source sentence representation based on the weight of the value vector of the first source sentence representation to obtain the cross-attention processing result corresponding to the sub-previous decoding network, and then perform a feed-forward processing on the cross-attention processing result, and continue to perform the above self-attention processing on the obtained hidden state vector with other previous words. Other previous words are previous words except the last previous word, such as the start symbol "B".

[0100] In step 102, the first position to be predicted with the first confidence lower than the confidence threshold is determined as the second position to be predicted of the second translation model, and the first pre-marked target word corresponding to the first position to be predicted with the first confidence not lower than the confidence threshold is determined as the context word corresponding to the second translation model.

[0101] As an example, for the three first positions to be predicted in the target word sequence "B where are you E", for the 1st first position to be predicted, the first confidence level corresponding to the first pre-marked target word "where" output by the previous decoding network is 0.4; for the 2nd first position to be predicted, the first confidence level corresponding to the first pre-marked target word "are" output by the previous decoding network is 0.8; for the 3rd first position to be predicted, the first confidence level corresponding to the first pre-marked target word "you" output by the previous decoding network is 0.3. If the confidence level threshold is 0.6, then the 1st and 3rd first positions to be predicted are determined as the second positions to be predicted of the second translation model, and the first pre-marked target word "are" corresponding to the 2nd first position to be predicted is used as the context word of the second translation model. This is equivalent to using the first pre-marked target word "are" at the 2nd first position to be predicted as the visible second pre-marked target word "are" at the 2nd second position to be predicted of the second translation model. That is, the task of the second translation model is to use the visible second pre-marked target word "are" at the 2nd second position to be predicted as a known condition to predict the translation results of the 1st and 3rd second positions to be predicted.

[0102] In step 103, the context word and the corpus sample are propagated forward in the second translation model to obtain the second confidence levels of the second pre-marked target words corresponding to each second position to be predicted.

[0103] As an example, when the corpus sample and the context word are propagated forward in the second translation model, they need to be processed by the encoding network and the previous decoding network successively. Among them, when decoding through the context decoding network, the confidence levels of the second pre-marked target words are predicted for each second position to be predicted simultaneously.

[0104] See Figure 3C , Figure 3C is a schematic flowchart of the training method of the translation model provided by the embodiment of the present application. The second translation model includes a second encoding network and a context decoding network; in step 103, the context word and the corpus sample are propagated forward in the second translation model to obtain the second confidence levels of the second pre-marked target words corresponding to each second position to be predicted, which can be implemented through Figure 3C the steps 1031 - 1033 shown.

[0105] In step 1031, each original word of the corpus sample and the original word vector corresponding to each original word are obtained, and the original word vectors corresponding to each original word are combined to obtain the original word vector sequence of the corpus sample.

[0106] As an example, in natural language processing tasks, we need to consider how words are represented in computers. There are usually two ways of representation, such as one-hot encoding and distributed encoding. The original word vector of each original word is obtained through one-hot encoding and distributed encoding. For example, for the source sentence "Where are you", there are four original words "you", "in", "where" and "in", as well as four original word vectors.

[0107] In step 1032, semantic encoding processing is performed on the original word vector sequence of the corpus sample through a second encoding network to obtain a second source sentence representation corresponding to the corpus sample.

[0108] As an example, the encoding process in step 1032 may refer to the specific implementation in step 1012, wherein the second encoding network and the first encoding network may have the same parameters or different parameters.

[0109] In step 1033, a corpus decoding process is performed on the second source sentence representation through a context decoding network to obtain a second confidence of a second pre-marked target word corresponding to each second position to be predicted.

[0110] As an example, the second confidence is generated based on context words corresponding to multiple second positions to be predicted. When generating the second confidence of the second pre-marked target word corresponding to each second position to be predicted, they are all generated based on the same context words, that is, they are all generated by the context words obtained based on the first translation model.

[0111] In some embodiments, in step 1033, the second source sentence representation is subjected to corpus decoding processing through the context decoding network to obtain the second confidence of the second pre-marked target word corresponding to each second position to be predicted, which can be achieved by the following technical scheme: the following processing is performed for each second position to be predicted output by the context decoding network: the context words and the second source sentence representation are subjected to semantic decoding processing through the context decoding network to obtain the second confidence of the second pre-marked target word decoded into the corresponding second position to be predicted.

[0112] In some embodiments, the context decoding network includes P cascaded sub-context decoding networks, where P is an integer greater than or equal to 2; the semantic decoding process of the context word and the second source statement representation by the context decoding network can be achieved through the following technical solutions: through the P cascaded sub-context decoding networks, perform semantic decoding processing on the context word set corresponding to the second position to be predicted and the source statement representation in the following manner: perform context self-attention processing on the input of the sub-context decoding network to obtain the context self-attention processing result corresponding to the sub-context decoding network; perform cross-attention processing on the context self-attention processing result to obtain the cross-attention processing result corresponding to the sub-context decoding network; perform hidden state mapping processing on the cross-attention processing result; where, among the P cascaded sub-context decoding networks, the input of the first sub-context decoding network includes: the context word set corresponding to the second position to be predicted and the second source statement representation: the hidden state mapping processing result of the P-th sub-context decoding network includes: the second confidence level that the second position to be predicted is decoded into the corresponding second pre-marked target word. As an alternative implementation, there may be only one sub-context decoding network for semantic decoding processing.

[0113] As an example, through the p-th sub-context decoding network among the P cascaded sub-context decoding networks, perform semantic decoding processing on the input of the p-th sub-context decoding network, and transmit the p-th semantic decoding processing result output by the p-th sub-context decoding network to the (p + 1)-th sub-context decoding network to continue semantic decoding processing to obtain the (p + 1)-th semantic decoding processing result; where, p is an integer variable starting from 1 and increasing, and the value range of p is 1 ≤ p < P. When p takes the value of 1, the input of the p-th sub-context decoding network is the second source statement representation of the second encoding network and the context word. When p takes the value of 2 ≤ p < P, the input of the p-th sub-context decoding network is the (p - 1)-th semantic decoding processing result output by the (p - 1)-th sub-context decoding network. When p takes the value of P - 1, the output of the (p + 1)-th sub-context decoding network is the second confidence level.

[0114] As an example, when performing semantic decoding on the input of the p-th sub-context decoding network through the p-th sub-context decoding network in P cascaded sub-context decoding networks, context self-attention processing is performed on the input of the p-th sub-context decoding network through the context self-attention layer of the p-th sub-context decoding network in P cascaded sub-context decoding networks to obtain the context self-attention processing result corresponding to the p-th sub-context decoding network; the context self-attention processing result corresponding to the p-th sub-context decoding network is transmitted to the cross-attention layer of the p-th sub-context decoding network for cross-attention processing to obtain the cross-attention processing result corresponding to the p-th sub-context decoding network; the cross-attention processing result corresponding to the p-th sub-context decoding network is transmitted to the feed-forward processing layer of the p-th sub-context decoding network for hidden state mapping processing to obtain the hidden state sequence corresponding to the p-th sub-context decoding network, which is used as the p-th semantic decoding processing result output by the p-th sub-context decoding network.

[0115] As an example, the following processing is performed for each context word. When the sub-previous decoding network is the first decoding network, the third intermediate vector is the word vector of each second pre-tokenized target word that is a previous word. When the sub-context decoding network is not the first decoding network, the third intermediate vector is the output of the previous sub-context decoding network. The third intermediate vectors correspond one-to-one with each second pre-tokenized target word that is a context word. Linear transformation processing is performed on the third intermediate vector corresponding to each context word to obtain the key vector and value vector of the context word. Linear transformation processing is performed on the third intermediate vector of the symbol at the second position to be predicted (which is marked with a special symbol because it is unknown) to obtain the query vector of the symbol. The query vector is dot-multiplied with the key vector of each context word, and the dot-multiplication result is normalized based on the maximum likelihood function to obtain the weight of the value vector of each context word; the value vectors of the context words are weighted based on the weights of the value vectors of each context word to obtain the masked attention processing result of the context word corresponding to the sub-context decoding network at the end.

[0116] As an example, the process of performing cross-attention processing on the context self-attention processing result to obtain the cross-attention processing result corresponding to the sub-context decoding network is similar to the processing method in the previous decoding network.

[0117] In step 104, based on the first confidence of the first pre-tokenized target word corresponding to each first position to be predicted and the second confidence of the second pre-tokenized target word corresponding to each second position to be predicted, the parameters of the first translation model and the second translation model are updated.

[0118] See Figure 3D , Figure 3DIt is a schematic flowchart of a method for training a translation model provided by an embodiment of the present application. In step 104, based on the first confidence of the first pre-labeled target word corresponding to each first position to be predicted and the second confidence of the second pre-labeled target word corresponding to each second position to be predicted, the parameters of the first translation model and the second translation model are updated, which can be achieved through Figure 3D the steps 1041-1044 shown.

[0119] In step 1041, based on the first confidence, a first loss corresponding to the first translation model is determined.

[0120] In some embodiments, there are multiple first positions to be predicted with corresponding multiple first pre-labeled target words one by one; in step 1041, based on the first confidence, determining the first loss corresponding to the first translation model can be achieved through the following technical solution: fusing the first confidence obtained for each first position to be predicted to obtain the first loss corresponding to the first translation model.

[0121] In step 1042, based on the first confidence lower than the confidence threshold and the second confidence of the second pre-labeled target word corresponding to each second position to be predicted, a second loss corresponding to the second translation model is determined.

[0122] In some embodiments, the second loss is used to represent the teaching loss of the second translation model to the first translation model, and there are multiple second positions to be predicted with corresponding multiple second pre-labeled target words one by one; in step 1042, based on the first confidence lower than the confidence threshold and the second confidence of the second pre-labeled target word corresponding to each second position to be predicted, determining the second loss corresponding to the second translation model can be achieved through the following technical solution: fusing the first confidence lower than the confidence threshold and the second confidence of the second pre-labeled target word corresponding to each second position to be predicted to obtain the second loss corresponding to the second translation model.

[0123] In step 1043, based on the aggregation parameters corresponding to the first loss and the second loss respectively, the first loss and the second loss are aggregated to obtain a joint loss.

[0124] As an example, through the first translation model, a partially visible sequence y of the second translation model is determined o , in the case of a given source sentence x and the partially visible sequence y o , the pre-trained second translation model predicts each word y m in the occluded target word subset y t to obtain the corresponding prediction probability distribution As the second confidence level, next, for the second position to be predicted at the target end of the second translation model, where the second position to be predicted is the first position to be predicted with a first confidence level lower than the confidence threshold, the knowledge distillation method is used to introduce bidirectional global context information (based on context words) for the first translation model in a targeted manner. The loss function for knowledge distillation is shown in Formula (1):

[0125]

[0126] Among them, KL(·) represents the Kullback–Leibler divergence, and α is a balance coefficient. The value-taking strategy for α is as follows. The summation is the second loss. As the number of training rounds increases, the value of α linearly decreases from 1 to 0. This can guide the first translation model to absorb more knowledge from the second translation model with bidirectional global context information in the early stage, and then gradually refocus on the prediction of the first pre-labeled target word, so as to be better trained. For other first pre-labeled target words that do not belong to y m For the first pre-labeled target word, the first loss function of the first translation model is still used for training. Therefore, the joint loss function can be seen in Formula (2):

[0127]

[0128] Among them, y t ∈y o \[M] represents the visible sequence of target words (the first pre-labeled target words with a first confidence level higher than the confidence threshold among multiple first pre-labeled target words) excluding all special symbols [M]. L CBKD (θ ne ,θ nd ) is the joint loss, and L kd (θ ne ,θ nd ) is the second loss. The summation result of and the sum of is the first loss. Through confidence-based knowledge distillation, bidirectional global context information is introduced for the first translation model at the target end in a targeted manner. At the same time, the second translation model only participates in the training process and does not participate in the inference stage of the first translation model.

[0129] In step 1044, update the parameters of the first translation model and the second translation model according to the joint loss.

[0130] As an example, to update the parameters of the two models according to the joint loss, the gradient descent algorithm can be used for update.

[0131] The exemplary application and implementation of the server 200 provided in the embodiments of the present application will be combined to illustrate the corpus translation method provided in the embodiments of the present application.

[0132] In some embodiments, in response to a translation request for a target corpus, the first translation model or the second translation model is called to perform translation processing on the target corpus to obtain a translation result for the target corpus; wherein, the first translation model and the second translation model are trained by the translation model training method provided in the embodiments of the present application.

[0133] As an example, there are differences between the application stage and the training stage of the first translation model. When the first translation model is applied, the corpus sample is encoded by the encoding network of the first translation model to obtain the first source sentence representation. Then, the decoding process is sequentially performed on each first position to be predicted step by step. The start symbol "B" and the first source sentence representation are decoded by the previous decoding network to obtain the word "where" with the highest confidence in the vocabulary as the translation result of the first position to be predicted. Then, the start symbol "B", "where", and the first source sentence representation are decoded by the previous decoding network to obtain the word with the highest confidence in the vocabulary as the translation result of the second first position to be predicted, and so on until the end symbol "E" is output. The application stage of the second translation model is the same as the training stage, that is, it is equivalent to performing a forward propagation.

[0134] Next, the exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0135] In some embodiments, a language translation function can be provided in the social client. For example, the social client running on terminal A receives text information sent by terminal B via the server. The text information is in English. In response to the language translation operation for the text information received by terminal A, terminal A calls the first translation model provided in the embodiments of the present application to perform translation processing on the text information belonging to the first language to obtain text information corresponding to the second language. Wherein, the first translation model is obtained by auxiliary training with the help of the second translation model, and when performing auxiliary training through the second translation model, only the first positions to be predicted with relatively low first confidence are jointly predicted.

[0136] In some embodiments, the embodiments of the present application provide a joint training framework that integrates bidirectional global context information based on confidence. The joint training framework includes a translation model (the first translation model) and a conditional masked language model (the second translation model). During the joint training process, based on the confidence of the first translation model in predicting the first pre-labeled target word, knowledge distillation is used to integrate bidirectional global context information (context words) for the first translation model by the second translation model. Wherein, the context words are unobscured target words. The joint training is divided into two stages: (1) pre-training the first translation model and the second translation model; (2) performing knowledge distillation based on confidence.

[0137] In some embodiments, the first translation model is pre-trained (i.e., trained separately). The first translation model and the second translation model have the same encoding network. The role of the encoding network is to encode the input source sentence (equivalent to a corpus sample) into a source semantic representation. The encoding network can be composed of L e identical sub-encoding networks, where L e is an integer greater than or equal to 1. Each layer includes two sub-layers: (1) a self-attention processing layer, and (2) a feed-forward processing layer. In the sub-encoding network, the self-attention processing layer takes the sequence of hidden state vectors output by the previous layer as input, and further maps the sequence of hidden state vectors through the self-attention mechanism, that is, performs multi-head self-attention operations. The self-attention processing layer can be formalized as the following formula (3):

[0138] c (l) = AN(SelfAtt(h (l-1) , h (l-1) , h (l-1) )) (3);

[0139] where c (l) is the intermediate vector calculated by the self-attention sub-layer, AN(·) represents layer normalization operation with residual connection, SelfAtt(·) represents multi-head self-attention operation, h (l) represents the sequence of hidden state vectors output by the l-th layer of the encoding network, h (l-1) represents the sequence of hidden state vectors output by the (l - 1)-th layer of the encoding network, and c (l) is mapped through the feed-forward processing layer to the sequence of hidden state vectors h (l) output by the l-th layer of the encoding network. See formula (4):

[0140] h (l) = AN(FFN(h (l-1) )) (4);

[0141] In some embodiments, h (0) of the encoding network is the sequence of word embedding vectors corresponding to the input source sentence, which is the final output source sentence representation of the encoding network.

[0142] In some embodiments, the previous decoding network can be composed of L dIt is composed of [[NUM]] identical sub-previous decoding networks. Each sub-previous decoding network has three layers: (1) Masked Self-Attention Processing Layer (MaskedSelfAtt), (2) Cross-Attention Processing Layer (CrossAtt), and (3) Feed-Forward Processing Layer (FFN). To ensure the autoregressive property of the first translation model, the MaskedSelfAtt layer uses an attention mask to block all subsequent words of the target word at each first position to be predicted, so that the first translation model only relies on previous words for prediction at the target end. Since the first translation model can only use the words generated before at each time step during the test phase, the MaskedSelfAtt layer has an attention mask during training to block all subsequent words of each first pre-tokenized target word. The MaskedSelfAtt layer can be formalized as Equation (5):

[0143] a (l) = AN(MaskedSelfAtt(s (l-1) , s (l-1) , s (l-1) )) (5);

[0144] Among them, s (l-1) represents the sequence of hidden state vectors of the (l - 1)-th sub-previous decoding network, and a (l) is the intermediate representation (masked attention processing result) after being mapped by the MaskedSelfAtt layer. AN(·) represents the layer regularization operation with residual connection, and FFN(·) represents the mapping operation of the feed-forward processing layer.

[0145] In some embodiments, the CrossAtt layer models a (l) and the source sentence representation output by the encoding network using the cross-attention mechanism. The calculation process is abstracted as Equation (6):

[0146]

[0147] Among them, z (l) is the intermediate representation (cross-attention processing result) after being mapped by the CrossAtt layer. AN(·) represents the layer regularization operation with residual connection, a (l) is the intermediate representation (masked attention processing result) after being mapped by the MaskedSelfAtt layer, is the source sentence representation output by the encoding network, and CrossAtt(·) is the cross-attention processing.

[0148] In some embodiments, z (l) is mapped by the FFN layer to the hidden state vector s (l) output by the sub-previous decoding network. The mapping process is abstracted as Equation (7):

[0149] s (l) = AN(FFN(z (l) )) (7);

[0150] Where s (l) represents the sequence of hidden state vectors of the l-th sub-previous decoding network, AN(·) represents the layer regularization operation with residual connections, FNN(·) is the feed-forward process, and z (l) is the intermediate representation (cross-attention processing result) after being mapped by the CrossAtt layer.

[0151] In some embodiments, given the source sentence x, the previous word sequence y <t corresponding to the t-th moment at the target side, and the hidden state vector s output by the highest layer of the decoding network, the first translation model predicts y t in the form of the following formula (8) as the probability distribution on the vocabulary:

[0152] p(y t |y <t ,x) = softmax(Ws t ) (8);

[0153] Where the moment can be understood as different first positions to be predicted, each first position to be predicted corresponds to a different time step, W represents the learnable linear transformation matrix, and s t is the hidden state vector corresponding to the t-th moment of the decoding network of the first translation model. For example, for the first pre-labeled target word y5 at the 5th first position to be predicted, it is based on the given source sentence x (e.g., "Where are you"), the given previous words y1 to y4, and the hidden state vector corresponding to the t-th moment output by the decoding network, to predict the probability p(y5|y <5 ,x) of being translated as the first pre-labeled target word y5 at the 5th first position to be predicted, where the hidden state vector corresponding to the t-th moment output by the decoding network is also obtained based on the given previous words y1 to y4.

[0154] In some embodiments, the first translation model and the second translation model can share the parameters of the encoding network, or can have independent encoding networks. The decoding networks of the first translation model and the second translation model can have some identical parameters, and the first translation model is additionally affected by the supervision signal from the context decoding network of the second translation model. The context decoding network of the second translation model has the target-side bidirectional global context information. Therefore, the first translation model after joint training has the ability to capture bidirectional global context information. The loss function of the first translation model is shown in formula (9):

[0155]

[0156] Among them, x is the source sentence, y t is the first pre-marked target word corresponding to the t-th moment (the t-th first position to be predicted) at the target end. P(y t |x, y <t ) is the probability distribution (the first probability) of y t on the vocabulary. The first translation model is trained in the standard teacher-guided manner.

[0157] In some embodiments, the second translation model is pre-trained (i.e., trained separately). In some embodiments, the decoding method of the context decoding network of the second translation model is different from that of the previous decoding network of the first translation model. In the second translation model, given a part of the visible input word sequence y o at the target end, the set of target words y m obscured at the target end is predicted. The decoding network of the second translation model consists of L d identical sub-context decoding networks. Each sub-context decoding network includes three layers: (1) a self-attention processing layer (SelfAtt); (2) a cross-attention processing layer (CrossAtt); (3) a feed-forward processing layer (FFN). The SelfAtt layer of the sub-context decoding network of the second translation model does not have a mask that obscures the subsequent words corresponding to each position when performing the self-attention mechanism, but can attend to all unobscured positions of the target-end input sequence (the visible input word sequence). After the same calculation process as the previous decoding network, given the source sentence x, the partially visible sequence y o at the target end, and the hidden state vector s′ output by the top layer of the context decoding network, the second translation model predicts the probability distribution of y t on the vocabulary through the following formula (10):

[0158] p(y t |y o , x) = softmax(W′s t ′) (10);

[0159] where W′ is a learnable linear transformation matrix, and s′ t is the hidden state vector corresponding to the t-th moment of the context decoding network. For example, for the second pre-marked target word y5 at the 5th second position to be predicted and the second pre-marked target word y3 at the 3rd second position to be predicted, based on the given source sentence x (e.g., "Do you like this flower?"), the given visible words y1, y2, and y4, and the hidden state vector corresponding to the t-th moment output by the decoding network, the probability p(y5|y 1,2,4, x), and the second pre-marked target word p(y3|y at the third second position to be predicted 1,2,4 , x), where the hidden state vector corresponding to the t-th moment output by the decoding network is also obtained based on the given visible words y1, y2, and y4. The visible words form a set of context words, thereby forming global context information.

[0160] In some embodiments, for the second translation model, first, a random integer v is generated between 1 and |y|, and then v words are randomly selected from the y second pre-marked target words. These words are replaced with a special character, thus splitting the second pre-marked target word sequence into the input sequence y of the observable context decoding network o (composed of visible second pre-marked target words) and the masked sequence y m (composed of invisible second pre-marked target words). The training objective of the second translation model can be expressed as the following formula (11):

[0161]

[0162] where x is the source language sentence, y t is the second pre-marked target word corresponding to the t-th moment (the t-th second position to be predicted) at the target end, P(y t |x, y o ) is the probability distribution (second probability) of y t on the vocabulary.

[0163] In some embodiments, the first translation model still predicts very low probabilities for many first pre-marked target words even when given completely correct previous words. See Figure 9 , Figure 9 is the confidence distribution diagram provided by the embodiments of the present application. Figure 9 The horizontal axis of Figure 9 is the confidence, and the vertical axis of

[0164] is the proportion of target words. For a first translation model that has been fully trained (the above pre-training process), when completely correct previous words are given for each first position to be predicted at the target end, the confidence distribution of the predicted first pre-marked target words. For example, there are 25.67% of the first pre-marked target words, and when given completely correct previous words, the predicted confidence is only 0.1. Figure 4 ,Figure 4 It is a schematic structural diagram of a joint training model for the training method of the translation model provided by an embodiment of the present application. First, the first translation model makes predictions based on the completely correct previous words corresponding to each first position to be predicted, and obtains the first probability distribution of each first pre-labeled target word at the corresponding first position to be predicted. Given a confidence threshold, the first positions to be predicted corresponding to the first pre-labeled target words with the first probability (first confidence) less than the confidence threshold are used as the subset y to be occluded and input into the second translation model subsequently. m The remaining first pre-labeled target words are used as the partially visible sequence y to be input into the second translation model. o This process can be represented by formula (12):

[0165]

[0166] Where represents the prediction probability (first confidence) of the first translation model for the first pre-labeled target word at the t-th moment, ε is the confidence threshold, t is the prediction time step, different first positions to be predicted are distinguished by the prediction time step, |y| is the number of first pre-labeled target words, and y t is the first pre-labeled target word corresponding to the t-th moment (the corresponding first position to be predicted).

[0167] The partially visible sequence y of the second translation model is determined by the first translation model. o Given the source sentence x and the partially visible sequence y o , the pre-trained second translation model makes predictions for each word y m in the occluded target word subset y t and obtains the corresponding prediction probability distribution as the second confidence. Next, for the second position to be predicted at the target end of the second translation model, where the second position to be predicted is the first position to be predicted with the first confidence lower than the confidence threshold, a knowledge distillation method is used to introduce two-way global context information (based on context words) for the first translation model in a targeted manner. The second loss function for knowledge distillation is shown in formula (13):

[0168]

[0169] Among them, KL(·) represents the Kullback–Leibler divergence, and α is a balance coefficient. The value-taking strategy for α is as follows: as the number of training rounds increases, the value of α linearly decreases from 1 to 0. This can guide the first translation model to absorb more knowledge from the second translation model with bidirectional global context information in the early stage, and then gradually refocus on the prediction of the first pre-marked target word, so as to be better trained. For other first pre-marked target words that do not belong to y m The first pre-marked target word is still trained using the first loss function of the first translation model. Therefore, the joint loss function can be seen in formula (14):

[0170]

[0171] Among them, y t ∈y o \[M] represents the visible sequence of target words (the first pre-marked target words among multiple first pre-marked target words with the first confidence higher than the confidence threshold) excluding all special symbols [M], and L CBKD (θ ne , θ nd ) is the joint loss, and L kd (θ ne , θ nd ) is the second loss. Through confidence-based knowledge distillation, bidirectional global context information is introduced for the first translation model at the target end in a targeted manner. At the same time, the second translation model only participates in the training process and does not participate in the inference stage of the first translation model.

[0172] Through the training method of the translation model provided by the embodiments of the present application, confidence-based knowledge distillation is performed, and bidirectional global context information is introduced for the first translation model at the target end for the first to-be-predicted positions with relatively low first confidence of the first pre-marked target words. Thus, when the first translation model makes predictions for each first to-be-predicted position, it not only utilizes the local context information of the corresponding previous words but also utilizes the global context information, thereby improving the translation performance of the first translation model.

[0173] Next, the implementation of the training device 255 of the translation model provided by the embodiments of the present application as an exemplary structure of software modules will be continued. In some embodiments, such as Figure 2As shown in the figure, the software modules stored in the training device 255 of the translation model in the memory 250 may include: a first task module 2551, configured to perform forward propagation of the corpus sample in the first translation model to obtain a first confidence level of the first pre-labeled target word corresponding to each first position to be predicted; a selection module 2552, configured to determine the second position to be predicted of the second translation model for the first position to be predicted with a first confidence level lower than the confidence threshold, and determine the first pre-labeled target word corresponding to the first position to be predicted with a first confidence level not lower than the confidence threshold as the context word corresponding to the second translation model; a second task module 2553, configured to perform forward propagation of the context word and the corpus sample in the second translation model to obtain a second confidence level of the second pre-labeled target word corresponding to each second position to be predicted; and an update module 2554, configured to update the parameters of the first translation model and the second translation model based on the first confidence level of the first pre-labeled target word corresponding to each first position to be predicted and the second confidence level of the second pre-labeled target word corresponding to each second position to be predicted.

[0174] In some embodiments, the update module 2554 is further configured to: determine a first loss corresponding to the first translation model based on the first confidence level; determine a second loss corresponding to the second translation model based on the first confidence level lower than the confidence threshold and the second confidence level of the second pre-labeled target word corresponding to each second position to be predicted, where the second loss is used to represent the teaching loss of the second translation model to the first translation model; perform an aggregation process on the first loss and the second loss based on the aggregation parameters corresponding to the first loss and the second loss respectively to obtain a joint loss; and update the parameters of the first translation model and the second translation model according to the joint loss.

[0175] In some embodiments, there are multiple first positions to be predicted with corresponding multiple first pre-labeled target words; the update module 2554 is further configured to: perform a fusion process on the first confidence levels obtained for each first position to be predicted to obtain a first loss corresponding to the first translation model.

[0176] In some embodiments, there are multiple second positions to be predicted with corresponding multiple second pre-labeled target words; the update module 2554 is further configured to: perform a fusion process on the first confidence level lower than the confidence threshold and the second confidence levels of the second pre-labeled target words corresponding to each second position to be predicted to obtain a second loss corresponding to the second translation model.

[0177] In some embodiments, the first translation model includes a first encoding network and a pre-order decoding network; the first task module 2551 is further configured to: determine each original word of the corpus sample and the original word vector corresponding to each original word, combine the original word vectors corresponding to each original word to obtain an original word vector sequence of the corpus sample; perform semantic encoding processing on the original word vector sequence of the corpus sample through the first encoding network to obtain a first source sentence representation corresponding to the corpus sample; perform corpus decoding processing on the first source sentence representation through the pre-order decoding network to obtain a first confidence level of the first pre-marked target word corresponding to each first position to be predicted; wherein the first confidence level is generated based on the previous words corresponding to each first position to be predicted.

[0178] In some embodiments, the encoding network includes N cascaded sub-encoding networks, where N is an integer greater than or equal to 2; the first task module 2551 is further configured to: perform semantic encoding processing on the original word vector sequence of the corpus sample in the following manner through the N cascaded first sub-encoding networks included in the first encoding network: perform self-attention processing on the input of the first sub-encoding network to obtain a self-attention processing result corresponding to the first sub-encoding network, perform hidden state mapping processing on the self-attention processing result to obtain a hidden state vector sequence corresponding to the first sub-encoding network, and use the hidden state vector sequence as the semantic encoding processing result of the first sub-encoding network; wherein, among the N cascaded first sub-encoding networks, the input of the first first sub-encoding network includes the original word vector sequence of the corpus sample, and the semantic encoding processing result of the Nth first sub-encoding network includes a first source sentence representation corresponding to the corpus sample.

[0179] In some embodiments, the first task module 2551 is further configured to: perform the following processing for each original word of the corpus sample: perform linear transformation processing on the first intermediate vector corresponding to the original word in the input of the first sub-encoding network to obtain a query vector, a key vector, and a value vector corresponding to the original word; perform dot product processing on the query vector of the original word and the key vector of each original word, and perform normalization processing on the dot product processing result based on the maximum likelihood function to obtain the weight of the value vector of the original word; perform weighted processing on the value vector of the original word based on the weight of the value vector of the original word to obtain the self-attention processing result of the sub-encoding network corresponding to each original word.

[0180] In some embodiments, the previous decoding network output for each of the first prediction position to be performed the following processing: obtaining from the corpus sample set corresponding to the corpus sample of the first pre-labeled target word sequence; extracting from the first pre-labeled target word sequence located before the first prediction position of the first pre-labeled target word, the first pre-labeled target word extraction as the previous word corresponding to the first prediction position; by the previous decoding network corresponding to the first prediction position of the previous word and the first source sentence representation for semantic decoding process, the first prediction position is decoded to obtain the corresponding first pre-labeled target word of the first confidence level.

[0181] In some embodiments, the previous decoding network includes M cascaded sub-previous decoding networks, M is an integer greater than or equal to 2; the first task module 2551, is further configured to: perform semantic decoding process in the following manner by each sub-previous decoding network: the input of the sub-previous decoding network for masked self-attention processing, to obtain the corresponding sub-previous decoding network masked self-attention processing result, the masked self-attention processing result for cross-attention processing, to obtain the corresponding sub-previous decoding network cross-attention processing result, the cross-attention processing result for hidden state mapping process; wherein, in the M cascaded sub-previous decoding networks, the input of the first sub-previous decoding network includes: the previous word corresponding to the first prediction position and the first source sentence representation; the M-th sub-previous decoding network hidden state mapping process result includes: the first prediction position is decoded to obtain the corresponding first pre-labeled target word of the first confidence level.

[0182] In some embodiments, the first task module 2551, is further configured to: perform a linear transformation process on the masked self-attention processing result, to obtain a query vector of the masked self-attention processing result; for each of the original word to perform the following processing: the first source sentence representation of the original word for linear transformation process, to obtain the first source sentence representation of the key vector and the value vector; the masked self-attention processing result of the query vector and the first source sentence representation of the key vector for dot product process, and the dot product process result for normalization process based on the maximum likelihood function, to obtain the first source sentence representation of the value vector weight; based on the first source sentence representation of the value vector weight of the first source sentence representation of the value vector for weighted processing, to obtain the corresponding sub-previous decoding network cross-attention processing result.

[0183] In some embodiments, the second translation model includes a second encoding network and a context decoding network; the second task module 2553 is further configured to: obtain each original word of the corpus sample and the original word vector corresponding to each original word, combine the original word vectors corresponding to each original word to obtain the original word vector sequence of the corpus sample; perform semantic encoding processing on the original word vector sequence of the corpus sample through the second encoding network to obtain the second source sentence representation corresponding to the corpus sample; perform corpus decoding processing on the second source sentence representation through the context decoding network to obtain the second confidence of the second pre-marked target word corresponding to each second position to be predicted; wherein, the second confidence is generated based on the context words corresponding to multiple second positions to be predicted.

[0184] In some embodiments, the second task module 2553 is further configured to: perform the following processing for each second position to be predicted output by the context decoding network: perform semantic decoding processing on the context words and the second source sentence representation through the context decoding network to obtain the second confidence of the second pre-marked target word decoded at the second position to be predicted.

[0185] In some embodiments, the context decoding network includes P cascaded sub-context decoding networks, where P is an integer greater than or equal to 2; the second task module 2553 is further configured to: perform semantic decoding processing on the context word set corresponding to the second position to be predicted and the source sentence representation through the P cascaded sub-context decoding networks in the following manner: perform context self-attention processing on the input of the sub-context decoding network to obtain the context self-attention processing result corresponding to the sub-context decoding network; perform cross-attention processing on the context self-attention processing result to obtain the cross-attention processing result corresponding to the sub-context decoding network; perform hidden state mapping processing on the cross-attention processing result; wherein, in the P cascaded sub-context decoding networks, the input of the first sub-context decoding network includes: the context word set corresponding to the second position to be predicted and the second source sentence representation; the hidden state mapping processing result of the P-th sub-context decoding network includes: the second confidence of the second pre-marked target word decoded at the second position to be predicted.

[0186] An embodiment of the present application provides a corpus translation device 256 for a translation model, including: an application module 2555, configured to, in response to a translation request for a target corpus, call the first translation model or the second translation model to perform translation processing on the target corpus to obtain a translation result for the target corpus; wherein, the first translation model and the second translation model are trained according to the training method of the translation model provided by the embodiment of the present application.

[0187] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the training method of the translation model and the corpus translation method of the embodiment of the present application.

[0188] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, where the executable instructions are stored. When the executable instructions are executed by a processor, the processor will be caused to execute the training method of the translation model provided by the embodiment of the present application. For example, as Figures 3A - 3D shown in the training method of the translation model.

[0189] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0190] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0191] As an example, the executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file storing other programs or data. For example, stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines, or code portions).

[0192] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.

[0193] In summary, according to the embodiments of the present application, the characteristics and confidence of the translation task of a neural network model are utilized to assist in the targeted training of another neural network model. Since the second translation model performs translation using the context word set, bidirectional global context information is effectively introduced through the context word set, and through the confidence threshold, the second translation model introduces bidirectional global context information based on context words for the first translation model at positions with relatively low confidence at the target end, so that the first translation model after joint training can not only utilize the local context information of the previous word corresponding to each position to be predicted during translation, but also can utilize the global context information in a targeted manner, thereby effectively improving the accuracy of translation by the first translation model.

[0194] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are all included in the protection scope of the present application.

Claims

1. A training method for a translation model, characterized in that, Including: Performing forward propagation on a corpus sample in a first translation model to obtain a first confidence level of a first pre-marked target word corresponding to each first position to be predicted; Determining a second position to be predicted of a second translation model for a first position to be predicted with a first confidence level lower than a confidence threshold, and determining a first pre-marked target word corresponding to a first position to be predicted with a first confidence level not lower than the confidence threshold as a context word corresponding to the second translation model; Performing forward propagation on the context word and the corpus sample in the second translation model to obtain a second confidence level of a second pre-marked target word corresponding to each second position to be predicted; Updating parameters of the first translation model and the second translation model based on the first confidence level of the first pre-marked target word corresponding to each first position to be predicted and the second confidence level of the second pre-marked target word corresponding to each second position to be predicted.

2. The method according to claim 1, wherein The updating the parameters of the first translation model and the second translation model based on the first confidence level of the first pre-marked target word corresponding to each first position to be predicted and the second confidence level of the second pre-marked target word corresponding to each second position to be predicted includes: Determining a first loss corresponding to the first translation model based on the first confidence level; Determining a second loss corresponding to the second translation model based on the first confidence level lower than the confidence threshold and the second confidence level of the second pre-marked target word corresponding to each second position to be predicted; wherein the second loss is used to represent a teaching loss of the second translation model for the first translation model; Performing an aggregation process on the first loss and the second loss based on aggregation parameters respectively corresponding to the first loss and the second loss to obtain a joint loss; Updating the parameters of the first translation model and the second translation model according to the joint loss.

3. The method according to claim 2, wherein There are a plurality of first positions to be predicted with a plurality of corresponding first pre-marked target words in one-to-one correspondence, and a plurality of second positions to be predicted with a plurality of corresponding second pre-marked target words in one-to-one correspondence; The determining a first loss corresponding to the first translation model based on the first confidence level includes: Performing a fusion process on the first confidence levels obtained for each first position to be predicted to obtain a first loss corresponding to the first translation model; The determining a second loss corresponding to the second translation model based on the first confidence level lower than the confidence threshold and the second confidence level of the second pre-marked target word corresponding to each second position to be predicted includes: Performing a fusion process on the first confidence level lower than the confidence threshold and the second confidence levels of the second pre-marked target words corresponding to each second position to be predicted to obtain a second loss corresponding to the second translation model.

4. The method according to claim 1, wherein The first translation model includes a first encoding network and a previous decoding network; The performing forward propagation on a corpus sample in a first translation model to obtain a first confidence level of a first pre-marked target word corresponding to each first position to be predicted includes: Determine each original word of the corpus sample and the original word vector corresponding to each original word, and combine the original word vectors corresponding to each original word to obtain the original word vector sequence of the corpus sample; Perform semantic encoding processing on the original word vector sequence of the corpus sample through the first encoding network to obtain the first source sentence representation corresponding to the corpus sample; Perform corpus decoding processing on the first source sentence representation through the previous decoding network to obtain the first confidence level of the first pre-marked target word corresponding to each first prediction position; Wherein, the first confidence level is generated based on the previous words corresponding to each first prediction position.

5. The method according to claim 4, characterized in that, The first encoding network includes N cascaded first sub-encoding networks, and N is an integer greater than or equal to 2; The performing semantic encoding processing on the original word vector sequence of the corpus sample through the first encoding network to obtain the first source sentence representation corresponding to the corpus sample includes: Performing semantic encoding processing on the original word vector sequence of the corpus sample in the following manner through the N cascaded first sub-encoding networks included in the first encoding network: Perform self-attention processing on the input of the first sub-encoding network to obtain the self-attention processing result corresponding to the first sub-encoding network, perform hidden state mapping processing on the self-attention processing result to obtain the hidden state vector sequence corresponding to the first sub-encoding network, and use the hidden state vector sequence as the semantic encoding processing result of the first sub-encoding network; Wherein, among the N cascaded first sub-encoding networks, the input of the first first sub-encoding network includes the original word vector sequence of the corpus sample, and the semantic encoding processing result of the Nth first sub-encoding network includes the first source sentence representation corresponding to the corpus sample.

6. The method according to claim 4, wherein The performing corpus decoding processing on the first source sentence representation through the previous decoding network to obtain the first confidence level of the first pre-marked target word corresponding to each first prediction position includes: Perform the following processing for each first prediction position output by the previous decoding network: Obtain the first pre-marked target word sequence corresponding to the corpus sample from the corpus sample set; Extract the first pre-marked target word located before the first prediction position from the first pre-marked target word sequence, and use the extracted first pre-marked target word as the previous word corresponding to the first prediction position; Perform semantic decoding processing on the previous word corresponding to the first prediction position and the first source sentence representation through the previous decoding network to obtain the first confidence level of being decoded as the corresponding first pre-marked target word at the first prediction position.

7. The method according to claim 6, characterized in that The previous decoding network includes M cascaded sub-previous decoding networks, and M is an integer greater than or equal to 2; The performing semantic decoding processing on the previous word corresponding to the first prediction position and the first source sentence representation through the previous decoding network includes: The semantic decoding process is performed by each of the sub-preamble decoding networks in the following manner: performing masked self-attention processing on the input of the sub-preamble decoding network to obtain the masked self-attention processing result corresponding to the sub-preamble decoding network, performing cross-attention processing on the masked self-attention processing result to obtain the cross-attention processing result corresponding to the sub-preamble decoding network, and performing hidden state mapping processing on the cross-attention processing result; Among them, in the M cascaded sub-preamble decoding networks, the input of the first sub-preamble decoding network includes: the preamble word corresponding to the first position to be predicted and the first source sentence representation; the hidden state mapping processing result of the Mth sub-preamble decoding network includes: the first confidence level that the first position to be predicted is decoded into the corresponding first pre-token target word.

8. The method according to claim 7, wherein The performing cross-attention processing on the masked self-attention processing result to obtain the cross-attention processing result corresponding to the sub-preamble decoding network includes: Performing linear transformation processing on the masked self-attention processing result to obtain the query vector of the masked self-attention processing result; Performing the following processing for each of the original words: Performing linear transformation processing on the first source sentence representation of the original word to obtain the key vector and value vector of the first source sentence representation; Performing dot product processing on the query vector of the masked self-attention processing result and the key vector of the first source sentence representation, and performing normalization processing based on the maximum likelihood function on the dot product processing result to obtain the weight of the value vector of the first source sentence representation; Performing weighted processing on the value vector of the first source sentence representation based on the weight of the value vector of the first source sentence representation to obtain the cross-attention processing result corresponding to the sub-preamble decoding network.

9. The method according to claim 1, wherein The second translation model includes a second encoding network and a context decoding network; The performing forward propagation of the context word and the corpus sample in the second translation model to obtain the second confidence level of the second pre-token target word corresponding to each second position to be predicted includes: Obtaining each original word of the corpus sample and the original word vector corresponding to each original word, and performing combination processing on the original word vectors corresponding to each original word to obtain the original word vector sequence of the corpus sample; Performing semantic encoding processing on the original word vector sequence of the corpus sample through the second encoding network to obtain the second source sentence representation corresponding to the corpus sample; Performing corpus decoding processing on the second source sentence representation through the context decoding network to obtain the second confidence level of the second pre-token target word corresponding to each of the second positions to be predicted; Among them, the second confidence level is generated based on the context words corresponding to multiple second positions to be predicted.

10. The method according to claim 9, wherein The performing corpus decoding processing on the second source sentence representation through the context decoding network to obtain the second confidence level of the second pre-token target word corresponding to each of the second positions to be predicted includes: Performing the following processing for each second position to be predicted output by the context decoding network: Performing semantic decoding processing on the context word and the second source sentence representation through the context decoding network to obtain a second confidence level that is decoded as a corresponding second pre-token target word at the second position to be predicted.

11. A corpus translation method, characterized in that, Including: In response to a translation request for a target corpus, calling a first translation model or a second translation model to perform translation processing on the target corpus to obtain a translation result for the target corpus; Wherein, the first translation model and the second translation model are trained according to the training method of the translation model according to any one of claims 1-10.

12. A training device for a translation model, characterized in that, Including: A first task module for propagating the corpus sample forward in the first translation model to obtain a first confidence level of the first pre-token target word corresponding to each first position to be predicted; A selection module for determining the first position to be predicted with the first confidence level lower than the confidence threshold as the second position to be predicted of the second translation model, and determining the first pre-token target word corresponding to the first position to be predicted with the first confidence level not lower than the confidence threshold as the context word corresponding to the second translation model; A second task module for propagating the context word and the corpus sample forward in the second translation model to obtain a second confidence level of the second pre-token target word corresponding to each second position to be predicted; An update module for updating the parameters of the first translation model and the second translation model based on the first confidence level of the first pre-token target word corresponding to each first position to be predicted and the second confidence level of the second pre-token target word corresponding to each second position to be predicted.

13. A corpus translation device, characterized in that, Including: An application module for, in response to a translation request for a target corpus, calling a first translation model or a second translation model to perform translation processing on the target corpus to obtain a translation result for the target corpus; wherein, the first translation model and the second translation model are trained according to the training method of the translation model according to any one of claims 1-10.

14. An electronic device, characterized in that, Including: A memory for storing executable instructions; A processor for, when executing the executable instructions stored in the memory, implementing the training method of the translation model according to any one of claims 1 to 10 or the corpus translation method according to claim 11.

15. A computer-readable storage medium, characterized in that, Stored with executable instructions for, when being executed by a processor, implementing the training method of the translation model according to any one of claims 1 to 10 or the corpus translation method according to claim 11.

Citation Information

Patent Citations

  • Pre-training method and device of intelligent translation model and storage medium

    CN111460838A

  • Text translation method and device, electronic equipment and storage medium

    CN111931517A