A text translation method, related device and storage medium
By multiplexing and updating the weight matrix in each layer of network structure of machine translation, the problems of high time complexity and spatial complexity in machine translation are solved, and the translation efficiency is improved and high performance is maintained.
Patent Information
- Application Number
- CN202110443463.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-04-23
AI Technical Summary
During machine translation, the multi-head self-attention mechanism based on dot product operations requires a lot of time complexity and spatial complexity, resulting in lower translation efficiency as sentence length and transformer layers grow.
By multiplexing the weight matrix in each layer's network structure and updating it at the next level, the need to dot product calculations for the Query vector and Key vector of each word is reduced, thereby reducing the time complexity and spatial complexity.
This method effectively reduces the time complexity and spatial complexity in machine translation, improves translation efficiency, and maintains high translation performance, such as BLEU values.
Smart Images

Figure CN113761949B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a text translation method, related devices and storage media. Background Art
[0002] Machine translation is the process of using computers to convert one natural language into another. It is widely used in all aspects of life. Among them, the transformer architecture commonly used in machine translation has a strong ability to express semantics and can capture long-term dependencies in text. Since its introduction, it has significantly surpassed previous models in a series of natural language processing tasks represented by translation. Pre-trained language models based on transformer models have also achieved very good results in question-answering systems, voice assistants and other fields.
[0003] The transformer model uses a multi-head self-attention mechanism (Dot-Product Multi-head Self-Attention) based on the dot product operation to calculate the weight matrix. In this multi-head self-attention mechanism, each word has three different vectors: Query vector, Key vector and Value vector. In the text translation process, each transformer layer needs to perform a dot product operation on the Query vector, Key vector and Value vector of each word with the Query vector, Key vector and Value vector of the word at other positions, so as to determine the semantic relationship between each word and the words at other positions and obtain the feature vector of each word.
[0004] In the multi-head self-attention mechanism based on dot product operations, since the word-to-word dot product operations require a lot of time complexity and space complexity, as the sentence length and the number of transformer layers increase, the time complexity and space complexity required in the machine translation process will increase significantly, resulting in lower efficiency of machine translation. Summary of the invention
[0005] In view of this, an embodiment of the present application provides a method, related device and storage medium for text translation, which are used to reduce the time complexity and space complexity in the machine translation process and improve the efficiency of machine translation.
[0006] On one hand, the present application provides a method for text translation, comprising:
[0007] Get the first embedding vector corresponding to the target text;
[0008] Obtain a first initial weight matrix;
[0009] The first initial weight matrix is processed using the first layer network structure to obtain a first weight matrix;
[0010] Using a first layer network structure, processing the first embedding vector through a first weight matrix to obtain a first eigenvector;
[0011] The first weight matrix is processed by using a second layer network structure to obtain a second weight matrix;
[0012] Using a second-layer network structure, the first eigenvector is processed by a second weight matrix to obtain a second eigenvector;
[0013] A first text translation result is obtained according to the second feature vector.
[0014] Another aspect of the present application provides a text translation device, comprising:
[0015] An acquisition unit, used to acquire a first embedding vector corresponding to the target text;
[0016] The acquisition unit is further used to acquire a first initial weight matrix;
[0017] A processing unit, configured to process the first initial weight matrix using the first layer network structure to obtain a first weight matrix;
[0018] The processing unit is further used to process the first embedding vector by using the first layer network structure through the first weight matrix to obtain a first eigenvector;
[0019] The processing unit is further used to process the first weight matrix using a second layer network structure to obtain a second weight matrix;
[0020] The processing unit is further used to process the first eigenvector by a second weight matrix using a second layer network structure to obtain a second eigenvector;
[0021] The acquiring unit is further configured to acquire a translation result of the first text according to the second feature vector.
[0022] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0023] The first layer of the network structure includes the first attention layer and the first feedforward layer.
[0024] The processing unit is specifically configured to use a first attention layer to process the first embedding vector through a first weight matrix to obtain a first attention vector;
[0025] The first feed-forward layer is used to process the first attention vector and the first embedding vector to obtain the first feature vector.
[0026] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0027] The first layer of the network structure includes the second attention layer, the encoder-decoder layer and the second feedforward layer.
[0028] The processing unit is specifically used to use the second attention layer to process the first embedding vector through the first weight matrix to obtain a second attention vector;
[0029] Obtaining a third feature vector from the encoder;
[0030] Using the Encoder-Decoder layer, the third feature vector is processed by the second attention vector to obtain the third attention vector;
[0031] The third attention vector and the third eigenvector are processed using the second feed-forward layer to obtain the first eigenvector.
[0032] In a possible design, in an implementation of another aspect of the embodiment of the present application, the text updating device further includes: an updating unit;
[0033] The acquisition unit is further used to acquire a second embedding vector corresponding to the training text;
[0034] The acquisition unit is further used to acquire a second initial weight matrix;
[0035] The processing unit is further used to process the second initial weight matrix using the first layer network structure to obtain a third weight matrix;
[0036] The processing unit is further used to process the second embedding vector by using the first layer network structure through the third weight matrix to obtain a fourth eigenvector;
[0037] The processing unit is further used to process the third weight matrix using the second layer network structure to obtain a fourth weight matrix;
[0038] The processing unit is further used to process the fourth eigenvector by using the second layer network structure through the fourth weight matrix to obtain a fifth eigenvector;
[0039] The acquiring unit is further used to acquire the second text translation result according to the fifth feature vector;
[0040] The updating unit is used to update the second initial weight matrix according to the second text translation result to obtain the first initial weight matrix.
[0041] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0042] The processing unit is further used to process the second feature vector using the third layer network structure to obtain a Query vector, a Key vector and a Value vector corresponding to the second feature vector;
[0043] The processing unit is further used to process the Query vector, the Key vector and the Value vector using the third layer network structure to obtain a fourth attention vector;
[0044] The processing unit is also used to use the third-layer network structure to process the fourth attention vector and the second eigenvector to obtain an updated second eigenvector.
[0045] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0046] A processing unit, specifically configured to process a first initial weight matrix using a first formula in a first layer network structure to obtain a first weight matrix;
[0047] The first formula is:
[0048] S1 = tanh(W*S0+b)+S0;
[0049] Among them, S0 is the first initial weight matrix, S1 is the first weight matrix, W and b are the model parameters of the first transformer layer, and tanh represents the hyperbolic tangent processing of W, S1 and b.
[0050] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0051] A processing unit, specifically configured to process the first weight matrix and the first embedding vector using a second formula in the first attention layer to obtain a first attention vector;
[0052] The second formula is:
[0053] C1=Softmax(S1)H0;
[0054] Among them, S1 is the first weight matrix, H0 is the first embedding vector, and C1 is the first attention vector.
[0055] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0056] The processing unit is also used to use the first attention layer to process the first weight matrix through the first embedding vector to obtain an updated first weight matrix.
[0057] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0058] A processing unit, specifically configured to process the first embedding vector and the first weight matrix in the first layer network structure using a third formula to obtain an updated first weight matrix;
[0059] The third formula is:
[0060] S1' = tanh(W1*S1+W2*H0+b);
[0061] Among them, S1 is the first weight matrix, H0 is the first embedding vector, W1, W2 and b are the model parameters of the first transformer layer, and S1' is the updated first weight matrix.
[0062] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0063] The processing unit is specifically used to use the first attention layer to process the first embedding vector through the updated first weight matrix to obtain the first attention vector.
[0064] In a possible design, in an implementation of another aspect of the embodiment of the present application,
[0065] A processing unit, specifically configured to process the updated first weight matrix and the first embedding vector in the first attention layer using the fourth formula to obtain a first attention vector;
[0066] The fourth formula is:
[0067] C1=Softmax(S1')H0;
[0068] Among them, S1' is the updated first weight matrix, H0 is the first embedding vector, and C1 is the first attention vector.
[0069] In a possible design, in an implementation of another aspect of the embodiment of the present application, the text translation device further includes: a generation unit.
[0070] The acquisition unit is also used to acquire the target speech;
[0071] A generating unit, used for generating a target text according to a target speech;
[0072] The generating unit is further used to generate a first embedding vector corresponding to the target text.
[0073] On the other hand, the present application provides a computer device, including: a memory, a processor and a bus system; the memory is used to store program codes; the processor is used to execute any of the above-mentioned text translation methods according to instructions in the program codes.
[0074] On the other hand, the present application provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is executed on a computer, the computer executes any one of the above-mentioned methods for text translation.
[0075] According to another aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the text translation method of any of the above aspects.
[0076] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0077] In an embodiment of the present application, a method for text translation is provided, and a first embedding vector corresponding to a target text is obtained, and a first initial weight matrix is obtained. The first initial weight matrix is processed by a first-layer network structure to obtain a first weight matrix, and then the first embedding vector is processed by the first weight matrix to obtain a first eigenvector, and input into a second-layer network structure. The first weight matrix obtained by processing the first-layer network structure is further input into the second-layer network structure for processing to obtain a second weight matrix, and the first eigenvector is processed by the second weight matrix to obtain a second eigenvector. According to the second eigenvector, the translation result of the first text is obtained. In the above manner, the weight matrix of each layer of the network structure is input into the next layer for updating, and the eigenvector corresponding to the target text is generated using the weight matrix, which reduces the time complexity and space complexity in the machine translation process and improves the efficiency of machine translation. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0079] Figure 1 A diagram of the network architecture that runs the text translation system;
[0080] Figure 2 A schematic diagram of a transformer framework provided in an embodiment of the present application;
[0081] Figure 3 A schematic diagram of an embodiment of the method for text translation in the embodiments of the present application;
[0082] Figure 4 A schematic diagram of processing the first initial weight matrix to obtain the first weight matrix in an embodiment of the present application;
[0083] Figure 5 A schematic diagram of obtaining an updated first weight matrix by processing the first weight matrix through the first embedding vector in an embodiment of the present application;
[0084] Figure 6 This is a schematic diagram of the structure of a neural network model in an embodiment of the present application;
[0085] Figure 7 This is a schematic diagram of the structure of a text translation device in an embodiment of the present application;
[0086] Figure 8 A schematic diagram of the structure of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0087] The embodiments of the present application provide a text translation method, related devices and storage medium for reducing the time complexity and space complexity in the machine translation process and improving the efficiency of machine translation.
[0088] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "corresponding to" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0089] It should be understood that this application uses natural language processing (NLP) technology based on artificial intelligence (AI) to achieve text translation. Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0090] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0091] Natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0092] Exemplarily, the text translation method provided in the embodiments of the present application can be applied to the following scenarios:
[0093] 1. Machine Translation:
[0094] In this application scenario, the machine translation model trained by the method provided in the embodiment of the present application can be applied to applications supporting translation functions such as electronic dictionary applications, e-book applications, web browsing applications, social applications, and image and text recognition applications. When the above-mentioned application receives the content to be translated, the trained machine translation model outputs the translation result according to the input content to be translated. Schematically, the content to be translated includes at least one of text-type content, picture-type content, audio-type content, and video-type content, wherein the picture-type content includes photos taken by the camera assembly of the terminal or pictures containing the content to be translated, which is not limited in the embodiment of the present application. By identifying the text content in the picture, the identified text content is input into the trained machine translation model, and the translation result is output and displayed to the user. Illustratively, by recognizing text in audio (such as converting the human voice contained in the audio into text) and inputting the recognized text content into a trained machine translation model, the translation result can also be output; by recognizing text in a video (such as recognizing subtitles in a video, or converting the human voice contained in the video into text) and inputting the recognized text content into a trained machine translation model, the translation result can also be output.
[0095] 2. Dialogue Q&A:
[0096] In this application scenario, the machine translation model trained by the method provided in the embodiment of the present application can be applied to smart devices such as smart terminals or smart homes. Taking the virtual assistant set in the smart terminal as an example, the automatic answer function of the virtual assistant is realized by the machine translation model after the above training. The user asks the virtual assistant questions about translation. When the virtual assistant receives the questions input by the user (the questions input by the user can be implemented in the form of voice or text input), the machine translation model outputs the translation result according to the input question. The translation result is converted into the form of voice or text, and fed back to the user with the help of the virtual assistant.
[0097] The above description only takes two application scenarios as examples. The method provided in the embodiments of the present application can also be applied to the application scenario of extracting text summaries. The embodiments of the present application do not limit the specific application scenarios.
[0098] Specifically, the text translation method provided in this application can be applied to Figure 1 In the network architecture shown in Figure 1 As shown in FIG. 1 , it is a network architecture diagram of the text translation system. As can be seen from the figure, the text translation system can provide a text translation process with multiple information sources, that is, through the server parsing and translating the information flow, the information that meets the user's needs is sent to the terminal side; it can be understood that Figure 1A variety of terminal devices are shown in FIG. 1 . The terminal devices may be computer devices. In actual scenarios, more or fewer types of terminal devices may participate in the process of text recognition based on location information. The specific number and type depend on the actual scenario and are not limited here. In addition, Figure 1 One server is shown in the figure, but in an actual scenario, multiple servers may be involved, and the specific number of servers depends on the actual scenario.
[0099] In this embodiment, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected through wired or wireless communication, and the terminal and the server can be connected to form a blockchain network, which is not limited in this application.
[0100] The text translation method provided in the embodiment of the present application can be applied to Figure 2 In the transformer framework shown in FIG, the encoder and decoder models are included, and the structures of the encoding layer and the decoding layer are similar or identical. Figure 2 As shown, both the encoder and the decoder include an embedding layer and at least one transformer layer, wherein each transformer layer includes an attention layer, a residual connection and normalization (add&norm) layer, a feed forward layer, and a residual connection and normalization layer that are adjacent in sequence. Taking the encoder as an example, in the embedding layer, the currently input target text is embedded to obtain an embedding vector for each word in the target text; in the attention layer, the current transformer layer can obtain the weight matrix and the feature vector of the target text from the previous transformer layer, and update the weight matrix from the previous layer based on the model parameters of the current transformer layer, and then use the updated weight matrix to process the feature vector from the previous layer to obtain the attention vector of the current transformer layer for the target text. Specifically, the weight matrix represents the attention weight corresponding to each word in the target text, and the attention vector corresponding to the target text can be calculated by combining the embedding vector of the target text and the weight matrix of each transformer layer.
[0101] It should be understood that the text translation method provided in the embodiment of the present application can input the weight matrix processed by each level of the transformer layer into the transformer layer of the next level, and the transformer layer of the next level continues to update and iterate the input weight matrix until the last transformer layer. Figure 2 The transformer framework shown in the figure, the text translation method provided in the present application can be applied to the processing of the weight matrix by the encoder alone, can also be applied to the processing of the weight matrix by the decoder, and can also be applied to the encoder and the decoder at the same time, which is not limited here.
[0102] The new weight matrix and attention vector processed by the attention layer are also input into the feedforward layer, which processes the new weight matrix and attention vector to calculate the feature vector of the current transformer layer for the target text.
[0103] Finally, the current transformer layer will output a new weight matrix and a new feature vector, which will serve as the input of the next transformer layer. The above steps of updating the weight matrix and feature vector will be repeated until the last transformer layer.
[0104] Similarly, the attention layer in the decoder also performs the same steps as the attention layer in the encoder until the last transformer layer, obtaining the feature vector corresponding to the target text, which is then processed by the output layer to generate the text translation result. I will not go into details here.
[0105] In combination with the above introduction, the text translation method provided by this application is introduced below. Figure 3 , Figure 3 This is a schematic diagram of an embodiment of the method for text translation in the embodiment of the present application. As shown in the figure, an embodiment of the method for text translation in the embodiment of the present application includes:
[0106] 101. Obtain a first embedding vector corresponding to the target text;
[0107] For easier understanding, see Figure 2 In the transformer framework shown in the figure, from the perspective of the encoder, the target text is first input into the embedding layer of the encoder and converted into the first feature vector. From the perspective of the decoder, the target text should be the start symbol, and it also needs to pass through the embedding layer of the decoder to convert the first feature vector.
[0108] It should be understood that the present application does not limit the language type of the target text, and the target text may be a Chinese text, an English text, or a text in another language.
[0109] After obtaining the target text, the embedding layer can embed each word in the target text to obtain the embedding vector of each word. Figure 2 The embedding layer in the target text may specifically include an input embedding layer and a positional encoding layer. In the input embedding layer, each word in the target text may be subjected to word embedding processing to obtain an embedding vector for each word. In the position encoding layer, the position of each word in the target text may be obtained, and then a position vector may be generated for the position of each word. In some examples, the position of each word may be the absolute position of each word in the target text. Taking the target text "How are you?" as an example, the position of "you" may be represented as the first position, the position of "good" may be represented as the second position, and the position of "may" may be represented as the third position. In some examples, the position of each word may be the relative position between each word. Still taking the target text "How are you?" as an example, the position of "you" may be represented as before "good", the position of "good" may be represented as after "you" and before "may", and the position of "may" may be represented as after "good". When the word vectors and position vectors of each word in the target text are obtained, the position vectors of each word and the corresponding word embedding vectors may be combined to obtain the embedding vectors of each word, that is, to obtain multiple embedding vectors corresponding to the target text, which are the first embedding vectors of the present application.
[0110] 102. Obtain a first initial weight matrix;
[0111] like Figure 2 As shown, it should be understood that both the encoder and the decoder include at least one embedding layer and N transformer layers. In the embodiment of the present application, initial weight matrices are configured for the encoder and the decoder, respectively, as the input of the first transformer layer. Specifically, the weight matrix represents the attention weight corresponding to each word in the target text, and the attention vector corresponding to the target text can be calculated by combining the embedding vector of the target text and the weight matrix of each transformer layer.
[0112] It should be understood that the attention mechanism of neural networks mimics the internal process of biological observation behavior, that is, a mechanism that aligns internal experience and external sensations to increase the fineness of observation of some areas, and can use limited attention resources to quickly filter out high-value information from a large amount of information. The attention mechanism can quickly extract important features of sparse data, and is therefore widely used in natural language processing tasks, especially machine translation. The self-attention mechanism is an improvement on the attention mechanism, which reduces dependence on external information and is better at capturing the internal correlation of data or features.
[0113] The weight matrix provided in this application is independent of the semantic relationship between each word in the target text. The attention weight value corresponding to each word is only related to its absolute position or relative position in the target text. Therefore, in the process of calculating the attention vector of each word in the target text, it is not necessary to consider the semantic relationship between each word, thereby improving the efficiency of machine translation.
[0114] 103. Processing the first initial weight matrix using the first layer network structure to obtain a first weight matrix;
[0115] The first initial weight matrix will be used as the input of the first layer of the network structure (hereinafter referred to as the first transformer layer), and the first transformer layer will process the first initial weight to obtain the first weight matrix.
[0116] For easier understanding, see Figure 4 , Figure 4 The first initial weight matrix is processed in the embodiment of the present application to obtain a schematic diagram of the first weight matrix. Among them, as the bottom layer (first transformer layer) of the transformer architecture, there will be an initial weight matrix first, and then the initial weight matrix can be updated to obtain the first weight matrix, and applied to the subsequent processing of the embedded vector and output to the next level of transformer layer.
[0117] Specifically, the first transformer layer includes a first attention layer. In the first attention layer, the first initial weight matrix is processed using the first formula to obtain a first weight matrix.
[0118] The first formula is:
[0119] S1 = tanh(W*S0+b)+S0;
[0120] Among them, S0 is the first initial weight matrix, S1 is the first weight matrix, W and b are the model parameters of the first transformer layer, and tanh represents the hyperbolic tangent processing of W, S0 and b.
[0121] Since the weight matrix provided in the embodiment of the present application is independent of the semantic relationship between each word in the target text, in order to improve the accuracy of text translation, in a possible implementation method, after processing the first initial weight matrix to obtain the first weight matrix, the obtained first weight matrix can be fused with the semantic information between each word to obtain an updated first weight matrix.
[0122] For easier understanding, see Figure 5 , Figure 5In the embodiment of the present application, the first weight matrix is processed by the first embedding vector to obtain a schematic diagram of the updated first weight matrix. Specifically, in the first attention layer, the first embedding vector and the first weight matrix can be processed by the third formula to obtain the updated first weight matrix;
[0123] The third formula is:
[0124] S1' = tanh(W1*S1+W2*H0+b);
[0125] Among them, S1 is the first weight matrix, H0 is the first embedding vector, W1, W2 and b are the model parameters of the first transformer layer, and S1' is the updated first weight matrix.
[0126] In this embodiment, the first weight matrix and the first embedding vector based on the position information are processed, and the updated first weight matrix can also integrate the semantic information between each word, thereby improving the accuracy of text translation.
[0127] 104. Using the first layer network structure, processing the first embedding vector by using the first weight matrix to obtain a first eigenvector;
[0128] After the first initial weight is processed by the first transformer layer to obtain a first weight matrix, the first embedding vector can be processed by the first weight matrix to obtain a first eigenvector.
[0129] For the encoder, the first transformer layer includes the first attention layer and the first feedforward layer. The specific steps to obtain the first feature vector are as follows:
[0130] Using a first attention layer, processing the first embedding vector through a first weight matrix to obtain a first attention vector;
[0131] The first feed-forward layer is used to process the first attention vector and the first embedding vector to obtain the first feature vector.
[0132] Specifically, in the first attention layer, the second formula may be used to process the first weight matrix and the first embedding vector to obtain a first attention vector;
[0133] The second formula is:
[0134] C1=Softmax(S1)H0;
[0135] Among them, S1 is the first weight matrix, H0 is the first embedding vector, and C1 is the first attention vector.
[0136] Further, in step 103, Figure 5 On the basis of the corresponding embodiment, after the first weight matrix is processed by the first embedding vector to obtain the updated first weight matrix, the updated first weight matrix at this time integrates the semantic information between each word. Then in step 103, the first embedding vector can be processed by the updated first weight matrix to obtain the first attention vector.
[0137] Specifically, in the first attention layer, the fourth formula may be used to process the updated first weight matrix and the first embedding vector to obtain a first attention vector;
[0138] The fourth formula is:
[0139] C1=Softmax(S1')H0;
[0140] Among them, S1' is the updated first weight matrix, H0 is the first embedding vector, and C1 is the first attention vector.
[0141] The first embedding vector is processed with the updated first weight matrix to obtain the first attention vector. Since the updated first weight matrix integrates the semantic information between each word, the generated first feature vector can be made more accurate, thereby improving the accuracy of machine translation.
[0142] For the decoder, the first transformer layer includes the second attention layer, the encoder-decoder layer, and the second feedforward layer. The specific steps to obtain the first feature vector are as follows:
[0143] Using a second attention layer, the first embedding vector is processed by the first weight matrix to obtain a second attention vector;
[0144] Obtaining a third feature vector from the encoder;
[0145] Using the Encoder-Decoder layer, the third feature vector is processed by the second attention vector to obtain the third attention vector;
[0146] The third attention vector and the third eigenvector are processed using the second feed-forward layer to obtain the first eigenvector.
[0147] In the embodiment of the present application, the process of the decoder calculating and generating the second attention vector is similar to the implementation method of the encoder calculating and generating the first attention vector, and the details will not be repeated here.
[0148] The decoder is different from the encoder in that, on the one hand, the second attention vector generated by the second attention layer of the decoder is output to the Encoder-Decoder layer, rather than directly to the feedforward layer; the Encoder-Decoder layer also receives the third feature vector from the encoder, which is the output of the last transformer layer of the encoder. The Encoder-Decoder layer can process the third feature vector through the second attention vector to obtain the third attention vector.
[0149] The process of calculating and generating the first feature vector by the decoder is similar to the implementation method of calculating and generating the first feature vector by the encoder, and the details are not repeated here.
[0150] In an embodiment of the present application, in the process of calculating the feature vector corresponding to each word, it is not necessary to perform dot product calculation on the Query vector and Key vector of each word. Instead, the weight matrix of each layer is reused to the next level for updating, thereby generating a feature vector in combination with the weight matrix, thereby reducing the time complexity and space complexity of the machine translation process and improving the efficiency of machine translation.
[0151] 105. Process the first weight matrix using a second layer network structure to obtain a second weight matrix;
[0152] The first weight matrix obtained by the first transformer layer will be used as the input of the next level (the second transformer layer) and processed by the second transformer layer to obtain the second weight matrix. The specific implementation method is similar to step 103 and will not be repeated here.
[0153] 106. Using a second-layer network structure, processing the first eigenvector by a second weight matrix to obtain a second eigenvector;
[0154] After the second weight matrix is generated, the implementation method of obtaining the second eigenvector in step 106 is similar to the implementation method of obtaining the first eigenvector in step 104, and will not be described in detail here.
[0155] 107. Obtain a translation result of the first text according to the second feature vector;
[0156] In the embodiment of the present application, on the one hand, for the encoder, since the second transformer layer is the last transformer layer of the encoder, the generated second feature vector needs to be input to the Encoder-Decoder layer of the decoder, and the decoder uses the second feature vector as the input of the Encoder-Decoder layer, and the decoder performs the subsequent translation process. Specifically, regarding the cyclic iteration method of each transformer layer in the decoder for the weight matrix and feature vector, reference can be made to the description of the aforementioned steps 101 to 106, which will not be repeated here.
[0157] On the other hand, for the decoder, after being processed by the last transformer layer (the second transformer layer), the processed second feature vector is passed to Figure 2 The output layer shown is used to obtain a first translation text.
[0158] It should be noted that the text translation method provided in the present application can be applied to an encoder and a decoder with at least two transformer layers. In this embodiment and subsequent embodiments, only an encoder and a decoder with two transformer layers are used as an example for explanation, wherein the first transformer layer is the bottom transformer layer, and the second transformer layer is the last transformer layer. In practical applications, the number of transformer layers in the encoder and the decoder is not limited, and the number of transformer layers can be set according to actual needs.
[0159] For easier understanding, see Figure 6 , Figure 6 This is a schematic diagram of the structure of a neural network model in an embodiment of the present application. Figure 6 The neural network model shown may be a neural network model of an encoder or decoder in an embodiment of the present application. Figure 6 As shown, the encoder or decoder may include multiple transformer layers connected in sequence. The Nth transformer layer receives the weight matrix and feature vector from the N-1th transformer layer. When the Nth transformer layer completes the update of the input weight matrix and feature vector, it continues to send the updated weight matrix and feature vector to the N+1th transformer layer, and repeats this cycle until the last transformer layer, and outputs the feature vector processed by the last transformer layer.
[0160] In an embodiment of the present application, a method for text translation is provided, and a first embedding vector corresponding to a target text is obtained, and a first initial weight matrix is obtained. The first initial weight matrix is processed by a first-layer network structure to obtain a first weight matrix, and then the first embedding vector is processed by the first weight matrix to obtain a first eigenvector, and input into a second-layer network structure. The first weight matrix obtained by processing the first-layer network structure is further input into the second-layer network structure for processing to obtain a second weight matrix, and the first eigenvector is processed by the second weight matrix to obtain a second eigenvector. According to the second eigenvector, the translation result of the first text is obtained. In the above manner, the weight matrix of each layer of the network structure is input into the next layer for updating, and the eigenvector corresponding to the target text is generated using the weight matrix, which reduces the time complexity and space complexity in the machine translation process and improves the efficiency of machine translation.
[0161] Furthermore, verification on the public data set NIST12 (ZH-EN) shows that compared with the existing dot product-based transformer architecture, the text translation method of this application reduces 5M parameters, increases the calculation speed by 8%, and reaches a BLEU value of 47.96, while the existing dot product-based transformer architecture has a corresponding BLEU value of 48.17. Therefore, the text translation method provided by this application can maintain high performance while improving the efficiency of machine translation.
[0162] Optionally, in the above Figure 3 On the basis of the corresponding embodiment, in another optional embodiment of the text translation method provided in the embodiment of the present application, before obtaining the first embedding vector corresponding to the target text, the following steps may also be included:
[0163] Obtain the second embedding vector corresponding to the training text;
[0164] Obtain a second initial weight matrix;
[0165] The second initial weight matrix is processed using the first layer network structure to obtain a third weight matrix;
[0166] Using the first layer network structure, the second embedding vector is processed by the third weight matrix to obtain a fourth eigenvector;
[0167] The third weight matrix is processed by using the second layer network structure to obtain a fourth weight matrix;
[0168] The fourth eigenvector is processed by a fourth weight matrix using a second layer network structure to obtain a fifth eigenvector;
[0169] According to the fifth eigenvector, obtain the second text translation result
[0170] According to the second text translation result, the second initial weight matrix is updated to obtain the first initial weight matrix.
[0171] In this embodiment, the training process of the text translation method provided in the embodiment of the present application is introduced. Specifically, the second embedding vector corresponding to the training text is processed, and the process of finally generating the fifth feature vector is similar to step 101 to step 106, and will not be described in detail here. After obtaining the training result (i.e., the second text translation result), the model parameters of the encoder and the decoder can be updated according to the second text translation result, and the initial weight matrix in the encoder and the decoder is updated, so as to improve the accuracy of the initial weight matrix.
[0172] Optionally, in the above Figure 3 On the basis of the corresponding embodiment, in another optional embodiment of the text translation method provided in the embodiment of the present application, before obtaining the first text translation result according to the second feature vector, a second layer network structure is used to process the first feature vector through a second weight matrix, and after obtaining the second feature vector, the following steps may also be included:
[0173] The third-layer network structure is used to process the second feature vector to obtain the Query vector, Key vector and Value vector corresponding to the second feature vector;
[0174] The third-layer network structure is used to process the Query vector, Key vector and Value vector to obtain the fourth attention vector;
[0175] The third-layer network structure is used to process the fourth attention vector and the second eigenvector to obtain an updated second eigenvector.
[0176] In this embodiment, after the second transformer layer processes and obtains the second feature vector, it can be output to the next level (third transformer layer), and the third transformer layer calculates the Query vector, Key vector and Value vector corresponding to the second feature vector according to the traditional method of calculating the attention vector, thereby further calculating the attention vector. In the above manner, the text translation method provided by the present application is combined with the traditional text translation method, which improves the flexibility of the solution implementation.
[0177] Optionally, in the above Figure 3 On the basis of the corresponding embodiment, in another optional embodiment of the text translation method provided in the embodiment of the present application, obtaining the first embedding vector corresponding to the target text specifically includes the following steps:
[0178] Get the target voice;
[0179] Generate target text according to target speech;
[0180] Generate the first embedding vector corresponding to the target text.
[0181] In this embodiment, after acquiring the target speech, the target speech can be recognized to generate a target text, and then the text translation method provided in this application is executed, thereby improving the feasibility of the solution.
[0182] Optional, in the above Figure 3 On the basis of the corresponding embodiment, in another optional embodiment of the text translation method provided in the embodiment of the present application, a text translation system for performing a text translation function may be composed of multiple servers, wherein the multiple servers may be composed of a blockchain, and the server is a node on the blockchain.
[0183] In order to better implement the above solution of the embodiment of the present application, the following also provides related devices for implementing the above solution. Figure 7 , Figure 7 A structural diagram of a text translation device provided in an embodiment of the present application, the text translation device includes:
[0184] An acquisition unit 201 is used to acquire a first embedding vector corresponding to the target text;
[0185] The acquisition unit 201 is further used to acquire a first initial weight matrix;
[0186] The processing unit 202 is further configured to process the first initial weight matrix using the first layer network structure to obtain a first weight matrix;
[0187] The processing unit 202 is further configured to use a first layer network structure to process the first embedding vector through a first weight matrix to obtain a first eigenvector;
[0188] The processing unit 202 is further configured to process the first weight matrix using a second layer network structure to obtain a second weight matrix;
[0189] The processing unit 202 is further configured to use a second layer network structure to process the first eigenvector through a second weight matrix to obtain a second eigenvector;
[0190] The acquisition unit 201 is further configured to acquire a translation result of the first text according to the second feature vector.
[0191] Optionally, in the above Figure 7 Based on the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application,
[0192] The first layer of the network structure includes the first attention layer and the first feedforward layer.
[0193] A processing unit 202 is specifically configured to use a first attention layer to process the first embedding vector through a first weight matrix to obtain a first attention vector;
[0194] The first feed-forward layer is used to process the first attention vector and the first embedding vector to obtain the first feature vector.
[0195] Optionally, in the above Figure 7 Based on the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application,
[0196] The first layer of the network structure includes the second attention layer, the encoder-decoder layer and the second feedforward layer.
[0197] A processing unit 202 is specifically configured to use a second attention layer to process the first embedding vector through a first weight matrix to obtain a second attention vector;
[0198] Obtaining a third feature vector from the encoder;
[0199] Using the Encoder-Decoder layer, the third feature vector is processed by the second attention vector to obtain the third attention vector;
[0200] The third attention vector and the third eigenvector are processed using the second feed-forward layer to obtain the first eigenvector.
[0201] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application, the text updating device further includes: an updating unit 203;
[0202] The acquisition unit 201 is further used to acquire a second embedding vector corresponding to the training text;
[0203] The acquisition unit 201 is further used to acquire a second initial weight matrix;
[0204] The processing unit 202 is further configured to process the second initial weight matrix using the first layer network structure to obtain a third weight matrix;
[0205] The processing unit 202 is further configured to use the first layer network structure to process the second embedding vector through the third weight matrix to obtain a fourth eigenvector;
[0206] The processing unit 202 is further configured to process the third weight matrix using the second layer network structure to obtain a fourth weight matrix;
[0207] The processing unit 202 is further configured to use the second layer network structure to process the fourth eigenvector by using a fourth weight matrix to obtain a fifth eigenvector;
[0208] The acquiring unit 201 is further configured to acquire a second text translation result according to the fifth feature vector;
[0209] The updating unit 203 is used to update the second initial weight matrix according to the second text translation result to obtain the first initial weight matrix.
[0210] Optionally, in the above Figure 7 Based on the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application,
[0211] The processing unit 202 is further used to process the second feature vector using the third layer network structure to obtain a Query vector, a Key vector and a Value vector corresponding to the second feature vector;
[0212] The processing unit 202 is further used to process the Query vector, the Key vector and the Value vector using the third layer network structure to obtain a fourth attention vector;
[0213] The processing unit 202 is also used to use the third-layer network structure to process the fourth attention vector and the second feature vector to obtain an updated second feature vector.
[0214] Optionally, in the above Figure 7 Based on the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application,
[0215] The processing unit 202 is specifically configured to process the first initial weight matrix using the first formula in the first layer network structure to obtain a first weight matrix;
[0216] The first formula is:
[0217] S1 = tanh(W*S0+b)+S0;
[0218] Among them, S0 is the first initial weight matrix, S1 is the first weight matrix, W and b are the model parameters of the first transformer layer, and tanh represents the hyperbolic tangent processing of W, S1 and b.
[0219] Optionally, in the above Figure 7 Based on the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application,
[0220] The processing unit 202 is specifically configured to process the first weight matrix and the first embedding vector in the first attention layer using the second formula to obtain a first attention vector;
[0221] The second formula is:
[0222] C1=Softmax(S1)H0;
[0223] Among them, S1 is the first weight matrix, H0 is the first embedding vector, and C1 is the first attention vector.
[0224] Optionally, in the above Figure 7 Based on the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application,
[0225] The processing unit 202 is also used to use the first attention layer to process the first weight matrix through the first embedding vector to obtain an updated first weight matrix.
[0226] Optionally, in the above Figure 7 Based on the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application,
[0227] The processing unit 202 is specifically configured to process the first embedding vector and the first weight matrix in the first layer network structure using a third formula to obtain an updated first weight matrix;
[0228] The third formula is:
[0229] S1' = tanh(W1*S1+W2*H0+b);
[0230] Among them, S1 is the first weight matrix, H0 is the first embedding vector, W1, W2 and b are the model parameters of the first transformer layer, and S1' is the updated first weight matrix.
[0231] Optionally, in the above Figure 7 Based on the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application,
[0232] The processing unit 202 is specifically configured to use a first attention layer to process the first embedding vector through an updated first weight matrix to obtain a first attention vector.
[0233] Optionally, in the above Figure 7 Based on the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application,
[0234] The processing unit 202 is specifically configured to process the updated first weight matrix and the first embedding vector in the first attention layer using the fourth formula to obtain a first attention vector;
[0235] The fourth formula is:
[0236] C1=Softmax(S1')H0;
[0237] Among them, S1' is the updated first weight matrix, H0 is the first embedding vector, and C1 is the first attention vector.
[0238] Optionally, in the above Figure 7 On the basis of the corresponding embodiment, in one embodiment of the text translation device provided by the embodiment of the present application, the text translation device further includes: a generation unit 204.
[0239] The acquisition unit 201 is also used to acquire the target speech;
[0240] A generating unit 204, configured to generate a target text according to the target speech;
[0241] The generating unit 204 is further configured to generate a first embedding vector corresponding to the target text.
[0242] In this embodiment, the text translation device can perform the above Figure 3 The operation of the illustrated embodiment will not be described in detail here.
[0243] The present application also provides a computer device for executing Figure 3 The operation of the corresponding embodiment. Figure 8 , Figure 8 It is a structural diagram of a computer device 300 in an embodiment of the present application. As shown in the figure, the computer device 300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 322 (for example, one or more processors) and memory 332, and one or more storage media 330 (for example, one or more mass storage devices) storing application programs 342 or data 344. Among them, the memory 332 and the storage medium 330 can be short-term storage or permanent storage. The program stored in the storage medium 430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the computer device. Furthermore, the central processing unit 322 can be configured to communicate with the storage medium 330 to execute a series of instruction operations in the storage medium 330 on the computer device 300.
[0244] The computer device 300 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input and output interfaces 358, and / or one or more operating systems 341, such as Windows Server 2008, ... TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM etc.
[0245] The steps performed in the above embodiments can be based on the Figure 8 The structure of the computer device shown.
[0246] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the methods described in the aforementioned embodiments.
[0247] The present application also provides a computer program product or a computer program in an embodiment, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the aforementioned Figure 3 The text translation method of the illustrated embodiment.
[0248] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0249] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0250] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0251] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0252] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, an interactive video management device, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0253] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for text translation, characterized in that: Applicable to an encoder and / or decoder including a first network structure and a second network structure; the first network structure is a first transformer layer, and the second network structure is a second transformer layer; The first network structure and the second network structure both include an attention layer and a feedforward layer; the method includes: Get the first embedding vector corresponding to the target text; Obtain a first initial weight matrix; Using the attention layer in the first network structure to process the first initial weight matrix to obtain a first weight matrix; Using the attention layer and the feedforward layer in the first network structure, the first embedding vector is processed by the first weight matrix to obtain a first feature vector; Using the attention layer in the second network structure to process the first weight matrix to obtain a second weight matrix; Using the attention layer and the feedforward layer in the second network structure, processing the first eigenvector by the second weight matrix to obtain a second eigenvector; A first text translation result is obtained according to the second feature vector.
2. The method according to claim 1, characterized in that The first network structure includes a first attention layer and a first feedforward layer. The first embedding vector is processed by the first weight matrix using the attention layer and the feedforward layer in the first network structure to obtain a first feature vector including: Using the first attention layer, processing the first embedding vector by the first weight matrix to obtain a first attention vector; The first feed-forward layer is used to process the first attention vector and the first embedding vector to obtain a first feature vector.
3. The method according to claim 1, characterized in that The first network structure includes a second attention layer, an encoder-decoder layer, and a second feedforward layer. The first embedding vector is processed by the first weight matrix using the attention layer and the feedforward layer in the first network structure to obtain a first feature vector including: Using the second attention layer, processing the first embedding vector by the first weight matrix to obtain a second attention vector; Obtaining a third feature vector from the encoder; Using the Encoder-Decoder layer, processing the third feature vector by the second attention vector to obtain a third attention vector; The second feed-forward layer is used to process the third attention vector and the third eigenvector to obtain a first eigenvector.
4. The method according to any one of claims 1 to 3, characterized in that Before obtaining the first embedding vector corresponding to the target text, the method further includes: Obtain the second embedding vector corresponding to the training text; Obtain a second initial weight matrix; The second initial weight matrix is processed by using the first layer network structure to obtain a third weight matrix; Using the first layer network structure, processing the second embedding vector by using the third weight matrix to obtain a fourth eigenvector; Processing the third weight matrix using a second layer network structure to obtain a fourth weight matrix; Using the second layer network structure, processing the fourth eigenvector by using the fourth weight matrix to obtain a fifth eigenvector; Obtaining a second text translation result according to the fifth feature vector; According to the second text translation result, the second initial weight matrix is updated to obtain a first initial weight matrix.
5. The method according to any one of claims 1 to 3, characterized in that Before obtaining the first text translation result according to the second feature vector, the method further includes: using the attention layer and the feedforward layer in the second network structure to process the first feature vector through the second weight matrix to obtain the second feature vector. Using a third-layer network structure, the second feature vector is processed to obtain a Query vector, a Key vector, and a Value vector corresponding to the second feature vector; Using a third-layer network structure, the Query vector, the Key vector, and the Value vector are processed to obtain a fourth attention vector; The third-layer network structure is used to process the fourth attention vector and the second feature vector to obtain an updated second feature vector.
6. The method according to any one of claims 1 to 3, characterized in that The first initial weight matrix is processed by the attention layer in the first network structure to obtain a first weight matrix, including: In the attention layer of the first network structure, the first initial weight matrix is processed by the first formula to obtain a first weight matrix; The first formula is: S1=tanh(W*S0+b)+S0; Among them, S0 is the first initial weight matrix, S1 is the first weight matrix, W and b are model parameters of the first layer network structure, and tanh represents hyperbolic tangent processing of W, S0 and b.
7. The method according to claim 2, characterized in that: The adopting the first attention layer and processing the first embedding vector by the first weight matrix to obtain the first attention vector includes: In the first attention layer, the first weight matrix and the first embedding vector are processed using a second formula to obtain a first attention vector; The second formula is: C1=Softmax(S1)H0; Among them, the S1 is the first weight matrix, the H0 is the first embedding vector, and the C1 is the first attention vector.
8. The method according to claim 2, characterized in that: Before the first attention layer is used to process the first embedding vector by the first weight matrix to obtain the first attention vector, after the attention layer in the first layer of the network structure is used to process the first initial weight matrix to obtain the first weight matrix, the method further includes: The first attention layer is used to process the first weight matrix through the first embedding vector to obtain an updated first weight matrix.
9. The method according to claim 8, characterized in that The first attention layer is used to process the first weight matrix through the first embedding vector to obtain an updated first weight matrix, which includes: In the first layer network structure, the first embedding vector and the first weight matrix are processed using a third formula to obtain an updated first weight matrix; The third formula is: S1'=tanh(W1*S1+W2*H0+b); Among them, the S1 is the first weight matrix, the H0 is the first embedding vector, the W1, the W2 and the b are model parameters of the first layer network structure, and the S1' is the updated first weight matrix.
10. The method according to claim 8 or 9, characterized in that: The adopting the first attention layer and processing the first embedding vector by the first weight matrix to obtain the first attention vector includes: The first attention layer is used to process the first embedding vector through the updated first weight matrix to obtain a first attention vector.
11. The method according to claim 10, characterized in that The first attention layer is used to process the first embedding vector through the updated first weight matrix to obtain the first attention vector, including: In the first attention layer, the updated first weight matrix and the first embedding vector are processed using the fourth formula to obtain a first attention vector; The fourth formula is: C1=Softmax(S1')H0; Among them, the S1' is the updated first weight matrix, the H0 is the first embedding vector, and the C1 is the first attention vector.
12. The method according to claim 1, 2, 3, 7, 8 or 9, characterized in that: The obtaining of a first embedding vector corresponding to the target text includes: Get the target voice; Generating a target text according to the target speech; Generate a first embedding vector corresponding to the target text.
13. A text translation device, characterized in that: Applicable to an encoder and / or decoder including a first network structure and a second network structure; the first network structure is a first transformer layer, and the second network structure is a second transformer layer; The first network structure and the second network structure both include an attention layer and a feedforward layer; the device includes: An acquisition unit, used to acquire a first embedding vector corresponding to the target text; The acquisition unit is further used to acquire a first initial weight matrix; A processing unit, configured to process the first initial weight matrix using the attention layer in the first layer network structure to obtain a first weight matrix; The processing unit is further configured to use the attention layer and the feedforward layer in the first network structure to process the first embedding vector through the first weight matrix to obtain a first eigenvector; The processing unit is further configured to process the first weight matrix using the attention layer in the second layer network structure to obtain a second weight matrix; The processing unit is further configured to use the attention layer and the feedforward layer in the second network structure to process the first eigenvector through the second weight matrix to obtain a second eigenvector; The acquisition unit is further configured to acquire a first text translation result according to the second feature vector.
14. The device according to claim 13, characterized in that The first network structure includes a first attention layer and a first feedforward layer, and the processing unit is specifically used for: Using the first attention layer, processing the first embedding vector by the first weight matrix to obtain a first attention vector; The first feed-forward layer is used to process the first attention vector and the first embedding vector to obtain a first feature vector.
15. The device according to claim 13, characterized in that The first network structure includes a second attention layer, an encoder-decoder layer, and a second feedforward layer, and the processing unit is specifically used to: Using the second attention layer, processing the first embedding vector by the first weight matrix to obtain a second attention vector; Obtaining a third feature vector from the encoder; Using the Encoder-Decoder layer, processing the third feature vector by the second attention vector to obtain a third attention vector; The second feed-forward layer is used to process the third attention vector and the third eigenvector to obtain a first eigenvector.
16. The device according to any one of claims 13 to 15, characterized in that The device further comprises: an updating unit; The acquisition unit is further used to acquire a second embedding vector corresponding to the training text; The acquisition unit is further used to acquire a second initial weight matrix; The processing unit is further configured to process the second initial weight matrix using the first layer network structure to obtain a third weight matrix; The processing unit is further configured to use the first layer network structure to process the second embedding vector through the third weight matrix to obtain a fourth eigenvector; The processing unit is further configured to process the third weight matrix using a second layer network structure to obtain a fourth weight matrix; The processing unit is further configured to use the second layer network structure to process the fourth eigenvector through the fourth weight matrix to obtain a fifth eigenvector; The acquisition unit is further configured to acquire a second text translation result according to the fifth feature vector; The updating unit is used to update the second initial weight matrix according to the second text translation result to obtain a first initial weight matrix.
17. The device according to any one of claims 13 to 15, characterized in that The processing unit is further used to process the second feature vector using a third-layer network structure to obtain a Query vector, a Key vector, and a Value vector corresponding to the second feature vector; The processing unit is further configured to process the Query vector, the Key vector and the Value vector using a third-layer network structure to obtain a fourth attention vector; The processing unit is also used to use a third-layer network structure to process the fourth attention vector and the second feature vector to obtain an updated second feature vector.
18. The device according to any one of claims 13 to 15, characterized in that The processing unit is specifically used to process the first initial weight matrix using the first formula in the attention layer of the first network structure to obtain the first weight matrix; The first formula is: S1=tanh(W*S0+b)+S0; Among them, S0 is the first initial weight matrix, S1 is the first weight matrix, W and b are model parameters of the first layer network structure, and tanh represents hyperbolic tangent processing of W, S0 and b.
19. The device according to claim 14, characterized in that The processing unit is specifically configured to process the first weight matrix and the first embedding vector using a second formula in the first attention layer to obtain a first attention vector; The second formula is: C1=Softmax(S1)H0; Among them, the S1 is the first weight matrix, the H0 is the first embedding vector, and the C1 is the first attention vector.
20. The device according to claim 14, characterized in that The processing unit is also used to use the first attention layer to process the first weight matrix through the first embedding vector to obtain an updated first weight matrix.
21. The device according to claim 20, characterized in that The processing unit is specifically configured to process the first embedding vector and the first weight matrix using a third formula in the first layer network structure to obtain an updated first weight matrix; The third formula is: S1'=tanh(W1*S1+W2*H0+b); Among them, the S1 is the first weight matrix, the H0 is the first embedding vector, the W1, the W2 and the b are model parameters of the first layer network structure, and the S1' is the updated first weight matrix.
22. The device according to claim 20 or 21, characterized in that The processing unit is specifically used to use the first attention layer to process the first embedding vector through the updated first weight matrix to obtain a first attention vector.
23. The device according to claim 22, characterized in that The processing unit is specifically configured to process the updated first weight matrix and the first embedding vector in the first attention layer using a fourth formula to obtain a first attention vector; The fourth formula is: C1=Softmax(S1')H0; Among them, the S1' is the updated first weight matrix, the H0 is the first embedding vector, and the C1 is the first attention vector.
24. The device according to claim 13, 14, 15, 19, 20 or 21, characterized in that The device further comprises: a generating unit; The acquisition unit is further used to acquire the target speech; The generating unit is used to generate a target text according to the target speech; The generating unit is further configured to generate a first embedding vector corresponding to the target text.
25. A computer device, characterized in that: The computer device comprises a processor and a memory: The memory is used to store program codes; the processor is used to execute the text translation method according to any one of claims 1 to 12 according to instructions in the program codes.
26. A computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, and when the computer-readable storage medium is executed on a computer, the computer is enabled to execute the text translation method according to any one of claims 1 to 12.
27. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text translation method described in any one of claims 1 to 12 above.
Citation Information
Patent Citations
Text translation method and device, readable storage medium and computer equipment
CN111382584A
Data processing method and related equipment
CN112288075A