Network representation generation, encoding method and device in neural network

By extracting and calculating the key vector sequence of local elements and their correlation in the neural network and generating weight coefficients, the problem of weakening of local information in network representation generation in neural networks is solved, and the local context capture capability and network representation quality are improved.

CN110163339BActive Publication Date: 2025-06-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910167405.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-03-06
Publication Date
2025-06-17
Estimated Expiration
2039-11-12

AI Technical Summary

Technical Problem

In neural networks, networks represent that local information is weakened, resulting in insufficient ability to capture local contexts.

Method used

By obtaining the spatial representation of elements in the input sequence, a key-value pair vector sequence and a request vector sequence are generated, a key-vector sequence of local elements is extracted, the correlation between the request vector sequence and the set of key vector sequences is calculated, and the weight coefficient is obtained, which is used to generate a network representation of the output of the neural network to elements.

Benefits of technology

It strengthens the capture of local information, improves the capture ability of neural networks to local contexts, and thus improves the quality of network representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110163339B_ABST
    Figure CN110163339B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for generating and encoding network representations in a neural network, an encoder, and a machine device. The method includes: obtaining a spatial representation corresponding to an element in an input sequence; generating a key-value pair vector sequence and a request vector sequence mapped to the key vector sequence in the key-value pair vector sequence by encoding the spatial representation; extracting a key vector sequence of local elements relative to the element to obtain a set of key vector sequences centered on the element; obtaining a weight coefficient of the set of value vector sequences corresponding to the set of key vector sequences by calculating the correlation between the request vector sequence and the set of key vector sequences; and generating a network representation of the output of the element through the weight coefficient and the set of value vector sequences corresponding to the set of key vector sequences. Since the weight coefficient is no longer a scattered weight distribution and is obtained based on the correlation of the request vector sequence of the element, local information is strengthened, and the local context capture ability of the neural network is correspondingly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer application technologies, and particularly relates to a method and apparatus for generating and encoding network representations in a neural network, an encoder, and a machine device. Background Art

[0002] With the application and development of neural networks in various fields, the attention mechanism is increasingly introduced into neural networks as a basic module to dynamically and on-demand select relevant representations of input sequences in neural networks, thereby obtaining better output quality compared to recurrent neural networks (RNNs for short).

[0003] In a neural network, each upper-layer network representation will establish a direct connection with all lower-layer network representations, that is, the output of each layer serves as the input of the next layer, and this is repeated multiple times until the network representation is encoded.

[0004] Through the attention mechanism introduced in the neural network, all network representations in each layer will be fully considered, and a weighted summation operation will be performed on each network representation. This will, to a certain extent, disperse the distribution of the obtained attention weights, and for the elements in the input sequence, the information of its adjacent elements will also be weakened. In other words, there is a limitation in that local information is weakened in the network representation generated by the neural network for encoding elements.

[0005] How to strengthen local information in the implementation of network representation generation in a neural network and improve the ability of the neural network to capture local context has become a technical problem that needs to be solved urgently. Summary of the Invention

[0006] In order to strengthen local information in the implementation of network representation generation in a neural network and improve the ability of the neural network to capture local context in related technologies, the present invention provides a method and apparatus for generating and encoding network representations in a neural network, an encoder, and a machine device

[0007] A method for generating a network representation in a neural network, the method comprising:

[0008] Obtaining a spatial representation corresponding to an element in an input sequence, where the input sequence is the source input for generating a network representation by the neural network;

[0009] Generating a key-value pair vector sequence corresponding to the element and a request vector sequence mapped to the key vector sequence in the key-value pair vector sequence through encoding of the spatial representation corresponding to the element;

[0010] Extracting a key vector sequence of local elements relative to the element to obtain a set of key vector sequences centered on the element;

[0011] By calculating the correlation between the sequence of request vectors and the set of key vector sequences, the weight coefficients of the set of value vector sequences corresponding to the set of key vector sequences are obtained;

[0012] Based on the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences, the network representation output by the neural network for the element is generated.

[0013] A method for implementing neural network encoding, the method comprising:

[0014] The processor obtains the spatial representation corresponding to the element in the input sequence, and the input sequence is the source input for the neural network to generate the network representation;

[0015] Through the encoding of the network layers stacked with each other in the neural network by the spatial representation corresponding to the element, the key-value pair vector sequence of the element in the network layer and the sequence of request vectors mapped to the key vector sequence in the key-value pair vector sequence are generated;

[0016] Extract the key vector sequences of local elements in the network layer relative to the element to obtain a set of key vector sequences centered on the element in the network layer;

[0017] The processor connects the corresponding arithmetic components to calculate the correlation between the sequence of request vectors generated by the encoding of the element in the network layer and the set of key vector sequences, and obtains the weight coefficients of the set of value vector sequences corresponding to the set of key vector sequences in the network layer;

[0018] Based on the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences, each network layer generates the network representation output by the neural network for the encoding of the element.

[0019] A method for implementing neural network encoding, the method comprising:

[0020] Obtain the spatial representation corresponding to the element in the input sequence, and the input sequence is the source input for the neural network to generate the network representation;

[0021] Through the encoding of the network layers stacked with each other in the neural network by the network representation corresponding to the element, the key-value pair vector sequence of the element in the network layer and the sequence of request vectors mapped to the key vector sequence in the key-value pair vector sequence are generated;

[0022] Extract the key vector sequences of local elements in the network layer and the surrounding network layers relative to the element to obtain a set of key vector sequences centered on the element in the network layer;

[0023] Calculate the relevance between the sequence of request vectors encoded by the element in the network layer and the set of key vector sequences, and obtain the weight coefficients of the set of key vector sequences corresponding to the set of value vector sequences in the network layer;

[0024] Generate the network representation output by the neural network for the element through each network layer based on the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences.

[0025] A network representation generation device in a neural network, the device includes:

[0026] A spatial representation acquisition module, configured to acquire the spatial representation corresponding to the element in the input sequence, where the input sequence is the source input for the neural network to generate the network representation;

[0027] An encoding module, configured to generate the key-value pair vector sequence of the element and the sequence of request vectors mapped to the key vector sequence in the key-value pair vector sequence through encoding the spatial representation corresponding to the element;

[0028] An extraction module, configured to extract the key vector sequence of local elements relative to the element, and obtain a set of key vector sequences centered on the element;

[0029] A relevance calculation module, configured to calculate the relevance between the sequence of request vectors and the set of key vector sequences, and obtain the weight coefficients of the set of key vector sequences corresponding to the set of value vector sequences;

[0030] A network representation generation module, configured to generate the network representation output by the neural network for the element through the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences.

[0031] An encoder in a neural network, the encoder includes:

[0032] An input module, configured to acquire the spatial representation corresponding to the element in the input sequence, where the input sequence is the source input for the neural network to generate the network representation;

[0033] A network layer encoding module, configured to generate the key-value pair vector sequence of the element in the network layer and the sequence of request vectors mapped to the key vector sequence in the key-value pair vector sequence through encoding the spatial representation corresponding to the element in the mutually stacked network layers in the neural network;

[0034] An element extraction module, configured to extract the key vector sequence of local elements in the network layer relative to the element, and obtain a set of key vector sequences centered on the element in the network layer;

[0035] A one-dimensional correlation calculation module, configured to calculate the correlation between the request vector sequence encoded and generated by the element in the network layer and the set of key vector sequences, and obtain the weight coefficients of the set of value vector sequences corresponding to the set of key vector sequences in the network layer;

[0036] An encoding output module, configured to generate, by each network layer, a network representation of the encoded output of the neural network for the element through the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences.

[0037] An encoder in a neural network, the encoder comprising:

[0038] An encoding input module, configured to obtain a spatial representation corresponding to an element in an input sequence, where the input sequence is a source input for the neural network to generate a network representation;

[0039] An encoding generation module, configured to generate, through encoding in network layers stacked with each other in the neural network by the network representation corresponding to the element, a key-value pair vector sequence of the element in the network layer and a request vector sequence mapped to the key vector sequence in the key-value pair vector sequence;

[0040] A two-dimensional sequence extraction module, configured to extract a set of key vector sequences centered on the element in the network layer by relatively extracting key vector sequences of local elements in the network layer and surrounding network layers with respect to the element;

[0041] A two-dimensional correlation calculation module, configured to calculate the correlation between the request vector sequence encoded and generated by the element in the network layer and the set of key vector sequences, and obtain the weight coefficients of the set of value vector sequences corresponding to the set of key vector sequences in the network layer;

[0042] An output module, configured to generate, by each network layer, a network representation of the encoded output of the neural network for the element through the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences.

[0043] A machine device, comprising:

[0044] A processor; and

[0045] A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method described above is implemented.

[0046] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0047] For a given input sequence, to generate its network representation in a neural network, first obtain the spatial representations corresponding to the elements in the input sequence. The input sequence is the source input for the neural network to generate the network representation. Then, through the encoding of the spatial representations corresponding to the elements, generate a sequence of key-value pair vectors for the element and a sequence of request vectors mapped to the sequence of key vectors in the sequence of key-value pair vectors. Extract the sequence of key vectors of local elements relative to the element to obtain a set of sequences of key vectors centered on the element. Use the set of sequences of value vectors corresponding to this set of sequences of key vectors as local information. By calculating the correlation between the sequence of request vectors and the set of sequences of key vectors, obtain the weight coefficients for the set of sequences of value vectors corresponding to the set of sequences of key vectors. The weight coefficients will be used to characterize the importance of the sequence of value vectors corresponding to the local element to the element. Furthermore, through the weight coefficients and the set of sequences of value vectors corresponding to the set of sequences of key vectors, generate the network representation output by the neural network for the element. Since the weight coefficients are no longer a scattered weight distribution and are obtained based on the correlation of the request vector sequence of the element, the local information corresponding to the element is strengthened, and the local context capture ability of the neural network is correspondingly improved, thereby improving the quality of the network representation output by the neural network.

[0048] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.

[0050] Figure 1 is a schematic hardware structure diagram of the implementation environment related to the present invention;

[0051] Figure 2 is a flowchart of a method for generating a network representation in a neural network shown according to an exemplary embodiment;

[0052] Figure 3 is according to Figure 2 a flowchart describing step 250 shown according to the corresponding embodiment;

[0053] Figure 4 is according to Figure 3 a flowchart describing step 251 shown according to the corresponding embodiment;

[0054] Figure 5 is according to Figure 2 a flowchart describing step 270 shown according to the corresponding embodiment;

[0055] Figure 6 is according to Figure 2Flowchart describing step 290 shown in the corresponding embodiment;

[0056] Figure 7 is according to Figure 6 Flowchart describing step 291 shown in the corresponding embodiment;

[0057] Figure 8 Flowchart of a method for implementing neural network encoding shown in an exemplary embodiment;

[0058] Figure 9 is according to Figure 8 Flowchart describing step 650 shown in the corresponding embodiment;

[0059] Figure 10 Flowchart of a method for implementing neural network encoding shown in an exemplary embodiment;

[0060] Figure 11 Flowchart describing step 750 shown in an exemplary embodiment;

[0061] Figure 12 Schematic diagram of the architecture of the machine translation system implemented by the present invention in an exemplary embodiment;

[0062] Figure 13 Schematic diagram of the implementation principle of encoder 1 shown in an exemplary embodiment;

[0063] Figure 14 Schematic diagram of the implementation principle of encoder 2 shown in an exemplary embodiment;

[0064] Figure 15 Schematic diagram of the translation quality of encoder 1 and encoder 2 at different phrase lengths in a test;

[0065] Figure 16 Block diagram of a network representation generation device in a neural network shown in an exemplary embodiment;

[0066] Figure 17 Block diagram of an encoder in a neural network shown in an exemplary embodiment;

[0067] Figure 18 Block diagram of an encoder in a neural network shown in another exemplary embodiment. Detailed implementation manners

[0068] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0069] Figure 1 It is a schematic diagram of the hardware structure of the implementation environment related to the present invention. In an exemplary embodiment, this implementation environment may be a machine device on which a neural network is deployed. It should be understood that, according to the implementation of the neural network deployment, this machine device may be a mobile device, a server, or even a server cluster and other devices with excellent computing capabilities.

[0070] As Figure 1 described in the implementation environment, the implementation environment related to the present invention can be implemented through a server. It should be noted that the server 100 is only an example adapted to the present invention and cannot be considered as providing any limitation to the scope of use of the present invention. The server 100 cannot be interpreted as requiring dependence on or necessarily having Figure 1 one or more components in the exemplary server 100 shown in

[0071] The hardware structure of the server 100 may vary greatly due to different configurations or performances. As Figure 1 shown, the server 100 includes: a power supply 110, an interface 130, at least one storage medium 150, and at least one central processing unit (CPU) 170.

[0072] Among them, the power supply 110 is used to provide working voltage for each hardware device on the server 100.

[0073] The interface 130 includes at least one wired or wireless network interface 131, at least one serial-parallel conversion interface 133, at least one input / output interface 135, and at least one USB interface 137, etc., for communicating with external devices.

[0074] The storage medium 150 serves as a carrier for resource storage and can be a random access storage medium, a disk, an optical disc, etc. The resources stored thereon include an operating system 151, application programs 153, data 155, etc. The storage method can be transient storage or permanent storage. Among them, the operating system 151 is used to manage and control each hardware device and application program 153 on the server 100 to enable the central processing unit 170 to perform calculations and processing on the massive data 155. It can be Windows ServerTM, MacOS XTM, UnixTM, LinuxTM, FreeBSDTM, etc. The application program 153 is a computer program that completes at least one specific task based on the operating system 151. It can include at least one module ( Figure 1 not shown in the figure), and each module can separately contain a series of operation instructions for the server 100. The data 155 can be photos, pictures, etc. stored on the disk.

[0075] The central processing unit 170 can include one or more processors and is set to communicate with the storage medium 150 through a bus for operating on and processing the massive data 155 in the storage medium 150.

[0076] As described in detail above, the server 100 applicable to the present invention will perform audio recognition by reading a series of operation instructions stored in the storage medium 150 through the central processing unit 170.

[0077] Figure 2 It is a flowchart of a method for generating a network representation in a neural network shown according to an exemplary embodiment. In an exemplary embodiment, the method for generating a network representation in the neural network, as Figure 2 shown, includes at least the following steps.

[0078] In step 210, the processor obtains the spatial representation corresponding to the elements in the input sequence, and the input sequence is the source input for the neural network to generate the network representation.

[0079] Among them, this neural network will be used to perform network representation generation tasks in many fields such as machine translation, text processing, speech recognition, and image processing. This neural network can be deployed in machines of types such as servers and terminals, as well as in machine combinations formed by servers and terminals. Therefore, the machine on which the neural network is deployed can be any one of a server and a terminal, or a combination of the two. The machine on which the neural network is deployed will, under the action of the processor, complete the generation of the network representation in the neural network through the devices inside the processor and the connected components.

[0080] The input sequence includes a number of elements. The input sequence will serve as the source target for the neural network to generate a network representation. Correspondingly, the network representation generated by the neural network is the output target of the neural network. Depending on the task performed by the neural network, the corresponding input sequences are also different. For example, in the sentence processing task of natural language processing, the input sequence is a sentence, and the elements in the input sequence are the words in the sentence. In other words, the sequence of words forming the sentence is the input sequence. The neural network will generate a corresponding network representation for each element.

[0081] The spatial representation corresponding to the elements in the input sequence is obtained by converting the discrete elements in the input sequence into a continuous spatial representation through the network layers in the neural network. In this way, the spatial representation (embedding) corresponding to the elements in the input sequence can be obtained, which is also called the source-side vector representation.

[0082] It should be understood that the obtained spatial representation corresponding to the elements is essentially a process of numericalizing the elements by mapping them to a space, so that the neural network can recognize and understand the elements in the input sequence.

[0083] In step 230, through the encoding of the spatial representation corresponding to the elements, a sequence of key-value pair vectors for the element and a sequence of request vectors mapped to the key vector sequence in the sequence of key-value pair vectors are generated.

[0084] Among them, the generated sequence of key-value pair vectors by encoding is the encoded representation of the corresponding element. That is to say, the input sequence is regarded as being composed of a series of sequences of key-value pair vectors, and each sequence of key-value pair vectors numerically describes the corresponding element. And the generated sequence of request vectors by encoding the element has a mapping relationship with the key vector sequence in the corresponding sequence of key-value pair vectors. For an element, it can be mapped from the sequence of request vectors to the key vector sequence in the sequence of key-value pair vectors. The key vector sequence in the sequence of key-value pair vectors is mapped to the value vector sequence, and there is a key-value pair correspondence between the key vector sequence and the value vector sequence.

[0085] It should be understood that for the generated sequence of request vectors, sequence of key-value pair vectors, key vector sequence, and value vector sequence by encoding the spatial representation corresponding to the elements, they are combined into one and point to the same element in the input sequence. For example, when the input sequence is an input sentence, the key vector sequence and the value vector sequence constitute the semantic encoding corresponding to the words in the input sentence. Through multiple linear mappings between the sequence of request vectors and the key vector sequence, the mapping between the sequence of request vectors and the sequence of key-value pair vectors can be obtained. The generated sequence of key-value pair vectors and sequence of request vectors for the element will be written into the internal register by the processor for subsequent quick access.

[0086] The encoding of the element corresponding spatial representation is achieved through the performed linear transformation. In an exemplary embodiment, step 230 includes: for each element, performing a linear transformation of the element corresponding spatial representation through the parameter matrix learned by the neural network to respectively encode and generate the key-value pair vector sequence of the element, and the request vector sequence mapped to the key vector sequence in the key-value pair vector sequence.

[0087] Among them, the parameter matrices used to encode and generate the key vector sequence, value vector sequence, and request vector sequence are three different learnable parameter matrices. Through the linear transformations performed by these three different learnable parameter matrices, the element corresponding spatial representation is respectively encoded to generate the key vector sequence, value vector sequence, and request vector sequence.

[0088] The key-value pair vector sequence and the request vector sequence are the encoded outputs of a network layer in the neural network. In an exemplary embodiment, this network layer is stacked with other network layers. In other words, the neural network includes multiple stacked network layers. Correspondingly, each network layer will encode the spatial representation corresponding to the element based on its configured parameter matrix to output the key-value pair vector sequence and the request vector sequence at this network layer.

[0089] In this neural network, the spatial representation corresponding to the element will be encoded through multiple stacked network layers, and subsequent processes will be executed on this basis until the network representation of the element is output. It can be understood that this neural network introduces a multi-head mechanism, which can use multiple network layers in parallel, and the multiple stacked network layers are independent of each other in the generation of the network representation of the element.

[0090] In summary, the key-value pair vector sequence of the element in the input sequence and the request vector sequence mapped to the key vector sequence in the key-value pair vector sequence will be encoded and generated in each of the multiple stacked network layers.

[0091] In step 250, the key vector sequence of the local element is extracted relative to the element to obtain a set of key vector sequences centered on the element.

[0092] Among them, through the aforementioned step 230, the key-value pair vector sequence and the request vector sequence corresponding to each element in the input sequence can be encoded and generated. Discretely distributed elements can all obtain their corresponding key-value pair vector sequences and request vector sequences.

[0093] In the input sequence, each element has other elements related to itself, and these elements are the local elements of the element. The local elements will carry the local information closely related to the element.

[0094] For local elements, their key vector sequences exist as indices (Keys). Therefore, for capturing the local information related to an element, the key vector sequence of the corresponding local element of this element will be extracted first.

[0095] In the input sequence, each element has a corresponding local element, and each local element has a key vector sequence. Therefore, as the key vector sequence of the local element of this element is extracted, a set of key vector sequences centered on this element will be obtained.

[0096] In an exemplary embodiment, the local elements of an element are a number of elements determined forward and backward in the input sequence centered on this element. The local elements correspond to the context of this element and thus carry context local information.

[0097] For a neural network introducing the multi-head mechanism, its multi-layer stacked network layers will all extract the key vector sequences of local elements for elements. And so on, each network layer will obtain a set of key vector sequences for each element in the input sequence.

[0098] It should be noted that the extraction of the key vector sequence of the local element for the current element can be one-dimensional, that is, the defined local range can be a one-dimensional element axis, or two-dimensional, that is, extended from the one-dimensional element axis to a two-dimensional region. For example, the key vector extraction of local elements in multi-layer stacked network layers is not limited here.

[0099] In step 270, the processor connects the corresponding arithmetic components to calculate the relevance between the request vector sequence and the set of key vector sequences, and obtains the weight coefficients of the set of key vector sequences corresponding to the set of value vector sequences.

[0100] Among them, in the execution stage of network representation generation, the corresponding arithmetic components connected by the processor, such as the Arithmetic Logic Unit (ALU), will perform the required calculations.

[0101] For an element, the execution of the foregoing steps obtains a request vector sequence for the element in the input sequence and a set of key vector sequences related to the corresponding local elements. The set of key vector sequences contains the key vector sequences of each local element. From this, the key vector sequence can be mapped to the value vector sequence of the local element. Therefore, correspondingly, it can also be mapped to the set of value vector sequences corresponding to it, that is, the set composed of the value vector sequences of local elements.

[0102] For each element in the input sequence, during the execution of step 270, the calculation of the correlation between the request vector sequence and the set of key vector sequences is performed. Here, the request vector sequence referred to herein is the request vector sequence of the current element, and the set of key vector sequences also belongs to the current element and is a set of key vector sequences extracted from the local elements of the current element.

[0103] Calculate the correlation between the request vector sequence of the current element and the key vector sequences of the local elements. In this calculation, the greater the correlation between the request vector sequence and the key vector sequence, the more relevant or similar the local element corresponding to the key vector sequence is to the current element, and the greater the weight coefficient of the corresponding local element value vector sequence; conversely, the smaller the weight coefficient of the local element value vector sequence.

[0104] The calculation of the correlation between the request vector sequence and the set of key vector sequences is essentially the calculation of the correlation or similarity between the request vector sequence of the current element and the key-value pair vector sequences of the local elements. In an exemplary embodiment, this similarity can be measured by logical similarity.

[0105] In a multi-layer stacked network layer, each network layer calculates the correlation between the request vector sequence and the key-value pair vector sequences of each local element corresponding to itself for the elements in the input sequence. Here, taking the calculation of logical similarity as an example, a detailed description is given.

[0106] In a network layer, that is, in the h-th head of the neural network, the spatial representation corresponding to the element in the input sequence is first linearly transformed into a request (query) vector sequence Q h 、key (key) vector sequence K h and value (value) vector sequence V h .

[0107] For the i-th element, the request vector sequence q of the i-th element relative to the input sequence i h , calculate the set of key vector sequences of the M + 1 (M <= I) elements centered on i, that is, the set of key vector sequences formed by the key vector sequences of the local elements, which can be expressed as:

[0108]

[0109] where I is the number of elements in the input sequence.

[0110] Calculate the request vector sequence of the element and the set of key vector sequences to obtain the logical similarity e between the request vector sequence and each key-value pair, that is, the key-value pair vector sequence of each local element i h :

[0111]

[0112] Among them, d is the dimension of the spatial representation corresponding to the element.

[0113] By analogy, the relevance calculation between the elements in the input sequence and each of their local elements can be completed at each network layer, and then converted into the weight coefficients between the request vector sequence and each key-value pair. The weight coefficients will simply describe the dependence relationship between two elements, that is, the current element and the corresponding local element.

[0114] In step 290, based on the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences, a network representation of the element output by the neural network is generated.

[0115] Among them, as described above, each element in the input sequence has obtained its own weight coefficient relative to the local element value vector sequence. Therefore, it is possible to fuse the value vector sequences between the local elements and the current element, and then generate a network representation of the current element output by the neural network.

[0116] Through the above-described exemplary embodiments, the implemented neural network is able to capture local information, and the network representation generated for the element also refers to the local information closely related to itself, that is, the context local information. Therefore, the quality of the network representation output for the element is improved, and the performance of the neural network is also optimized.

[0117] The neural network implemented by the above-described exemplary embodiments will be applied to many fields such as machine translation, text processing, speech recognition, and image processing, and then used to complete network representation generation tasks such as machine translation, natural language inference, speech models, and image annotation. Thanks to the deployment of the neural network implemented by the present invention, local information can be strengthened, and thus it performs excellently in the execution of various tasks.

[0118] In summary, the neural network implemented by the present invention is simple and effective and can be applied to all models that need to strengthen local information and model discrete sequences.

[0119] Figure 3 is according to Figure 2 The flowchart for describing step 250 shown in the corresponding embodiment. In an exemplary embodiment, as Figure 3 shown, this step 250 includes:

[0120] In step 251, determine the local elements corresponding to the elements in the input sequence. The local elements are other elements corresponding to a preset number centered on the element in the input sequence.

[0121] Among them, for an element in the input sequence, there is a corresponding local element, and the other elements in the input sequence surrounding this element are the local elements corresponding to this element.

[0122] In an exemplary embodiment, the number of local elements is preset, and taking the current element as the center and according to the preset number, the local elements corresponding to the current element can be determined from the input sequence.

[0123] For example, the preset number is M, taking the current element i as the center, the and elements behind in the input sequence are all local elements of the current element i.

[0124] And so on, the local elements corresponding to each element in the input sequence can be determined.

[0125] In step 253, for this element, a key vector sequence extraction of the corresponding local elements is performed for the multi-layer stacked network layers in the neural network, and a key vector sequence of each local element is obtained for this element in this network layer.

[0126] Among them, it should be noted first that in this exemplary embodiment, the neural network includes multi-layer stacked network layers. For the key vector sequence extraction of the local elements, it is performed in multiple heads, and a key vector sequence of the corresponding local elements for the elements in the input sequence is extracted for each mutually stacked network layer.

[0127] The key vector sequence extraction of the local elements for the multi-layer stacked network layers in the neural network includes the key vector sequence extraction with a one-dimensional element axis as the defined local range and the key vector sequence extraction with a two-dimensional region as the defined local range.

[0128] Whether it is with a one-dimensional element axis as the defined local range or a two-dimensional region as the defined local range, a key vector sequence generated by encoding the corresponding spatial representation can be extracted for each element according to the determined local elements.

[0129] It should be noted that the key vector sequence extraction is performed for the element and the network layer where it is located, and it is also to provide the required key vector sequence set for the subsequent correlation calculation between the element request vector sequence and the key vector sequence set. Therefore, correspondingly, the key vector sequence extraction is performed for the network layer that encodes and generates the current element request vector sequence.

[0130] For each element of the input sequence, the key vector sequences of the corresponding local elements are extracted along the one-dimensional element axis or in a two-dimensional region in each of the mutually stacked network layers, so as to obtain the set of key vector sequences of the element in this network layer.

[0131] Specifically, the extraction of the key vector sequence for local elements along the one-dimensional element axis is to extract the key vector sequence of each local element in the network layer that encodes and generates the request vector sequence for the request vector sequence of the current element, and these key vector sequences together with the key vector sequence of the current element form the set of key vector sequences.

[0132] The extraction of the key vector sequence for local elements in a defined two-dimensional region is to take the network layer that encodes and generates the request vector sequence as the current network layer for the request vector sequence of the current element, and extract the key vector sequences of local elements in the current network layer and the surrounding network layers of the current network layer. These key vector sequences together with the key vector sequences of the current element in the current network layer and the surrounding network layers form the set of key vector sequences.

[0133] In step 255, the key vector sequences of each local element and the key vector sequence of the element are sequentially formed into the set of key vector sequences of the element in this network layer.

[0134] Among them, the key vector sequences extracted through step 253 will be sequentially formed into the set of key vector sequences of the current element in this network layer according to the corresponding element or elements and the network layer where they are located.

[0135] For the key vector sequences extracted along the one-dimensional element axis, they will form the set of key vector sequences of the current element in a network layer together with the key vector sequence of the current element in the order of the corresponding elements.

[0136] For the key vector sequences extracted in the same network layer in a two-dimensional region, after arranging them in the order of the corresponding elements with the key vector sequence of the current element in this network layer, they are then spliced in the order of the corresponding network layers, so as to sequentially form the set of key vector sequences.

[0137] In an exemplary embodiment, Figure 3 The shown step 251 includes: for the elements in the input sequence, taking the key vector sequences corresponding to the local elements in the network layer as the one-dimensional element axis, extracting the key vector sequences of the local elements in the network layer, and the extracted key vector sequences are used to form the set of key vector sequences of the element in the network layer.

[0138] Among them, this process is the acquisition of the one-dimensional key vector sequence set. For each element in the input sequence, the key vector sequences of local elements will be extracted in each of the mutually stacked network layers, so as to obtain the set of key vector sequences of the element in this network layer.

[0139] And so on, a set of key vector sequences will be obtained for each element in each network layer, with no interaction between network layers and they are independent of each other.

[0140] Figure 4 It is based on Figure 3 The flowchart for describing step 251 shown in the corresponding embodiment. In an exemplary embodiment, as Figure 4 shown, step 251 includes:

[0141] In step 301, for the request vector sequence generated by encoding an element in the input sequence in a network layer, other network layers in the neural network are determined as the peripheral network layers of this element with this network layer as the center. The peripheral network layers and the local element are used to determine the two-dimensional region of this element relative to this network layer.

[0142] Among them, for the extraction of the key vector sequence of the local element, an element in the input sequence is used as the current element, and a stacked network layer is used as the current network layer. After determining the local element and the peripheral network layer respectively relative to the current element and the current network layer, a limited local range, that is, a limited two-dimensional region, is obtained.

[0143] The peripheral network layers are the network layers determined according to a set number with the current network layer as the center. For example, for the current network layer h, when the set number is N, the determined peripheral network layers are network layer to network layer Of course, the current network layer h is excluded.

[0144] For each network layer, its peripheral network layer is determined in this way, and then based on this, the two-dimensional region for extracting the key vector sequence of the local element for the encoded request vector sequence of this network layer is determined.

[0145] In step 303, in the two-dimensional region, the key vector sequences of the corresponding local elements are respectively extracted from the network layer as the center and each peripheral network layer of this element. The extracted key vector sequences are used to form the key vector sequence set of this element in this network layer with the key vector sequence of this element in the two-dimensional region.

[0146] Among them, in the limited two-dimensional region, on the one hand, the key vector sequence is extracted according to the local element, and on the other hand, the extraction of the key vector sequence for the local element is carried out separately in each network layer in the two-dimensional region, so as to extract the required key vector sequence for the current network layer and the current element.

[0147] An element i, that is, the i-th element in the input sequence, defines a two-dimensional region of size (N + 1) × (M + 1) (N <= M) centered on the element i and the network layer that encodes and generates the request vector sequence of the element i in multiple mutually stacked network layers, and then obtains a set of key vector sequences containing (N + 1) × (M + 1) key vector sequences. Here, N is the set number of surrounding network layers, and M is the set number of local elements.

[0148] At this time, the set of key vector sequences can be expressed as:

[0149]

[0150] Among them, is obtained by arranging the local element key vector sequence and the current element key vector sequence of the current network layer in order.

[0151] Correspondingly, Figure 3 Step 255 shown includes: After arranging the key vector sequences extracted from the same network layer and the key vector sequence of this element in this network layer according to the corresponding elements, perform splicing of the key vector sequences in the order of the network layers where they are located to obtain the set of key vector sequences of this element in this network layer.

[0152] Among them, as described above, the set of key vector sequences extracted from the two-dimensional region will form the set of key vector sequences of the current element in the current network layer in order together with the key vector sequences of the corresponding element, the corresponding network layer, and the current element in the two-dimensional region. The current element and the current network layer are closely related to the two-dimensional region and are the center of the two-dimensional region.

[0153] Figure 5 is based on Figure 2 The flowchart corresponding to the embodiment shows the description of step 270. In an exemplary embodiment, as Figure 5 shown, this step 270 includes:

[0154] In step 271, calculate the correlation between the request vector sequence encoded by the element in the network layer and the set of extracted key vector sequences. The correlation is characterized by the correlation degree and similarity between the request vector sequence and the key vector sequence.

[0155] Among them, is the request vector sequence encoded by the element in each network layer, and is the set of key vector sequences obtained for this element in this network layer. Therefore, the correlation between the request vector sequence and the set of key vector sequences in this network layer can be calculated for this element.

[0156] The calculation of the correlation between the request vector sequence generated by encoding an element in the network layer and the set of extracted key vector sequences is performed for this element with respect to this network layer, and is essentially the calculation of the correlation between the request vector sequence and each key vector sequence in the set of key vector sequences.

[0157] In step 273, the correlation of the non-linear transformation calculation is used to obtain the weight coefficient of the value vector sequence in the set of value vector sequences corresponding to the set of key vector sequences in this network layer.

[0158] Among them, after the correlation calculation is completed to obtain the correlation of the request vector sequence with each key-value pair, such as logical similarity, it is converted into the weight coefficient of the value vector sequence corresponding to the key-value pair.

[0159] In an exemplary embodiment, the SoftMax function is applied for non-linear transformation to convert the logical similarity into the weight relationship between the request vector sequence and each key-value pair, that is:

[0160] α i h =softmax(e i h )

[0161] Among them, α i h is the weight relationship between the request vector sequence of the i-th element in the network layer h and each key-value pair, and e i h is the logical similarity.

[0162] Figure 6 is the flowchart describing step 290 according to the corresponding embodiment shown. In an exemplary embodiment, as Figure 2 shown, this step 290 at least includes: Figure 6 shown, this step 290 at least includes:

[0163] In step 291, in the network layers stacked in multiple layers in the neural network, a set of value vector sequences corresponding to the set of key vector sequences are extracted for the elements of the input sequence.

[0164] Among them, for the set of key vector sequences of an element in a network layer, the extraction of the set of value vector sequences is correspondingly performed, and the set of value vector sequences includes the value vector sequences corresponding to the key vector sequences in the set of key vector sequences.

[0165] Therefore, the acquisition of the set of value vector sequences corresponds to the extraction of the key vector sequences in the set of key vector sequences. There is a key-value pair mapping relationship between the key vector sequence and the value vector sequence, and therefore, it can be mapped from the key vector sequences in the set of key vector sequences.

[0166] In addition, the value vector sequence can also be extracted through the one-dimensional element axis and two-dimensional region corresponding to the key vector sequence set, and then the corresponding value vector sequence set can be obtained.

[0167] In one exemplary embodiment, the key vector set corresponds to a one-dimensional element axis. Then, step 291 includes: in the network layers stacked in multiple layers in the neural network, extracting the value vector sequence of the corresponding local elements for the elements in the input sequence. The extracted value vector sequence forms a value vector sequence set corresponding to the key vector sequence set of the ground scarlet in the network layer.

[0168] Figure 7 is based on Figure 6 The flowchart describing step 291 is shown in the corresponding embodiment. In another exemplary embodiment, as Figure 7 shown, this step 291 includes:

[0169] In step 501, according to the key vector sequence set of the element in the network layer, the value vector sequence of the element and the corresponding local elements is extracted in the network layer and each surrounding network layer, and the value vector sequence encoded and generated by the element and the local elements in this network layer and each surrounding network layer is obtained.

[0170] In step 503, the value vector sequences are concatenated in the element order and the network layer order to obtain a value vector sequence set corresponding to the key vector sequence set.

[0171] In step 293, a weighted sum is performed between the weight coefficient and the value vector sequences in the value vector sequence set to obtain the output representation of the element in this network layer.

[0172] Among them, the weight coefficient calculated through the foregoing steps indicates the importance degree of the value vector sequence of the local element relative to the current element. Therefore, the value vector sequence of the local element and the value vector sequence of the current element will be fused through the weight coefficient to obtain the output representation of the current element in the current network layer.

[0173] Each element obtains the calculated weight coefficient and value vector sequence set in a network layer. Through the weighted sum between the weight coefficient and the value vector sequences in the value vector sequence set, the output representation of the element in this network layer can be obtained.

[0174] For example, for the key vector sequence set extracted from the one-dimensional element axis, that is The weighted sum calculation of the weight coefficient and the value vector sequences in the value vector sequence set is completed through the performed dot product calculation to obtain the output of the i-th element in network layer h, that is:

[0175]

[0176] Thus, the outputs of multiple network layers can be connected to obtain the network representation of the output of the neural network for the i-th element, that is:

[0177] O = [O 1 ,..., O H

[0178] For another example, for the set of key vector sequences extracted from a two-dimensional region, that is The dot product calculation is also performed on the weight coefficients and the value vector sequence to obtain the output of the i-th element in network layer h, that is:

[0179]

[0180] Connect the outputs of the element in multiple network layers to obtain the network representation of the neural network for this element:

[0181] O = [O 1 ,..., O H

[0182] In step 295, splice the output representations of the element in each network layer to generate the network representation of the output of the neural network for this element.

[0183] Through the above-described exemplary embodiments, it can be applied to a neural network model and exist as a network layer therein, and then process discrete sequences that require strengthening of local information.

[0184] As described above, the generation of the network representation realizes the extraction of the set of key vector sequences through the introduction of a two-dimensional region, which enables the information interaction of the network layers stacked in multiple layers in the neural network, that is, the network layer interacts with the surrounding network layers, so that the output calculation of the network layer is no longer limited to its own subspace, avoiding the neglect of information in different subspaces, and the performance of the neural network is further enhanced.

[0185] In the above-described exemplary embodiments, through the generation of the network representation by the neural network, an attention mechanism is introduced to dynamically select relevant representations in the network as needed, that is, select the relevant representations of local elements to generate the network representation for the current element, and through the introduced multi-head mechanism, that is, parallel application of the attention network to parallelly use and focus on different information, thus through the method described above, a neural network model with stacked multi-head self-attention is realized.

[0186] ​​The method for generating network representations in the neural network as described above will be a general idea for generating network representations and does not need to rely on a specific framework. Of course, for the Nncoder-Decoder framework applicable to machine translation, the neural network model implemented by the present invention can also be attached to this framework.

[0187] That is to say, the method for generating network representations in the neural network implemented by the present invention can be applied to the encoder in the neural network and can also be applied to the decoder in the neural network.

[0188] The following is a method embodiment of the encoder implemented by the present invention. On the other hand, based on the method for generating network representations in the aforementioned neural network, a method for implementing neural network encoding can also be provided. Figure 8 It is a flowchart of a method for implementing neural network encoding shown according to an exemplary embodiment.

[0189] In an exemplary embodiment, as Figure 8 shown, a method for implementing neural network encoding includes:

[0190] In step 610, the processor obtains the spatial representation corresponding to the element in the input sequence, and the input sequence is the source input for the neural network to generate network representations.

[0191] In step 630, through the network layers stacked with each other in the neural network by the spatial representation corresponding to the element, a key-value pair vector sequence of the element in this network layer and a request vector sequence mapped to the key vector sequence in the key-value pair vector sequence are generated.

[0192] In step 650, the key vector sequence of the local elements in this network layer is extracted relative to the element, and a set of key vector sequences centered on the element in the network layer is obtained.

[0193] In step 670, the processor connects the corresponding arithmetic components to calculate the relevance between the request vector sequence and the set of key vector sequences generated by the encoding of the element in this network layer, and obtains the weight coefficients of the set of value vector sequences corresponding to the set of key vector sequences in the network layer.

[0194] In step 690, through the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences, each network layer generates the network representation output by the neural network for encoding the element.

[0195] In this exemplary embodiment, the spatial representation corresponding to the element is encoded, and then the network representation output by encoding this element is obtained.

[0196] In natural language processing based on neural networks, especially in sentence processing tasks, the encoder plays a crucial role. Through the encoder, discrete sentences are encoded into one or more continuous high-dimensional space representations for further operations.

[0197] And through the above-described exemplary embodiments, it can become an encoder block in the encoder, and then form an encoder together with other neural network models. That is to say, through the above-described exemplary embodiments, the self-attention neural network layer in the encoder can be realized.

[0198] Exemplarily, in the encoding performed, for the network layers stacked with each other in the neural network, a key vector sequence corresponding to local elements will be extracted for each element along the one-dimensional element axis to obtain a set of key vector sequences, and then subsequent steps will be executed. That is to say, through the above-described exemplary embodiments, a one-dimensional convolutional self-attention neural network encoder can be realized.

[0199] In this encoding implementation, as the extraction of the local element key vector sequence is carried out, a scope of attention can be defined for the elements, and then under the action of the weight coefficients, the local information to be concerned can be obtained according to the dependence relationship between the local elements and the current element, avoiding discreteness.

[0200] Figure 9 It is according to Figure 8 The flowchart for describing step 650 shown in the corresponding embodiment. In an exemplary embodiment, as Figure 9 shown, this step 650 includes:

[0201] In step 651, determine the local elements corresponding to the elements in the input sequence. The local elements are other elements corresponding to a preset number with the element as the center in the input sequence.

[0202] In step 653, for this element, extract the key vector sequence corresponding to the local elements in the network layer along the one-dimensional element axis of the key vector sequence corresponding to the local elements in the network layer.

[0203] In step 655, form a set of key vector sequences of the element in the network layer by arranging the key vector sequence of the local element and the key vector sequence of the element in the network layer in order.

[0204] This exemplary embodiment is the realization of the extraction of the key vector sequence via the one-dimensional element axis, so as to define the scope of attention for the introduced attention mechanism, that is, the one-dimensional element axis of the local elements in this network layer.

[0205] The following is a method embodiment of the encoder implemented by the present invention. On the other hand, based on the network representation generation method in the foregoing neural network, a method for realizing neural network encoding can also be provided.Figure 10 It is a flowchart of a method for implementing neural network encoding shown according to an exemplary embodiment.

[0206] In one exemplary embodiment, as Figure 10 shown, a method for implementing neural network encoding includes the following steps:

[0207] In step 710, the processor obtains the spatial representation corresponding to the elements in the input sequence, and the input sequence is the source input for the neural network to generate the network representation.

[0208] In step 730, through the network layers stacked with each other in the neural network by the network representation corresponding to the element, a key-value pair vector sequence of the element in this network layer and a request vector sequence mapped to the key vector sequence in the key-value pair vector sequence are generated.

[0209] In step 750, the key vector sequence of the local elements in the network layer and the surrounding network layers is extracted relative to the element, and a set of key vector sequences centered on the element in this network layer is obtained.

[0210] In step 770, the processor connects the corresponding arithmetic components to calculate the relevance between the request vector sequence generated by the encoding of the element in the network layer and the set of key vector sequences, and obtains the weight coefficients of the set of value vector sequences corresponding to the set of key vector sequences in the network layer.

[0211] In step 790, through the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences, each network layer generates the network representation of the neural network's encoded output for the element.

[0212] In natural language processing based on neural networks, especially in sentence processing tasks, the encoder is used to encode the source language into a high-dimensional vector sequence, and then the decoder decodes it into the corresponding target language.

[0213] And through the above-described exemplary embodiment, it forms an encoder block together with other types of neural networks, and then multiple encoder blocks are connected in series to form an encoder.

[0214] Exemplarily, in the encoding performed, for the network layers stacked with each other in the neural network, the key vector sequence corresponding to the local elements will be extracted for each element according to the defined two-dimensional region, so as to obtain the set of key vector sequences. That is to say, through the above-described exemplary embodiment, a two-dimensional convolutional self-attention neural network encoder is implemented.

[0215] Figure 11 It is a flowchart for describing step 750 shown according to an exemplary embodiment. In one exemplary embodiment, step 750, as Figure 11 shown, at least includes the following steps.

[0216] In step 751, determine the local elements corresponding to the elements in the input sequence. The local elements are the elements in the input sequence centered on the element, corresponding to a preset number of other elements.

[0217] In step 753, for the request vector sequence generated by encoding the elements in the input sequence in a network layer, determine the surrounding network layers in the neural network for this element with this network layer as the center. The surrounding network layers and the local elements are used to determine the two-dimensional region of this element relative to this network layer.

[0218] In step 755, in the two-dimensional region, extract the key vector sequences corresponding to the local elements of the network layer as the center and each surrounding network layer of the element respectively. The extracted key vector sequences are used to form the key vector sequence set of this element in this network layer with the key vector sequence of this element in the two-dimensional region.

[0219] In step 757, after arranging the key vector sequences extracted from the same network layer and the key vector sequence of this element in this network layer according to the corresponding elements, splice the key vector sequences in the order of the network layers where they are located to obtain the key vector sequence set of this element in this network layer.

[0220] This exemplary embodiment is realized by the extraction of the key vector sequence via the two-dimensional region. While defining the attention range for the introduced attention mechanism, it also realizes the interaction between network layers. In this two-dimensional region, the local range is restricted to adjacent network layers, that is, adjacent subspaces. Since multi-head parallel processing is performed, no limitation is imposed on the combination between network layers.

[0221] Taking the implementation of a machine translation system as an example, it is described in combination with the above method implementation. Through the implementation of the above method, the self-attention network layer of the encoder and the self-attention network layer of the decoder in the original machine translation system will be improved.

[0222] Figure 12 It is a schematic diagram of the architecture of the machine translation system implemented in an exemplary embodiment of the present invention. The implemented machine translation system adopts the Encoder-Decoder framework, that is, it includes two main parts: the encoder and the decoder.

[0223] As Figure 12 shown, overall, the left side is the implementation architecture of the encoder, and the right side is the implementation architecture of the decoder. In both the encoder and the decoder, the multi-headed attention mechanism (Multi-headed self-attention) is used, which is the implementation of the method of the present invention.

[0224] In a network block of the encoder, it consists of a multi-head attention sub-layer 810 and a feed-forward neural network sub-layer 830, and the entire encoder stack is built with N blocks.

[0225] Similar to the encoder, a multi-head attention layer 910 is added to a network block of the decoder. In the right decoder, a decoder block is composed of a self-attention neural network layer, a source-to-target neural network layer, and a feed-forward neural network 930, and multiple decoder blocks are connected in series to form the decoder.

[0226] Residual connections and layer normalization (Add&Norm) are used throughout the network.

[0227] In the existing implementation of the self-attention neural network, attention weights are calculated for each element of the discrete sequence. Therefore, compared with traditional sequence modeling methods, such as RNN (Recurrent Neural Network), it can capture the dependencies between elements more directly regardless of their distance, and thus can achieve significantly better translation quality than the machine translation system using RNN for modeling in the translation tasks of multiple language pairs.

[0228] Both the encoder and the decoder are based on the self-attention neural network and are improved and implemented through the network representation generation method in the above neural network of the present invention.

[0229] Here, first, an example of the implementation of the encoder in the machine translation scenario is used for illustration.

[0230] Figure 13 It is the implementation schematic diagram of the encoder 1 shown according to an exemplary embodiment. The original self-attention neural network will consider the vector representations of all elements in each network layer completely, and then perform a weighted sum operation on all elements, which will disperse the distribution of weights to a certain extent, and thus weaken the information of adjacent elements, and this information often plays a key role in many tasks.

[0231] For example, in the natural language processing task of a machine translation system, when word A corresponds to word B, the words around word B often correspond to word A. Take the sentence "Bush held a talk with Sharon" as an example. If "Bush" has a strong correlation with "held" and obtains a high weight, it is expected that the self-attention neural network can focus more on the words "a talk" around "held". In this way, the information of the phrase "held a talk" can be captured and associated with the subject "Bush", while the original self-attention neural network cannot achieve this. The discrete weights and the weakening of local information have become the main problems of the self-attention neural network.

[0232] In this exemplary embodiment, as Figure 13 shown, the locality of attention is modeled by restricting the scope of attention, i.e., the local scope as previously indicated, such that the phrase "held a talk" can be captured.

[0233] The encoder 1, also referred to as the "one-dimensional convolutional self-attention neural network encoder", is specifically implemented as follows:

[0234] 1. For a given input sentence x = {x1,..., x I}, the first layer of the one-dimensional relational self-attention neural network converts discrete words into continuous spatial representations;

[0235] 2. The output of the upper layer is used as the input of the current layer. At the h-th head, i.e., the h-th stacked network layer, the spatial representation is linearly transformed by three different learnable parameter matrices into a sequence of query vectors Q h , a sequence of key vectors K h and a sequence of value vectors V h ;

[0236] 3. For the i-th element, calculate the i-th element q h in Q i h with the set of M + 1 (M <= I) keys centered around i, i.e., the set of key vector sequences as previously indicated, which is represented as:

[0237]

[0238] 4. Take the dot product of q i h and to obtain the logical similarity e i h between the sequence of query vectors and each key-value pair, i.e.:

[0239]

[0240] Subsequently, apply the SoftMax non-linear transformation to convert the logical similarity into the weight coefficient α i h between the sequence of query vectors and the key-value pairs. This weight coefficient can be simply understood as the dependency relationship between two words, i.e.:

[0241] α i h = softmax(e i h )

[0242] 5. Weighted Average of Values: According to the weight coefficients obtained in the previous step, the output vector of the current element is obtained by weighted summation of the M+1 (M <= I) value sets centered around i, that is, the value vector sequence sets as described above. The M+1 value vector sequence sets around the i-th element in the h-th head are represented as follows:

[0243]

[0244] In actual calculation, the dot product of the weight coefficient and the value is calculated. Finally, the output of the i-th element in the h-th head is:

[0245]

[0246] The final output is represented as the concatenation of multiple heads along the last dimension, that is:

[0247] O = [O 1 ,..., O H

[0248] Figure 14 It is the schematic diagram of the implementation of the encoder 2 shown according to an exemplary embodiment. This encoder 2 can also be called a two-dimensional convolutional self-attention neural network with the attention scope extended to multiple heads. That is to say, the defined local scope will be extended from the one-dimensional element axis of the encoder to a two-dimensional area, that is, a rectangular area of element × head, so as to Figure 14 as shown, enabling each head to interact with multiple adjacent heads during attention calculation. Its specific implementation process is as follows:

[0249] 1. For the given input sequence x = {x1,..., x I}, the first layer of the two-dimensional convolutional self-attention neural network converts the words into continuous spatial representations;

[0250] 2. The output of the upper layer is used as the input of the current layer and is linearly transformed into the query vector sequence Q h , key vector sequence K h and value vector sequence V h by three different learnable parameter matrices in the h-th head;

[0251] 3. For the i-th element, calculate the i-th element q h in Q i h with the (N+1)×(M+1) (N <= H) key sets centered around i, that is, the key vector sequence sets. This key set can be represented as:

[0252]

[0253] 4. For q i ​h Perform a dot product with to obtain the logical similarity e between the request vector sequence and each key-value pair i h , that is:

[0254]

[0255] Subsequently, apply SoftMax for non-linear transformation to convert the logical similarity into the weight coefficient α between the request vector sequence and each key-value pair i h , that is:

[0256] α i h = softmax(e i h )

[0257] 5. Weighted average of values. According to the weight coefficients obtained in the previous step, the output vector of the current element is obtained by weighted summation of the (N + 1) × (M + 1) (N <= H) value sets centered on i, that is, the value vector sequence sets. The M + 1 value vector sequence sets around the i-th element in the h-th head are represented as follows:

[0258]

[0259] In actual calculation, a dot product calculation is performed on the weight coefficients and values. Finally, the output of the i-th element in the h-th head is:

[0260]

[0261] The final output is the concatenation of multiple heads along the last dimension, that is:

[0262] O = [O 1 ,..., O H

[0263] It can be seen that the implementation process described above is simple and easy to model. No additional components and parameters are required in the neural network, and the calculation speed will not be reduced. In a machine translation system, the present invention can significantly improve the translation quality and perform well in the translation of longer phrases and longer sentences.

[0264] For example, in the development set test of the WMT2017 Chinese-English machine translation task, applying the method of the present invention as described above significantly improved the translation quality. As shown in Table 1, a significant improvement is generally an increase of more than 0.5 points in BLEU. The Δ in this column refers to the absolute value of the increase. The unit of the number of parameters is million (M), and the unit of the training speed is the number of iterations per second.

[0265] ​

[0266] Table 1

[0267] Figure 15 It is a schematic diagram of the translation quality of Encoder 1 and Encoder 2 at different phrase lengths in a test. In the figure, the vertical coordinate is the BLEU difference between the encoder and the reference model. The upper broken line represents Encoder 2, and the lower broken line represents Encoder 1. The horizontal coordinate represents the phrase length.

[0268] From Figure 15 , it can be seen that the method provided by the present invention will perform excellently in the translation of larger phrases and longer sentences.

[0269] The following is an embodiment of the device of the present invention for implementing the network representation generation method embodiment in the above neural network of the present invention. For the details not disclosed in the embodiment of the device of the present invention, please refer to the embodiment of the network representation generation method in the neural network of the present invention.

[0270] Figure 16 It is a block diagram of a network representation generation device in a neural network shown according to an exemplary embodiment. In an exemplary embodiment, as Figure 16 shown, the network representation generation device in the neural network includes, but is not limited to: a spatial representation acquisition module 1010, an encoding module 1030, an extraction module 1050, a relevance calculation module 1070, and a network representation generation module 1090.

[0271] The spatial representation acquisition module 1010 is used to acquire the spatial representation corresponding to the elements in the input sequence, and the input sequence is the source input for the neural network to generate the network representation.

[0272] The encoding module is used to generate a key-value pair vector sequence of the element and a request vector sequence mapped to the key vector sequence in the key-value pair vector sequence through the encoding of the spatial representation corresponding to the element.

[0273] The extraction module 1050 is used to extract the key vector sequence of local elements relative to the element to obtain a set of key vector sequences centered on the element.

[0274] The relevance calculation module 1070 is used to obtain the weight coefficients of the value vector sequence set corresponding to the key vector sequence set by calculating the relevance between the request vector sequence and the key vector sequence set.

[0275] The network representation generation module 1090 is used to generate the network representation output by the neural network for the element through the weight coefficients and the value vector sequence set corresponding to the key vector sequence set.

[0276] Figure 17It is a block diagram of an encoder in a neural network shown according to an exemplary embodiment. In one exemplary embodiment, as Figure 17 shown, the encoder includes: an input module 1110, a network layer encoding module 1130, an element extraction module 1150, a one-dimensional correlation calculation module 1170, and an encoding output module 1190.

[0277] The input module 1110 is used to obtain the spatial representation corresponding to the elements in the input sequence, and the input sequence is the source input for the neural network to generate the network representation;

[0278] The network layer encoding module 1130 is used to generate a key-value pair vector sequence of the elements in the network layer and a request vector sequence mapped to the key vector sequence in the key-value pair vector sequence through the network layer encoding of the spatial representations corresponding to the elements stacked with each other in the neural network;

[0279] The element extraction module 1150 is used to extract the key vector sequence of the local elements in the network layer for the elements, and obtain a set of key vector sequences centered on the elements in the network layer;

[0280] The one-dimensional correlation calculation module 1170 is used to calculate the correlation between the request vector sequence encoded by the elements in the network layer and the set of key vector sequences, and obtain the weight coefficient of the set of value vector sequences corresponding to the set of key vector sequences in the network layer;

[0281] The encoding output module 1190 is used to generate the network representation encoded by the neural network for the elements through the weight coefficient and the set of value vector sequences corresponding to the set of key vector sequences in each network layer.

[0282] Figure 18 It is a block diagram of an encoder in a neural network shown according to another exemplary embodiment. In another exemplary embodiment, as Figure 18 shown, the encoder includes: an encoding input module 1210, an encoding generation module 1230, a two-dimensional sequence extraction module 1250, a two-dimensional correlation calculation module 1270, and an output module 1290.

[0283] The encoding input module 1210 is used to obtain the spatial representation corresponding to the elements in the input sequence, and the input sequence is the source input for the neural network to generate the network representation;

[0284] The encoding generation module 1230 is used to generate a key-value pair vector sequence of the elements in the network layer and a request vector sequence mapped to the key vector sequence in the key-value pair vector sequence through the network layer encoding of the network representations corresponding to the elements stacked with each other in the neural network;

[0285] A two-dimensional sequence extraction module 1250 is configured to extract key vector sequences of local elements in the network layer and the surrounding network layers relative to the element, and obtain a set of key vector sequences centered on the element in the network layer;

[0286] A two-dimensional correlation calculation module 1270 is configured to calculate the correlation between the request vector sequence encoded by the element in the network layer and the set of key vector sequences, and obtain the weight coefficients of the set of value vector sequences corresponding to the set of key vector sequences in the network layer;

[0287] An output module 1290 is configured to generate a network representation output by the neural network for encoding the element through each network layer based on the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences.

[0288] Optionally, the present invention further provides a machine device, which can be used in Figure 1 the shown implementation environment to execute Figure 2 、 Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 、 Figure 7 、 Figure 8 、 Figure 9 、 Figure 10 and Figure 11 all or part of the steps of any of the shown methods. The device includes:

[0289] A processor;

[0290] A memory for storing instructions executable by the processor;

[0291] Wherein, the processor is configured to execute the method as described above.

[0292] The specific manner in which the processor of the device in this embodiment performs operations has been described in detail in the relevant foregoing embodiments, and will not be elaborated here.

[0293] It should be understood that the present invention is not limited to the exact structure described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A coding method for a machine translation task, characterized in that, The method includes: The processor converts discrete words in the input sequence into continuous spatial representations. The input sequence serves as the source input for the neural network to generate network representations, and the input sequence is the input sentence in the machine translation sentence processing task. By encoding the continuous spatial representations converted from the discrete words, a key-value pair vector sequence of the words and a request vector sequence mapped to the key vector sequence in the key-value pair vector sequence are generated. The key vector sequence in the key-value pair vector sequence is mapped to a value vector sequence, and the key vector sequence and the value vector sequence constitute the semantic encoding corresponding to the words in the input sentence. Extract the key vector sequence of local words relative to the word to obtain a set of key vector sequences centered on the word. The local words are other words corresponding to a preset number centered on the word in the input sentence. The processor connects the corresponding computing components to calculate the relevance between the request vector sequence and the set of key vector sequences, and obtains the weight coefficients of the set of key vector sequences corresponding to the set of value vector sequences. The weight coefficients are the dependency relationships between two words. Through the weight coefficients and the set of value vector sequences corresponding to the set of key vector sequences, the network representation output by the neural network for the word is generated.

2. The method according to claim 1, characterized in that, The neural network includes multiple stacked network layers, and the key-value pair vector sequence of the word and the request vector sequence mapped to the key vector sequence in the key-value pair vector sequence are both encoded and generated in each of the multiple stacked network layers.

3. The method according to claim 1, characterized in that, The extracting the key vector sequence of local words relative to the word to obtain a set of key vector sequences centered on the word includes: Determine the local words corresponding to the word in the input sequence. Extract the key vector sequence of the corresponding local words for the word facing the multiple stacked network layers in the neural network, and obtain the key vector sequence of each local word relative to the word in the network layer. Form the set of key vector sequences of the word in the network layer by arranging the key vector sequences of each local word and the key vector sequence of the word in order.

4. The method according to claim 3, characterized in that, The extracting the key vector sequence of the corresponding local words for the word facing the multiple stacked network layers in the neural network and obtaining the key vector sequence of each local word relative to the word in the network layer includes: For the words in the input sequence, taking the key vector sequence corresponding to the local words in the network layer as the one-dimensional element axis, extract the key vector sequence of the local words in the network layer. The extracted key vector sequence is used to form the set of key vector sequences of the word in the network layer.

5. The method according to claim 3, characterized in that, The words in the input sequence are encoded and generate a key-value pair vector sequence and a request vector sequence in each network layer. The extracting the key vector sequence of the corresponding local words for the word facing the multiple stacked network layers in the neural network and obtaining the key vector sequence of each local word relative to the word in the network layer includes: For the request vector sequence encoded by a word in a network layer in the input sequence, determine other network layers in the neural network as the peripheral network layers of the word with the network layer as the center. The peripheral network layers and the local words are used to determine the two-dimensional region of the word relative to the network layer. In the two-dimensional region, key vector sequences of corresponding local words are respectively extracted for the network layer serving as the center and each peripheral network layer of the word, and the extracted key vector sequences are used to form a key vector sequence set of the word in the network layer with the key vector sequence of the word in the two-dimensional region.

6. The method according to claim 1, characterized in that, Calculating the relevance between the request vector sequence and the key vector sequence set to obtain the weight coefficients of the value vector sequence corresponding to the key vector sequence set includes: Calculating the relevance between the request vector sequence generated by encoding the word in the network layer and the extracted key vector sequence set, where the relevance is characterized by the correlation or similarity between the request vector sequence and the key vector sequence set; Obtaining the weight coefficients of the value vector sequences in the value vector sequence set corresponding to the key vector sequence set in the network layer through non-linear transformation calculation of the relevance.

7. The method according to claim 1, characterized in that, Generating the network representation output by the neural network for the word through the weight coefficients and the value vector sequence set corresponding to the key vector sequence set includes: In the network layers stacked in multiple layers in the neural network, value vector sequence sets corresponding to the key vector sequence set are extracted for the words in the input sequence; Performing weighted summation between the weight coefficients and the value vector sequences in the value vector sequence set to obtain the output representation of the word in the network layer; Concatenating the output representations of the word in each network layer to generate the network representation output by the neural network for the word.

8. The method according to claim 7, characterized in that, The key vector set corresponds to a one-dimensional element axis. In the network layers stacked in multiple layers in the neural network, extracting value vector sequence sets corresponding to the key vector sequence set for the words in the input sequence includes: In the network layers stacked in multiple layers in the neural network, value vector sequences of corresponding local words are extracted for the words in the input sequence, and the extracted value vector sequences form a value vector sequence set corresponding to the key vector sequence set of the word in the network layer.

9. The method according to claim 7, characterized in that, The key vector set corresponds to a two-dimensional region. In the network layers stacked in multiple layers in the neural network, extracting value vector sequence sets corresponding to the key vector sequence set for the words in the input sequence includes: According to the key vector sequence set of the word in the network layer, value vector sequences are correspondingly extracted for the word and its corresponding local words in the network layer and each peripheral network layer, and value vector sequences generated by encoding the word and local words in the network layer and each peripheral network layer are obtained; The value vector sequences are concatenated in the order of the words and the order of the network layers where they are located to obtain a value vector sequence set corresponding to the key vector sequence set.

10. A coding method for implementing a machine translation task, characterized in that, The method includes: Converting the discrete words in the input sequence into continuous spatial representations, where the input sequence is the source input for the neural network to generate the network representation, and the input sequence is the input sentence in the machine translation sentence processing task. By encoding the continuous space representations converted from the discrete words in the mutually stacked network layers in the neural network, a sequence of key-value pair vectors of the words in the network layer and a sequence of request vectors mapped to the sequence of key vectors in the sequence of key-value pair vectors are generated. The sequence of key vectors in the sequence of key-value pair vectors is mapped to a sequence of value vectors, and the sequence of key vectors and the sequence of value vectors constitute the semantic encoding corresponding to the words in the input statement. Extract the sequence of key vectors of local words in the network layer relative to the word to obtain a set of sequences of key vectors centered on the word in the network layer. The local words are the words in the input statement centered on the word and corresponding to a preset number of other words. Calculate the correlation between the sequence of request vectors generated by encoding the word in the network layer and the set of sequences of key vectors to obtain the weight coefficients of the set of sequences of value vectors corresponding to the set of sequences of key vectors in the network layer. The weight coefficients are the dependency relationships between two words. Through the weight coefficients and the set of sequences of value vectors corresponding to the set of sequences of key vectors, each network layer generates the network representation output by the neural network for encoding the word.

11. A coding method for a machine translation task, characterized in that, The method includes: Convert the discrete words in the input sequence into continuous space representations. The input sequence is used as the source input for the neural network to generate the network representation, and the input sequence is the input statement in the machine translation sentence processing task. By encoding the continuous space representations converted from the discrete words in the mutually stacked network layers in the neural network, a sequence of key-value pair vectors of the words in the network layer and a sequence of request vectors mapped to the sequence of key vectors in the sequence of key-value pair vectors are generated. The sequence of key vectors in the sequence of key-value pair vectors is mapped to a sequence of value vectors, and the sequence of key vectors and the sequence of value vectors constitute the semantic encoding corresponding to the words in the input statement. Extract the sequence of key vectors of local words in the network layer and the surrounding network layers relative to the word to obtain a set of sequences of key vectors centered on the word in the network layer. The local words are the words in the input statement centered on the word and corresponding to a preset number of other words. Calculate the correlation between the sequence of request vectors generated by encoding the word in the network layer and the set of sequences of key vectors to obtain the weight coefficients of the set of sequences of value vectors corresponding to the set of sequences of key vectors in the network layer. The weight coefficients are the dependency relationships between two words. Through the weight coefficients and the set of sequences of value vectors corresponding to the set of sequences of key vectors, each network layer generates the network representation output by the neural network for encoding the word.

12. A coding device for a machine translation task, characterized in that, The device includes: A space representation acquisition module for converting the discrete words in the input sequence into continuous space representations. The input sequence is used as the source input for the neural network to generate the network representation, and the input sequence is the input statement in the machine translation sentence processing task. An encoding module, configured to generate a sequence of key-value pair vectors of the word and a sequence of request vectors mapped to the sequence of key vectors in the sequence of key-value pair vectors through encoding of the continuous space representation converted from the discrete word, wherein the sequence of key vectors in the sequence of key-value pair vectors is mapped to a sequence of value vectors, and the sequence of key vectors and the sequence of value vectors constitute the semantic encoding corresponding to the words in the input statement; An extraction module, configured to extract a sequence of key vectors of local words relative to the word, to obtain a set of sequences of key vectors centered on the word, where the local words are other words corresponding to a preset number of other words centered on the word in the input statement; A relevance calculation module, configured to obtain a weight coefficient of the set of sequences of value vectors corresponding to the set of sequences of key vectors by calculating the relevance between the sequence of request vectors and the set of sequences of key vectors, where the weight coefficient is the dependency relationship between two words; A network representation generation module, configured to generate a network representation output by the neural network for the word through the weight coefficient and the set of sequences of value vectors corresponding to the set of sequences of key vectors; 13. An encoder for a machine translation task, characterized in that, The encoder includes: An input module, configured to convert discrete words in an input sequence into a continuous space representation, where the input sequence serves as the source input for the neural network to generate a network representation, and the input sequence is the input statement in a machine translation sentence processing task; A network layer encoding module, configured to generate a sequence of key-value pair vectors of the word in the network layer and a sequence of request vectors mapped to the sequence of key vectors in the sequence of key-value pair vectors through stacked network layer encoding of the continuous space representation converted from the discrete word in the neural network, wherein the sequence of key vectors in the sequence of key-value pair vectors is mapped to a sequence of value vectors, and the sequence of key vectors and the sequence of value vectors constitute the semantic encoding corresponding to the words in the input statement; An element extraction module, configured to extract a sequence of key vectors of local words in the network layer relative to the word, to obtain a set of sequences of key vectors centered on the word in the network layer, where the local words are other words corresponding to a preset number of other words centered on the word in the input statement; A one-dimensional relevance calculation module, configured to calculate the relevance between the sequence of request vectors encoded and generated by the word in the network layer and the set of sequences of key vectors, to obtain a weight coefficient of the set of sequences of value vectors corresponding to the set of sequences of key vectors in the network layer, where the weight coefficient is the dependency relationship between two words; An encoding output module, configured to generate a network representation encoded and output by the neural network for the word by each network layer through the weight coefficient and the set of sequences of value vectors corresponding to the set of sequences of key vectors; 14. An encoder for a machine translation task, characterized in that, The encoder includes: An encoding input module, configured to convert discrete words in an input sequence into a continuous space representation, where the input sequence serves as the source input for the neural network to generate a network representation, and the input sequence is the input statement in a machine translation sentence processing task; An encoding generation module, configured to generate a key-value pair vector sequence of the word in the network layer and a request vector sequence mapped to the key vector sequence in the key-value pair vector sequence by encoding the continuous space representations converted from the discrete words in the mutually stacked network layers in the neural network, where the key vector sequence in the key-value pair vector sequence is mapped to a value vector sequence, and the key vector sequence and the value vector sequence constitute the semantic encoding corresponding to the words in the input statement; A two-dimensional sequence extraction module, configured to extract the key vector sequence of local words in the network layer and the surrounding network layers relative to the word, to obtain a set of key vector sequences centered on the word in the network layer, where the local words are other words corresponding to a preset number of words centered on the word in the input statement; A two-dimensional correlation calculation module, configured to calculate the correlation between the request vector sequence generated by encoding the word in the network layer and the set of key vector sequences, to obtain a weight coefficient of the set of value vector sequences corresponding to the set of key vector sequences in the network layer, where the weight coefficient is the dependency relationship between two words; An output module, configured to generate a network representation output by the neural network for encoding the word by each network layer through the weight coefficient and the set of value vector sequences corresponding to the set of key vector sequences; 15. A machine device, characterized in that, Comprising: A processor; And A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Mongolian-Chinese neural translation method based on convolutional neural network

    CN108681539A

  • Method, apparatus, storage medium and apparatus for generating network representation of neural network

    CN109034378A