Information Processing Apparatus, Information Processing Method, and Method for Generating Learning Model

The information processing apparatus addresses the challenge of extracting knowledge from documents by using a multi-layer encoder and decoder configuration, enhancing answer accuracy and reducing computational load through efficient attention and sum-product operations across layers.

JP7696736B2Active Publication Date: 2025-06-23KIOXIA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021048635
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-23
Publication Date
2025-06-23
Estimated Expiration
2041-03-23

AI Technical Summary

Technical Problem

Existing language models struggle to efficiently extract knowledge from documents and generate accurate answers to questions, particularly when dealing with large datasets and complex queries.

Method used

The information processing apparatus employs a multi-layer encoder and decoder configuration, where the encoder generates multiple keys, values, and queries across layers based on input data, and the decoder uses these outputs to generate answers by performing attention operations and sum-product operations across layers.

Benefits of technology

This approach improves answer accuracy in inference operations by leveraging knowledge extracted from multiple layers of the encoder, reducing computational load, and enabling more efficient data processing compared to methods that rely solely on the output of the final encoder layer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696736000005
    Figure 0007696736000005
  • Figure 0007696736000006
    Figure 0007696736000006
  • Figure 0007696736000007
    Figure 0007696736000007
Patent Text Reader

Abstract

To extract knowledge from a document.SOLUTION: An information processing device includes: an encoder including a first layer and a second layer which are coupled in series; and a decoder. The encoder is configured to: generate, based on first data, a first key and a first value in the first layer, and a second key and a second value in the second layer; and generate, based on second data different from the first data, a first query in the first layer, and a second query in the second layer. The decoder is configured to generate third data which is included in the first data and is not included in the second data, based on the first key, the first value, the first query, the second key, the second value, and the second query.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments relate to an information processing apparatus, an information processing method, and a method for generating a learning model.

Background Art

[0002] As a technique for processing information such as natural language, a language model is known. A language model is constructed, for example, by deep learning that uses a neural network and takes a large number of documents as input. The language model obtained by deep learning may include knowledge contained in the large number of documents used during learning.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0004] Extract knowledge from documents.

Means for Solving the Problems

[0005] The information processing apparatus according to the embodiment includes an encoder including a first layer and a second layer connected in series, and a decoder. The encoder is configured to generate a first key and a first value in the first layer and a second key and a second value in the second layer based on first data, and generate a first query in the first layer and a second query in the second layer based on second data different from the first data. The decoder includes the first key, the first value,and the first query the operation result based on , the second key, the second value, and the second query the operation result based on, configured to generate third data that is included in the first data and not included in the second data based on the above.

Brief Description of Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Embodiments for Carrying Out the Invention

[0007] Hereinafter, embodiments will be described with reference to the drawings. In the description, components having substantially the same functions and configurations are denoted by the same reference numerals. Also, the embodiments shown below illustrate the technical idea. Various changes can be made to the embodiments.

[0008] 1. Embodiment 1.1 Configuration First, the configuration of the embodiment will be described.

[0009] 1.1.1 Information Processing Apparatus FIG. 1 is a block diagram showing an example of the hardware configuration of the information processing apparatus according to the embodiment. The information processing apparatus 1 is an apparatus that converts information such as natural language into data and processes it. The information processing apparatus 1 is, for example, a personal computer or a smartphone. The information processing apparatus 1 includes a control circuit 11, a memory 12, a storage 13, and a user interface 14.

[0010] The control circuit 11 is a circuit that controls the entire information processing apparatus 1. The control circuit 11 includes a CPU (Central Processing Unit), a ROM (Read Only Memory), and a RAM (Random Access Memory). The control circuit 11 may include a GPU (Graphics Processing Unit). The control circuit 11 executes various operations by loading a program stored in the ROM into the RAM in response to a request from an external user. The various operations include, for example, a learning operation based on an information source and an inference operation for inferring an answer to a question.

[0011] The memory 12 is the main memory of the information processing apparatus 1. The memory 12 is, for example, a DRAM (Dynamic Random Access Memory). The memory 12 temporarily stores data related to various operations executed by the control circuit 11.

[0012] The storage 13 is a storage device of the information processing apparatus 1. The storage 13 is, for example, an SSD (Solid State Drive) or an HDD (Hard Disk Drive). The SSD may include a NAND flash memory. The storage 13 stores data related to various operations executed by the control circuit 11 in a non-volatile manner.

[0013] The user interface 14 is a device that controls communication between the user and the control circuit 11. The user interface 14 includes an input device and an output device. The input device includes, for example, a touch panel, a keyboard, and operation buttons. The output device includes, for example, a display or a printer. The user interface 14 inputs a request for executing various operations from the user to the control circuit 11 via the input device. The user interface 14 provides the execution result of various operations to the user via the output device.

[0014] FIG. 2 is a block diagram showing an example of an outline of a functional configuration of an information processing apparatus according to an embodiment. As shown in FIG. 2, the information processing apparatus 1 has functions as an encoder 15 and a decoder 16. The encoder 15 and the decoder 16 are realized by the control circuit 11 executing operations based on a program using the memory 12. Thereby, the information processing apparatus 1 is configured to output an answer 23 to the input question 22 based on the information source 21. Further, the information processing apparatus 1 is configured to generate a re-question 22R as an intermediate product. The encoder 15 and the decoder 16 are realized using a neural network having a plurality of layers.

[0015] The information source 21, the question 22, the re-question 22R, and the answer 23 correspond to natural languages including one or more sentences. A sentence includes one or more words. A word includes one or more sub-words. A sub-word corresponds to a token. A token is a unit of data when treating a natural language as data.

[0016] The information source 21 includes information for deriving the answer 23 from the question 22 and the re-question 22R. The information source 21 may include unnecessary information for deriving the answer 23 from the question 22 and the re-question 22R. The question 22 and the re-question 22R are, for example, sentences including a masked part at the end of the sentence. The masked part includes one or more sub-words. The answer 23 is a sentence in which the masked part in the question 22 is replaced with one or more correct words.

[0017] The encoder 15 is a language model that converts an input natural language into a vector according to the context for each token. The encoder 15 generates a key 24 and a value 25 based on the information source 21. The encoder 15 stores the generated key 24 and value 25 in association with each other in the storage 13. The key 24 is data for identifying the value 25. The value 25 is data representing a sub-word included in the information source 21. The key 24 and the value 25 are associated one-to-one.

[0018] Further, based on the question 22 or the follow-up question 22R, the encoder 15 generates a query 26. The encoder 15 transmits the generated query 26 to the decoder 16. The query 26 is data for retrieving the key 24.

[0019] Based on the output from the encoder 15, the decoder 16 generates a new natural language. Based on the key 24 and the value 25 in the storage 13, and the query 26 from the encoder 15, the decoder 16 generates a follow-up question 22R and an answer 23. The decoder 16 transmits the follow-up question 22R to the encoder 15. The decoder 16 outputs the answer 23.

[0020] 1.1.2 Encoder Next, the configuration of the encoder 15 according to the embodiment will be described. Hereinafter, the functional configuration of the encoder 15 will be described separately for the case of processing the information source 21 and the case of processing the question 22 or the follow-up question 22R.

[0021] (Information source processing function) First, the functional configuration of the encoder 15 in the case of processing the information source 21 will be described.

[0022] FIG. 3 is a block diagram showing an example of the information source processing functional configuration of the encoder according to the embodiment. As shown in FIG. 3, the encoder 15 includes a receiving unit 15_s and N layers (the first layer 15_1,..., the nth layer 15_n,..., the Nth layer 15_N) (N is an integer of 3 or more, and n is an integer greater than 1 and less than N). The receiving unit 15_s and the N layers 15_1 to 15_N are connected in series. The key 24 includes keys 24_1,..., 24_n,..., 24_N. The value 25 includes values 25_1,..., 25_n,..., 25_N.

[0023] When receiving the information source 21, the receiving unit 15_s generates data 21_0 based on the information source 21. When the number of tokens of the information source 21 is L D , the data 21_0 is a multi-dimensional array in which L D d-dimensional vectors are arranged (L Dand d are natural numbers). The receiving unit 15_s transmits the data 21_0 to the first layer 15_1 of the encoder 15. In the following description, the size of the data 21_0 may be represented as [L D , d].

[0024] Based on the data 21_0, the first layer 15_1 generates data 21_1. The data 21_1 has a size of [L D , d]. Also, the first layer 15_1 generates a key 24_1 and a value 25_1 as intermediate products. Each of the key 24_1 and the value 25_1 has a size of [L D , d]. The first layer 15_1 outputs the data 21_1, the key 24_1, and the value 25_1.

[0025] The nth layer 15_n of the encoder 15 generates data 21_n based on the data 21_(n - 1). Each of the data 21_(n - 1) and the data 21_n has a size of [L D , d]. Also, the nth layer 15_n generates a key 24_n and a value 25_n as intermediate products. Each of the key 24_n and the value 25_n has a size of [L D , d]. The nth layer 15_n outputs the data 21_n, the key 24_n, and the value 25_n. The description regarding the nth layer 15_n holds for all (N - 2) layers connected in series between the first layer 15_1 and the Nth layer 15_N of the encoder 15.

[0026] The Nth layer 15_N generates data 21_N based on the data 21_(N - 1). Each of the data 21_(N - 1) and the data 21_N has a size of [L D , d]. Also, the Nth layer 15_N generates a key 24_N and a value 25_N as intermediate products. Each of the key 24_N and the value 25_N has a size of [L D , d]. The Nth layer 15_N outputs the data 21_N, the key 24_N, and the value 25_N.

[0027] With the above configuration, each of the N layers 15_1 to 15_N in the encoder 15 generates N pairs of keys and values 24_1 and 25_1 to 24_N and 25_N based on the information source 21.

[0028] Each of the N layers 15_1 to 15_N in the encoder 15 has an equivalent configuration. Hereinafter, the configuration of the n-th layer 15_n will be described on behalf of the N layers 15_1 to 15_N. The descriptions of the configurations of the other (N - 1) layers 15_1 to 15_(n - 1) and 15_(n + 1) to 15_N will be omitted.

[0029] FIG. 4 is a block diagram showing an example of the configuration of the information source processing function of the n-th layer of the encoder according to the embodiment. As shown in FIG. 4, the n-th layer 15_n of the encoder 15 includes a self-attention sub-layer SA_n and a neural network sub-layer NL1_n. The self-attention sub-layer SA_n includes a query conversion unit 30_n, a key conversion unit 31_n, a value conversion unit 32_n, a similarity calculation unit 33_n, a weighted sum calculation unit 34_n, a residual connection unit 35_n, and a normalization unit 36_n. The neural network sub-layer NL1_n includes a feed-forward network 37_n, a residual connection unit 38_n, and a normalization unit 39_n.

[0030] The query conversion unit 30_n generates a query q Dn based on the data 21_(n - 1). The query q Dn has a size [L D , d]. The query conversion unit 30_n transmits the query q Dn to the similarity calculation unit 33_n.

[0031] The key conversion unit 31_n generates a key k Dn based on the data 21_(n - 1). The key k Dn has a size [L D , d]. The key k Dn is equal to the key 24_n. The key conversion unit 31_n transmits the key k Dn to the similarity calculation unit 33_n and the storage 13.

[0032] Value conversion unit 32_n generates value v based on data 21_(n - 1). Dn Value v Dn has a size of [L D , d]. Value v Dn is equal to value 25_n. Value conversion unit 32_n sends value v Dn to weighted sum calculation unit 34_n and storage 13. Storage 13 stores key k Dn and value v Dn in association with each other.

[0033] Similarity calculation unit 33_n performs a similarity operation based on query q Dn and key k Dn . The similarity operation is an operation for calculating an attention weight. The similarity operation is, for example, dot-product attention. The calculated attention weight is sent to weighted sum calculation unit 34_n.

[0034] Weighted sum calculation unit 34_n performs a weighted sum operation based on value v Dn and the attention weight. By the weighted sum operation, elements of the part corresponding to key k Dn similar to query q Dn in value v Dn are extracted. The output from weighted sum calculation unit 34_n is sent to residual connection unit 35_n.

[0035] Note that the similarity operation and the weighted sum operation are also referred to as attention operations. The attention operation in the n-th layer 15_n when processing information source 21 is expressed as the following formula (1).

[0036] [Number]

[0037] The n-th layer 15_n includes query q Dn , key k Dn , and value vDn They are generated from the same information source 21. Therefore, when processing the information source 21, the attention operation in the n-th layer 15_n is self-attention based on the information source 21 and not based on the question 22.

[0038] The residual connection part 35_n executes a residual connection by adding the data 21_(n - 1) to the output from the weighted sum calculation part 34_n. A residual connection is a process of converting the output from a target component (e.g., Attention(q Dn , k Dn , v Dn )) into a desired output based on the input to the target component (e.g., the data 21_(n - 1)). A residual connection is executed when the target component is configured to output the residual of the desired output with respect to the input to the target component.

[0039] The normalization part 36_n performs layer normalization on the output from the residual connection part 35_n. The output from the normalization part 36_n becomes the output of the self-attention sublayer SA_n.

[0040] The feed-forward network 37_n performs a product-sum operation on the output from the self-attention sublayer SA_n using a weight tensor and a bias term. The weight tensor and the bias term are parameters that determine the characteristics of the n-th layer 15_n of the encoder 15. In this embodiment, the weight tensors and bias terms of all the feed-forward networks within the encoder 15 are assumed to be fixed values regardless of the learning operation, inference preparation operation, and inference operation described hereinafter.

[0041] The residual connection part 38_n executes a residual connection by adding the output from the self-attention sublayer SA_n to the output from the feed-forward network 37_n.

[0042] The normalization unit 39_n performs layer normalization on the output from the residual connection unit 38_n. The output from the normalization unit 39_n becomes the output of the neural network sublayer NL1_n. The output of the neural network sublayer NL1_n is transmitted as data 21_n to the (n + 1)-th layer 15_(n + 1) of the encoder 15.

[0043] As described above, the n-th layer 15_n of the encoder 15 generates data 21_n based on data 21_(n - 1) and transmits it to the (n + 1)-th layer 15_(n + 1) of the encoder 15.

[0044] (Question processing function) Next, the functional configuration of the encoder 15 when processing the question 22 and the follow-up question 22R will be described.

[0045] FIG. 5 is a block diagram showing an example of the question processing functional configuration of the encoder according to the embodiment. FIG. 5 corresponds to FIG. 3. As shown in FIG. 5, when processing the question 22 and the follow-up question 22R, similar to the case of processing the information source 21, the encoder 15 includes a receiving unit 15_s and N layers 15_1 to 15_N. Also, the query 26 includes queries 26_1,..., 26_n,..., 26_N.

[0046] When receiving the question 22 or the follow-up question 22R, the receiving unit 15_s generates data 22_0 based on the question 22 or the follow-up question 22R. When receiving the question 22, the receiving unit 15_s converts the question 22 into data 22_0 in the form of a d-dimensional vector for each token. The masked part in the question 22 is one special token <mask>It is converted. When the re-question 22R is received, the receiving unit 15_s outputs the re-question 22R as the data 22_0.

[0047] When the number of tokens of the question 22 and the re-question 22R is L Q in this case, the data 22_0 is a multi-dimensional array in which L Q d-dimensional vectors are arranged (L Q is a natural number smaller than L D ). That is, the data 22_0 generated based on the question 22 and the re-question 22R has a size of [L Q , d]. The receiving unit 15_s transmits the data 22_0 to the first layer 15_1 of the encoder 15.

[0048] The first layer 15_1 generates the data 22_1 based on the data 22_0. The data 22_1 has a size of [L Q , d]. Also, the first layer 15_1 generates the query 26_1 as an intermediate product. The query 26_1 has a size of [1, d]. The query 26_1 is a special token <mask>It is a vector of dimension d corresponding thereto. The first layer 15_1 outputs data 22_1 and query 26_1.

[0049] The n-th layer 15_n of the encoder 15 generates data 22_n based on data 22_(n - 1). Each of data 22_(n - 1) and data 22_n has a size of [L Q , d]. Also, the n-th layer 15_n generates a query 26_n as an intermediate product. The query 26_n has a size of [1, d]. The query 26_n is a special token <mask>It is a vector of dimension d corresponding thereto. The n-th layer 15_n outputs data 22_n and query 26_n. The description of the n-th layer 15_n of the encoder holds for all (N - 2) layers connected in series between the first layer 15_1 and the N-th layer 15_N of the encoder 15.

[0050] The N-th layer 15_N generates data 22_N based on data 22_(N - 1). Each of data 22_(N - 1) and data 22_N has a size of [L Q , d]. Also, the N-th layer 15_N generates query 26_N as an intermediate product. Query 26_N has a size of [1, d]. Query 26_N is a special token <mask>It is a vector of dimension d corresponding thereto. The Nth layer 15_N outputs data 22_N and query 26_N.

[0051] Configured as described above, each of the N layers 15_1 to 15_N in the encoder 15 generates N queries 26_1 to 26_N based on the question 22 and the re-question 22R.

[0052] FIG. 6 is a block diagram showing an example of the configuration of the question processing function of the nth layer of the encoder according to the embodiment. FIG. 6 corresponds to FIG. 4. In FIG. 6, similar to FIG. 4, the configuration of the nth layer 15_n will be described representing the N layers 15_1 to 15_N.

[0053] The query conversion unit 30_n generates a query q of size [L Q , d] based on the data 22_(n - 1). The query conversion unit 30_n transmits the query q Qn to the similarity calculation unit 33_n. Also, the query conversion unit 30_n of the query q Qn among the special tokens Qn <mask>The query q for the corresponding part Mn (= Query 26_n) is sent to the decoder 16.

[0054] Based on the data 22_(n - 1), the key conversion unit 31_n generates a key k of size [L Q , d]. The key conversion unit 31_n sends the key k Qn to the similarity calculation unit 33_n. Qn

[0055] Based on the data 22_(n - 1), the value conversion unit 32_n generates a value v of size [L Q , d]. The value conversion unit 32_n sends the value v Qn to the weighted sum calculation unit 34_n. Qn

[0056] The similarity calculation unit 33_n executes a similarity operation based on the query q Qn and the key k Qn . The attention weights calculated by the similarity operation are sent to the weighted sum calculation unit 34_n.

[0057] The weighted sum calculation unit 34_n executes a weighted sum operation based on the value v Qn and the attention weights received from the similarity calculation unit 33_n. By the weighted sum operation, among the values v Qn , the elements corresponding to the key k Qn similar to the query q Qn are extracted. The output from the weighted sum calculation unit 34_n is sent to the residual connection unit 35_n.

[0058] Note that the attention operation in the n-th layer 15_n of the encoder 15 when processing the question 22 and the re-question 22R is expressed as the following formula (2).

[0059]

Equation

[0060] The n-th layer 15_n receives the query q Qn , the key k Qn and value v Qn are generated from the same question 22 or re-question 22R. Therefore, when processing the question 22 and the re-question 22R, the attention operation in the n-th layer 15_n is self-attention based on the question 22 and the re-question 22R and not based on the information source 21.

[0061] The residual connection part 35_n executes a residual connection by adding the output from the weighted sum calculation part 34_n to the data 22_(n - 1).

[0062] The normalization part 36_n performs layer normalization on the output from the residual connection part 35_n. The output from the normalization part 36_n becomes the output of the self-attention sublayer SA_n.

[0063] The functional configuration of the neural network sublayer NL1_n is the same as when processing the information source 21. That is, the weight tensor and bias term of the feed-forward network 37_n are equal to those when processing the information source 21.

[0064] As described above, the n-th layer 15_n of the encoder 15 generates the data 22_n based on the data 22_(n - 1) and transmits it to the (n + 1)-th layer 15_(n + 1) of the encoder 15.

[0065] 1.1.3 Decoder Next, the configuration of the decoder 16 according to the embodiment will be described.

[0066] FIG. 7 is a block diagram showing an example of the functional configuration of the decoder according to the embodiment. As shown in FIG. 7, the decoder 16 includes N layers (the first layer 16_1,..., the n-th layer 16_n,..., the N-th layer 16_N), and a determination unit 16_e. The N layers 16_1 to 16_N and the determination unit 16_e are connected in series.

[0067] The first layer 16_1 of the decoder 16 generates data 23_1 based on the key 24_1, the value 25_1, and the query 26_1. The data 23_1 has a size of [1, d]. The data 23_1 is a vector of dimension d corresponding to one token. The first layer 16_1 transmits the generated data 23_1 to the second layer 16_2 of the decoder 16.

[0068] When the nth layer 16_n of the decoder 16 receives the data 23_(n - 1) from the (n - 1)th layer 16_(n - 1) of the decoder 16, it generates the data 23_n based on the data 23_(n - 1), the key 24_n, the value 25_n, and the query 26_n. Each of the data 23_(n - 1) and the data 23_n has a size of [1, d]. The data 23_n is a vector of dimension d corresponding to one token. The first layer 16_n transmits the generated data 23_n to the (n + 1)th layer 16_(n + 1) of the decoder 16. The description regarding the nth layer 16_n holds for all (N - 2) layers connected in series between the first layer 16_1 and the Nth layer 16_N of the decoder 16.

[0069] The Nth layer 16_N generates the data 23_N based on the data 23_(N - 1), the key 24_N, the value 25_N, and the query 26_N. Each of the data 23_(N - 1) and the data 23_N has a size of [1, d]. The data 23_N is a vector of dimension d corresponding to one token. The Nth layer 16_N transmits the generated data 23_N to the determination unit 16_e.

[0070] Based on the data 23_N, the determination unit 16_e determines whether the process for generating the answer 23 has ended. If it is determined that the process for generating the answer 23 has not ended, the determination unit 16_e generates a re - question 22R. If it is determined that the process for generating the answer 23 has ended, the determination unit 16_e generates the answer 23. The determination process of the determination unit 16_e will be described later.

[0071] With the above configuration, each of the N layers 16_1 to 16_N in the decoder 16 generates data 23_1 to 23_N based on at least the sets of key 24_1, value 25_1, and query 26_1 to key 24_N, value 25_N, and query 26_N, respectively.

[0072] Each of the N layers 16_1 to 16_N in the decoder 16 has an equivalent configuration. Hereinafter, the configuration of the n-th layer 16_n will be described on behalf of the N layers 16_1 to 16_N. The descriptions of the configurations of the other (N - 1) layers 16_1 to 16_(n - 1) and 16_(n + 1) to 16_N are omitted.

[0073] FIG. 8 is a block diagram showing an example of the functional configuration of the n-th layer of the decoder according to the embodiment. As shown in FIG. 8, the n-th layer 16_n of the decoder 16 includes a source-target attention sublayer STA_n and a neural network sublayer NL2_n. The source-target attention sublayer STA_n includes a residual connection unit 40_n, a similarity calculation unit 41_n, a weighted sum calculation unit 42_n, a residual connection unit 43_n, and a normalization unit 44_n. The neural network sublayer NL2_n includes a feed-forward network 45_n, a residual connection unit 46_n, and a normalization unit 47_n.

[0074] The residual connection unit 40_n adds the data 23_(n - 1), which is the output from the (n - 1)-th layer 16_(n - 1) of the decoder 16, to the query q Mn (= query 26_n) to obtain a query q'. Mn The data 23_(n - 1) means the hidden state transmitted from the (n - 1)-th layer 16_(n - 1). Note that the residual connection unit 40_1 of the first layer 16_1 of the decoder 16 does not need to add any data to the query q M1 (= query 26_1).

[0075] The similarity calculation unit 41_n calculates the similarity between the query q' Mn , and the key k Dn Based on [=Key 24_n], a similarity operation is performed. The similarity operation in the similarity calculation unit 41_n is an inner product attention, similar to the similarity operation in the similarity calculation unit 33_n. The attention weights calculated by the similarity operation of the similarity calculation unit 33_n are sent to the weighted sum calculation unit 42_n.

[0076] The weighted sum calculation unit 42_n performs a weighted sum operation based on the value v Dn (=Value 25_n), and the attention weights received from the similarity calculation unit 41_n. Through the weighted sum operation, the elements corresponding to the part of the key k Dn similar to the query q’ Mn among the values v Dn are extracted. The output from the weighted sum calculation unit 42_n is sent to the residual connection unit 43_n.

[0077] Note that the attention operation in the n-th layer 16_n of the decoder 16 is represented by the following formula (3).

[0078]

Equation

[0079] Here, the key k Dn , and v Dn are generated based on the information source 21. The query q’ Mn is generated based on the question 22 or the re-question 22R. Therefore, the attention operation in the n-th layer 16_n is source-target attention.

[0080] The residual connection unit 43_n performs a residual connection by adding the data 23_(n - 1) to the output from the weighted sum calculation unit 42_n.

[0081] The normalization unit 44_n performs layer normalization on the output from the residual connection unit 43_n. The output from the normalization unit 44_n becomes the output of the source-target attention sublayer STA_n.

[0082] The feed-forward network 45_n performs a sum-of-products operation on the output from the source-target attention sub-layer STA_n using a weight tensor and a bias term. The weight tensor and the bias term are parameters that determine the characteristics of the n-th layer 16_n. In this embodiment, the parameters of all the feed-forward networks within the decoder 16 are determined by the learning operation described hereinafter. Hereinafter, the parameters of all the feed-forward networks within the decoder 16 are collectively referred to as the learning model as well.

[0083] The feed-forward network 45_n has, for example, one hidden layer. Let the data output from the source-target attention sub-layer STA_n be x n , the weight tensors be W A and W B , and the bias terms be b A and b B . Then, the output FFN(x n ) from the feed-forward network 45_n is expressed as in the following equation (4).

[0084]

Equation

[0085] The residual connection part 46_n performs a residual connection by adding the output FFN(x n ) from the feed-forward network 45_n to the output x n from the source-target attention sub-layer STA_n.

[0086] The normalization part 47_n performs layer normalization on the output from the residual connection part 46_n. The output from the normalization part 47_n becomes the output of the neural network sub-layer NL2_n. The output of the neural network sub-layer NL2_n is transmitted as data 23_n to the (n + 1)-th layer 16_(n + 1) of the decoder 16.

[0087] As described above, the n-th layer 16_n of the decoder 16 generates data 23_n based on data 23_(n-1) and transmits it to the (n+1)-th layer 16_(n+1) of the decoder 16.

[0088] 1.2 Operations The operations of the embodiment will be described.

[0089] 1.2.1 Inference Preparation Operations First, the inference preparation operations in the information processing apparatus 1 according to the embodiment will be described.

[0090] The inference preparation operations are operations for storing the key 24 and the value 25 in the storage 13. The inference preparation operations are executed before the inference operations.

[0091] FIG. 9 is a flowchart showing an example of the inference preparation operations in the information processing apparatus according to the embodiment.

[0092] As shown in FIG. 9, when the information source 21 is input (start), the encoder 15 encodes the information source 21 and generates N keys 24_1 to 24_N and N values 25_1 to 25_N (S101).

[0093] The encoder 15 stores the generated N keys 24_1 to 24_N and N values 25_1 to 25_N in the storage 13 (S102).

[0094] When the process of S102 is completed, the inference preparation operations end (end).

[0095] 1.2.2 Inference Operations Next, the inference operations in the information processing apparatus 1 according to the embodiment will be described.

[0096] FIG. 10 is a flowchart showing an example of the inference operations in the information processing apparatus according to the embodiment.

[0097] As shown in FIG. 10, when question 22 is input (start), decoder 16 loads the N keys 24_1 to 24_N and the N values 25_1 to 25_N stored in storage 13 in the inference preparation operation (S111).

[0098] Encoder 15 encodes question 22 to generate N queries 26_1 to 26_N (S112). Encoder 15 transmits the generated N queries 26_1 to 26_N to decoder 16.

[0099] Decoder 16 decodes the N keys 24_1 to 24_N and the N values 25_1 to 25_N loaded in the process of S111, and the N queries 26_1 to 26_N generated in the process of S112 (S113). As a decoding result, decoder 16 generates data 23_N corresponding to question 22.

[0100] The determination unit 16_e of decoder 16 determines whether the process for generating answer 23 has ended based on data 23_N (S114). Specifically, determination unit 16_e determines whether the token corresponding to data 23_N is a special token. The special token is a token indicating the end of a sentence. If the token corresponding to data 23_N is not a special token, determination unit 16_e determines that the process for generating answer 23 has not ended. If the token corresponding to data 23_N is a special token, determination unit 16_e determines that the process for generating answer 23 has ended.

[0101] If it is determined that the process for generating answer 23 has not ended (S114; no), determination unit 16_e generates a re-question 22R (S115). Specifically, determination unit 16_e is the special token in question 22 or re-question 22R used for the generation of data 23_N <mask>Immediately before [[ID=]], a new follow-up question 22R is generated by inserting a token corresponding to data 23_N. The determination unit 16_e transmits the generated follow-up question 22R to the reception unit 15_s of the encoder 15. As a result, the encoding of the follow-up question 22R generated in the process of S115 is started.

[0102] The encoder 15 encodes the follow-up question 22R generated in the process of S115 and generates N queries 26_1 to 26_N (S116).

[0103] After the process of S116, the decoder 16 decodes the N keys 24_1 to 24_N and the N values 25_1 to 25_N loaded in the process of S111, and the N queries 26_1 to 26_N generated in the process of S116 (S113). As a decoding result, the decoder 16 generates data 23_N corresponding to the follow-up question 22R. By such an operation, the data 23_N is updated until it is determined that the process for generating the answer 23 in the process of S114 is completed.

[0104] When it is determined that the process for generating the answer 23 is completed (S114; yes), the determination unit 16_e generates the answer 23. As a result, the inference operation ends (ends).

[0105] FIG. 11 is a diagram showing an example of determination processing in the information processing apparatus according to the embodiment. In FIG. 11, as the question 22, "Bernhard Fries was born in <mask>A specific example of the loop of the determination process until it is determined that the process for generating Answer 23 has ended when "」" is input is shown. In this case, the Answer 23 to be generated shall be "Bernhard Fries was born in Heidelberg.". Here, it is assumed that the word "Heidelberg" is composed of three sub-words (tokens) "He", "idel", and "berg".

[0106] As shown in FIG. 11, in the first loop, the decoder 16 generates "He" as the token corresponding to the data 23_N. The determination unit 16_e determines that the decoding result of the decoder 16 is not a special token. Therefore, the inference operation proceeds to the second loop.

[0107] In the second loop, the determination unit 16_e uses, as the re-question 22R, "Bernhard Fries was born in He" <mask>Generate. The encoder 15 is "Bernhard Fries was born in He <mask>Encode 」. Accordingly, the decoder 16 generates "idel" as the token corresponding to the data 23_N. The determination unit 16_e determines that the decoding result of the decoder 16 is not a special token. Therefore, the inference operation proceeds to the third loop.

[0108] In the third loop, the determination unit 16_e uses, as the re-question 22R, "Bernhard Fries was born in Heidel" <mask>Generate "". The encoder 15 is "Bernhard Fries was born in Heidel <mask>Encode "」". Accordingly, the decoder 16 generates "berg" as the token corresponding to the data 23_N. The determination unit 16_e determines that the decoding result of the decoder 16 is not a special token. Therefore, the inference operation proceeds to the fourth loop.

[0109] In the fourth loop, the determination unit 16_e sets, as the re-question 22R, "Bernhard Fries was born in Heidelberg" <mask>Generate ". The encoder 15 is "Bernhard Fries was born in Heidelberg <mask>Encode 「」. Accordingly, the decoder 16 generates ". (period)" as the token corresponding to the data 23_N. The determination unit 16_e determines that the decoding result of the decoder 16 is not a special token. Therefore, the inference operation proceeds to the fifth loop.

[0110] In the fifth loop, the determination unit 16_e sets the re-question 22R to "Bernhard Fries was born in Heidelberg. <mask>Generate "". The encoder 15 is "Bernhard Fries was born in Heidelberg. <mask>Encode 「」. Accordingly, the decoder 16 generates a special token as the token corresponding to the data 23_N. The determination unit 16_e determines that the decoding result of the decoder 16 is a special token. Therefore, the inference operation ends in the fifth loop. As a result, the determination unit 16_e can generate 「Bernhard Fries was born in Heidelberg.」 as the answer 23.

[0111] 1.2.3 Learning operation Next, the learning operation in the information processing apparatus 1 according to the embodiment will be described.

[0112] The learning operation is an operation of generating a learning model by determining the parameters in the decoder 16. The learning operation is executed before the inference preparation operation and the inference operation. In the learning operation, a set of the information source D, the question Q, and the correct label L is used as teacher data (D, Q, L). By performing the learning operation on a large amount of teacher data (D, Q, L), a learning model with high answer accuracy can be generated.

[0113] The correct label L is the subword that the decoder 16 should answer. That is, the correct label L corresponds to one token. The question Q is that the token corresponding to the correct label L is a special token <mask>It is a masked sentence. In question Q, the special token <mask>is located at the end of the sentence. The information source D includes at least two sentences: a sentence containing information for deriving the correct label L from the question Q, and a sentence containing information unnecessary for deriving the correct label L from the question Q.

[0114] In the following description, the case where the learning operation is executed by the information processing apparatus 1 will be described, but it is not limited thereto. That is, the learning operation may be executed in a hardware configuration that functions as the encoder 15 and the decoder 16, and does not necessarily have to be executed in the same hardware configuration as the information processing apparatus 1. When the learning operation is executed in a hardware configuration different from that of the information processing apparatus 1, the configuration corresponding to the control circuit 11 may have a processor (for example, TPU: Tensor Processing Unit) that can execute calculations faster than the control circuit 11. When the learning operation is executed in a hardware configuration different from that of FIG. 1, the learning model generated by the learning operation is appropriately stored in the memory 12, the storage 13, etc. within the information processing apparatus 1.

[0115] (Flowchart) FIG. 12 is a flowchart showing an example of a learning operation in the information processing apparatus according to the embodiment. In FIG. 12, an example of a learning operation using a set of teacher data (D, Q, L) is shown.

[0116] As shown in FIG. 12, when the teacher data (D, Q, L) is input (start), the control circuit 11 initializes the loop count i to, for example, 1 (S201). The loop count i is an integer greater than or equal to 1 and less than or equal to a specified value imax. The specified value imax is the number of loops executed for a set of teacher data (D, Q, L).

[0117] The control circuit 11 determines whether to execute data augmentation processing (S202). The data augmentation processing is a method for pseudo-increasing the number of teacher data when the number of teacher data is small. The control circuit 11 can probabilistically determine whether to execute the data augmentation processing. For example, the control circuit 11 may determine to execute the data augmentation processing with a probability of 50% among the specified value imax loops.

[0118] When it is determined that the data augmentation process is to be executed (S202; yes), the control circuit 11 executes the data augmentation process (S203). As a result, in the process of the loop number i, instead of the teacher data (D, Q, L), the pseudo-expanded teacher data (D’, Q, L’) is used. Details of the data augmentation process will be described later. When it is determined that the data augmentation process is not to be executed (S202; no), in the process of the loop number i, the process of S203 is omitted.

[0119] The encoder 15 encodes the information source D or D’ to generate N keys k D1 ~k DN and N values v D1 ~v DN (S204).

[0120] The encoder 15 encodes the question Q to generate N queries q M1 ~q MN (S205).

[0121] The decoder 16 generates an answer A based on the N keys k D1 ~k DN , N values v D1 ~v DN , and N queries q M1 ~q MN generated in the processes of S204 and S205 (S206). The answer A is one token corresponding to the correct label L. Note that during the learning operation, the determination unit 16_e generates the answer A without determining whether the process for generating the answer A has been completed. That is, the determination unit 16_e does not generate the re-question 22R.

[0122] The control circuit 11 calculates a loss function based on the answer A and the correct label L generated in the process of S206 (S207). For example, cross-entropy loss is used as the loss function.

[0123] Based on the loss function calculated in the process of S207, the control circuit 11 updates at least one parameter of the feed-forward network in the decoder 16 (S208). For example, the backpropagation method is used for parameter update.

[0124] The control circuit 11 determines whether the loop count i has reached the specified value imax (S209).

[0125] If the loop count i has not reached the specified value imax (S209; no), the control circuit 11 increments the loop count i (S210). After incrementing the loop count i, the processes of S202 to S209 are executed again. In this way, parameter update based on the teacher data (D, Q, L) or (D’, Q, L’) is repeatedly executed until the loop count i reaches the specified value imax.

[0126] If the loop count i has reached the specified value imax (S209; yes), the learning operation ends (ends).

[0127] As described above, in the learning operation, the decoder 16 does not generate the re-question 22R. Therefore, the learning operation assuming each loop in the inference operation is executed individually. Specifically, for example, "Nico Gardener was born in <mask>To answer the question "」" with "Nico Gardener was born in Riga.", the following four pieces of teacher data (1) to (4) are prepared individually. Here, it is assumed that the word "Riga" is composed of two sub-words (tokens) "R" and "iga". ·(1):(Q,L)=("Nico Gardener was born in <mask>”,"R”) ·(2):(Q,L)=("Nico Gardener was born in R <mask>”,"iga”) ·(3):(Q,L)=("Nico Gardener was born in Riga <mask>”," (period) · (4): (Q, L) = ("Nico Gardener was born in Riga. <mask>”,"”) The learning operations using these four pieces of teacher data (1) to (4) do not need to be executed continuously. Note that a common information source D can be used for the teacher data (1) to (4).

[0128] As a result, the situations corresponding to each loop in the inference operation can be learned independently. Therefore, it is possible to perform highly generalizable learning that does not depend on the preceding and succeeding loops.

[0129] (Data augmentation process) Next, the data augmentation process in the information processing apparatus 1 according to the embodiment will be described. FIG. 13 is a diagram showing an example of teacher data used in the data augmentation process in the information processing apparatus according to the embodiment.

[0130] In the example of FIG. 13, when the data augmentation process is not executed, "Nico Gardener (1908 - 1989) was a British international bridge player born in Riga Latvia (then part of Imperial Russia)." is input to the encoder 15 as the information source D. Also, "Nico Gardener was born in <mask>"] is input. The correct place name for this is "Riga".

[0131] When performing data augmentation processing on this, the same question Q as when not performing data augmentation processing is input to the encoder 15, and a different information source D' from the information source D is input. The information source D' is generated by randomly replacing the part of the place name in the information source D that matches the correct place name ("Riga") with another place name ("Heidelberg", "Lyon", "Hawaii",...). At this time, the correct label L is also replaced with the correct label L' of the replaced place name ("Heidelberg", "Lyon", "Hawaii",...).

[0132] Note that the learning operation is not for the purpose of learning facts, but for the purpose of learning a method of deriving the correct label L corresponding to the question Q from the information source D. For this reason, even if the information source D' becomes incorrect content due to word replacement in the data augmentation processing. Therefore, more teacher data can be prepared from a small dataset.

[0133] 1.3 Effects of the Present Embodiment According to the embodiment, each of the N layers 15_1 to 15_N of the encoder 15 generates a set of a key 24_1 and a value 25_1 to a set of a key 24_N and a value 25_N based on the information source 21. The N layers 15_1 to 15_N of the encoder 15 generate queries 26_1 to 26_N based on the question 22. The decoder 16 generates the data 23_N based on the keys 24_1 to 24_N, the values 25_1 to 25_N, and the queries 26_1 to 26_N. Thereby, when generating the answer 23, the decoder 16 can use the information generated by the N layers 15_1 to 15_N of the encoder 15. For this reason, the answer accuracy in the inference operation can be improved compared to a method that uses only the output of the final layer of the encoder 15 (for example, the Dual-Encoder method).

[0134] Specifically, the values of the key 24, value 25, and query 26 generated by the encoder 15 are different for each of the N layers 15_1 to 15_N. This indicates that the information contained in the key 24, value 25, and query 26 is different for each generated layer. That is, the keys 24_1 to 24_(N - 1), values 25_1 to 25_(N - 1), and queries 26_1 to 26_(N - 1) may have information not included in the key 24_N, value 25_N, and query 26_N. Here, the information input from the encoder 15 to the decoder 16 is knowledge obtained from the context of the information source 21. Specifically, for example, the knowledge includes the relationship between two place names (e.g., the relationship where the two place names are a country name and the capital name of that country, etc.). On the other hand, the decoder 16 can learn how to generate the answer 23 to the question 22 through a learning operation, but the above-described knowledge cannot be learned by the decoder 16 alone.

[0135] According to this embodiment, the decoder 16 executes an inference operation using the information from the N layers 15_1 to 15_N of the encoder 15. Thereby, the decoder 16 can generate the answer 23 while making maximum use of the knowledge collected by the encoder 15 from the information source 21. Therefore, the accuracy of the answer in the inference operation can be improved.

[0136] In addition, the encoder 15 independently executes the generation of the key 24 and value 25 and the generation of the query 26. Thereby, when generating the answer 23, the key 24 and value 25 can be loaded from the storage 13. Therefore, when generating the answer 23, the computational load required for generating the key 24 and value 25 can be omitted. Thus, the load required for extracting the knowledge from the information source 21 can be reduced.

[0137] Regarding this effect, a supplementary explanation will be given with reference to FIG. 14. FIG. 14 is a diagram showing an example of the amount of calculation required for the inference operation in the information processing apparatus according to the embodiment. In the example of FIG. 14, "Obama was born in Hawaii. He was a president of USA." is input as the information source 21, and "Obama was born in <mask>An example where "」" is input is shown. And in FIG. 14, the amount of calculation required for the inference operation is represented by the size of the area determined by the token sequences arranged in the vertical and horizontal directions of the paper surface.

[0138] The amount of calculation by the encoder 15 and the decoder 16 is dominated by the amount of calculation of source-target attention and self-attention. In the case of a method of encoding the information source and the question together in the encoder (for example, the BERT method), the amount of calculation is O((the number of tokens of the information source + the number of tokens of the question)^2). The amount of calculation O((the number of tokens of the information source + the number of tokens of the question)^2) corresponds to the area S load _comp in FIG. 14.

[0139] On the other hand, according to the present embodiment, the amount of calculation of the encoder 15 is O((the number of tokens of the information source 21)^2) + O((the number of tokens of the question 22)^2). The amount of calculation O((the number of tokens of the information source 21)^2) is the amount of calculation required for the process of S101 in FIG. 9, and corresponds to the area S load _101 in FIG. 14. The amount of calculation O((the number of tokens of the question 22)^2) is the amount of calculation required for the process of S112 in FIG. 10, and corresponds to the area S load _112 in FIG. 14. Also, the amount of calculation of the decoder 16 is O(the number of tokens of the information source). The amount of calculation O((the number of tokens of the information source 21)) is the amount of calculation required for the process of S113 in FIG. 10, and corresponds to the area S load _113 in FIG. 14.

[0140] Thus, according to the present embodiment, the amount of calculation can be reduced compared to the method of encoding the information source and the question together in the encoder. In addition, among the processes in the present embodiment, the process related to the information source 21 can be completed in advance before the inference operation. As a result, the above-mentioned amount of calculation O((the number of tokens of the information source 21)^2) can be omitted during the inference operation. That is, the amount of calculation in the inference operation can be substantially reduced to O((the number of tokens of the question 22)^2) + O(the number of tokens of the information source). Therefore, the requirement for the computing performance of the control circuit 11 can be relaxed.

[0141] 2. Variations and the like Note that the above-described embodiments can be variously modified.

[0142] 2.1 First Variation For example, in the above-described embodiment, the case where the information source 21 and the question 22 are encoded by one encoder 15 has been described, but it is not limited to this. For example, the information source 21 and the question 22 may be encoded by different encoders.

[0143] FIG. 15 is a block diagram showing an example of the functional configuration of an information processing apparatus according to the first variation. As shown in FIG. 15, the information processing apparatus 1a may include encoders 15-1 and 15-2.

[0144] The encoder 15-1 has a functional configuration equivalent to the functional configuration shown in FIGS. 3 and 4 in the embodiment. That is, based on the information source 21, the encoder 15-1 generates the key 24 and the value 25. The encoder 15-1 stores the generated key 24 and value 25 in the storage 13. The encoder 15-1 has an N-layer structure. That is, the encoder 15-1 generates N keys 24-1 to 24-N and N values 25-1 to 25-N. The dimensionality of each of the keys 24-1 to 24-N generated by the encoder 15-1 is d.

[0145] The encoder 15-2 has a functional configuration equivalent to the functional configuration shown in FIGS. 5 and 6 in the embodiment. That is, based on the question 22 or the re-question 22R, the encoder 15-2 generates the query 26. The encoder 15-2 transmits the generated query 26 to the decoder 16. The encoder 15-2 has an N-layer structure. That is, the encoder 15-2 generates N queries 26-1 to 26-N. The dimensionality of each of the queries 26-1 to 26-N generated by the encoder 15-2 is d.

[0146] Thus, the encoders 15-1 and 15-2 are each configured to generate keys 24 and queries 26 of the same dimension number d. On the other hand, the parameters set in the feed-forward network in the encoder 15-1 and the parameters set in the feed-forward network in the encoder 15-2 may be the same or different from each other. When the parameters set in the feed-forward network in the encoder 15-1 and the parameters set in the feed-forward network in the encoder 15-2 are the same, the encoders 15-1 and 15-2 generate the same keys, queries, and values based on the same input. When the parameters set in the feed-forward network in the encoder 15-1 and the parameters set in the feed-forward network in the encoder 15-2 are different from each other, the encoders 15-1 and 15-2 generate different keys, queries, and values based on the same input.

[0147] FIG. 16 is a flowchart showing an example of an inference operation in the information processing apparatus according to the first modification. FIG. 16 corresponds to FIGS. 9 and 10 in the embodiment.

[0148] As shown in FIG. 16, when the question 22 is input (start), the encoder 15-1 encodes the information source 21 to generate N keys 24_1 to 24_N and N values 25_1 to 25_N (S121). The encoder 15-1 transmits the generated N keys 24_1 to 24_N and N values 25_1 to 25_N to the decoder 16.

[0149] The encoder 15-2 encodes the question 22 to generate N queries 26_1 to 26_N (S122). The encoder 15-2 transmits the generated N queries 26_1 to 26_N to the decoder 16.

[0150] The processes of S121 and S122 can be executed in parallel.

[0151] The decoder 16 decodes the N keys 24_1 to 24_N and the N values 25_1 to 25_N generated in the process of S121, and the N queries 26_1 to 26_N generated in the process of S122 (S123). As a decoding result, the decoder 16 generates the data 23_N corresponding to the question 22.

[0152] The processes of S124 to S126 are equivalent to the processes of S114 to S116 in FIG. 10. That is, after the processes of S124 to S126, the decoder 16 decodes the N keys 24_1 to 24_N and the N values 25_1 to 25_N generated in the process of S121, and the N queries 26_1 to 26_N generated based on the re-question 22R in the process of S126 (S123). As a decoding result, the decoder 16 generates the data 23_N corresponding to the re-question 22R. Thereby, the data 23_N is updated until it is determined that the process for generating the answer 23 in the process of S124 is completed.

[0153] When it is determined that the process for generating the answer 23 is completed (S124; yes), the determination unit 16_e of the decoder 16 generates the answer 23. Thereby, the inference operation is completed (end).

[0154] According to the first modification example, the key 24 and the value 25, and the query 26 are generated by different encoders 15-1 and 15-2, respectively. Thereby, during the inference operation, the generation of the key 24 and the value 25 and the generation of the query 26 can be executed in parallel. For this reason, the generation time of the key 24 and the value 25 can be shortened without performing the inference preparation operation.

[0155] 2.2 Second Modification Example Also, for example, in the above-described embodiment, in the n-th layer 16_n of the decoder 16, the case where the data 23_(n-1) from the (n-1)-th layer 16_(n-1) of the decoder 16 is residual-connected to the query 26_n has been described, but it is not limited thereto. In the n-th layer 16_n of the decoder 16, the data 23_(n-1) may not be residual-connected to the query 26_n.

[0156] FIG. 17 is a block diagram showing an example of the functional configuration of the n-th layer of the decoder according to the second modification. FIG. 17 corresponds to FIG. 8 in the embodiment. As shown in FIG. 17, the source-target attention sublayer STAa_n included in the n-th layer 16a_n of the decoder 16a may not include the residual connection part 40_n.

[0157] That is, the similarity calculation unit 41_n executes a similarity operation based on the query q Mn (=query 26_n), and the key k Dn (=key 24_n). The attention weights calculated by the similarity operation of the similarity calculation unit 41_n are sent to the weighted sum calculation unit 42_n.

[0158] The configurations of the weighted sum calculation unit 42_n, the residual connection unit 43_n, the normalization unit 44_n, the feed-forward network 45_n, the residual connection unit 46_n, and the normalization unit 47_n are the same as those in FIG. 8, so the description thereof is omitted.

[0159] With the configuration as described above, the decoder 16a can also use the information generated in the N layers 15_1 to 15_N of the encoder 15 when generating the answer 23. Therefore, the answer accuracy of the inference operation can be improved compared to the method of using only the output of the final layer of the encoder 15. Accordingly, the same effects as those of the embodiment can be achieved.

[0160] Also, in the n-th layer 16a_n, the data 23_(n-1) is not residual-connected to the query 26_n. For this reason, the amount of calculation in the decoder 16a is reduced. Therefore, the time required for the inference operation can be shortened.

[0161] 2.3 Third Modification Example Also, for example, in the above-described embodiment, the case where data 23_(n - 1) from the (n - 1)-th layer 16_(n - 1) of the decoder 16 is residually connected to the output of the weighted sum calculation unit 42_n in the n-th layer 16_n of the decoder 16 has been described. However, the present invention is not limited to this. In the n-th layer 16_n of the decoder 16, data 23_(n - 1) may not be residually connected to the output of the weighted sum calculation unit 42_n.

[0162] FIG. 18 is a block diagram showing an example of the functional configuration of the n-th layer of the decoder according to the third modification example. FIG. 18 corresponds to FIG. 8 in the embodiment. As shown in FIG. 18, the source-target attention sublayer STAb_n included in the n-th layer 16b_n of the decoder 16b may not include the residual connection unit 43_n.

[0163] That is, the weighted sum calculation unit 42_n Dn (= value v (= value 25_n)), and based on the attention weights received from the similarity calculation unit 41_n, executes a weighted sum operation. The output from the weighted sum calculation unit 42_n is transmitted to the normalization unit 44_n.

[0164] The configurations of the residual connection unit 40_n, the similarity calculation unit 41_n, the normalization unit 44_n, the feed-forward network 45_n, the residual connection unit 46_n, and the normalization unit 47_n are the same as those in FIG. 8, and thus the description thereof is omitted.

[0165] With the configuration as described above, the decoder 16 can also use the information generated in the N layers 15_1 to 15_N of the encoder 15 when generating the answer 23. Therefore, the inference operation answer accuracy can be improved more than the method of using only the output of the final layer of the encoder 15. Therefore, the same effects as those of the embodiment can be achieved.

[0166] Further, for the n-th layer 16b_n, data 23_(n - 1) is not residual-connected to the output of the weighted sum calculation unit 42_n. Therefore, the amount of calculation in the decoder 16b is reduced. Thus, the time required for the inference operation can be shortened.

[0167] 2.4 Fourth Modification Example Also, for example, in the above-described embodiment, the case where the N layers 16_1 to 16_n of the decoder 16 are configured to use the data output from the previous layer by being connected in series has been described, but the present invention is not limited thereto. The N layers 16_1 to 16_N of the decoder 16 may be configured not to use the data output from other layers.

[0168] FIG. 19 is a block diagram showing an example of the functional configuration of a decoder according to a fourth modification example. FIG. 19 corresponds to FIG. 7 in the embodiment. As shown in FIG. 19, the decoder 16c includes N layers 16c_1 to 16c_N instead of the N layers 16_1 to 16_N. Further, the decoder 16c further includes a feed-forward network 16_f in addition to the N layers 16c_1 to 16c_N and the determination unit 16_e.

[0169] The n-th layer 16c_n of the decoder 16c generates data 23_n based on the key 24_n, the value 25_n, and the query 26_n. The n-th layer 16c_n transmits the generated data 23_n to the feed-forward network 16_f. The description regarding the n-th layer 16c_n holds for all N layers of the decoder 16c.

[0170] The feed-forward network 16_f takes as input the data 23_1 to 23_N output from the N layers 16c_1 to 16c_N, and performs a sum-of-products operation using the weight tensor and the bias term. The weight tensor and the bias term are parameters that determine the characteristics of the decoder 16c. The parameters of the feed-forward network 16_f, together with all the other N feed-forward networks 45_1 to 45_N within the decoder 16c, are determined by the above-described learning operation. The output from the feed-forward network 16_f is sent to the determination unit 16_e. That is, the determination unit 16_e processes the output from the feed-forward network 16_f as data equivalent to the data 23_N in the embodiment.

[0171] FIG. 20 is a block diagram showing an example of the functional configuration of the n-th layer of the decoder 16c according to the fourth modification. FIG. 20 corresponds to FIG. 8 according to the embodiment. As shown in FIG. 20, the source-target attention sublayer STAc_n included in the n-th layer 16c_n of the decoder 16c does not include the residual connection parts 40_n and 43_n.

[0172] That is, the similarity calculation unit 41_n Mn (= query 26_n), and the key k Dn (= key 24_n), and performs a similarity operation. The attention weights calculated by the similarity operation of the similarity calculation unit 41_n are sent to the weighted sum calculation unit 42_n.

[0173] The weighted sum calculation unit 42_n Dn (= value 25_n), and based on the attention weights received from the similarity calculation unit 41_n, performs a weighted sum operation. The output from the weighted sum calculation unit 42_n is sent to the normalization unit 44_n.

[0174] The configurations of the normalization unit 44_n, the feed-forward network 45_n, the residual connection unit 46_n, and the normalization unit 47_n are the same as those in FIG. 8, and thus the description thereof is omitted.

[0175] Even with the configuration as described above, when generating the response 23, the decoder 16 can use the information generated in the N layers 15_1 to 15_N of the encoder 15. Therefore, the accuracy of the inference operation can be improved compared to the method of using only the output of the final layer of the encoder 15. Thus, effects equivalent to those of the embodiments can be achieved.

[0176] 2.5 Others In each of the above embodiments, for example, as shown in FIGS. 4 and 6, the case where the normalization units 36_n and 39_n are provided after the similarity calculation unit 33_n, the weighted sum calculation unit 34_n, and the feed-forward network 37_n in the n-th layer 15_n of the encoder 15 has been described, but it is not limited to this. For example, the normalization units 36_n and 39_n may be provided before the similarity calculation unit 33_n, the weighted sum calculation unit 34_n, and the feed-forward network 37_n, respectively. Similarly, for example, as shown in FIG. 8, the case where the normalization units 44_n and 47_n are provided after the similarity calculation unit 41_n, the weighted sum calculation unit 42_n, and the feed-forward network 45_n in the n-th layer 16_n of the decoder 16 has been described, but it is not limited to this. For example, the normalization units 44_n and 47_n may be provided before the similarity calculation unit 41_n, the weighted sum calculation unit 42_n, and the feed-forward network 45_n, respectively.

[0177] Also, in each of the above embodiments, for example, as shown in FIG. 4, in the n-th layer 15_n of the encoder 15, the similarity calculation unit 33_n and the weighted sum calculation unit 34_n use the query q Dn , key k Dn , and value v Dn collectively in the attention operation has been described, but it is not limited to this. For example, the similarity calculation unit 33_n and the weighted sum calculation unit 34_n may use the query q Dn , key k Dn , and value v Dn separately for h heads in the attention operation (h is an integer of 2 or more). In this case, for each of the h heads, the query q Dn , key k Dn , and value v Dn each has a size of [L D , d / h]. Similarly, for example, as shown in FIG. 8, in the n-th layer 16_n of the decoder 16, the similarity calculation unit 41_n and the weighted sum calculation unit 42_n perform a query q' of dimension d Dn , key k Dn , and value v Dn are used together in the attention operation. However, it is not limited to this. For example, the similarity calculation unit 41_n and the weighted sum calculation unit 42_n may use the query q' of dimension d Mn , key k Dn , and value v Dn separately for the attention operation in h heads. In this case, for each of the h heads, the query q' Mn , key k Dn , and value v Dn have sizes of [1, d / h], [L D , d / h], and [L D , d / h], respectively. Such an attention operation is also referred to as a multi-head attention operation. In a form that includes both the attention operation and the multi-head attention operation in each of the above embodiments, the dimension d in the above formulas (1) to (3) is extended to d / H (H is an integer of 1 or more).

[0178] Also, in each of the above embodiments, for example, as shown in FIGS. 4 and 6, in the n-th layer 15_n of the encoder 15, the residual connection units 35_n and 38_n perform a residual connection by addition processing. However, it is not limited to this. For example, the residual connection units 35_n and 38_n may perform a residual connection by subtraction processing, multiplication processing, concatenation processing, and inner product processing. Similarly, for example, as shown in FIG. 8, in the n-th layer 16_n of the decoder 16, the residual connection units 43_n and 46_n perform a residual connection by addition processing. However, it is not limited to this. For example, the residual connection units 43_n and 46_n may perform a residual connection by subtraction processing, multiplication processing, concatenation processing, and inner product processing.

[0179] Also, in each of the above embodiments, the decoder 16 has been described as reading all of the keys 24 and values 25 stored in the storage 13 and performing an attention operation, but the present invention is not limited to this. For example, the decoder 16 may cooperate with the memory 12 to search for a portion of the tokens L' with high similarity (i.e., a portion of size [L', d]) among the keys 24 and values 25 of size [L, d]. The decoder 16 may read the keys 24 and values 25 of size [L', d] extracted by the search from the storage 13 and perform an attention operation. Thereby, the computational complexity of the attention operation by the decoder 16 can be further reduced. D , d] of the keys 24 and values 25, among which the number of tokens L' with high similarity D ' (i.e., a portion of size [L D ', d]) may be searched. D ', d] of the keys 24 and values 25 and perform an attention operation.

[0180] Also, in each of the above embodiments, the case where the encoder 15 and the decoder 16 have a configuration of three or more layers has been described, but the present invention is not limited to this. For example, the encoder 15 and the decoder 16 may have a two-layer configuration.

[0181] Also, in each of the above embodiments, the case where the question 22 with the end masked is input to the encoder 15 has been described, but the present invention is not limited to this. For example, the question 22 with the beginning or the middle masked may be input to the encoder 15.

[0182] Also, in each of the above embodiments, the case where the information processing apparatus 1 performs question answering as an inference operation has been described, but the present invention is not limited to this. For example, the information processing apparatus 1 may perform reading comprehension as an inference operation.

[0183] Also, in each of the above embodiments, the case where the information processing apparatus 1 converts natural language into data in an inference operation has been described, but the present invention is not limited to this. For example, the information processing apparatus 1 may convert information other than natural language, such as an image, into data in an inference operation.

[0184] Note that some or all of the above embodiments may also be described as follows, but are not limited thereto.

[0185] [Appendix 1] An encoder including a first layer and a second layer connected in series, and a decoder. The encoder is configured to generate a first key and a first value in the first layer and a second key and a second value in the second layer based on first data, and generate a first query in the first layer and a second query in the second layer based on second data different from the first data. The decoder is configured to generate third data included in the first data and not included in the second data based on the first key, the first value, the first query, the second key, the second value, and the second query. An information processing apparatus.

[0186] [Appendix 2] The decoder includes a first attention layer, a first neural network layer, a second attention layer, and a second neural network layer. The first attention layer is configured to execute a first attention operation based on the first query, the first key, and the first value to generate fourth data. The first neural network layer is configured to execute a first sum-product operation based on the fourth data to generate fifth data. The second attention layer is configured to execute a second attention operation based on the second query, the second key, and the second value to generate sixth data. The second neural network layer is configured to execute a second sum-product operation based on the sixth data to generate the third data. The information processing apparatus according to Appendix 1.

[0187] [Appendix 3] Each of the first neural network layer and the second neural network layer is configured to use a feedforward network, The information processing apparatus according to Supplementary Note 2.

[0188] [Supplementary Note 4] The first attention operation and the second attention operation are source-target attention operations, The information processing apparatus according to Supplementary Note 2.

[0189] [Supplementary Note 5] The encoder includes a first encoder and a second encoder, The first encoder includes a third layer and a fourth layer connected in series, the third layer is the first layer, and the fourth layer is the second layer, The second encoder includes a fifth layer and a sixth layer connected in series, the fifth layer is the first layer, and the sixth layer is the second layer, The first encoder is configured to generate the first key and the first value in the third layer and generate the second key and the second value in the fourth layer based on the first data, The second encoder is configured to generate the first query in the fifth layer and generate the second query in the sixth layer based on the second data, The information processing apparatus according to Supplementary Note 1.

[0190] [Supplementary Note 6] The first encoder is configured to generate a third query in the third layer and generate a fourth query in the fourth layer based on the second data, The third query is the same as the first query, The fourth query is the same as the second query, The information processing apparatus according to Supplementary Note 5.

[0191] [Supplementary Note 7] The first encoder is configured to generate a third query in the third layer and a fourth query in the fourth layer based on the second data, The third query is different from the first query, The fourth query is different from the second query, The information processing apparatus according to Supplementary Note 5.

[0192] [Supplementary Note 8] Generating a first key, a first value, a second key, and a second value based on the first data, Generating a first query and a second query based on second data different from the first data, Generating third data that is included in the first data and not included in the second data based on the first key, the first value, the first query, the second key, the second value, and the second query, An information processing method comprising:

[0193] [Supplementary Note 9] Generating the third data includes: Executing a first attention operation based on the first query, the first key, and the first value to generate fourth data, Executing a first sum-product operation based on the fourth data to generate fifth data, Executing a second attention operation based on the second query, the second key, and the second value to generate sixth data, Executing a second sum-product operation based on the sixth data to generate the third data, Including: The information processing method according to Supplementary Note 8.

[0194] Although some embodiments of the present invention have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention and are included in the invention described in the claims and the equivalent scope thereof.

Explanation of Reference Numerals

[0195] 1, 1a... Information processing device 11... Control circuit 12... Memory 13... Storage 14... User interface 15, 15-1, 15-2... Encoder 16, 16a, 16b, 16c... Decoder 15_1, 16_1, 16c_1... First layer 15_n, 16_n, 16c_n... nth layer 15_N, 16_N, 16c_N... Nth layer 15_s... Receiver 16_e... Judgment unit 21... Information source 22... Question 22R... Re-question 23... Answer 24, 24_1, 24_n, 24_N... Key 25, 25_1, 25_n, 25_N... Value 26, 26_1, 26_n, 26_N... Query 30_n... Query conversion unit 31_n... Key conversion unit 32_n... Value conversion unit 33_n, 41_n... Similarity calculation unit 34_n, 42_n... Weighted sum calculation unit 35_n, 38_n, 40_n, 43_n, 46_n... Residual connection unit 36_n, 39_n, 44_n, 47_n... Normalization unit 16_f, 37_n, 45_n... Feedforward network< / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask> < / mask>

Claims

1. An encoder including a first layer and a second layer connected in series, A decoder, Comprising, The encoder is configured to: Based on the first data, generate a first key and a first value in the first layer, and generate a second key and a second value in the second layer; Based on second data different from the first data, generate a first query in the first layer and generate a second query in the second layer. Configured as such, The decoder is configured to generate third data that is included in the first data and not included in the second data based on the operation result based on the first key, the first value, and the first query, and the operation result based on the second key, the second value, and the second query. An information processing apparatus.

2. The decoder includes a third layer and a fourth layer connected in series, The third layer includes a first attention layer and a first neural network layer connected in series, The fourth layer includes a second attention layer and a second neural network layer connected in series, The first attention layer is configured to execute a first attention operation based on the first query, the first key, and the first value to generate fourth data; The first neural network layer is configured to execute a first sum-product operation based on the fourth data using a first parameter to generate fifth data; The second attention layer is configured to execute a second attention operation based on the second query, the second key, and the second value to generate sixth data; The second neural network layer is configured to execute a second sum-product operation based on the sixth data using a second parameter to generate the third data. The information processing apparatus according to Claim 1.

3. The second attention layer is configured to execute a third query based on the fifth data and the second query, and a second attention operation based on the second key and the second value to generate the sixth data. The information processing apparatus according to claim 2.

4. The second attention layer is configured to generate the third query by performing a residual connection on the fifth data and the second query. The information processing apparatus according to claim 3.

5. The second neural network layer is configured to execute a second sum-product operation based on seventh data based on the fifth data and the sixth data to generate the third data. The information processing apparatus according to claim 3.

6. The second attention layer is configured to generate the seventh data by performing a residual connection on the fifth data and the sixth data. The information processing apparatus according to claim 5.

7. The decoder includes a third neural network layer, and a third layer and a fourth layer connected to the third neural network layer. The third layer includes a first attention layer connected in series, and a first neural network layer. The fourth layer includes a second attention layer connected in series, and a second neural network layer. The first attention layer is configured to execute a first attention operation based on the first query, the first key, and the first value to generate fourth data. The first neural network layer is configured to execute a first sum-product operation based on the fourth data using a first parameter to generate fifth data. The second attention layer is configured to execute a second attention operation based on the second query, the second key, and the second value to generate sixth data. The second neural network layer is configured to perform a second sum-of-products operation based on the sixth data using second parameters to generate seventh data. The seventh data is independent of the fifth data. The third neural network layer is configured to perform a third sum-of-products operation based on the fifth data and the seventh data to generate the third data. The information processing apparatus according to claim 1.

8. The encoder is configured to perform a third attention operation in the first layer based on the first data to generate the first key and the first value, and perform a fourth attention operation in the second layer to generate the second key and the second value. Based on the second data, perform a fifth attention operation in the first layer to generate a first query, and perform a sixth attention operation in the second layer to generate the second query. configured as such. The information processing apparatus according to claim 1.

9. The third attention operation, the fourth attention operation, the fifth attention operation, and the sixth attention operation are self-attention operations. The information processing apparatus according to claim 8.

10. The apparatus further comprises a storage configured to store the first key and the first value in association with each other and store the second key and the second value in a non-volatile manner in association with each other. The decoder is configured to load the first key, the first value, the second key, and the second value from the storage. The information processing apparatus according to claim 1.

11. The encoder includes a first encoder and a second encoder. The first encoder includes a third layer and a fourth layer connected in series, the third layer being the first layer and the fourth layer being the second layer. The second encoder includes a fifth layer and a sixth layer connected in series, the fifth layer being the first layer and the sixth layer being the second layer, The first encoder is configured to generate the first key and the first value in the third layer and generate the second key and the second value in the fourth layer based on the first data, The second encoder is configured to generate the first query in the fifth layer and generate the second query in the sixth layer based on the second data, The information processing apparatus according to claim 1.

12. The first key, the second key, the first query, and the second query have the same number of dimensions, The information processing apparatus according to claim 11.

13. An information processing method by an information processing apparatus including an encoder including a first layer and a second layer connected in series and a decoder, generating, by the encoder, a first key and a first value in the first layer and generating a second key and a second value in the second layer based on first data; generating, by the encoder, a first query in the first layer and generating a second query in the second layer based on second data different from the first data; generating, by the decoder, third data included in the first data and not included in the second data based on an operation result based on the first key, the first value, and the first query and an operation result based on the second key, the second value, and the second query; An information processing method comprising:

14. After executing the information processing method according to claim 13, by the information processing apparatus, calculating a loss function based on the third data and correct data; Updating the parameters of the decoder used to generate the third data based on the calculated loss function; Generating the first key, the first value, the second key, and the second value based on the updated parameters, generating the first query and the second query, generating the third data, performing the calculation, and performing the update, repeating a first number of times; A method for generating a learning model, comprising:

15. Further comprising, in at least one of the first number of repetitions, generating the first key, the first value, the second key, and the second value based on data obtained by modifying a portion of the first data corresponding to the third data; The generation method according to claim 14.

Citation Information

Patent Citations

  • Interactive response system, model learning device and interactive device

    JP2019040574A

  • Inquiry answering apparatus and computer program

    JP2020004045A

  • Neural network-based translation of natural language queries into database queries

    JP2020520516A

  • Data Processing System

    JP6772692B2