Device, Method, and Program
A shared input and modular output layer model architecture for machine reading comprehension systems addresses resource inefficiencies by enabling efficient multi-format response output and selection.
Patent Information
- Application Number
- JP2024047724
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-06-29
AI Technical Summary
Existing machine reading comprehension systems require multiple models for different output formats, leading to excessive resource consumption and inefficiency in devices with resource constraints.
A single model architecture with shared input and modular output layers, combined with a probability distribution calculation for response selection, allows for outputting responses in multiple formats while minimizing resource usage.
Enables efficient output of responses in multiple formats without excessive resource consumption, ensuring appropriate response selection.
Smart Images

Figure 0007704245000001 
Figure 0007704245000002 
Figure 0007704245000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to Device , Method and Program .
Background Art
[0002] Machine reading comprehension is known as a question-and-answer technology that automatically answers questions input by a user in natural language. Machine reading comprehension is a technology that inputs a question by a user and a related document (referred to as a "passage") described in natural language, and outputs an answer to the input question based on information extracted from the passage.
[0003] When outputting an answer to a question by the machine reading comprehension, the output format is various. As an example, · A format for outputting an answer sentence generated by sentence generation based on information extracted from a passage, · A format for outputting a label such as YES / NO based on information extracted from a passage, · A format for outputting a question (a question for narrowing down an answer) generated based on information extracted from a passage, etc. can be mentioned.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
[0005] Here, in order to realize machine reading comprehension that can be output in any of the above output formats, it is conceivable to combine machine reading comprehension models corresponding to the respective output formats.
[0006] However, if the configuration is such that a number of machine reading models corresponding to the number of output formats are deployed and executed in the memory, a large amount of resources in the device will be consumed, and it will be less feasible in a device with resource constraints.
[0007] An object of the present disclosure is to provide a question-and-answer device, a question-and-answer method, and a question-and-answer program that can output responses to questions in a plurality of output formats when outputting responses to questions by machine reading.
Means for Solving the Problem
[0008] According to one aspect of the present disclosure, Device is, A plurality of models each having a plurality of layers, configured by a common layer on the input side of the plurality of models corresponding to the number of output formats when outputting a response to a first text described in natural language a calculation unit, Composed of layers on the output side of each of the plurality of models a plurality of output units, Having When the first text described in natural language and the output of the calculation unit when the second text used when responding to the first text are input to the calculation unit, the responses in the corresponding output formats are output from the plurality of output units when the output of the calculation unit is input to each of the plurality of output units. Perform learning processing on the calculation unit and the plurality of output units .
Effect of the Invention
[0009] According to the present disclosure, it is possible to provide a question-and-answer device, a question-and-answer method, and a question-and-answer program that can output responses to questions in a plurality of output formats when outputting responses to questions by machine reading.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
[0011] Hereinafter, each embodiment will be described with reference to the accompanying drawings. In the present specification and drawings, components having substantially the same functional configuration are denoted by the same reference numerals, and redundant descriptions are omitted.
[0012] [First Embodiment] <Hardware Configuration of Question-and-Answer Device> First, the hardware configuration of the question-and-answer device according to the first embodiment will be described. FIG. 1 is a diagram showing an example of the hardware configuration of the question-and-answer device.
[0013] As shown in FIG. 1, the question-and-answer device 100 includes a processor 101, a memory 102, an auxiliary storage device 103, an I / F (Interface) device 104, a communication device 105, and a drive device 106. Each hardware of the question-and-answer device 100 is connected to each other via a bus 107.
[0014] The processor 101 has various arithmetic devices such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit). The processor 101 reads out various programs (not shown) installed in the auxiliary storage device 103 onto the memory 102 and executes them.
[0015] The memory 102 has main storage devices such as a ROM (Read Only Memory) and a RAM (Random Access Memory). The processor 101 and the memory 102 form a so-called computer, and the computer realizes various functions by the processor 101 executing various programs read onto the memory 102.
[0016] For example, in the present embodiment, the computer formed by the processor 101 and the memory 102 realizes the question-and-answer unit 110 by the processor 101 executing the question-and-answer program read onto the memory 102. As will be described later, the question-and-answer unit 110 is configured to suppress the consumption of computer resources in the question-and-answer device 100.
[0017] The auxiliary storage device 103 stores various programs and various data used when the various programs are executed by the processor 101. For example, in the present embodiment, the auxiliary storage device 103 has a learning data set storage unit 120 and a message storage unit 130, and stores various data (learning data sets and messages (related documents described in natural language)).
[0018] The I / F device 104 is a connection device that connects the input device 140 and the output device 141 to the question-and-answer device 100. The I / F device 104 receives questions for the question-and-answer device 100 via the input device 140. Also, the I / F device 104 outputs the response generated by the question-and-answer device 100 for the input question via the output device 141. The input device 140 mentioned here includes devices that convert the input question into voice data, devices that convert it into text data, and the like. Similarly, the output device 141 mentioned here includes devices that output responses using voice data, devices that output responses using text data, and the like.
[0019] The communication device 105 is a communication device for communicating with other devices via a network.
[0020] The drive device 106 is a device for setting the recording medium 142. The recording medium 142 mentioned here includes media that optically, electrically, or magnetically record information, such as CD-ROMs, flexible disks, magneto-optical disks, etc. Also, the recording medium 142 may include semiconductor memories that electrically record information, such as ROMs, flash memories, etc.
[0021] Note that various programs installed in the auxiliary storage device 103 are installed, for example, when the distributed recording medium 142 is set in the drive device 106 and the various programs recorded on the recording medium 142 are read by the drive device 106. Alternatively, various programs installed in the auxiliary storage device 103 may be installed by being downloaded from a network via the communication device 105.
[0022] Similarly, various data stored in each storage unit of the auxiliary storage device 103 are stored, for example, when the distributed recording medium 142 is set in the drive device 106 and various data recorded on the recording medium 142 are read out by the drive device 106. Alternatively, various data stored in each storage unit of the auxiliary storage device 103 may be stored by being downloaded from a network via the communication device 105.
[0023] <Functional Configuration of Question-Answering Unit> Next, the detailed functional configuration of the question-answering unit 110 will be described. In the description, in order to clarify the features of the functional configuration of the question-answering unit 110, as a comparative example, first, the functional configuration of the question-answering unit constructed by combining a number of machine reading models corresponding to the number of output formats will be described.
[0024] (1) Functional Configuration of Question-Answering Unit in Comparative Example FIG. 2 is a diagram showing the functional configuration of the question-answering unit in the comparative example. When outputting an answer to a question by machine reading so as to be able to output in a plurality of output formats, the question-answering unit 200 in the comparative example includes · an input unit 210, · a plurality of machine reading models (in the example of FIG. 2, the first machine reading model 221 to the third machine reading model 223), · a selection unit 230, and has.
[0025] The input unit 210 inputs the input question and a message (related document described in natural language) to each of the plurality of machine reading models.
[0026] The first machine reading model 221 outputs, as a first response, an answer sentence generated by sentence generation based on the information extracted from the message.
[0027] The second machine reading model 222 outputs, as a second response, a label such as YES / NO generated based on the information extracted from the message.
[0028] The third machine reading comprehension model 223 outputs, as a third response, a question (a question generated to narrow down the answer (referred to as a revised question)) generated based on the information extracted from the message.
[0029] The selection unit 230 selects and outputs a predetermined number of responses from among the responses output from each of the first machine reading comprehension model 221 to the third machine reading comprehension model 223.
[0030] As shown in the question-and-answer unit 200 of the comparative example, by combining a number of machine reading comprehension models corresponding to the number of output formats, responses to questions can be output in a plurality of output formats. On the other hand, in the case of the question-and-answer unit 200 of the comparative example, there are the following problems. ·Due to the configuration in which a number of machine reading comprehension models corresponding to the number of output formats are expanded and executed in the memory, a large amount of computer resources in the question-and-answer device 100 will be consumed. ·When selecting a predetermined number of responses from among the responses output from each of the first machine reading comprehension model 221 to the third machine reading comprehension model 223, an appropriate response cannot be selected. This is because it is impossible to determine the superiority or inferiority of the responses by simply comparing the first response to the third response, and it is necessary to calculate some selection index.
[0031] In contrast, the question-and-answer unit 110 of the question-and-answer device 100 according to the first embodiment has a configuration for solving these problems. This will be described in detail below.
[0032] (2) Functional configuration of the question-and-answer unit FIG. 3 is a diagram showing an example of the functional configuration of the question-and-answer unit of the question-and-answer device according to the first embodiment. When outputting a response to a question by machine reading comprehension, the question-and-answer unit 110 is configured to be able to output in a plurality of output formats while suppressing the consumption of computer resources and being able to select an appropriate response from among the responses in a plurality of output formats for the question. Therefore, ·An input unit 310, ·An understanding layer 320 that functions as an input layer, ·The first output layer 321, the second output layer 322, and the third output layer 323 that function as output layers, ·The output determination layer 324 that functions as an output layer, ·The selection unit 330, and has.
[0033] Among these, the input unit 310 inputs the input question and the message (related document described in natural language) into the understanding layer 320.
[0034] Also, the understanding layer 320 is an example of a calculation unit. Taking the question and the message as inputs, it calculates information indicating the relevance between the question and the message on the vector of deep learning, and outputs a state vector or a state tensor. Note that the understanding layer 320 can adopt any structure as long as it can input the question and the message and calculate information indicating the relevance between the question and the message.
[0035] For example, the understanding layer 320 can adopt BiDAF using RNN() (see Non-Patent Document 1) or BERT based on Transformer (see Non-Patent Document 2), etc. However, the understanding layer 320 needs to have a structure that outputs in a form of a state vector or a state tensor that is input to the output layers (the first output layer 321 to the third output layer 323, the output determination layer 324) at the subsequent stage of the understanding layer 320.
[0036] The first output layer 321 to the third output layer 323 are examples of output units. Taking as input the information (state vector or state tensor) indicating the relevance between the question and the message output from the understanding layer 320, they output a response (the output result of machine reading comprehension).
[0037] In the case of FIG. 3, the first output layer 321 outputs the answer sentence generated by sentence generation as the first response. Also, the second output layer 322 outputs labels such as YES / NO as the second response. Further, the third output layer 323 outputs a revised question for narrowing down the answer to the input question as the third response.
[0038] Note that when the first output layer 321 to the third output layer 323 output the first response to the third response, they may have any deep learning structure. Also, the output formats output by the first output layer 321 to the third output layer 323 are not limited to answer sentences, labels, and revised questions, and responses in other output formats may be output.
[0039] Also, in the example of FIG. 3, the configuration is such that the first to third output layers are installed. However, the number of output layers to be installed is arbitrary, and a plurality of output layers having the same structure may be installed. For example, two structures of a sentence generation decoder may be installed. One output layer may be an output layer learned to generate an answer sentence by sentence generation, and the other output layer may be an output layer learned to generate a revised question by sentence generation.
[0040] The output determination layer 324 is an example of an index calculation unit, and calculates the probability distribution of each response output from the first output layer 321 to the third output layer 323.
[0041] Specifically, the output determination layer 324 has a softmax layer with a number of dimensions corresponding to the number N of output formats (in the example of FIG. 3, N = 3). The softmax layer with N dimensions receives the state vector or state tensor from the understanding layer 320 and calculates the probability distribution of each response.
[0042] Note that such a configuration is effective when the number N of output formats is fixed. On the other hand, when it is necessary to add an output format, since the number of dimensions of the softmax layer cannot cope, it is necessary to perform a relearning process for the entire question-and-answer unit 110.
[0043] The selection unit 330 selects the top M responses (M is a predetermined number, M = 1 in the example of FIG. 3) from the responses output from each of the first output layer 321 to the third output layer 323, the probability distribution calculated by the output determination layer 324, and outputs them to the user via the output device 141.
[0044] In this way, in the question-and-answer unit 110, · Install a shared understanding layer for multiple output formats (share the input layer), · Separate the input layer (understanding layer) and the output layer (the first output layer 321 to the third output layer 323, output judgment layer 324), and modularize each layer. · Using the probability distribution of each response calculated by the output judgment layer 324 installed in the output layer as a selection index, the selection unit 330 selects the response to be finally output. It is configured in this way. Thus, according to the question answering unit 110, when outputting a response to a question by machine reading comprehension, it can be output in multiple output formats, and · By sharing the input layer, the consumption of computer resources is suppressed, and · By calculating the selection index, an appropriate response is selected. This becomes possible.
[0045] <Learning method of the question answering unit> Next, the learning method of the question answering unit 110 will be described. In the learning phase of performing learning processing on the question answering unit 110, the question answering unit 110 of the question answering device 100 according to the present embodiment first installs a comparison / modification unit 410 instead of the selection unit 330. Subsequently, the question answering unit 110 of the question answering device 100 according to the present embodiment performs learning processing on each layer (understanding layer 320, the first output layer 321 to the third output layer 323, output judgment layer 324) installed in the question answering unit 110.
[0046] At this time, in the question answering unit 110 of the question answering device 100, the learning dataset used when performing learning processing on individual machine reading comprehension models (the first machine reading comprehension model 221 to the third machine reading comprehension model 223) is used.
[0047] Specifically, the question answering unit 110 of the question answering device 100 first sets a flag indicating for which output layer the learning dataset is used for learning processing.
[0048] Subsequently, the question-and-answer unit 110 of the question-and-answer device 100 stores the correct data of the output determination layer 324 in the training data set. For example, the question-and-answer unit 110 of the question-and-answer device 100 stores vector data in which the value of the dimension corresponding to the set flag is "1" and the values of the other dimensions are "0" as the correct data of the output determination layer 324 in the training data set.
[0049] Subsequently, the question-and-answer unit 110 of the question-and-answer device 100 calculates the learning loss between the response output from the output layer corresponding to the set flag and the corresponding correct data, and updates the parameters of the output layer and the understanding layer corresponding to the set flag based on the calculated learning loss. At this time, the question-and-answer unit 110 of the question-and-answer device 100 ignores the responses output from the output layers other than the output layer corresponding to the set flag.
[0050] Also, the question-and-answer unit 110 of the question-and-answer device 100 calculates the learning loss between the vector data output from the output determination layer and the corresponding correct data, and updates the parameters of the output determination layer based on the calculated learning loss.
[0051] FIG. 4 is a first diagram showing an operation example in the learning phase of the question-and-answer device according to the first embodiment. In the case of FIG. 4, it is assumed that a flag indicating that the training data set 400 = "the training data set used for the learning process of the first output layer 321" is set by the question-and-answer unit 110.
[0052] As shown in FIG. 4, the training data set 400 includes "input data", "correct data of the first output layer", and "correct data of the output determination layer" as information items. · "Input data" stores questions and messages. · "Correct data of the first output layer" stores the correct data of the answer sentence generated by sentence generation based on the information extracted from the corresponding message. · The "correct data of the output judgment layer" includes "the first dimension" to "the third dimension", and vector data with the value of the first dimension being "1" and the values of the second and third dimensions being "0" is stored.
[0053] In FIG. 4, when the input unit 310 inputs the input data (a set of a question and a message) of the learning data set 400 to the understanding layer 320, the first response to the third response are output as the output results of machine reading comprehension from the first output layer 321 to the third output layer 323. Also, vector data of N dimensions (N = 3 in the example of FIG. 4) is output from the output judgment layer 324.
[0054] In the comparison / modification unit 410, a learning loss is calculated between the first response (answer sentence) output from the first output layer 321 and the answer sentence stored in the "correct data of the first output layer" of the learning data set 400. Also, in the comparison / modification unit 410, based on the calculated learning loss, the parameters of the first output layer 321 and the understanding layer 320 are updated.
[0055] Similarly, in the comparison / modification unit 410, a learning loss is calculated between the vector data of N (N = 3 in the example of FIG. 4) dimensions output from the output judgment layer 324 and the vector data of the first dimension to the third dimension stored in the "correct data of the output judgment layer" of the learning data set 400. Also, in the comparison / modification unit 410, based on the calculated learning loss, the parameters of the output judgment layer 324 are updated.
[0056] On the other hand, FIG. 5 is a second diagram showing an operation example in the learning phase of the question-answering device according to the first embodiment. In the case of FIG. 5, it is assumed that a flag indicating that the learning data set 500 = "the learning data set used for the learning process of the second output layer 322" is set by the question-answering unit 110.
[0057] As shown in FIG. 5, the learning data set 500 includes, as information items, "input data", "correct data of the second output layer", and "correct data of the output judgment layer". · The "input data" stores questions and messages. · The "correct data for the second output layer" stores the correct data of labels such as YES / NO, which is generated based on the information extracted from the corresponding message. · The "correct data for the output judgment layer" includes "the first dimension" to "the third dimension", and stores vector data with the value of the second dimension being "1" and the values of the first dimension and the third dimension being "0".
[0058] In FIG. 5, when the input unit 310 inputs the input data (a set of questions and messages) of the learning dataset 500 to the understanding layer 320, the first response to the third response are output from the first output layer 321 to the third output layer 323 as the output results of machine reading comprehension. Also, vector data of N dimensions (N = 3 in the example of FIG. 5) is output from the output judgment layer 324.
[0059] In the comparison / modification unit 410, the learning loss is calculated between the second response (label) output from the second output layer 322 and the label stored in the "correct data for the second output layer" of the learning dataset 500. Also, in the comparison / modification unit 410, based on the calculated learning loss, the parameters of the second output layer 322 and the understanding layer 320 are updated.
[0060] Similarly, in the comparison / modification unit 410, the learning loss is calculated between the vector data of N dimensions (N = 3 in the example of FIG. 5) output from the output judgment layer 324 and the vector data of the first dimension to the third dimension stored in the "correct data for the output judgment layer" of the learning dataset 500. Also, in the comparison / modification unit 410, based on the calculated learning loss, the parameters of the output judgment layer 324 are updated.
[0061] On the other hand, FIG. 6 is a third diagram showing an operation example in the learning phase of the question-answering apparatus according to the first embodiment. In the case of FIG. 6, it is assumed that a flag indicating that the learning dataset 600 = "the learning dataset used for the learning process of the third output layer 323" is set by the question-answering unit 110.
[0062] As shown in FIG. 6, the learning dataset 600 includes, as information items, "input data", "correct data for the third output layer", and "correct data for the output determination layer". · The "input data" stores a question and a message. · The "correct data for the third output layer" stores the correct data for the revised question generated based on the information extracted from the corresponding message. · The "correct data for the output determination layer" includes "first dimension" to "third dimension", and stores vector data with the value of the third dimension being "1" and the values of the first and second dimensions being "0".
[0063] In FIG. 6, when the input unit 310 inputs the input data (a set of a question and a message) of the learning dataset 600 to the understanding layer 320, the first response to the third response are output as the output results of machine reading comprehension from the first output layer 321 to the third output layer 323. Also, vector data of N dimensions (N = 3 in the example of FIG. 6) is output from the output determination layer 324.
[0064] In the comparison / modification unit 410, a learning loss is calculated between the third response (revised question) output from the third output layer 323 and the revised question stored in the "correct data for the third output layer" of the learning dataset 600. Also, in the comparison / modification unit 410, based on the calculated learning loss, the parameters of the third output layer 323 and the understanding layer 320 are updated.
[0065] Similarly, in the comparison / modification unit 410, a learning loss is calculated between the vector data of N dimensions (N = 3 in the example of FIG. 6) output from the output determination layer 324 and the vector data of the first dimension to the third dimension stored in the "correct data for the output determination layer" of the learning dataset 600. Also, in the comparison / modification unit 410, based on the calculated learning loss, the parameters of the output determination layer 324 are updated.
[0066] In this way, in the question-and-answer unit 110 of the question-and-answer device 100, learning processes are sequentially performed on the understanding layer 320, the first output layer 321 to the third output layer 323, and the output determination layer 324 using the learning data sets 400 to 600.
[0067] <Flow of question-and-answer processing> Next, the flow of question-and-answer processing by the question-and-answer device 100 will be described. FIG. 7 is a flowchart showing the flow of question-and-answer processing by the question-and-answer device according to the first embodiment. Among these, steps S701 to S703 represent the processing in the learning phase, and steps S704 to S707 represent the processing in the response phase.
[0068] In step S701, the question-and-answer unit 110 performs a learning process on the understanding layer 320, the first output layer 321, and the output determination layer 324 using the learning data set 400.
[0069] In step S702, the question-and-answer unit 110 performs a learning process on the understanding layer 320, the second output layer 322, and the output determination layer 324 using the learning data set 500.
[0070] In step S703, the question-and-answer unit 110 performs a learning process on the understanding layer 320, the third output layer 323, and the output determination layer 324 using the learning data set 600.
[0071] In step S704, the input unit 310 of the question-and-answer unit 110 receives the input of the question and the message, and inputs the input question and message to the understanding layer 320.
[0072] In step S705, the first output layer 321 to the third output layer 323 take the state vector output from the understanding layer 320 as an input and output the first response to the third response.
[0073] In step S706, the output determination layer 324 of the question-and-answer unit 110 takes the state vector output from the understanding layer 320 as input, calculates the probability distributions of the first to third responses, and outputs a selection index.
[0074] In step S707, the selection unit 330 of the question-and-answer unit 110 selects the top M predetermined responses based on the selection index output from the output determination layer 324, and outputs the selected responses.
[0075] <Summary> As is clear from the above description, the question-and-answer device 100 according to the first embodiment has: · An understanding layer that takes a question and a message as input and calculates information indicating the relevance between the question and the message. · A first output layer to a third output layer that take the information indicating the relevance calculated by the understanding layer as respective inputs and output the first to third responses having different output formats. · An output determination layer that calculates the probability distributions of the first to third responses based on the information indicating the relevance output by the understanding layer. Further, it has a selection unit that selects a predetermined number of responses using the probability distributions of the first to third responses calculated by the output determination layer as a selection index.
[0076] Thus, according to the question-and-answer device 100 according to the first embodiment, it is possible to output responses to a question in a plurality of output formats, suppress the consumption of computer resources, and select appropriate responses.
[0077] That is, according to the first embodiment, when outputting a response to a question by machine reading comprehension, it is possible to provide a highly feasible question-and-answer device, question-and-answer method, and question-and-answer program that can output in a plurality of output formats.
[0078] [Second Embodiment] In the first embodiment described above, on the premise that the number of output formats is fixed, when adding a new output format, the question-and-answer unit 110 is configured to perform relearning processing on the entire question-and-answer unit.
[0079] In contrast, in the second embodiment, even when a new output format is added, the question-and-answer unit is configured so that it is not necessary to perform relearning processing on the entire question-and-answer unit 110. Hereinafter, the second embodiment will be described centering on the differences from the first embodiment.
[0080] <Functional Configuration of Question-and-Answer Unit> First, the functional configuration of the question-and-answer unit of the question-and-answer device according to the second embodiment will be described. FIG. 8 is a diagram showing an example of the functional configuration of the question-and-answer unit of the question-and-answer device according to the second embodiment.
[0081] The difference from the functional configuration shown in FIG. 3 is that, in the case of the question-and-answer unit 800 in FIG. 8, for each of the first output layer 321 to the third output layer 323, as a selection index, the first output determination layer 801 to the third output determination layer 803 that calculate individual scores are installed. Also, in the case of the question-and-answer unit 800 in FIG. 8, the function of the selection unit 810 is different from the function of the selection unit 330 in FIG. 3.
[0082] The first output determination layer 801 is an example of an index calculation unit, and has a logit layer that receives the state vector of the first output layer 321 and calculates a scalar value of 0 to 1.0 as the first score.
[0083] Similarly, the second output determination layer 802 is an example of an index calculation unit, and has a logit layer that receives the state vector of the second output layer 322 and calculates a scalar value of 0 to 1.0 as the second score.
[0084] Similarly, the third output determination layer 803 is an example of an index calculation unit, and has a logit layer that receives the state vector of the third output layer 323 and calculates a scalar value of 0 to 1.0 as the third score.
[0085] The selection unit 810 selects and outputs a response corresponding to the top M scores determined in advance based on the first score to the third score calculated by the first output determination layer 801 to the third output determination layer 803.
[0086] As described above, in the question-and-answer device 100 according to the second embodiment, the first output determination layer 801 to the third output determination layer 803 that calculate individual scores for each of the first output layer 321 to the third output layer 323 are provided. Accordingly, according to the second embodiment, even when a new output format is added, it is sufficient to perform learning processing on the added new output layer and output determination layer and the understanding layer, and there is no need to perform re-learning processing on the already learned output layer and output determination layer.
[0087] <Learning method of the question-and-answer unit> Next, a learning method of the question-and-answer unit 800 will be described. In a learning phase in which learning processing is performed on the question-and-answer unit 110, the question-and-answer unit 800 of the question-and-answer device 100 according to the present embodiment first installs a comparison / modification unit 910 instead of the selection unit 810. Subsequently, the question-and-answer unit 800 of the question-and-answer device 100 according to the present embodiment performs learning processing on each layer (understanding layer 320, first output layer 321 to third output layer 323, first output determination layer 801 to third output determination layer 803) installed in the question-and-answer unit 800.
[0088] At this time, in the question-and-answer device 100, similar to the first embodiment, a learning data set used when performing learning processing on individual machine reading models (first machine reading model 221 to third machine reading model 223) is used.
[0089] Specifically, the question-and-answer unit 800 of the question-and-answer device 100 first sets a flag indicating for which output layer the learning data set is used for learning processing.
[0090] Subsequently, the question-and-answer unit 800 of the question-and-answer device 100 stores the correct answer data of the first score to the third score calculated by the first output determination layer 801 to the third output determination layer 803 in the learning data set. For example, the question-and-answer unit 800 of the question-and-answer device 100 stores the correct answer data with the score corresponding to the set flag being "1.0" in the learning data set.
[0091] Subsequently, the question-and-answer unit 800 of the question-and-answer device 100 calculates the learning loss between the response output from the output layer corresponding to the set flag and the corresponding correct answer data, and updates the parameters of the output layer and the understanding layer corresponding to the set flag based on the calculated learning loss. At this time, the question-and-answer unit 800 of the question-and-answer device 100 ignores the responses output from the output layers other than the output layer corresponding to the set flag.
[0092] Also, the question-and-answer unit 800 of the question-and-answer device 100 calculates the learning loss between the score output from the output determination layer corresponding to the set flag and the corresponding correct answer data, and updates the parameters of the output determination layer corresponding to the set flag based on the calculated learning loss. At this time, in the question-and-answer device 100, the scores output from the output determination layers other than the output determination layer corresponding to the set flag are ignored.
[0093] FIG. 9 is a first diagram showing an operation example in the learning phase of the question-and-answer device according to the second embodiment. In the case of FIG. 9, it is assumed that a flag indicating that the learning data set 900 = "the learning data set used for the learning process of the first output layer 321 and the first output determination layer 801" is set by the question-and-answer unit 800.
[0094] As shown in FIG. 9, the learning data set 900 includes "input data", "correct answer data of the first output layer", and "correct answer data of the output determination layer" as information items. · "Input data" stores questions and messages. · Based on the information extracted from the corresponding message, the correct data of the answer sentence generated by sentence generation is stored in the "correct data of the first output layer". · The "correct data of the output judgment layer" stores the correct data of the score (first score) output from the first output judgment layer 801.
[0095] In FIG. 9, when the input unit 310 inputs the input data (a set of a question and a message) of the learning data set 900 to the understanding layer 320, the first response to the third response are output from the first output layer 321 to the third output layer 323 as the output results of machine reading comprehension. Further, the first score to the third score are output from the first output judgment layer 801 to the third output judgment layer 803.
[0096] In the comparison / modification unit 910, a learning loss is calculated between the first response (answer sentence) output from the first output layer 321 and the answer sentence stored in the "correct data of the first output layer" of the learning data set 900. Further, in the comparison / modification unit 910, based on the calculated learning loss, the parameters of the first output layer 321 and the understanding layer 320 are updated.
[0097] Similarly, in the comparison / modification unit 910, a learning loss is calculated between the first score output from the first output judgment layer 801 and the value stored in the first score of the "correct data of the output judgment layer" of the learning data set 900. Further, in the comparison / modification unit 910, based on the calculated learning loss, the parameters of the first output judgment layer 801 are updated.
[0098] On the other hand, FIG. 10 is a second diagram showing an operation example in the learning phase of the question-answering apparatus according to the second embodiment. In the case of FIG. 10, it is assumed that a flag indicating that the learning data set 1000 = "the learning data set used for the learning process of the second output layer 322 and the second output judgment layer 802" is set by the question-answering unit 800.
[0099] Note that, as shown in FIG. 10, the learning dataset 1000 includes, as information items, "input data", "correct data for the second output layer", and "correct data for the output judgment layer". · In the "input data", questions and messages are stored. · In the "correct data for the second output layer", correct data of labels such as YES / NO, which is generated based on information extracted from the corresponding message, is stored.
[0100] In the "correct data for the output judgment layer", correct data of the score (second score) output from the second output judgment layer 802 is stored.
[0101] In FIG. 10, when the input unit 310 inputs the input data (a set of a question and a message) of the learning dataset 1000 to the understanding layer 320, the first response to the third response are output from the first output layer 321 to the third output layer 323 as the output results of machine reading comprehension. Also, the first score to the third score are output from the first output judgment layer 801 to the third output judgment layer 803.
[0102] In the comparison / modification unit 910, a learning loss is calculated between the second response (label) output from the second output layer 322 and the label stored in the "correct data for the second output layer" of the learning dataset 1000. Also, in the comparison / modification unit 910, based on the calculated learning loss, the parameters of the second output layer 322 and the understanding layer 320 are updated.
[0103] Similarly, in the comparison / modification unit 910, a learning loss is calculated between the second score output from the second output judgment layer 802 and the value stored in the second score of the "correct data for the output judgment layer" of the learning dataset 1000. Also, in the comparison / modification unit 910, based on the calculated learning loss, the parameters of the second output judgment layer 802 are updated.
[0104] On the other hand, FIG. 11 is a third diagram showing an operation example in the learning phase of the question-and-answer device according to the second embodiment. In the case of FIG. 11, it is assumed that a flag indicating that the learning dataset 1100 = "a learning dataset used for the learning process for the third output layer 323 and the third output determination layer 803" is set by the question-and-answer unit 800.
[0105] As shown in FIG. 11, the learning dataset 1100 includes, as information items, "input data", "correct data for the third output layer", and "correct data for the output determination layer". · "Input data" stores questions and messages. · "Correct data for the third output layer" stores the correct data for the revised questions generated based on the information extracted from the corresponding messages. · "Correct data for the output determination layer" stores the correct data for the scores (third scores) output from the third output determination layer 803.
[0106] In FIG. 11, when the input unit 310 inputs the input data (a set of questions and messages) of the learning dataset 1100 to the understanding layer 320, the first response to the third response are output from the first output layer 321 to the third output layer 323 as the output results of machine learning. Also, the first score to the third score are output from the first output determination layer 801 to the third output determination layer 803.
[0107] The comparison / modification unit 910 calculates the learning loss between the third response (revised question) output from the third output layer 323 and the revised question stored in the "correct data for the third output layer" of the learning dataset 1100. Also, the comparison / modification unit 910 updates the parameters of the third output layer 323 and the understanding layer 320 based on the calculated learning loss.
[0108] Similarly, in the comparison / modification unit 910, a learning loss is calculated between the third score output from the third output determination layer 803 and the value stored in the third score of the "correct data of the output determination layer" of the learning dataset 1100. Also, in the comparison / modification unit 910, based on the calculated learning loss, the parameters of the third output determination layer 803 are updated.
[0109] <Question and Answer Processing Flow> Next, the flow of the question and answer process by the question and answer device 100 according to the second embodiment will be described. FIG. 12 is a flowchart showing the flow of the question and answer process by the question and answer device according to the second embodiment. The differences from the flowchart described with reference to FIG. 7 in the above first embodiment are steps S1201 to S1203.
[0110] In step S1201, the question and answer unit 110 performs learning processing on the understanding layer 320, the first output layer 321, and the first output determination layer 801 using the learning dataset 900.
[0111] In step S1202, the question and answer unit 110 performs learning processing on the understanding layer 320, the second output layer 322, and the second output determination layer 802 using the learning dataset 1000.
[0112] In step S1203, the question and answer unit 110 performs learning processing on the understanding layer 320, the third output layer 323, and the third output determination layer 803 using the learning dataset 1100.
[0113] <Summary> As is clear from the above description, the question and answer device 100 according to the second embodiment has · An understanding layer that calculates information indicating the relevance between a question and a message with the question and the message as inputs. · First to third output layers that output first to third responses, which are different output formats, using the information indicating the relevance calculated by the understanding layer as respective inputs. · It has first to third output determination layers that receive the state vectors of the first to third output layers and calculate individual scores (first to third scores) for the first to third output layers. Further, it has a selection unit that selects a predetermined number of responses using the first to third scores calculated by the first to third output determination layers as selection indicators.
[0114] Accordingly, according to the question-and-answer device 100 according to the second embodiment, similar to the first embodiment, responses to questions can be output in a plurality of output formats, computer resource consumption can be suppressed, and appropriate responses can be selected. In addition, according to the question-and-answer device 100 according to the second embodiment, even when a new output format is added, it is not necessary to perform relearning processing on the entire question-and-answer unit.
[0115] That is, according to the second embodiment, when outputting responses to questions by machine reading comprehension, it is possible to provide a more feasible question-and-answer device, question-and-answer method, and question-and-answer program that can output in a plurality of output formats.
[0116] [Other Embodiments] In the above first and second embodiments, cases where different output determination layers are installed have been described. However, the decision of whether to install the output determination layer in the first embodiment or the output determination layer in the second embodiment is arbitrary. For example, it may be determined in consideration of the tasks and purposes to be set, the system configuration of the question-and-answer device, and the like.
[0117] Also, in the above first and second embodiments, it has been described that the learning phase and the response phase are executed in the same question-and-answer device 100. However, the learning phase and the response phase may be configured to be executed by separate devices. In this case, the device that executes the response phase does not need to have the learning dataset storage unit 120, and the comparison / change units 410 and 910 are not installed either.
[0118] Note that the configurations and the like described in the above embodiments are not limited to the configurations shown here, such as combinations with other elements. Regarding these points, it is possible to make changes without departing from the spirit of the present invention, and it can be appropriately determined according to the application form.
Explanation of Signs
[0119] 100: Question and Answer Device 110: Question and Answer Unit 120: Learning Data Set Storage Unit 130: Message Storage Unit 310: Input Unit 320: Understanding Layer 321: First Output Layer 322: Second Output Layer 323: Third Output Layer 324: Output Judgment Layer 330: Selection Unit 400~600: Learning Data Set 801: First Output Judgment Layer 802: Second Output Judgment Layer 803: Third Output Judgment Layer 810: Selection Unit 900~1100: Learning Data Set
Claims
Claim 1. A plurality of models each having a plurality of layers, comprising: a calculation unit constituted by a common layer on the input side of the plurality of models corresponding to the number of output formats when outputting a response to a first text described in natural language; a plurality of output units constituted by layers on the output side of each of the plurality of models; and a device that performs learning processing on the calculation unit and the plurality of output units so that when an output of the calculation unit when a first text described in natural language and a second text used when responding to the first text are input is input to each of the plurality of output units, responses in corresponding output formats are output from the plurality of output units. Claim 2. It has an index calculation unit that calculates an index value for selecting any one of the responses in each output format output from each of the plurality of output units, The device according to claim 1, wherein learning processing is performed on the calculation unit and the index calculation unit so that when an output of the calculation unit when the first text and the second text used when responding to the first text are input is input to the index calculation unit, the index value for selecting any one of the responses is calculated by the index calculation unit. Claim 3. The plurality of output units have a plurality of index calculation units that receive information calculated when outputting responses in each output format, and calculate index values for each of the plurality of output units, The device according to claim 1, wherein learning processing is performed on the plurality of index calculation units so that when an output of the calculation unit when the first text and the second text used when responding to the first text are input is input to each of the plurality of output units, and the plurality of index calculation units receive the information calculated when each of the plurality of output units outputs each response, the index values are calculated by each of the plurality of index calculation units. Claim 4. The device according to claim 1, having a selection unit that selects any one of the responses in each output format output from each of the plurality of output units. Claim 5. The device according to claim 2 or 3, having a selection unit that selects any one of the responses in each output format output from each of the plurality of output units based on the index value. Claim 6. A computer A plurality of models each having a plurality of layers, a calculation unit composed of a common layer on the input side of the plurality of models according to the number of output formats when outputting a response to a first text described in natural language, functioning as a plurality of output units each composed of a layer on the output side of each of the plurality of models, a method of performing learning processing on the calculation unit and the plurality of output units such that when the output of the calculation unit when a first text described in natural language and a second text used when responding to the first text are input to the calculation unit is input to each of the plurality of output units, responses in corresponding output formats are output from the plurality of output units.
7. A computer, A plurality of models each having a plurality of layers, a calculation unit composed of a common layer on the input side of the plurality of models according to the number of output formats when outputting a response to a first text described in natural language, functioning as a plurality of output units each composed of a layer on the output side of each of the plurality of models, a program for performing learning processing on the calculation unit and the plurality of output units such that when the output of the calculation unit when a first text described in natural language and a second text used when responding to the first text are input to the calculation unit is input to each of the plurality of output units, responses in corresponding output formats are output from the plurality of output units.
8. A plurality of machine-learned models corresponding to the number of output formats, machine-learned using a first text described in natural language, a second text used when responding to the first text, and responses in each output format to the first text, each of the machine-learned models having a plurality of layers, a calculation unit composed of a common layer on the input side of the machine-learned models, and a plurality of output units each composed of a layer on the output side of each of the plurality of machine-learned models, wherein the plurality of output units output responses in corresponding output formats when the output of the calculation unit when a first text described in natural language and a second text used when responding to the first text are input to the calculation unit is input thereto.
9. An index calculation unit that calculates an index value for selecting any one of the responses in each output format output from each of the plurality of output units, where the output of the calculation unit when a first text described in natural language and a second text used when responding to the first text are input is input to each of the plurality of output units. The apparatus according to claim 8, further comprising the same.
10. The apparatus according to claim 8, further comprising a plurality of index calculation units that receive information calculated when each of the plurality of output units outputs each response and calculate index values for each of the plurality of output units.
11. The apparatus according to claim 8, further comprising a selection unit that selects any one of the responses in each output format output from each of the plurality of output units.
12. The apparatus according to claim 9 or 10, comprising a selection unit that selects any one of the responses in each output format output from each of the plurality of output units based on the index value.
13. A computer A plurality of machine-learned models corresponding to the number of output formats, learned using a first text described in natural language, a second text used when responding to the first text, and responses in each output format for the first text, the calculation unit being composed of a common layer on the input side of the machine-learned models each having a plurality of layers, Functioning as a plurality of output units composed of layers on the output side of each of the plurality of machine-learned models, The plurality of output units output responses in corresponding output formats by inputting the output of the calculation unit when a first text described in natural language and a second text used when responding to the first text are input to the calculation unit.
14. A computer A plurality of machine-learned models corresponding to the number of output formats, learned using a first text described in natural language, a second text used when responding to the first text, and responses in each output format for the first text, the calculation unit being composed of a common layer on the input side of the machine-learned models each having a plurality of layers, Functioning as a plurality of output units composed of layers on the output side of each of the plurality of machine-learned models. The plurality of output units are programs that output responses in corresponding output formats when the outputs of the calculation unit are respectively input when the first text described in natural language and the second text used when responding to the first text are input to the calculation unit.
Citation Information
Patent Citations
Dynamic Mutual Attention Networks for Question Answering
JP2020501229A
JPP6649536B
Data Processing Method, Apparatus and Electronic Device
US20190108273A1
CRF-based span prediction for fine machine learning comprehension
US20200175015A1
Question generation device, question generation method, and program
WO2019235103A1