Learning device, generation device, learning method, generation method, program, method for manufacturing a learned machine learning model, learning system, and generation system

The learning device enhances machine reading comprehension by incorporating visual information, addressing the limitations of conventional technologies that ignore document layout, thereby improving answer generation accuracy in complex document formats.

JP7704190B2Active Publication Date: 2025-07-08NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023216105
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-28
Filing Date
2023-12-21
Publication Date
2025-07-08
Estimated Expiration
2040-12-09

AI Technical Summary

Technical Problem

Conventional machine reading technologies fail to consider visual information in documents, such as position and size of text, leading to incomplete understanding when processing formats like HTML or PDF.

Method used

A learning device that utilizes a neural network model to incorporate visual information, including position and size of text, through model parameters, to generate accurate answers by learning from input data and correct answers.

Benefits of technology

Enables machine reading comprehension that considers visual information, improving the accuracy of answering questions in documents with layouts like HTML or PDF by generating answers that include visual context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704190000012
    Figure 0007704190000012
  • Figure 0007704190000013
    Figure 0007704190000013
  • Figure 0007704190000014
    Figure 0007704190000014
Patent Text Reader

Abstract

To achieve machine reading comprehension that takes visual information into consideration.SOLUTION: A learning apparatus includes: a generation unit which receives data including a visual region and first information related to the data, and generates, by using a model parameter of a machine learning model, second information corresponding to the first information, from information representing features of the region; and a learning unit which learns the model parameter on the basis of the second information and third information representing a correct answer of the second information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning device, a generation device , learning a learning method, a generation method, a program , learning a method for manufacturing a learned machine learning model , learning system, and generation system and relates thereto.

Background Art

[0002] If artificial intelligence can accurately perform "machine reading comprehension" that generates answers to questions based on a given set of documents, it can be applied to a wide range of services such as question answering and intelligent agent dialogue. Machine reading comprehension is divided into extraction type and generation type. As a conventional technique for performing generation type machine reading comprehension, for example, the technique disclosed in Non-Patent Document 1 is known.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, conventional machine reading technologies only handle text and cannot handle visual information such as the position and size of text in a document. For this reason, when understanding a document in which multiple texts are laid out (for example, an HTML (HyperText Markup Language) document, a PDF (Portable Document Format) document, etc.) by machine reading, all information other than the content of the text is handled in a state where it is missing.

[0005] One embodiment of the present invention has been made in view of the above points, and an object thereof is to realize machine reading that takes visual information into consideration.

Means for Solving the Problems

[0006] To achieve the above object, a learning device according to one embodiment uses model parameters of a machine learning model with data including a visual area and first information related to the data as inputs, and from information representing the characteristics of the area, A generation unit that generates second information corresponding to the first information; and a learning unit that learns the model parameters based on the second information and third information representing the correct answer of the second information.

Effects of the Invention

[0007] Machine reading that takes visual information into consideration can be realized.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

MODE FOR CARRYING OUT THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described.

[0010] ·First Embodiment In this embodiment, a question-and-answer device 10 will be described which can generate an answer text considering visual information (e.g., the position and size of text in the image, etc.) in an image containing text when the image containing text and a question text related to this image are given. Further, the question-and-answer device 10 according to this embodiment can generate an answer text considering not only the position and size of text in the image but also visual information such as graphs and photos included in the image (in other words, auxiliary information that helps understand the text).

[0011] Note that, as described above, although it is assumed that an image containing text is given to the question-and-answer device 10, this is not limiting, and this embodiment can be similarly applied when any data containing text is given. Therefore, for example, without depending on formats such as HTML and PDF, any data containing text can be similarly applied. Examples of data containing text include, for example, an HTML document (web page) containing text, a PDF document containing text, a landscape image containing a description, document data, and the like.

[0012] Here, the question-and-answer device 10 according to this embodiment realizes machine reading by a neural network model. Therefore, for the question-and-answer device 10 according to this embodiment, there are a learning time for learning the parameters of this neural network model (hereinafter also referred to as "model parameters") and an inference time for performing machine reading by the neural network model using the learned model parameters. Therefore, hereinafter, the learning time and the inference time of the question-and-answer device 10 will be described.

[0013] [Learning time] First, the learning time will be described. At the learning time of the question-and-answer device 10, a set of training data (training data set) including an image containing text, a question text related to this image, and a correct answer text representing the correct answer to this question text is input.

[0014] <Overall Configuration of the Question-Answering Device 10 During Learning> The overall configuration of the question - answering device 10 during learning will be described with reference to FIG. 1. FIG. 1 is a diagram showing an example of the overall configuration (during learning) of the question - answering device according to the first embodiment.

[0015] As shown in FIG. 1, the question - answering device 10 during learning includes a feature region extraction unit 101, a text recognition unit 102, a text analysis unit 103, a language understanding unit with visual effects 104, an answer text generation unit 105, a parameter learning unit 106, and a parameter storage unit 107.

[0016] The feature region extraction unit 101 extracts a feature region from the input image. The text recognition unit 102 performs text recognition on the feature regions containing text among the feature regions extracted by the feature region extraction unit 101 and outputs the text. The text analysis unit 103 divides the text output by the text recognition unit 102 and the input question text into token sequences respectively. Also, the text analysis unit 103 divides the correct answer text into a token sequence.

[0017] The language understanding unit with visual effects 104 is realized by a neural network and encodes the token sequence obtained by the text analysis unit 103 using the in - learning model parameters stored in the parameter storage unit 107. Thereby, an encoded sequence considering visual information is obtained. That is, language understanding considering the visual effects in the image is obtained.

[0018] The answer text generation unit 105 is realized by a neural network and calculates a probability distribution representing the generation probability of the answer text from the encoded sequence obtained by the language understanding unit with visual effects 104 using the in - learning model parameters stored in the parameter storage unit 107.

[0019] The parameter learning unit 106 updates the in-training model parameters stored in the parameter storage unit 107 using the loss between the answer text generated by the answer text generation unit 105 and the input correct answer text. Thereby, the model parameters are learned.

[0020] The parameter storage unit 107 stores the in-training model parameters (i.e., the model parameters to be learned) of the neural network model that realizes the visual effect-enhanced language understanding unit 104 and the answer text generation unit 105. Note that the in-training model parameters mean the model parameters that have not been learned yet.

[0021] <Learning Process> Next, the learning process according to the present embodiment will be described with reference to FIG. 2. FIG. 2 is a flowchart showing an example of the learning process according to the first embodiment. Hereinafter, as an example, the case of learning the in-training model parameters by the stochastic gradient descent method will be described. However, the in-training model parameters may be learned by other optimization methods other than the stochastic gradient descent method.

[0022] First, the parameter learning unit 106 initializes a variable n e representing the number of epochs to 1 (step S101).

[0023] Next, the parameter learning unit 106 divides the input training data set into mini-batches each containing a maximum of N b training data (step S102). N b is a preset value and can be set to any value. For example, it is conceivable that N b = 60 or the like.

[0024] Next, the question-and-answer device 10 executes model parameter update processing for each mini-batch (step S103). The details of the model parameter update processing will be described later.

[0025] Next, the parameter learning unit 106 sets n e>N e Determine whether it is -1 (step S104). N e is the number of epochs set in advance, and any value can be set. For example, N e = 15 or the like can be considered.

[0026] In step S104 above, n e >N e If it is determined that it is -1, the parameter learning unit 106 ends the learning process. Thereby, the learning of the model parameters stored in the parameter storage unit 107 is completed.

[0027] On the other hand, in step S104 above, if it is determined that n e >N e is not -1, the parameter learning unit 106 adds 1 to n e (step S105) and returns to step S102 above. Thereby, steps S102 and S103 above are repeatedly executed for the number of epochs N e times.

[0028] ≪Model Parameter Update Process≫ Next, the details of the model parameter update process in step S103 above will be described with reference to FIG. 3. FIG. 3 is a flowchart showing an example of the model parameter update process according to the first embodiment. Hereinafter, the model parameter update process for a certain mini-batch will be described.

[0029] First, the parameter learning unit 106 reads one piece of training data in the mini-batch (step S201).

[0030] Next, the feature region extraction unit 101 extracts K feature regions from the images included in the read training data (step S202). A feature region is a region based on visual features, and in this embodiment, it is assumed to be represented by a rectangular region. Also, the k-th feature region is an image token i having position information (a total of 7 dimensions) including the upper left coordinate, the lower right coordinate, the width, the height, and the area, a rectangular image representation (D dimensions), and a region type (C types). k It is assumed to be represented as such. However, any information may be used as long as the position information can specify the position of the feature region (for example, at least one of the width, height, and area information may be absent, or instead of the upper left coordinate and the lower right coordinate, the upper right coordinate and the lower left coordinate may be used, or the center coordinate may be used). Also, either the rectangular image representation or the region type information may be absent. For example, when the feature region is a polygon (polygonal region), the rectangular region surrounding this polygon may be used as the feature region again.

[0031] Here, in this embodiment, as the region type, for example, nine types such as "image", "data (chart)", "paragraph / text", "sub - data", "heading / title", "caption", "sub - title / author name", "list", and "other text" are handled. Also, those other than "image" and "data (chart)" are region types including text. However, these region types are just examples, and other region types may be set. For example, a region type "image information" that combines "image" and "data (chart)" into one may be set, or a region type "text information" that combines "paragraph / text", "sub - data", "heading / title", "caption", "sub - title / author name", "list", and "other text" into one may be set. Thus, at least two types of region types, that is, a region type indicating that no text is included in the feature region and a region type indicating that text is included in the feature region, may be set for the region type.

[0032] An example of the extraction of the feature regions by the feature region extraction unit 101 is shown in FIG. 4. In the example shown in FIG. 4, a case is shown where five feature regions, i.e., a feature region 1100, a feature region 1200, a feature region 1300, a feature region 1400, and a feature region 1500, are extracted from an image 1000 including text. Also, in the example shown in FIG. 4, the region type of the feature region 1100 is "image", the region type of the feature region 1200 is "paragraph / text", the region type of the feature region 1300 is "heading / title", the region type of the feature region 1400 is "list", and the region type of the feature region 1500 is "list".

[0033] Note that for the extraction of such feature regions, for example, Faster R-CNN described in Reference 1 "Shaoqing Ren, Kaiming He, Ross B. Girshick, Jian Sun: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. NIPS 2015: 91-99" can be used. However, as long as it is a method capable of extracting regions based on visual features, other methods (such as object recognition technologies, etc.) can also be used. In addition to this, for example, the feature regions may be extracted manually from the input image (that is, for example, image tokens with the upper left coordinates, lower right coordinates, region type, etc. set manually are created).

[0034] Next, the text recognition unit 102 performs text recognition on the feature regions of the region type indicating that text is included among the feature regions extracted in the above step S202, and outputs the text (step S203). Note that for this text recognition, for example, Tesseract described in Reference 2 "Google: Tesseract Manual. 2018. Internet <URL:https: / / github.com / tesseract-ocr / tesseract / blob / master / doc / tesseract.1.asc>" can be used.

[0035] Next, the text analysis unit 103 divides the text output in step S203 above into a text token sequence (step S204). Hereinafter, assuming that a certain k-th feature region contains text, the text token sequence obtained by dividing this text is

[0036]

Number

[0037] Note that by using the above Byte-level BPE, the text is divided into a subword token sequence. Instead of subword tokens, for example, a sequence of words separated by blanks or the like may be used as the text token sequence.

[0038] Next, in the same manner as in step S204 above, the text analysis unit 103 divides the question text included in the loaded training data into a question token sequence (x1 q , x2 q , ···, x J q ). J is the number of tokens of the question text. Note that the question token sequence is a subword token sequence.

[0039] Next, the question-and-answer device 10 executes a language understanding process with visual effects to obtain an encoded sequence considering visual information (step S206). Here, the details of the language understanding process with visual effects will be described with reference to FIG. 5. FIG. 5 is a flowchart showing an example of the language understanding process with visual effects according to the first embodiment.

[0040] First, the language understanding unit 104 with visual effects uses the image token, the text token sequence, and the question token sequence to create the following input token sequence

[0041]

Number

[0042]

Number

[0043] Hereinafter, the length of the input token sequence is denoted as L. This L is generally adjusted to a predetermined length (for example, L = 512, etc.). If the length of the input token sequence exceeds L, among the texts included in each feature region, the longest text is deleted, or each text is evenly deleted, etc., so that the length L of the input token sequence becomes the predetermined length. On the other hand, if the length L of the input token sequence is less than the predetermined length, it may be padded with special tokens.

[0044] Next, the visual effect-enhanced language understanding unit 104 sets the first token in the input token sequence as the processing target (step S302).

[0045] Next, the visual effect-enhanced language understanding unit 104 determines whether the token set as the processing target is a text token (step S303). Here, a text token refers to a token included in the question token sequence, a token included in the text token sequence, and special tokens such as [CLS], [SEP], [EOS], etc. (that is, sub-word tokens).

[0046] When it is determined in step S303 above that the processing target token is a text token, the visual effect-enhanced language understanding unit 104 encodes the processing target token (step S304). Here, in this embodiment, assuming that the visual effect-enhanced language understanding unit 104 is implemented by a neural network model including BERT (Bidirectional Encoder Representations from Transformers), the visual effect-enhanced language understanding unit 104 encodes the processing target token as follows.

[0047] h = LayerNorm(TokenEmb(x)+PositionEmb(x)+SegmentEmb(x)) x represents the processing target token (that is, the sub-word token), and h represents the processed target token after encoding. For BERT, for example, refer to Reference 4, "Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, 'BERT: Pre-training of Deep Bidirectional Transformers for Language'".

[0048] TokenEmb is a process of converting subword tokens into corresponding G-dimensional vectors by a neural network model. In this embodiment, the pre-trained embedding vectors (G = 1024) according to Reference 5 "Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, Luke Zettlemoyer: BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension, arXiv, 2019." are used as the initial values of the model parameters of the neural network model and are used as the model parameters to be learned. Note that the parameters of pre-trained language models other than Reference 5 may also be used as the objects to be learned.

[0049] PositionEmb is a process of converting, by a neural network model, the position of a token to be processed in an input token sequence into a G-dimensional vector. In this embodiment, the method described in Reference 6 "Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pp. 5998-6008, 2017." is used.

[0050] SegmentEmb is a process that converts a segment of a token sequence to be processed in the input token sequence into a G-dimensional vector. In this embodiment, without distinguishing segments, the vector after conversion is treated as a G-dimensional zero vector. A segment is information for distinguishing the text input to BERT. In this embodiment, since image token i k serves as a segment, SegmentEmb does not distinguish segments. Since SegmentEmb is required in BERT, SegmentEmb is also used in this embodiment. However, when BERT is not used, SegmentEmb is unnecessary.

[0051] LayerNorm takes a G-dimensional vector as input and outputs a G-dimensional vector by the normalization method described in Reference 7 "Jimmy Lei Ba, Jamie Ryan Kiros, Geoffrey E. Hinton: Layer Normalization. Arxiv, 2016."

[0052] On the other hand, when it is determined in step S303 that the token to be processed is not a text token (that is, when the token to be processed is an image token), the visual effect-enhanced language understanding unit 104 encodes the token to be processed as follows (step S305).

[0053] h = LayerNorm(ImgfEmb(i) + LocationEmb(i) + SegmentEmb(i)) i represents the token to be processed (that is, the image token), and h represents the token to be processed after encoding. Also, SegmentEmb and LayerNorm are as described in step S304 above.

[0054] ImgfEmb is a process of converting the rectangular image representation included in the image token from D dimensions to G dimensions by a feed-forward network model composed of a fully connected layer. In this embodiment, assuming a feed-forward network model composed of a single fully connected layer is used, the model parameters of this feed-forward network model are set as the model parameters to be learned.

[0055] LocationEmb is a process of converting the position information included in the image token from 7 dimensions to D dimensions by a feed-forward network model composed of a fully connected layer. In this embodiment, assuming a feed-forward network model composed of a single fully connected layer is used, the model parameters of this feed-forward network model are set as the model parameters to be learned.

[0056] Following the above step S304 or step S305, the visual effect-enhanced language understanding unit 104 determines whether the processing target token is the last token of the input token sequence (step S306).

[0057] If it is determined in the above step S306 that the processing target token is not the last token, the visual effect-enhanced language understanding unit 104 sets the token next to the current processing target token in the input token sequence as the processing target (step S307), and returns to the above step S303. As a result, each token in the input token sequence is encoded, and an encoded sequence H=(h1,h2,···,h L ) is obtained. Note that h r is the encoding of the r-th (r = 1,2,···,L) token in the input token sequence.

[0058] On the other hand, when it is determined in step S306 above that the token to be processed is the final token, the visual effect-enhanced language understanding unit 104 converts the encoded sequence H of the input token sequence into H' by the M-layer Transformer Encoder (step S308). That is, the visual effect-enhanced language understanding unit 104 sets H' = TransformerEncoder(H). For the Transformer Encoder, refer to, for example, the above reference 5. In this embodiment, M = 12, and the Transformer Encoder that has been learned according to the above reference 5 is used as the initial value, and the parameters of this Transformer Encoder are used as the model parameters to be learned.

[0059] Return to FIG. 3. Following step S206, the question-and-answer device 10 executes a generation probability calculation process for the answer text to calculate a probability distribution representing the generation probability of the answer text (step S207). Here, the details of the generation probability calculation process for the answer text will be described with reference to FIG. 6. FIG. 6 is a flowchart showing an example of the generation probability calculation process for the answer text according to the first embodiment.

[0060] First, the text analysis unit 103, in the same manner as in step S204 above, divides the correct answer text included in the loaded training data into a correct answer token sequence

[0061]

Number

[0062] Next, the visual effect-enhanced language understanding unit 104 uses the correct answer token sequence obtained in step S401 above to generate a correct output token sequence

[0063] [Number] Create it (step S402). Thereafter, this correct output token sequence is denoted as Y * .

[0064] Next, the visually enhanced language understanding unit 104 sets t = 2 and sets the leading token in the correct output token sequence Y * as the processing target (the (t - 1)-th processing target token) (step S403).

[0065] Next, the visually enhanced language understanding unit 104 encodes the processing target token as follows in the same manner as step S304 in FIG. 5 (step S404).

[0066] h y* = LayerNorm(TokenEmb(y*) + PositionEmb(y*) + SegmentEmb(y*)) y* is the processing target token (i.e., the subword token), and h y* represents the encoded processing target token.

[0067] Thereafter, the encoding sequence representing the encoding results of the processing target tokens up to the (t - 1)-th is denoted as H y* = (h1 y* , h2 y* , ···, h t-1 y* ).

[0068] Next, the visually enhanced language understanding unit 104 converts the encoding sequence H' obtained in the visually enhanced language understanding process and the encoding sequence H y* = (h1 y* , h2 y* , ···, h t-1 y* ) obtained in step S404 above by the M-layer Transformer Decoder (step S405). That is, the visually enhanced language understanding unit 104 sets h t y = TransformerDecoder(Hy* , let it be H').

[0069] Regarding the Transformer Decoder, for example, refer to the above-mentioned reference 5 etc. In this embodiment, with M = 12 and the Transformer Decoder that has been learned according to the above-mentioned reference 5 as the initial value, the parameters of this Transformer Decoder are used as the model parameters to be learned. Note that the parameters of a pre-trained language model other than reference 5 may also be used as the object of learning.

[0070] Next, the answer text generation unit 105 calculates the probability distribution p(y t |y <t *) of the t-th word being generated (step S406). The probability distribution of the word y t in the preset output vocabulary (vocabulary size V) is calculated by p(y t |y <t *) = softmax(Wh t y + b). Here, W ∈ R V×G , b ∈ R V are the model parameters to be learned. Note that as the vocabulary size V, any value can be set, but for example, it can be considered to be 50257 etc.

[0071] Next, the visual effect - enhanced language understanding unit 104 determines whether the correct word y t * is the last word (step S407). The last word refers to the word at the t = L T -th position or the special word [EOS] indicating the end of the sentence. Note that if the correct word y t * is a word indicating the end of the sentence and t < L T , the answer text generation unit 105 pads from t + 1 to L T with special words.

[0072] In step S407 described above, for the correct word y t *When it is determined that it is not the last word, the visually enhanced language understanding unit 104 adds 1 to t (step S408) and returns to step S404. As a result, steps S404 to S406 described above are repeatedly executed.

[0073] Return to Figure 3. Following step S207, the parameter learning unit 106 calculates the loss Loss using the probability distribution p(y t |y <t *) calculated in step S406 of Figure 6 (step S208). The parameter learning unit 106 may calculate the loss Loss, for example, as follows.

[0074]

Equation

[0075] Next, the parameter learning unit 106 determines whether all the training data in the mini-batch has been read (step S209).

[0076] When it is determined in step S209 described above that there is still unread training data in the mini-batch, the parameter learning unit 106 reads one piece of unread training data (step S210) and returns to step S202 described above. As a result, steps S202 to S208 described above are repeatedly executed for each training data included in the mini-batch.

[0077] On the other hand, when it is determined in step S209 above that all the training data during the mini-batch has been read, the parameter learning unit 106 updates the model parameters using the loss Loss calculated in step S208 above for each training data (step S211). That is, the model parameters are updated by a known optimization method so that the loss Loss is minimized.

[0078] As described above, in the question-and-answer device 10 according to the present embodiment, when an image including text and a question text related to this image are given, the model parameters are learned so that an answer text considering the visual information in this image is generated. That is, the model parameters are learned so that machine reading comprehension considering visual information is possible.

[0079] [During inference] Next, the inference process will be described. During inference, test data including an image containing text and a question text related to this image is input to the question-and-answer device 10.

[0080] <Overall configuration of the question-and-answer device 10 during inference> The overall configuration of the question-and-answer device 10 during inference will be described with reference to FIG. 7. FIG. 7 is a diagram showing an example of the overall configuration (during inference) of the question-and-answer device according to the first embodiment.

[0081] As shown in FIG. 7, the question-and-answer device 10 during inference includes a feature region extraction unit 101, a text recognition unit 102, a text analysis unit 103, a vision-augmented language understanding unit 104, an answer text generation unit 105, and a parameter storage unit 107. Among these units, the feature region extraction unit 101, the text recognition unit 102, and the text analysis unit 103 are the same as those during learning. On the other hand, the vision-augmented language understanding unit 104 and the answer text generation unit 105 use the learned model parameters stored in the parameter storage unit 107. Also, the answer text generation unit 105 generates an answer text using the probability distribution calculated from the encoded sequence obtained by the vision-augmented language understanding unit 104.

[0082] <Inference processing> Next, the inference processing according to the present embodiment will be described with reference to FIG. 8. FIG. 8 is a flowchart showing an example of the inference processing according to the first embodiment. Hereinafter, it is assumed that the test data given to the question-and-answer device 10 has been read.

[0083] First, the feature region extraction unit 101 extracts K feature regions from the image included in the read test data in the same manner as in step S202 of FIG. 3 (step S501).

[0084] Next, the text recognition unit 102 performs text recognition on the feature regions of the region type indicating that text is included among the feature regions extracted in step S501 above, in the same manner as in step S203 of FIG. 3, and outputs the text (step S502).

[0085] Next, the text analysis unit 103 divides the text output in step S502 above into a text token sequence in the same manner as in step S204 of FIG. 3 (step S503).

[0086] Next, the text analysis unit 103 divides the question text included in the read test data into a question token sequence in the same manner as in step S205 of FIG. 3 (step S504).

[0087] Next, the question-and-answer device 10 executes a language understanding process with visual effects to obtain an encoded sequence in which visual information is considered (step S505). Since the language process with visual effects is the same as step S206 of FIG. 3, the description thereof is omitted. Hereinafter, the description will continue assuming that the encoded sequence H' has been obtained.

[0088] Next, the question-and-answer device 10 executes an answer text generation process to generate an answer text (step S506). Here, the details of the answer text generation process will be described with reference to FIG. 9. FIG. 9 is a flowchart showing an example of the answer text generation process according to the first embodiment.

[0089] First, the answer text generation unit 105 sets the first token of the output token sequence as [CLS] (step S601). At this point, the tokens included in the output token sequence are only [CLS].

[0090] Next, the vision effect-enhanced language understanding unit 104 sets t = 2 and sets the first token of the output token sequence as the processing target (the (t - 1)-th processing target token) (step S602).

[0091] Next, the vision effect-enhanced language understanding unit 104 encodes the processing target token as follows in the same way as step S404 in FIG. 6 (step S603).

[0092] h y =LayerNorm(TokenEmb(y)+PositionEmb(y)+SegmentEmb(y)) y is the processing target token (i.e., the sub-word token), and h y represents the processed target token after encoding.

[0093] Hereafter, the encoded sequence H y =(h1 y ,h2 y ,···,h t-1 y ) represents the encoded results of the processing target tokens up to the (t - 1)-th one.

[0094] Next, the vision effect-enhanced language understanding unit 104 converts the encoded sequence H' obtained by the vision effect-enhanced language understanding process and the encoded sequence H obtained in step S603 above by the M-layer Transformer Decoder (step S604). That is, the vision effect-enhanced language understanding unit 104 sets h y t y y =TransformerDecoder(H y ,H'). Thereby, H y '= (h1 y ,h2 y, ···, h t-1 y , h t y ) is obtained.

[0095] Next, the answer text generation unit 105 calculates the probability distribution p(y t |y <t ) (step S605). The probability distribution of the word y in the preset output vocabulary (vocabulary size V) is p(y t |y t |y <t ) = softmax(Wh t y + b). Here, W ∈ R V×G , b ∈ R V are learned model parameters.

[0096] Next, the answer text generation unit 105 generates the t-th word based on the probability distribution p(y t |y <t ) calculated in step S605 above (step S606). The answer text generation unit 105 may generate the word with the maximum probability as the t-th word, or may generate the t-th word by sampling according to the probability distribution.

[0097] Next, the answer text generation unit 105 concatenates the t-th word generated in step S606 above to the end of the output token sequence (step S607).

[0098] Next, the visual effect-augmented language understanding unit 104 determines whether the t-th word generated in step S606 above is the last word (step S608). The last word refers to the special word [EOS] indicating the end of the sentence.

[0099] If it is determined in step S608 above that the t-th word is not the last word, the visual effect-augmented language understanding unit 104 adds 1 to t (step S609) and returns to step S603. As a result, steps S603 to S607 above are repeatedly executed to obtain a word sequence.

[0100] As described above, in the question-and-answer device 10 according to this embodiment, when an image including text and a question text related to this image are given, an answer text (word sequence) considering the visual information in this image can be generated.

[0101] ·Second Embodiment In this embodiment, a case will be described in which the answer text is generated in consideration of whether the feature region extracted by the feature region extraction unit 101 is information necessary for answering the question.

[0102] In this embodiment, mainly, the differences from the first embodiment will be described, and the description of the same components as those in the first embodiment will be omitted.

[0103] [During learning] First, the learning process will be described. The training data input to the question-and-answer device 10 during learning is assumed to include, in addition to an image including text, a question text, and a correct answer, a set of correct feature regions. The set of correct feature regions is a set of feature regions necessary for obtaining the correct answer among the feature regions extracted from the image.

[0104] <Overall configuration of the question-and-answer device 10 during learning> The overall configuration of the question-and-answer device 10 during learning will be described with reference to FIG. 10. FIG. 10 is a diagram showing an example of the overall configuration (during learning) of the question-and-answer device 10 according to the second embodiment.

[0105] As shown in FIG. 10, the question-and-answer device 10 during learning includes a feature region extraction unit 101, a text recognition unit 102, a text analysis unit 103, a vision-augmented language understanding unit 104, an answer text generation unit 105, a parameter learning unit 106, a relevant feature region determination unit 108, and a parameter storage unit 107. In the second embodiment, mainly, the difference from the first embodiment is that the question-and-answer device 10 has a relevant feature region determination unit 108.

[0106] The related feature region determination unit 108 is implemented by a neural network, and uses the in-training model parameters stored in the parameter storage unit 107 to calculate the probability indicating whether the feature region extracted by the feature region extraction unit 101 is information necessary for answering the question. Therefore, the in-training model parameters stored in the parameter storage unit 107 also include the in-training model parameters of the neural network model that implements the related feature region determination unit 108.

[0107] Also, the parameter learning unit 106 calculates a loss using the probability calculated by the related feature region determination unit 108 and the set of correct feature regions, and updates the in-training model parameters stored in the parameter storage unit 107.

[0108] <Learning Process> Next, the learning process according to the present embodiment will be described. Since the overall flow of the learning process may be the same as the learning process described with reference to FIG. 2, hereinafter, the details of the model parameter update process in step S103 of FIG. 2 will be described. However, the number of epochs N e and the maximum number of training data N included in the mini-batch b may be different from those in the first embodiment. For example, the maximum number of training data N b included in the mini-batch may be set to N b = 32 or the like.

[0109] ≪Model Parameter Update Process≫ Details of the model parameter update process in step S103 of FIG. 2 will be described with reference to FIG. 11. FIG. 11 is a flowchart showing an example of the model parameter update process according to the second embodiment. Hereinafter, the model parameter update process for a certain mini-batch will be described.

[0110] First, the parameter learning unit 106 reads one piece of training data in the mini-batch (step S701).

[0111] Next, the feature region extraction unit 101 extracts K feature regions from the images included in the read training data (step S702). In the present embodiment, as in the first embodiment, assuming that the feature region is represented by a rectangular region, the k-th feature region has position information (a total of 4 dimensions) including the upper left coordinates and the lower right coordinates, a rectangular image representation (D dimensions), and a region type (C types). However, any information may be used as long as the position information can specify the position of the feature region, and either the rectangular image representation or the region type information may be absent. In addition, as in the first embodiment, for example, when the feature region is a polygon (polygonal region), the rectangular region surrounding this polygon may be re-set as the feature region.

[0112] Also, the region type shall handle 9 types similar to those in the first embodiment. However, it goes without saying that these 9 types of region types are just an example, and other region types may be set. Also in this embodiment, as in the first embodiment, at least two types of region types are required to be set: a region type indicating that the feature region does not contain text, and a region type indicating that the feature region contains text.

[0113] Note that for the extraction of the feature region, as in the first embodiment, for example, Faster R-CNN described in the above reference 1 may be used. Also in this embodiment, for example, D = 2048 or the like is set.

[0114] Next, the text recognition unit 102 performs text recognition on the feature regions of the region type indicating that text is included among the feature regions extracted in the above step S702, and outputs a word region series composed of word regions that are regions including the words that are the results of the text recognition (step S703). Hereinafter, each word region is assumed to be a rectangular region, and has position information (a total of 4 dimensions) including the upper left coordinate and the lower right coordinate of the word region, and the word obtained by text recognition. For text recognition, as in the first embodiment, for example, Tesseract or the like described in the above reference 2 may be used. Note that the word region is a partial region of the feature region that includes the word that is the result of the text recognition.

[0115] Next, the feature region extraction unit 101 outputs a rectangular image representation (D dimensions) of each word region obtained in the above step S704 for each word region (step S704). Note that this rectangular image representation may be output in the same manner as when the rectangular image representation of the feature region was obtained in the above step S702. As a result, each word region will have position information (a total of 4 dimensions) including the upper left coordinate and the lower right coordinate of the word region, the word obtained by text recognition, and the rectangular image representation (D dimensions) of the word region.

[0116] Next, the text analysis unit 103 divides the word region series obtained in the above step S704 into a sub-word token series (step S705). Hereinafter, the sub-word token series obtained by dividing the word region series obtained from a certain k-th feature region is represented as

[0117]

Number

[0118] In addition, when the words in one word region are divided into a plurality of sub-words, the word regions of each sub-word shall be the same as the word region of the word before division.

[0119] Next, the text analysis unit 103 divides the question text included in the read training data into a sub-word token sequence (x1 q , x2 q , ···, x J q ). J is the number of sub-word tokens of the question text.

[0120] Next, the question answering device 10 executes a language understanding process with visual effects to obtain an encoded sequence considering visual information (step S707). Here, the details of the language understanding process with visual effects will be described with reference to FIG. 12. FIG. 12 is a flowchart showing an example of the language understanding process with visual effects according to the second embodiment.

[0121] First, the language understanding unit 104 with visual effects uses the sub-word token sequence of the word region sequence and the sub-word token sequence of the question text to generate the following input token sequence

[0122]

Number

[0123] In addition, when the k-th feature region does not contain text, the length of the sub-word token sequence of the word region sequence obtained from the k-th feature region is 0 (that is, L k = 0).

[0124] Hereinafter, similar to the first embodiment, the length of the input token sequence is set to L. If the length of the input token sequence exceeds L, among the texts included in each feature region, the longest text is deleted, or each text is evenly deleted, etc., so that the length L of the input token sequence becomes a predetermined length. On the other hand, when the length L of the input token sequence is less than the predetermined length, it may be padded with a special token.

[0125] Next, the visual effect-enhanced language understanding unit 104 encodes each token (sub-word token) in the input token sequence (step S802). Here, in this embodiment, the visual effect-enhanced language understanding unit 104 encodes each token x as follows.

[0126] h = LayerNorm(TokenEmb(x)+PositionEmb(x)+SegmentEmb(x)+ROIEmb(x)+LocationEmb(x)) TokenEmb is a process of converting a sub-word token (including a special token) into a corresponding G-dimensional vector. In this embodiment, similar to the first embodiment, the pre-trained embedding vector (G = 1024) according to the above reference 5 is used as the initial value and is used as the model parameters to be learned. Note that the parameters of a pre-trained language model other than reference 5 may also be used as the learning target. However, for special tokens that have not been learned, they are initialized with random numbers following a normal distribution N(0, 0.02).

[0127] PositionEmb is a process of converting into a G-dimensional vector according to the positions of sub-word tokens in the input token sequence. In this embodiment, the embedding vector (G = 1024) learned according to the above reference 5 is used as the initial value and is used as the model parameters to be learned. However, similar to the first embodiment, the method described in the above reference 6 may be used to convert into a G-dimensional vector.

[0128] SegmentEmb is a process of converting into a G-dimensional vector according to the segment to which the sub-word token belongs. In this embodiment, 9 types of region types and 10 types in total including questions are used as segments. Then, after preparing embedding vectors (G = 1024) for each segment, they are initialized with random numbers following the normal distribution N(0, 0.02) and are used as the model parameters to be learned.

[0129] ROIEmb is a process of converting from a rectangular image representation corresponding to a sub-word token into a G-dimensional vector. The rectangular image representation is a D-dimensional vector obtained by inputting a certain rectangular region in the input image into the neural network that realizes the feature region extraction unit 101. When the sub-word token is the region token i k it is the rectangular image representation of the k-th feature region, and when the sub-word token is the document token x j k it is the rectangular image representation of the i-th word region obtained from the k-th feature region. On the other hand, when the sub-word token is the document token x j k it is assumed that the output of ROIEmb is a G-dimensional zero vector. In this embodiment, with D = 2048, ROIEmb is converted into a G-dimensional (G = 1024) vector by a feed-forward network composed of a fully connected layer. Also, in this embodiment, the feed-forward network is composed of one layer of fully connected layer, and its parameters are initialized with random numbers following the normal distribution N(0, 0.02) and are used as the model parameters to be learned.

[0130] LocationEmb is a process that converts the position information of a region (feature region or word region) corresponding to a subword token (either a region token or a document token) into a vector of G dimensions (G = 1024) from 4 dimensions by a feed-forward network composed of a fully connected layer. LocationEmb normalizes the x coordinate of the position information of the region by dividing it by the width of the input image, and normalizes the y coordinate of the position information of the region by dividing it by the height of the image, and then inputs it to the feed-forward network. In this embodiment, the feed-forward network is composed of a single fully connected layer, and its parameters are initialized with random numbers following a normal distribution N(0, 0.02) and used as the model parameters to be learned. Note that when the subword token is other than the region token and the document token, the output of LocationEmb is assumed to be a G-dimensional zero vector.

[0131] Similar to the first embodiment, LayerNorm takes a G-dimensional vector as input and outputs a G-dimensional vector by the normalization method described in the above reference 7.

[0132] As described above, if the encoded r-th subword token in the input token sequence is represented as h r , the encoded sequence H=(h1, h2, ···, h L ) can be obtained. Since each h r is a G-dimensional vector, H is a sequence of vectors.

[0133] Next, the language understanding unit 104 with visual effects converts the encoded sequence H obtained in step S802 above into a vector sequence H' using an M-layer Transformer Encoder (step S803). That is, the language understanding unit 104 with visual effects sets H' = TransformerEncoder(H). For the Transformer Encoder, refer to, for example, the above reference 5. In this embodiment, M = 12, and the Transformer Encoder that has been learned according to the above reference 5 is used as the initial value, and the parameters of this Transformer Encoder are used as the model parameters to be learned.

[0134] Next, the relevant feature region determination unit 108 calculates the probability indicating whether the feature region is a region necessary for answer generation (step S804). That is, if the element of H' corresponding to the sub-word token x (where x is either a region token or a document token) in the input token sequence is h', the relevant feature region determination unit 108 calculates the probability that the feature region corresponding to the sub-word token x is necessary for the correct answer as follows.

[0135] p = sigmoid(w1 τ h'+b1) Here, w1 ∈ R G , b1 ∈ R are the model parameters to be learned, and τ represents transpose.

[0136] Next, the relevant feature region determination unit 108 uses the probability obtained in step S804 above to convert the vector sequence H' into a vector sequence H'' (step S805). That is, the relevant feature region determination unit 108 converts the vector sequence H' into a vector sequence H'' using h r '' = h r 'a r . Here, h r '' is the r-th element of the vector sequence H'', and h r ' is the r-th element of the vector sequence H'. Also, a r is a weight, and its value is such that the r-th sub-word token in the input token sequence is the region token i kor document token x j k If p k、 In other cases, it is set to 1.0. k is the region token i k is the probability calculated in step S804 above.

[0137] Returning to Fig. 11. Following step S707, the question answering device 10 executes an answer text generation probability calculation process to calculate a probability distribution representing the answer text generation probability (step S708). Here, the answer text generation probability calculation process will be described in detail with reference to Fig. 13. Fig. 13 is a flowchart showing an example of an answer text generation probability calculation process according to the second embodiment.

[0138] First, similarly to step S401 in FIG. 6, the text analysis unit 103 converts the correct answer text included in the loaded training data into a correct answer token sequence

[0139]

number

[0140] Next, the visual effect added language understanding unit 104 sets an index indicating the number of repetitions as t and initializes t to 0 (step S902). In the following, the process at the tth repetition will be described.

[0141] The visual effect language understanding unit 104 calculates the decoder input token sequence y <t is created (step S903).

[0142] y <t=([CLS], y1 * , ···, y t-1 * ) = (y0, y1, ···, y t-1 ) However, when t = 0, let y <t = ([CLS]). Also, at the final step when t = L T + 1, let y t = [EOS].

[0143] Next, the visual effect - enhanced language understanding unit 104 encodes each sub - word token y included in the decoder input token sequence y <t as follows (step S904).

[0144] h y = LayerNorm(TokenEmb(y) + PositionEmb(y)) Thus, if the encoded form of the sub - word token y t is h t y , then the encoded sequence H y = (h0 y , h1 y , ···, h t-1 y ) can be obtained.

[0145] Next, the visual effect - enhanced language understanding unit 104 converts the encoded sequence H y obtained in step S904 above into H y ' using an M - layer Transformer Decoder (step S905). That is, the visual effect - enhanced language understanding unit 104 sets H y ' = TransformerDecoder(H y , H''). Thus, H y ' = (h0 y , h1 y , ···, h t-1 y ) can be obtained.

[0146] Regarding the Transformer Decoder, for example, refer to the above-mentioned Reference 5 etc. In this embodiment, with M = 12, the Transformer Decoder pre-trained according to the above-mentioned Reference 5 is used as the initial value, and the parameters of this Transformer Decoder are used as the model parameters to be learned. Note that the parameters of a pre-trained language model other than Reference 5 may also be used as the learning target.

[0147] Next, the answer text generation unit 105 calculates the probability distribution p(y t |y <t ) (step S906). The probability distribution of the word y t in the preset output vocabulary (vocabulary size V) is calculated by p(y t |y <t ) = softmax(Wh t-1 y '+b). Here, W ∈ R V×G , b ∈ R V are the model parameters to be learned. Note that any value can be set for the vocabulary size V, and for example, it can be considered to be 50257 etc.

[0148] Next, the vision effect enhanced language understanding unit 104 determines whether t = L T +1 (step S907).

[0149] If it is determined in step S907 above that t is not equal to L T +1, the vision effect enhanced language understanding unit 104 adds 1 to t (step S908) and returns to step S903. As a result, the above steps S903 to S906 are repeatedly executed for t = 0, 1, ···, L T +1.

[0150] Returning to FIG. 11. Following step S708, the parameter learning unit 106 uses the probability distribution p(y t |y <tUsing the read training data and the set of correct feature regions included therein, calculate the loss Loss (step S709). The parameter learning unit 106 may calculate the loss Loss, for example, as follows.

[0151]

Number

[0152] Since the subsequent steps S710 to S712 are the same as steps S209 to S211 in FIG. 3, the description thereof is omitted.

[0153] As described above, in the question-and-answer device 10 according to this embodiment, when an image including text, a question text related to this image, and a set of correct feature regions are given, the model parameters are learned so that an answer text considering the visual information in this image is generated. That is, the model parameters are learned so that machine reading comprehension considering visual information is possible.

[0154] [During inference] Next, the inference process will be described. During inference, test data including an image including text and a question text related to this image is input to the question-and-answer device 10.

[0155] <Overall configuration of the question-and-answer device 10 during inference> The overall configuration of the question-and-answer device 10 during inference will be described with reference to FIG. 14. FIG. 14 is a diagram showing an example of the overall configuration (during inference) of the question-and-answer device 10 according to the second embodiment.

[0156] As shown in FIG. 14, at the time of inference, the question-and-answer device 10 includes a feature region extraction unit 101, a text recognition unit 102, a text analysis unit 103, a language understanding unit with visual effects 104, an answer text generation unit 105, a related feature region determination unit 108, and a parameter storage unit 107. Among these units, the feature region extraction unit 101, the text recognition unit 102, and the text analysis unit 103 are the same as those during learning. On the other hand, the language understanding unit with visual effects 104, the answer text generation unit 105, and the related feature region determination unit 108 use the learned model parameters stored in the parameter storage unit 107. Further, the answer text generation unit 105 generates an answer text using the probability distribution calculated from the encoded sequence obtained by the language understanding unit with visual effects 104. Note that the related feature region determination unit 108 may output a score (related feature region score) calculated or determined from the probability indicating whether the feature region extracted by the feature region extraction unit 101 is information necessary to answer the question.

[0157] <Inference process> Next, the inference process according to the present embodiment will be described with reference to FIG. 15. FIG. 15 is a flowchart showing an example of the inference process according to the second embodiment. Hereinafter, it is assumed that the test data given to the question-and-answer device 10 has been read.

[0158] First, the feature region extraction unit 101 extracts K feature regions from the image included in the read test data in the same manner as in step S702 of FIG. 11 (step S1001).

[0159] Next, the text recognition unit 102 performs text recognition on the feature regions of the region type indicating that text is included among the feature regions extracted in step S1001 above, in the same manner as in step S703 of FIG. 11, and outputs a word region sequence (step S1002).

[0160] Next, in the same manner as step S704 in FIG. 11, the feature region extraction unit 101 outputs the rectangular image representation (D dimensions) of each word region obtained in step S1002 above for each word region (step S1003). As a result, a series of word regions having position information (a total of 4 dimensions) including the upper left coordinate and the lower right coordinate, the word obtained by text recognition, and the rectangular image representation (D dimensions) is obtained.

[0161] Next, in the same manner as step S705 in FIG. 11, the text analysis unit 103 divides the word region series obtained in step S1003 above into a sub-word token series (step S1004).

[0162] Next, in the same manner as step S706 in FIG. 11, the text analysis unit 103 divides the question text included in the read test data into a sub-word token series (x1 q , x2 q , ···, x J q ). (Step S1005).

[0163] Next, the question answering device 10 executes a language understanding process with visual effects to obtain an encoded series in which visual information is considered (step S1006). Since the language understanding process with visual effects is the same as step S707 in FIG. 11, its description is omitted. Hereinafter, the description will continue on the assumption that the vector series H'' has been obtained.

[0164] Next, the question answering device 10 executes an answer text generation process to generate an answer text (step S1007). Here, the details of the answer text generation process will be described with reference to FIG. 16. FIG. 16 is a flowchart showing an example of the answer text generation process according to the second embodiment.

[0165] First, the language understanding unit 104 with visual effects initializes t, which represents the number of repetitions, to 0 (step S1101). Hereinafter, the processing at the t-th repetition will be described.

[0166] The language understanding unit 104 with visual effects initializes the decoder input token sequence as y <t = ([CLS]) (step S1102). That is, the language understanding unit 104 with visual effects sets the decoder input token sequence y <t at t = 0 to a sequence containing only [CLS].

[0167] Hereinafter, the processing at a certain t-th iteration will be described.

[0168] The language understanding unit 104 with visual effects encodes each sub-word token y included in the decoder input token sequence y <t in the same way as in step S904 of FIG. 13 (step S1103).

[0169] h y = LayerNorm(TokenEmb(y)+PositionEmb(y)) Thus, if the encoded sub-word token y t is h t y , the encoded sequence H y =(h0 y ,h1 y ,···,h t-1 y ) can be obtained.

[0170] Next, the language understanding unit 104 with visual effects, in the same way as in step S905 of FIG. 13, converts the encoded sequence H y obtained in step S1103 above into H y ' using the M-layer Transformer Decoder (step S1104). That is, the language understanding unit 104 with visual effects sets H y ' = TransformerDecoder(H y ,H''). Thus, H y '=(h0 y ',h1 y ',···,h t-1 y ) can be obtained.

[0171] Next, the response text generation unit 105 calculates the probability distribution p(y t |y <t ) in the same manner as step S906 in FIG. 13 (step S1105). The probability distribution of the word y t in the preset output vocabulary (vocabulary size V) is calculated as p(y t |y <t ) = softmax(Wh t-1 y '+b). Here, W ∈ R V×G , b ∈ R V are learned model parameters.

[0172] Next, the response text generation unit 105 generates the t-th word based on the probability distribution p(y t |y <t ) calculated in step S1105 above (step S1106). The response text generation unit 105 may generate the word with the maximum probability as the t-th word, or may generate the t-th word by sampling according to the probability distribution.

[0173] Next, the response text generation unit 105 concatenates the t-th word generated in step S1106 above to the end of the decoder input token sequence y <t (step S1107).

[0174] Next, the vision effect-enhanced language understanding unit 104 determines whether the t-th word generated in step S1106 above is the last word (step S1108).

[0175] If it is determined in step S1108 above that the t-th word is not the last word, the vision effect-enhanced language understanding unit 104 adds 1 to t (step S1109) and returns to step S1103. As a result, steps S1103 to S1107 above are repeatedly executed to obtain a word sequence.

[0176] As described above, in the question-and-answer device 10 according to the present embodiment, when an image including text and question text related to this image are given, an answer text (word sequence) considering the visual information in this image can be generated.

[0177] [Evaluation of the Present Embodiment] Next, an evaluation regarding considering whether the feature region is information necessary for answering a question will be described.

[0178] To evaluate the present embodiment, a baseline and performance comparison were performed. As models of the present embodiment, a model using BART described in the above Reference 5 as a pre-trained model and a model using T5 described in Reference 8 "Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. J. Mach. Learn. Res. 21(140): 1-67." as a pre-trained model were used. Hereinafter, the model using BART is referred to as "LayoutBART", and the model using T5 is referred to as "LayoutT5". In particular, those using LARGE as BERT are respectively denoted as "LayoutBART LARGE ", "LayoutT5 LARGE ".

[0179] In addition, as the baseline, the model called M4C described in Reference 9 "Hu, R.; Singh, A.; Darrell, T.; and Rohrbach, M. 2020. Iterative Answer Prediction with Pointer-Augmented Multi-modal Transformers for TextVQA. In CVPR, 9992-10002." was adopted. M4C is a model that generates an answer to a question by taking as input the tokens of the question text, the feature region, and the OCR tokens (corresponding to the document tokens in this embodiment), and it has been confirmed that high performance can be achieved.

[0180] As evaluation metrics, five metrics were used: BLEU, METEOR, ROUGE-L, CIDEr, and BERTscore. After training the model using a pre-prepared training dataset for experiments, the above four evaluation metrics were calculated using the test data. The results are shown in Table 1 below.

[0181]

Table 1

[0182] <Hardware Configuration> Finally, the hardware configuration of the question answering device 10 according to the first and second embodiments will be described with reference to FIG. 17. FIG. 17 is a diagram showing an example of the hardware configuration of the question answering device 10 according to an embodiment.

[0183] As shown in FIG. 17, the question-and-answer device 10 according to one embodiment is realized by a general computer or computer system, and includes an input device 201, a display device 202, an external I / F 203, a communication I / F 204, a processor 205, and a memory device 206. Each of these hardware components is communicably connected via a bus 207.

[0184] The input device 201 is, for example, a keyboard, a mouse, a touch panel, or the like. The display device 202 is, for example, a display or the like. Note that the question-and-answer device 10 may not have at least one of the input device 201 and the display device 202.

[0185] The external I / F 203 is an interface with an external device. Examples of the external device include a recording medium 203a. The question-and-answer device 10 can read from and write to the recording medium 203a via the external I / F 203. The recording medium 203a may store one or more programs that implement each functional unit (feature region extraction unit 101, text recognition unit 102, text analysis unit 103, visually enhanced language understanding unit 104, answer text generation unit 105, parameter learning unit 106, and related feature region determination unit 108) of the question-and-answer device 10.

[0186] Examples of the recording medium 203a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), a USB (Universal Serial Bus) memory card, and the like.

[0187] The communication I / F 204 is an interface for connecting the question-and-answer device 10 to a communication network. Note that one or more programs that implement each functional unit of the question-and-answer device 10 may be acquired (downloaded) from a predetermined server device or the like via the communication I / F 204.

[0188] The processor 205 is various arithmetic units such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), for example. Each functional unit of the question-and-answer device 10 is realized by processing in which one or more programs stored in the memory device 206 are executed by the processor 205, for example.

[0189] The memory device 206 is various storage devices such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), a RAM (Random Access Memory), a ROM (Read Only Memory), a flash memory, etc., for example. The parameter storage unit 107 included in the question-and-answer device 10 can be realized using the memory device 206, for example. Note that the parameter storage unit 107 may be realized using a storage device (such as a database server, etc.) connected to the question-and-answer device 10 via a communication network.

[0190] The question-and-answer device 10 according to the first and second embodiments can realize the above-described learning process and inference process by having the hardware configuration shown in FIG. 17. Note that the hardware configuration shown in FIG. 17 is an example, and the question-and-answer device 10 may have other hardware configurations. For example, the question-and-answer device 10 may have a plurality of processors 205, or may have a plurality of memory devices 206.

[0191] Regarding the above embodiments, the following supplementary notes are further disclosed.

[0192] (Supplementary Note 1) A memory, At least one processor connected to the memory, Including, Taking data including a visual area and first information related to the data as input, using model parameters of a machine learning model, generating second information corresponding to the first information from information representing the characteristics of the area, A learning device that learns the model parameters based on the second information and third information representing the correct answer of the second information. (Appendix 2) The processor The learning device according to Appendix 1, which creates a feature amount between information representing the feature of the region and the first information, and generates the second information from the feature amount. (Appendix 3) The learning device according to Appendix 1 or 2, wherein the region at least includes an image or a chart. (Appendix 4) The learning device according to any one of Appendices 1 to 3, wherein the first information is text information representing the content related to the data. (Appendix 5) A memory, At least one processor connected to the memory, Including, Using the data including a visual region and the first information related to the data as inputs, and using the model parameters of a machine learning model, calculates the degree of association between the region and the second information corresponding to the first information. A learning device that learns the model parameters based on the degree of association and information representing the correct answer of the degree of association. (Appendix 6) A memory, At least one processor connected to the memory, Including, Using the data including a visual region and the first information related to the data as inputs, and using the model parameters of a learned machine learning model, generates the second information corresponding to the first information from the information representing the feature of the region. A generating device. (Appendix 7) A memory, At least one processor connected to the memory, Including, An output device that takes as input data including a visual area and first information related to the data, and uses model parameters of a learned machine learning model to output a predetermined evaluation value based on the degree of association between the area and second information corresponding to the first information with respect to the area. (Appendix 8) A non-transitory storage medium storing a program executable by a computer to execute a learning process, wherein the learning process takes as input data including a visual area and first information related to the data, and uses model parameters of a machine learning model to generate second information corresponding to the first information from information representing features of the area, and learns the model parameters based on the second information and third information representing the correct answer of the second information. (Appendix 9) A non-transitory storage medium storing a program executable by a computer to execute a generation process, wherein the generation process takes as input data including a visual area and first information related to the data, and uses model parameters of a learned machine learning model to generate second information corresponding to the first information from information representing features of the area.

[0193] The present invention is not limited to the specifically disclosed above embodiments, and various modifications, changes, combinations with known technologies, etc. are possible without departing from the description of the claims.

[0194] This application is based on a basic application PCT / JP2020 / 008390 filed in Japan on February 28, 2020, the entire contents of which are incorporated herein by reference.

Description of Reference Numerals

[0195] 10 Question-and-Answer Device 101 Feature Region Extraction Unit 102 Text Recognition Unit 103 Text analysis unit 104 Visual effect - added language understanding unit 105 Response text generation unit 106 Parameter learning unit 107 Parameter memory unit 108 Related feature region determination unit

Claims

1. A generation unit that generates second information corresponding to the first information using model parameters of a machine learning model, with data including at least a visual area containing an image area and a character area, and the first information related to the data as inputs; A learning unit that learns the model parameters based on the second information and third information representing the correct answer of the second information; It has, The generation unit, When the area extracted from the data contains text, generates text tokens based on the information representing the text, Generates image tokens for the area extracted from the data, Generates the second information based on a token sequence including the generated image tokens. A learning device characterized by this.

2. A generation unit that generates second information corresponding to the first information using model parameters of a machine learning model, with image data including text and the first information related to the image data as inputs; A learning unit that learns the model parameters based on the second information and third information representing the correct answer of the second information; It has, The generation unit, When the area extracted from the image data contains text, generates text tokens based on the information representing the text, Generates image tokens for the area extracted from the image data, Generates the second information based on a token sequence including the generated image tokens. A learning device characterized by this.

3. The learning device according to claim 1, wherein the area further includes at least a chart area.

4. The learning device according to claim 1 or 3, wherein the first information is text information representing the content related to the data.

5. It has a calculation unit that calculates the degree of association between the area and the second information corresponding to the first information using the model parameters, with the data and the first information as inputs. The learning unit, Learns the model parameters based on the degree of association and information representing the correct answer of the degree of association. The learning device according to claim 1, 3 or 4.

6. A generation unit that generates second information corresponding to the first information using model parameters of a learned machine learning model, with data including at least a visual area containing an image area and a character area, and the first information related to the data as inputs, It has, The generation unit, When the area extracted from the data contains text, text tokens are generated based on the information representing the text, image tokens are generated for the area extracted from the data, and the second information is generated based on a token sequence including the generated image tokens. A generation device characterized by the above.

7. The generation device according to claim 6, further comprising an output unit that outputs a predetermined evaluation value based on the degree of association between the area and the second information corresponding to the first information with respect to the area.

8. Using the model parameters of a machine learning model, a generation procedure for generating second information corresponding to the first information, with data including a visual area including at least an image area and a character area and first information related to the data as inputs, a learning procedure for learning the model parameters based on the second information and third information representing the correct answer of the second information, wherein the computer executes, the generation procedure includes, when the area extracted from the data contains text, text tokens are generated based on the information representing the text, image tokens are generated for the area extracted from the data, and the second information is generated based on a token sequence including the generated image tokens. A learning method characterized by the above.

9. Using the model parameters of a machine learning model, a generation procedure for generating second information corresponding to the first information, with image data including text and first information related to the image data as inputs, a learning procedure for learning the model parameters based on the second information and third information representing the correct answer of the second information, wherein the computer executes, the generation procedure includes, when the area extracted from the image data contains text, text tokens are generated based on the information representing the text, image tokens are generated for the area extracted from the image data, and the second information is generated based on a token sequence including the generated image tokens. A learning method characterized by the above.

10. A generation procedure for generating second information corresponding to the first information using the model parameters of a learned machine learning model, with data including a visual area including at least an image area and a character area and first information related to the data as inputs, wherein the computer executes, the generation procedure includes, When the area extracted from the data contains text, text tokens are generated based on the information representing the text, image tokens are generated for the area extracted from the data, and the second information is generated based on a token sequence including the generated image tokens. A generation method characterized by this.

11. Using the model parameters of a machine learning model with data including a visual area including at least an image area and a character area and first information related to the data as inputs, a generation unit that generates second information corresponding to the first information, a learning unit that learns the model parameters based on the second information and third information representing the correct answer of the second information, functioning a computer as, The generation unit, When the area extracted from the data contains text, text tokens are generated based on the information representing the text, image tokens are generated for the area extracted from the data, and the second information is generated based on a token sequence including the generated image tokens. A program characterized by this.

12. Using the model parameters of a machine learning model with image data including text and first information related to the image data as inputs, a generation unit that generates second information corresponding to the first information, a learning unit that learns the model parameters based on the second information and third information representing the correct answer of the second information, functioning a computer as, The generation unit, When the area extracted from the image data contains text, text tokens are generated based on the information representing the text, image tokens are generated for the area extracted from the image data, and the second information is generated based on a token sequence including the generated image tokens. A program characterized by this.

13. Using the model parameters of a learned machine learning model with data including a visual area including at least an image area and a character area and first information related to the data as inputs, a generation unit that generates second information corresponding to the first information, functioning a computer as, The generation unit, When the area extracted from the data contains text, text tokens are generated based on the information representing the text, image tokens are generated for the area extracted from the data, A program that generates the second information based on a token sequence including the generated image tokens. **Claim 14** Using the model parameters of a machine learning model, a generation procedure that takes as input data including a visual region including at least an image region and a character region, and first information related to the data, and generates second information corresponding to the first information; A learning procedure for learning the model parameters based on the second information and third information representing the correct answer of the second information; wherein a computer executes the following: The generation procedure includes: when the region extracted from the data includes text, generating text tokens based on the information representing the text; generating image tokens for the region extracted from the data; A method for manufacturing a trained machine learning model, characterized by generating the second information based on a token sequence including the generated image tokens. **Claim 15** A generation unit that takes as input data including a visual region including at least an image region and a character region, and first information related to the data, and generates second information corresponding to the first information using the model parameters of a machine learning model; A learning unit that learns the model parameters based on the second information and third information representing the correct answer of the second information; having: The generation unit includes: when the region extracted from the data includes text, generating text tokens based on the information representing the text; generating image tokens for the region extracted from the data; A learning system, characterized by generating the second information based on a token sequence including the generated image tokens. **Claim 16** A generation unit that takes as input image data including text and first information related to the image data, and generates second information corresponding to the first information using the model parameters of a machine learning model; A learning unit that learns the model parameters based on the second information and third information representing the correct answer of the second information; having: The generation unit includes: when the region extracted from the image data includes text, generating text tokens based on the information representing the text; generating image tokens for the region extracted from the image data; A learning system that generates the second information based on a token sequence including the generated image tokens. **Claim 17** A generation unit that generates second information corresponding to the first information using model parameters of a learned machine learning model, with data including a visual region including at least an image region and a character region and the first information related to the data as inputs. comprising the generation unit generates text tokens based on information representing the text when the region extracted from the data contains text, generates image tokens for the region extracted from the data, A generation system that generates the second information based on a token sequence including the generated image tokens.

Citation Information

Patent Citations

  • Multilingual image question answering

    JP2017534956A

  • Question answering device, question answering method and program

    JP2019191827A