Inference device, training device, inference method, and program

By incorporating visual information extraction into document image processing, the accuracy of LLM-based task execution is enhanced, addressing the limitations of conventional methods that neglect layout and visual features in document images.

JP2025137328APending Publication Date: 2025-09-19NIPPON TELEGRAPH & TELEPHONE CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024091632
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-08
Filing Date
2024-06-05
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Conventional techniques fail to utilize visual information related to text in document images, leading to insufficient accuracy in task execution, particularly when using pre-trained large language models (LLM) for document image processing.

Method used

A feature generation unit that extracts visual information from document images and integrates it with text-based features, using a trained model to enhance the generation of output text.

Benefits of technology

Improves task execution accuracy by explicitly considering visual information, enhancing the performance of LLMs in document image understanding tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025137328000001_ABST
    Figure 2025137328000001_ABST
Patent Text Reader

Abstract

To generate, from data including texts, a feature based explicitly on visual information related to the data.SOLUTION: An inference device comprises a feature generation unit that generates a feature on the basis of a trained model from visual information that is extracted from data including texts, and that is related to the texts, and a first text, and outputs the feature to be input to a generation unit. The generation unit outputs a second text on the basis of the feature and the first text.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technique for outputting text based on a document image. [Background technology]

[0002] There are known techniques for executing tasks targeting document images in which text and images are arranged in various positions. Examples of such tasks include a question-answering task in which a document image is used as a knowledge source to generate answer text to a question, and an information extraction task in which specific information is extracted from a document image.

[0003] The technology disclosed in Non-Patent Document 1 is known as one of the conventional technologies for executing tasks targeting natural images, which are images of natural scenery, etc. The technology disclosed in Non-Patent Document 1 uses a pre-trained large language model (LLM). [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. ICML23 Summary of the Invention [Problem to be solved by the invention]

[0005] However, in the conventional technology disclosed in Non-Patent Document 1, when a document image is input, it is not possible to use features that explicitly take into account visual information, such as layout information related to the text, in the document image. Note that a document image is an example of "data containing text."

[0006] The present invention has been made in consideration of the above points, and aims to provide a technique for generating features from data including text that are explicitly based on visual information related to the data. [Means for solving the problem]

[0007] According to the disclosed technology, a feature generation unit is provided that generates features from visual information related to the text and a first text based on a trained model, and outputs the features to be input to a generation unit; The generation unit outputs a second text based on the feature and the first text. A reasoning device is provided. [Effects of the Invention]

[0008] The disclosed technology provides a technique for generating features from data that includes text that are explicitly based on visual information about the data. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 10 is a diagram illustrating a first example of a task. [Figure 2] FIG. 10 is a diagram illustrating a second example of a task. [Figure 3] FIG. 10 is a diagram illustrating a third example of a task. [Figure 4] FIG. 10 is a diagram illustrating a fourth example of a task. [Figure 5] FIG. 10 is a diagram illustrating a fifth example of a task. [Figure 6] FIG. 1 illustrates an example of the configuration of a learning device 100. [Figure 7] 10 is a flowchart showing an example of the operation of the learning device 100. [Figure 8] FIG. 2 is a diagram illustrating an example of the internal configuration of a generation unit 140. [Figure 9] FIG. 10 is a diagram showing parameters to be updated and parameters not to be updated. [Figure 10] FIG. 2 is a diagram illustrating an example of the configuration of an inference device 200. [Figure 11] FIG. 10 is a diagram showing parameters to be updated and parameters not to be updated. [Figure 12] 4 is a flowchart showing an example of the operation of the inference device 200. [Figure 13] FIG. 10 illustrates another exemplary configuration of the learning device 100. [Figure 14] FIG. 2 is a diagram illustrating another exemplary configuration of the inference device 200. [Figure 15] FIG. 2 illustrates an example of a hardware configuration of the apparatus. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, an embodiment of the present invention (the present embodiment) will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment.

[0011] With regard to the tasks targeting document images mentioned above, conventional techniques typically prepare a dataset for the target task (e.g., question answering) for each document image type (e.g., invoice, etc.), and then use the dataset for training to build a model that can be adapted to a specific document type and task. However, there are a wide variety of document images and tasks (user needs), and it is difficult to achieve all of these with a single model using general conventional techniques.

[0012] Furthermore, as mentioned above, the technique disclosed in Non-Patent Document 1 does not allow LLM to use features that explicitly consider visual information about the text in document images, and therefore the accuracy of task execution is insufficient.

[0013] In contrast, in this embodiment, a mechanism is provided for inputting features that explicitly consider visual information about text in document images into the LLM, thereby improving task execution accuracy. In other words, each device according to this embodiment provides specific improvements over conventional task execution techniques such as those disclosed in Non-Patent Document 1, and represents an advancement in the technical field of document image understanding.

[0014] The technology according to the present embodiment can be applied to any language. In addition, although learning device 100 and inference device 200 are described below as separate devices, learning device 100 and inference device 200 may be the same device.

[0015] (Definition of terms) In the present embodiment, a "document image" is an example of "data containing text." Examples of a "document image" include, but are not limited to, presentation slides, papers, PDF documents, and HTML documents.

[0016] A "document image" is assumed to include text (character string) information and image information. However, the data format of the "document image" may be an image format, an image and text format, or a text format.

[0017] The "image format" mentioned above refers to image data in which text information is embedded as image information. For example, the non-searchable PDF format falls into this category. The "image and text format" refers to image data in which text information is embedded as is. For example, the searchable PDF format falls into this category. The "text format" refers to data that contains only text, no figures or photos, but is treated as image data.

[0018] Examples of "data including text" that are not included in the general meaning of "document image" include landscape images including explanatory text, television screens such as news, etc. The technology according to this embodiment is applicable to "data including text" in general. Note that in this embodiment, "document image" also includes landscape images including explanatory text, television screens such as news, etc.

[0019] Furthermore, "visual information" in "data including text" refers to information that can be visually recognized in "data including text." "Visual information" may include features converted from visually recognized information.

[0020] Examples of "visual information" include the position of a text area, the size of the text area, the color of the text, the shape of the text, the color, type, shape, and size of the text font, the position of non-text areas, the size of non-text areas, the color of non-text areas, and the shape of non-text areas. Examples of non-text areas include graphs, photographs, landscapes, and backgrounds. Note that even non-text areas are related to the text. Other examples of "visual information" include depth information within an image and line information within an image. Depth information is, for example, information indicating that a certain area in an image is located further back than another area, and may also be referred to as context. Line information is, for example, information about lines used to separate tables or underline text. Note that "visual information" includes information such as the color and shape of text (visually recognizable information), but does not include linguistic information (semantic information) contained in the text strings.

[0021] In this specification, "visual information," "information indicating an area," "area information," and "position information" are used broadly, and "information indicating an area" is included in "visual information." In other words, "information indicating an area" is an example of "visual information." "Information indicating an area" and "area information" are synonymous. Furthermore, "position information" related to an area is included in "information indicating an area." In other words, "position information" related to an area is an example of "information indicating an area."

[0022] The "instruction sentence" is an example of the "first text." The "instruction sentence" may also be called the "instruction text." The "answer sentence" is an example of the "second text." The "answer sentence" may also be called the "answer text." Additionally, the text portion of the "data including text" is an example of the "third text."

[0023] The type of "instruction statement" used in this embodiment is not limited to a specific one. Examples of "instruction statements" include instructions, queries, questions, requests, and text for specifying output. However, "instruction statements" can be any text that has some meaning, and do not have to be text that expresses an "instruction" or a "question."

[0024] In this embodiment, both "token" and "token sequence" may be replaced with "text" or "word." Also, "word" may be replaced with "text."

[0025] (Task example) Learning device 100 and inference device 200 according to the present embodiment will be described in detail below, but before that, examples of tasks executed by learning device 100 / inference device 200 will be described to facilitate understanding of the operation of these devices. The examples described below may be considered to be tasks executed by learning device 100 in the learning phase, or may be considered to be tasks executed by inference device 200 in the inference phase, but here they will be described assuming that they are tasks executed by inference device 200.

[0026] <Example 1> Example 1 is shown in Figure 1. Example 1 uses data from WTQ (https: / / github.com / ppasupat / WikiTableQuestions). In Example 1, when the document image shown in (a) and the instruction statement shown in (b) are input to inference device 200, inference device 200 outputs the numerical value (example text) shown in (c). Note that the part of the document image in Figure 1 that is framed in bold has been added to make the contents of the document image easier to understand, and this framed part does not exist in the actual task.

[0027] As shown in Figure 1, the document image contains a multi-year forecast for Operations and Maintenance at the University of Minnesota, and the directive requests a forecast for 2010-2011. In response, reasoning apparatus 200 outputs accurate numerical values.

[0028] The same task as in Example 1 was performed using ChatGPT (registered trademark), an existing large-scale language model (LLM), and BLIP-2 disclosed in Non-Patent Document 1, but no appropriate answer was obtained.

[0029] <Example 2> Example 2 is shown in Figure 2. Example 2 uses data from Screen2Words (https: / / github.com / google-research-datasets / screen2words). In Example 2, when the document image shown in (a) and the directive shown in (b) are input to inference device 200, inference device 200 outputs the text shown in (c).

[0030] As shown in Figure 2, the document image contains an image of a guitar being played, a list of music genres, and an image of a search page. The directive shown in (b) is a query, limited in number of words, about what the UI image is. In response, the inference device 200 outputs the correct content for the document image.

[0031] As in Example 1, in Example 2, we confirmed that even when performing the same task as in Example 2 using ChatGPT (registered trademark) and BLIP-2, appropriate answers could not be obtained.

[0032] <Example 3> Example 3 is shown in Figure 3. Example 3 uses data from FUNSD (Form Understanding in Noisy Scanned Documents) (https: / / guillaumejaume.github.io / FUNSD).

[0033] In Example 3, when the document image shown in (a) and the directive shown in (b) are input to inference device 200, inference device 200 outputs the text shown in (c).

[0034] As shown in Figure 3, the document image shows an image titled "CASE FORM," and the instruction indicates answer options and asks about the category of "CASE FORM." In response, reasoning device 200 outputs the correct content.

[0035] <Example 4> Example 4 is shown in Figure 4. Example 4 uses data from SlideVQA (https: / / github.com / nttmdlab-nlp / SlideVQA).

[0036] In Example 4, when the two document images shown in (a) and the directive shown in (b) are input to inference device 200, inference device 200 outputs the text shown in (c).

[0037] As shown in Figure 4, the document image on the left displays a graph of the number of journalists by region, and the document image on the right displays percentage values ​​such as "Competition media" for each region. The instruction text includes a question and an answer style. In response, reasoning device 200 outputs the correct content.

[0038] <Example 5> Example 5 is shown in Figure 5. Example 5 uses data from SciCap (Scientific Figures Dataset) (https: / / github.com / tingyaohsu / SciCap).

[0039] In Example 5, when the document image shown in (a) and the directive shown in (b) are input to inference device 200, inference device 200 outputs the text shown in (c).

[0040] As shown in Figure 5, the document image shows a graph related to "Corruption Gaussian noise." The instruction sentence is a sentence requesting an explanation of the image. In response, inference device 200 outputs the correct content.

[0041] The configurations and operations of the learning device 100 and the inference device 200 will be described below.

[0042] (Configuration example of learning device 100) Fig. 6 shows an example configuration of the learning device 100. As shown in Fig. 6, the learning device 100 includes a text extraction unit 110, an image feature extraction unit 120, a feature conversion unit 130, a generation unit 140, and a parameter learning unit 150. The generation unit 140 includes an encoding unit 141 and a decoding unit 142.

[0043] The image feature extraction unit 120, the feature conversion unit 130, and the generation unit 140 are functional units realized by a neural network model, and parameters of the model are stored in a model DB 160. The learning device 100 (specifically, a computer) reads the parameters from the model DB 160 and executes the operation of each functional unit using the parameters. Note that functional units other than the "image feature extraction unit 120, the feature conversion unit 130, and the generation unit 140" may also be realized by a neural network model.

[0044] Furthermore, although the model parameters are updated by the parameter learning unit 150 of the learning device 100, there are also parameters that are not updated and remain fixed. Details of this will be described later.

[0045] (Operation Example of Learning Device 100) FIG. 7 is a flowchart showing an operation example of the learning device 100. Hereinafter, the operations of each part constituting the learning device 100 will be described in detail according to the procedure of the flowchart of FIG. 7.

[0046] <S101 (Step 101): Input of Learning Data> In S101, a learning data set is input to the learning device 100. The learning data set includes a plurality of learning data (which may also be called instances). Each learning data has one or more "instruction sentences, document images, and correct answers". Here, it is assumed that the process of executing S102 to S105 by one learning data is performed for each learning data included in the learning data set.

[0047] As shown in FIG. 6, the document image is input to each of the text extraction unit 110 and the image feature extraction unit 120, and the instruction sentence is input to each of the feature conversion unit 130 and the generation unit 140. Also, the correct answer is input to the parameter learning unit 150.

[0048] Note that the type of the learning data set (that is, the type of the task for learning) is not limited to a specific type. For example, a learning data set for tasks such as reading comprehension, summarization, information extraction, QA, dialogue, captioning, and illustration of regions can be used. Also, a learning data set of the type shown in FIGS. 1 to 5 (Example 1 to Example 5) may be used.

[0049] <S102: Information Extraction> In S102, the text extraction unit 110 extracts text information in the document image from the input document image and outputs the text information in the document image to the feature conversion unit 130. Also, the image feature extraction unit 120 extracts document image features from the input document image and outputs the document image features to the feature conversion unit 130. Hereinafter, the operations of each of the text extraction unit 110 and the image feature extraction unit 120 will be described in more detail.

[0050] <S102: Details of the Operation of the Text Extraction Unit 110> The text extraction unit 110 takes a document image as input and outputs text information in the document image. The text information in the document image includes a document text sequence (which may also be referred to as text) included in the input document image and information indicating a document text rectangular area (e.g., upper left coordinate, lower right coordinate) that is the area where the document text sequence exists.

[0051] Specifically, the text extraction unit 110 detects the area of each of one or more document text sequences included in the input document image, recognizes the document text sequence within the detected area, and outputs text information in the document image including this information (text and area). When there are multiple document text sequences, the text information in the document image includes the multiple document text sequences and their area information.

[0052] In the present embodiment, the area of the document text sequence is a rectangular area, and the rectangular area is represented by the upper left coordinates (x , y 1 ) and the lower right coordinates (x 2 , y 2 ). However, such a representation method is only an example, and other representation methods may be used. ​​​​​​​

[0055] A more detailed example of using "top left coordinates, bottom right coordinates" as information indicating an area will be described below.

[0056] A document text sequence is a set of M words, where {s i} M i=1 It is expressed as s i is the i-th word. In other words, the text extraction unit 110 extracts {s i} M i=1 The text extraction unit 110 also extracts a set of rectangular areas for each word {(x i 1 ,y i 1 ,x i 2 ,y i 2 )} M i=1 Extract and output (x i 1 ,y i 1 ), (x i 2 ,y i 2 ) are respectively, s i are the top left and bottom right coordinates of the

[0057] When outputting multiple document text sequences, the text extraction unit 110 rearranges the document text sequences in the order of "left to right" and "top to bottom" according to their coordinates in the document image. However, such rearrangement may not be performed.

[0058] As the text extraction unit 110, any method can be used as long as it can output a document text sequence and information indicating the region of the document text sequence from the document image. For example, OCR (Optical Character Recognition) can be used. Specifically, Google Vision API (https: / / cloud.google.com / vision) etc. can be used.

[0059] <S102: Details of the operation of the image feature extraction unit 120> The image feature extraction unit 120 takes a document image as input and outputs document image features. More specifically, the image feature extraction unit 120 has an image encoder, and the image encoder extracts a vector z representing the features of the input document image from the input document image vis and outputs it.

[0060] Any image encoder can be used as long as image features can be obtained from the document image. For example, a pre-trained model called CLIP (Contrastive Language-Image Pre-training) disclosed in "Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. ICML21" can be used as the image encoder.

[0061] <S103: Transformation of features> In S103, the feature conversion unit 130 receives as input the instruction sentence, the text information in the document image output from the text extraction unit 110, and the document image features output from the image feature extraction unit 120, converts the input information into document image conversion features, and outputs the document image conversion features. The output document image conversion features are input to the generation unit 140. The document image conversion features output from the feature conversion unit 130 are features that can be understood by the generation unit 140 (LLM). As described above, the text information in the document image includes the text itself and information on rectangular areas, which is visual information.

[0062] Document image conversion features are doc The feature conversion unit 130 converts h doc is calculated by the following steps S1 to S3.

[0063] S1) In S1, the feature conversion unit 130 converts the text information in the document image and the instruction sentence into an embedding sequence. i} M i=1 and {(x i 1 ,y i 1 ,x i 2 ,y i 2 )} M i=1 is input to the feature transformation unit 130.

[0064] The feature conversion unit 130 converts the i-th word s i Embedding z i ocr is calculated using the following formula: W s , W x , W y , W h , W w are learnable weights (specifically matrices), respectively.

[0065]

number

[0066] Similarly, the feature transformer 130 uses learnable weights W s Using this, we embed the directive into the sequence z ins For example, convert the words that make up the instruction into ins i Then, the feature conversion unit 130 converts W s (ins i ) and calculate W for each word. s (ins i ) are concatenated into z ins Let's say.

[0067] z ocr and z ins are vectors. The feature conversion unit 130 converts z ocr and z ins An input token sequence (vector) is constructed by vector concatenation of and.

[0068] S2) The feature transformation unit 130 has a cross-attention layer and a self-attention layer. The feature transformation unit 130 combines a trainable parameter sequence and document image features z vis The self-attention layer interacts with the learnable parameter sequence and the input token sequence (z ocr and z ins In S2, a learnable parameter sequence is obtained through these interactions.

[0069] More specifically, the cross-attention layer contains a learnable parameter sequence and document image features zz vis is input and the converted "learnable parameter sequence" is output. The self-attention layer receives the learnable parameter sequence and the input token sequence (z ocr and zins The input is a concatenation of the inputs, and the converted "trainable parameter sequence" is output.

[0070] In other words, using the input learnable parameter sequence as a pivot, the cross-attention layer interacts with document image features, and the self-attention layer interacts with the input token sequence, resulting in a "learnable parameter sequence" through these interactions.

[0071] As a variation, the learnable parameter sequence interacted with document image features in the cross-attention layer may be further interacted with an input token sequence in the self-attention layer to output a transformed "learnable parameter sequence." Also, the learnable parameter sequence interacted with an input token sequence in the self-attention layer may be further interacted with document image features in the cross-attention layer to output a transformed "learnable parameter sequence."

[0072] As a model having the above-mentioned cross-attention layer and self-attention layer, for example, the pre-trained BERT disclosed in "Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL19" can be used.

[0073] The trainable parameter sequence is a D×L dimensional vector, where D=1024 and L=32 in one embodiment.

[0074] S3) The feature transformation unit 130 further includes a feedforward network. In S3, the feature transformation unit 130 inputs the trainable parameter sequence obtained in S2 into the feedforward network, and outputs the document image transformation feature h docObtain it. The feature conversion unit 130 outputs the obtained document image conversion feature h doc Output it.

[0075] The dimension number of the document image conversion feature h doc is a dimension number suitable for the input to the generation unit 140 (specifically, the encoding unit 141) described in S104. Even if the dimension number of the series of learnable parameters is not a dimension number suitable for the input to the encoding unit 141, by passing the series of learnable parameters through the feed-forward network, the dimension number of the document image conversion feature h doc can be made a dimension number suitable for the input to the encoding unit 141.

[0076] Note that generating the document image conversion feature by the above method is just an example. Any method can be used as long as it can generate the feature reflecting "text information in the document image, document image feature, and instruction text". Input this feature to the generation unit 140.

[0077] <S104: Generation process> In S104, the generation unit 140 takes as input the instruction text, the document image conversion feature, the text information in the document image output from the text extraction unit 110 (specifically, the document text series), and the tokenized answer text output from the parameter learning unit 150, and outputs the probability distribution of the output series.

[0078] The generation unit 140 including the encoding unit 141 and the decoding unit 142 can be realized using a pre-trained LLM. As the LLM, for example, a model using the well-known Transformer can be used.

[0079] Specifically, as the pre-trained LLM, for example, a model called FlanT5 can be used. However, using FlanT5 is just an example. A pre-trained LLM other than FlanT5 may be used.

[0080] The internal configuration example of the generation unit 140 when using a model that uses a Transformer such as FlanT5 as an LLM is shown in FIG. 8. Note that the configuration of the generation unit 140 is the same during learning and inference, and FIG. 8 also shows the text generation unit 6 that is used only during inference, in addition to the functional units used during learning and inference.

[0081] As shown in FIG. 8, the encoding unit 141 in the generation unit 140 includes a document image conversion feature reception unit 1, an instruction sentence reception unit 2, a text information reception unit 3 in the document image, and a plurality of conversion layers 4. Further, the decoding unit 142 in the generation unit 140 includes a plurality of conversion layers 5 and a text generation unit 6. A Transformer is used as each conversion layer.

[0082] Hereinafter, the operations of the encoding unit 141 and the decoding unit 142 that constitute the generation unit 140 will be described.

[0083] <S104: Operation of the encoding unit 141> In the encoding unit 141, a document image conversion feature is input by the document image conversion feature reception unit 1, an instruction sentence is input by the instruction sentence reception unit 2, and a document text sequence is input by the text information reception unit 3 in the document image.

[0084] The conversion layer 4 (Transformer) converts the "instruction sentence, document image conversion feature, and document text sequence" into an encoded feature and outputs the encoded feature. The encoded feature is, for example, a K-dimensional vector, and in one embodiment, K = 1024.

[0085] <S105: Operation of the decoding unit 142 (partially, operation of the parameter learning unit 150)> The decoding unit 142 takes as inputs the encoded feature output from the encoding unit 141 and the tokenized response sentence output from the parameter learning unit 150, and outputs an output sequence probability distribution. More specifically, the encoding unit 142 during learning executes the following processes S1 to S3.

[0086] S1) In S1, the decoding unit 142 acquires a token sequence of the correct answer sentence output by the parameter learning unit 150 for each output step, and creates an output sequence (correct answer sentence) by combining the respective token sequences.

[0087] That is, during learning, the parameter learning unit 150 performs a tokenization process on the correct answer sentence, and inputs the resulting token sequence to the decoding unit 142. However, the tokenized sentence may also be input to the parameter learning unit 150.

[0088] During learning, the decoding unit 142 assigns a start symbol and a terminal symbol to the input token sequence. In this embodiment, as an example, the start symbol is <s> and assign it to the terminal symbol< / s> Assign a symbol.

[0089] The parameter learning unit 150 may use any method for tokenizing, but for example, it may use the Byte-level BPE disclosed in "Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever: Language models are unsupervised multitask learners. Technical report, OpenAI, 2019."

[0090] S2) In S2, the decoding unit 142 inputs the encoded features and the output sequence to the transformer layer 5, which then generates a representation H=[h1, . . . , h T ], where T is the length of the output token sequence, and as an example, T=128.

[0091] S3) In S3, the decoding unit 142 decodes the representation of the output token H=[h1, . . . , h TAfter performing a linear transformation on ]], the probability distribution p(y t │y <t ) of the t-th word is calculated using the softmax function. Here, t satisfies 0 ≤ t ≤ T.

[0092] <S105: Parameter Update> The parameter learning unit 150 takes as input the correct sentence (correct answer sentence) and the probability distribution of the output sequence output from the decoding unit 142, and outputs parameter update information and the tokenized correct sentence. As described above, the tokenized correct sentence may be input as the input to the parameter learning unit 150. Also, the tokenized correct sentence output from the parameter learning unit 150 is input to the decoding unit 141 as described above.

[0093] The parameter learning unit 150 calculates the loss L by the following formula based on the probability distribution of the output sequence and the tokenized correct sentence, and updates the parameters of the neural network model in the learning device 100 so that the loss L is minimized.

[0094]

Equation

[0095] In this embodiment, among the "image feature extraction unit 120, feature conversion unit 130, generation unit 140" configured using a neural network, only the parameters of the feature conversion unit 130 are updated, and the parameters of the other functional units, namely the "image feature extraction unit 120, generation unit 140", are fixed and not updated. Note that the parameters of the parameter update unit 150 may also be an update target.

[0096] This is shown in FIG. 9. As shown in FIG. 9, among the parameters stored in the model DB 160, the parameters of the model corresponding to the feature conversion unit 130 are the parameters 2 to be updated by learning. For parameters not updated by learning, pre-learned parameters may be used. Technically, any parameter may be updated as long as not all parameters are fixed. However, in the present embodiment, by updating only the parameters of the feature conversion unit 130, the learning cost is suppressed. The same applies to FIG. 11 described later.

[0097] <S106: Judgment> In the learning device 100, if there is unprocessed learning data, it returns to S101 and performs the processing of S102 to S105 using the next learning data. If there is no unprocessed learning data, the processing ends.

[0098] (Configuration example of the inference device 200) Subsequently, a configuration example and an operation example of the inference device 200 will be described. First, the configuration example of the inference device 200 will be described.

[0099] FIG. 10 shows a configuration example of the inference device 200. As shown in FIG. 10, the inference device 200 includes a text extraction unit 110, an image feature extraction unit 120, a feature conversion unit 130, and a generation unit 140. The generation unit 140 includes an encoding unit 141, a decoding unit 142, and a text generation unit 143.

[0100] The text extraction unit 110, the image feature extraction unit 120, and the feature conversion unit 130 in the inference device 200 are the same as the text extraction unit 110, the image feature extraction unit 120, and the feature conversion unit 130 in the learning device 100, respectively.

[0101] Also, the generation unit 140 is basically the same in the learning device 100 and the inference device 200. However, in the inference process, in addition to the encoding unit 141 and the decoding unit 142, the text generation unit 143 is used.

[0102] Also, as shown in FIG. 10, a model DB 160 is provided, and learned parameters are stored in the model DB 160.

[0103] As described above, the parameters of "the image feature extraction unit 120 and the generation unit 140" among "the image feature extraction unit 120, the feature conversion unit 130, and the generation unit 140" are parameters that have not been updated by learning in the learning device 100, and the parameters of "the feature conversion unit 130" are parameters that have been updated by learning in the learning device 100. This is shown in FIG. 11.

[0104] (Operation example of the inference device 200) FIG. 12 is a flowchart showing an operation example of the inference device 200. Hereinafter, the operation of the inference device 200 will be described according to the procedure of the flowchart in FIG. 12. Note that "the text extraction unit 110, the image feature extraction unit 120, the feature conversion unit 130, the encoding unit 141, and the decoding unit 142" in the inference device 100 are the same as those in the learning device 100, so only an outline of each operation is described.

[0105] <S201: Input of instruction text and document image> In S201, an instruction text and a document image are input to the inference device 200. As shown in FIG. 10, the document image is input to each of the text extraction unit 110 and the image feature extraction unit 120. Also, the instruction text is input to each of the generation unit 140 and the feature conversion unit 130.

[0106] <S202: Information extraction> In S202, the text extraction unit 110 extracts text information in the document image from the input document image and outputs the text in the document image to the feature conversion unit 130. Also, the image feature extraction unit 120 extracts document image features from the input document image and outputs the document image features to the feature conversion unit 130.

[0107] <S203: Feature conversion> In S203, the feature conversion unit 130 takes as input the instruction text, the text information in the document image output from the text extraction unit 110, and the document image features output from the image feature extraction unit 120, converts the input information into document image conversion features, and outputs the document image conversion features. The output document image conversion features are input to the generation unit 140.

[0108] Similar to the learning case, any method can be used as long as it can generate features reflecting "text information in the document image, document image features, and instruction text". The generated features are input to the generation unit 140.

[0109] <S204, S205: Generation of response text> In S204, the generation unit 140 takes as input the instruction text, the document image conversion features, and the text information in the document image (specifically, the document text sequence), and outputs a response text. More specifically, the following processing is executed. It is assumed that the generation unit 140 has the configuration shown in FIG. 8.

[0110] In the encoding unit 141, the document image conversion feature is input by the document image conversion feature reception unit 1, the instruction text is input by the instruction text reception unit 2, and the document text sequence is input by the text information reception unit 3 in the document image.

[0111] The conversion layer 4 (Transformer) converts the "instruction text, document image conversion features, and document text sequence" into encoded features and outputs the encoded features.

[0112] Subsequently, the decoding unit 142 and the text generation unit 143 execute the following processing S1 to S4.

[0113] S1) In S1, the decoding unit 142 receives the tokens recursively output for each output step from the text generation unit 143, and obtains the concatenated tokens as the output sequence. Also, the decoding unit 142 uses the start symbol of the sequence (e.g., <s>) is assigned to the output sequence.

[0114] S2) In S2, the decoding unit 142 inputs the encoded features and the output sequence to the transformer layer 5, which then generates a representation H=[h1, . . . , h T ]. As mentioned above, T is the length of the output token sequence, and as an example, T=128.

[0115] S3) In S3, the decoding unit 142 decodes the representation of the output token H=[h1, . . . , h T ] is linearly transformed and then the softmax function is used to obtain the probability distribution p(y t │y <t ) is calculated, where t satisfies 0≦t≦T.

[0116] S4) In S4, the text generation unit 143 receives the probability distribution of the output sequence output from the decoding unit 142, generates a response sentence based on the probability distribution, and outputs the response sentence.

[0117] Specifically, the text generation unit 143 recursively generates a token sequence of the answer sentence based on the probability distribution of the output sequence. More specifically, the text generation unit 143 selects a word that maximizes the probability distribution for each output step, or generates words by sampling according to the probability distribution, and ends generation when a sentence-end symbol is generated (Yes in S205 of FIG. 12).

[0118] (When inputting multiple document images) As shown in task example 4 (FIG. 4), inference device 200 can also handle the case where multiple document images are input along with instruction sentences.

[0119] When multiple document images are input, the image feature extraction unit 120, text extraction unit 110, and feature conversion unit 130 (excluding the feedforward network) perform the same processing as described above for each document image. Then, the features (trainable parameter series) for the multiple document images output from the feature conversion unit 130 (excluding the feedforward network) are averaged by mean-pooling processing in the feedforward network. In this way, the averaged document image conversion features obtained from the multiple document images are input to the generation unit 140.

[0120] Furthermore, the document text sequences obtained from each of the multiple document images by the text extraction unit 110 are concatenated for each of the multiple document images and input to the generation unit 140. The operation of the generation unit 140 is the same as that described above as the operation of the inference device 200.

[0121] The same applies to the operation during learning in learning device 100: when learning data containing multiple document images is input as learning data, image feature extraction unit 120, text extraction unit 110, and feature conversion unit 130 in learning device 100 perform the same operations as image feature extraction unit 120, text extraction unit 110, and feature conversion unit 130 in the above-mentioned inference device 200. Subsequent operations are the same as those of learning device 100 described above.

[0122] (About the experimental results) An experiment was conducted using an inference device 200 (herein referred to as the proposed method) equipped with a feature transformation unit 130 that uses model parameters learned by a learning device 100. For comparison, an experiment was also conducted using fine-tuned BLIP-2 (herein referred to as fine-tuned BLIP-2).

[0123] Experiments on various tasks using data not used in training showed that the average F1 score was 43.1 for the proposed method and 39.3 for fine-tuned BLIP-2. Furthermore, the average scores for evaluation indices other than F1 were 43.0 for the proposed method and 37.0 for fine-tuned BLIP-2. This demonstrates that the technology according to this embodiment improves task processing accuracy compared to conventional technologies.

[0124] (Variations in usage) Regarding the method of using the inference device 200, for example, when using a document image such as a receipt, a photo of the receipt may be taken immediately upon receiving the receipt and the image may be uploaded to the cloud (inference device 200).

[0125] Furthermore, for example, when using a task such as extracting only the total amount from a receipt, the user can be prevented from having to input an instruction each time. For example, when an image is received from a user who has set an instruction in advance for inference device 200, inference device 200 generates a response using the set instruction and the received image.

[0126] (Other configuration examples) The learning device 100 may be configured as shown in Fig. 13. As shown in Fig. 13, the learning device 100 includes an extraction unit 210, a feature generation unit 220, and a learning unit 230. In the example of Fig. 13, the generation unit 240 is provided outside the learning device 100, but the generation unit 240 may also be provided inside the learning device 100. Furthermore, the model DB may be provided inside or outside the learning device 100. Furthermore, the extraction unit 210 may also be provided outside the learning device 100.

[0127] Note that even when the generation unit 240 is provided outside the learning device 100, a system including the learning device 100 and the generation unit 240 is referred to as a "learning device." Furthermore, when the generation unit 240 is provided outside the learning device 100, the generation unit 240 is realized, for example, as a server connected to the learning device 100 via a network. The generation unit 240 provided outside the learning device 100 may also be referred to as a "generation device."

[0128] The extraction unit 210 extracts visual information related to text from data including the text (for example, a document image).

[0129] The feature generation unit 220 generates features (e.g., document-image conversion features) from the visual information and the first text (e.g., instructional sentence), and inputs the features to the generation unit 240. The extraction unit 210 may input the text to the generation unit 240.

[0130] The learning unit 230 uses information output from the generation unit 240 to which the features and the first text are input, and the data and a second text that is a correct answer to the first text, to learn model parameters of the neural network that constitutes the feature generation unit 220.

[0131] Note that the "text extraction unit 110 and image feature extraction unit 120" shown in FIG. 6 are examples of the extraction unit 210. The feature conversion unit 130 shown in FIG. 6 is an example of the feature generation unit 220. The generation unit 140 shown in FIG. 6 is an example of the generation unit 240. The parameter learning unit 150 shown in FIG. 6 is an example of the learning unit 130.

[0132] Inference device 200 may be configured as shown in Figure 14. As shown in Figure 14, inference device 200 includes extraction unit 210 and feature generation unit 220. In the example of Figure 14, generation unit 240 is provided outside inference device 200, but generation unit 240 may also be provided inside inference device 200. Furthermore, model DB may be provided inside or outside inference device 200. Furthermore, extraction unit 210 may also be provided outside inference device 200.

[0133] Note that even when generation unit 240 is provided outside inference device 200, a system including inference device 200 and generation unit 240 is referred to as an "inference device." Furthermore, when generation unit 240 is provided outside inference device 200, generation unit 240 is realized, for example, as a server connected to inference device 200 via a network. Generation unit 240 provided outside inference device 200 may also be referred to as a "generation device."

[0134] The extraction unit 210 extracts visual information related to text from data including the text (for example, a document image).

[0135] The feature generation unit 220 generates features from the visual information and a first text (e.g., an instructional sentence) and inputs the features to the generation unit 240. The extraction unit 210 may input the text to the generation unit 240. The generation unit 240, to which the features and the first text have been input, outputs a second text generated based on the features and the first text.

[0136] 9 is an example of the extraction unit 210. The feature conversion unit 130 shown in FIG. 9 is an example of the feature generation unit 220. The generation unit 140 shown in FIG. 9 is an example of the generation unit 240.

[0137] (Example of hardware configuration) Any of the devices described in this embodiment (learning device 100, inference device 200) can be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud.

[0138] That is, the device can be realized by executing a program corresponding to the processing performed by the device using hardware resources such as a CPU and memory built into a computer. The program can be recorded on a computer-readable recording medium (such as a portable memory) and stored or distributed. The program can also be provided via a network such as the Internet or email.

[0139] Fig. 15 is a diagram showing an example of the hardware configuration of the computer. The computer in Fig. 15 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, and the like, all of which are interconnected by a bus BS. The computer may further include a GPU.

[0140] A program for realizing processing on the computer is provided by a recording medium 1001 such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001, but may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files, data, etc.

[0141] The memory device 1003 reads and stores a program from the auxiliary storage device 1002 when an instruction to start the program is received. The CPU 1004 realizes the functions related to the device in accordance with the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network, etc. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, mouse, buttons, a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the results of calculations.

[0142] (Effects of the embodiment) As described above, the technology described in this embodiment makes it possible to accurately acquire information about data including text from the data.

[0143] More specifically, when a document image is used as the data, the document image can be understood as a knowledge source, and a response to a user's instruction can be obtained with high accuracy.

[0144] Furthermore, in the technology according to this embodiment, the parameters of the generation unit 140 (specifically, the LLM) are fixed, and the desired function can be achieved by training only the feature conversion unit 130, thereby shortening the training time and reducing the training cost.

[0145] The following additional notes are provided regarding the above-described embodiments.

[0146] <Additional Notes> (Additional note 1) 1. A reasoning apparatus that generates features of data including text based on a first text and the data to input to a sentence generation model that outputs a second text, Memory and at least one processor coupled to said memory; Including, The processor: generating the features based on a trained model from the visual information extracted from the data, the visual information relating to the text, and the first text, and outputting the features for input to the sentence generation model; Reasoning device. (Additional note 2) The visual information includes region information indicating regions of the text in the data and image features in the data, and the processor generates the features from the visual information, third text extracted from the data, and the first text. 2. The inference device according to claim 1. (Additional note 3) The processor generates the features using the weighted information on the text and the weighted information on the region. Item 2. An inference device according to claim 2. (Additional note 4) Memory and at least one processor coupled to said memory; Including, The processor: a feature generation process for generating features from visual information about the text extracted from data including the text and a first text, and outputting the features for input to a sentence generation model; a learning process for learning model parameters of a neural network that performs the feature generation process, using information output from the sentence generation model to which the features and the first text have been input, and the data and a second text that is a correct answer for the first text; A learning device that performs the following: (Additional note 5) An inference method executed by an inference device, comprising: visual information extracted from data including text, the method comprising: generating features from the visual information about the text and a first text based on a trained model; and outputting the features for input to a sentence generation model; The sentence generation model outputs a second text based on the features and the first text. Reasoning method. (Additional note 6) A non-transitory storage medium storing a program for causing a computer to function as the inference device described in any one of appendix 1 to 3. (Additional note 7) A non-transitory storage medium storing a program for causing a computer to function as the learning device described in appended claim 4.

[0147] Although the present embodiment has been described above, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims. [Explanation of symbols]

[0148] 100 Learning Device 110 Text Extraction Unit 120 Image feature extraction unit 130 Feature conversion unit 140 Generation part 141 Encoding section 1. Document image conversion feature reception unit 2 Instruction Reception Section 3. Text information reception section for document images 4. Conversion Layer 142 Decoding section 5. Conversion Layer 6, 143 Text Generation Section 150 Parameter Learning Unit 160 Model DB 200 Reasoning device 210 Extraction part 220 Feature Generation Unit 230 Learning Department 240 Generation part 1000 Drive Device 1001 Recording media 1002 Auxiliary storage 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input Device 1008 Output Device< / s>

Claims

1. visual information extracted from data including text, and a feature generation unit that generates features from the visual information about the text and a first text based on a trained model and outputs the features to be input to a generation unit; The generation unit outputs a second text based on the feature and the first text. Reasoning device.

2. The visual information includes area information indicating an area of ​​the text in the data and image features in the data, and the feature generation unit generates the features from the visual information, third text extracted from the data, and the first text based on the trained model. The inference device according to claim 1 .

3. The feature generation unit generates the features using information weighted on the text and information weighted on the region information. The inference device according to claim 2 .

4. visual information extracted from data including text, wherein the visual information about the text and a first text are used to generate features, and a feature generation unit is configured to output the features for input to the generation unit; a learning unit that learns model parameters of a neural network constituting the feature generation unit using information output from the generation unit to which the features and the first text are input, and the data and a second text that is a correct answer to the first text; A learning device comprising:

5. An inference method executed by an inference device, comprising: visual information extracted from data including text, the method comprising: generating features from the visual information about the text and a first text; and outputting the features for input to a generator; The generation unit outputs a second text based on the feature and the first text. Reasoning method.

6. A program for causing a computer to function as the feature generation unit in the inference device according to any one of claims 1 to 3.

7. A program for causing a computer to function as the feature generation unit and the learning unit in the learning device according to claim 4.