Information processing device, information processing method, and program

The information processing device addresses LLM answer inaccuracies by replacing area IDs with document objects, facilitating user verification and understanding through contextual linking.

JP7771457B1Active Publication Date: 2025-11-17KDDI CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025050052
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-11-17
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

Large Language Models (LLMs) often provide incorrect answers and make it difficult for users to determine the correctness and understand the content of the answers.

Method used

An information processing device that generates area identification information for documents, replaces area IDs in LLM outputs with corresponding objects, and transmits the replaced information to a terminal, allowing users to verify answers and understand the context through associated objects.

Benefits of technology

Enhances user confidence in answer correctness and comprehension by linking answers to relevant document sections, improving answer accuracy by preserving key points and maintaining document structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007771457000001_ABST
    Figure 0007771457000001_ABST
Patent Text Reader

Abstract

To make it easier for a user to judge whether an answer to a question is correct or not, and to make it easier for the user to understand the content of the answer. [Solution] The information processing device 2 has a document acquisition unit 231 that acquires a document, a generation unit 232 that generates area identification information for each area in the document that contains an object, and a reception unit 233 that receives a user's question, and the generation unit 232 generates a prompt that includes the question, multiple objects included in the document, and an instruction to present area identification information corresponding to an object related to the answer, and the information processing device 2 further has an output information acquisition unit 234 that acquires output information output by the large-scale language model by inputting the prompt into the large-scale language model, a replacement unit 235 that replaces the area identification information included in the output information with an object corresponding to the area identification information, and a transmission unit 236 that transmits the replaced output information to a terminal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] In recent years, large-scale language models (hereinafter referred to as LLMs (Large Language Models)) such as ChatGPT have become popular. Patent Document 1 discloses a technology that uses LLMs to provide answers to user questions while referring to documents. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2023-076413 Summary of the Invention [Problem to be solved by the invention]

[0004] However, LLM sometimes outputs incorrect answers, making it difficult for users to determine whether the answers provided are correct or not.In addition, the answers output by LLM are sometimes difficult for users to understand.

[0005] The present invention has been made in consideration of these points, and aims to make it easier for users to determine whether an answer to a question is correct, while also making it easier for them to understand the content of the answer. [Means for solving the problem]

[0006] an identification information generation unit that generates area identification information for identifying an area in each area of ​​the document acquired by the document acquisition unit, a reception unit that receives a user's question from a terminal, a prompt generation unit that generates a prompt to request an answer to the question received by the reception unit, the prompt including the question, the plurality of objects included in the document, and an instruction to present area identification information corresponding to the object related to the answer, an output information acquisition unit that acquires output information output by the large-scale language model by inputting the prompt generated by the prompt generation unit into the large-scale language model, a replacement unit that replaces the area identification information included in the output information acquired by the output information acquisition unit with the object corresponding to the area identification information, and a transmission unit that transmits the replaced output information, in which the replacement unit has replaced the area identification information included in the output information with the object, to the terminal.

[0007] The information processing device may further include a document generation unit that executes a process of generating a replacement document, which is the document in which, for each area containing a figure among the plurality of objects, the figure is replaced with the area identification information corresponding to the area, and the prompt generation unit may generate the prompt by treating the replacement document and the figure extracted from the document and associated with the corresponding area identification information as the plurality of objects contained in the document, and including an instruction to present the area identification information corresponding to the figure related to the answer as an instruction to present the area identification information corresponding to the object related to the answer.

[0008] The document generation unit may further execute a process of generating the replaced document by replacing the figure with a summary that summarizes the content of the figure for each of the areas including the figure.

[0009] The prompt generation unit may generate the prompt including an instruction to present the area identification information adjacent to the sentence related to the answer in the replacement document as an instruction to present the area identification information corresponding to the figure related to the answer.

[0010] The information processing device may further include a document generation unit that executes a process of generating one or more area documents that are the documents of hierarchical areas for each hierarchical level by repeatedly dividing the document into a plurality of areas and further dividing each area into a plurality of areas; the identification information generation unit may generate the area identification information corresponding to each hierarchical area for each hierarchical area; the prompt generation unit may treat the one or more area documents included in each of the plurality of hierarchies as the plurality of objects included in the document, and generate the prompt that includes an instruction to present the area identification information corresponding to the area document related to the answer as an instruction to present the area identification information corresponding to the object related to the answer; the replacement unit may replace the area identification information included in the output information acquired by the output information acquisition unit with the area document corresponding to the area identification information; and the transmission unit may transmit to the terminal the replaced output information in which the replacement unit has replaced the area identification information included in the output information with the area document.

[0011] The information processing device may further have a selection unit that selects an area document related to the question from the one or more area documents in each of the multiple hierarchies, and further selects at least one of an upper area document included in a hierarchical level above the hierarchical level that includes the selected area document, which is the selected area document, and a lower area document included in a hierarchical level below the hierarchical level that includes the selected area document, and which includes at least a portion of the selected area document, and the prompt generation unit may generate the prompt that includes the multiple area documents selected by the selection unit as the one or more area documents included in each of the multiple hierarchical levels.

[0012] The prompt generating unit may further generate the prompt including a summary sentence summarizing the domain document to be included in the prompt.

[0013] The document may include an image displaying a character string, and the prompt generation unit may further generate the prompt including a text sentence in which the character string of the image included in the area document to be included in the prompt is converted into text.

[0014] The prompt generating unit may further generate the prompt including information indicating a relationship between the hierarchy and the domain document.

[0015] The document generation unit may execute a process of generating the one or more area documents for each layer when the amount of information of the document is equal to or greater than a predetermined threshold.

[0016] An information processing method according to a second aspect of the present invention includes the steps of: acquiring a document composed of a plurality of objects including text and figures; generating region identification information for identifying each region in the acquired document that includes the objects; receiving a user's question from a terminal; generating a prompt for requesting an answer to the received question, the prompt including the question, the plurality of objects included in the document, and an instruction to present region identification information corresponding to the object related to the answer; acquiring output information output by the large-scale language model by inputting the generated prompt into a large-scale language model; replacing the region identification information included in the acquired output information with the object corresponding to the region identification information; and transmitting replaced output information in which the region identification information included in the output information has been replaced with the object to the terminal.

[0017] a reception unit for receiving user questions from a terminal; a prompt generation unit for generating a prompt for requesting an answer to the question received by the reception unit, the prompt including the question, the plurality of objects included in the document, and an instruction to present the area identification information corresponding to the object related to the answer; an output information acquisition unit for acquiring output information output by the large-scale language model by inputting the prompt generated by the prompt generation unit into the large-scale language model; a replacement unit for replacing the area identification information included in the output information acquired by the output information acquisition unit with the object corresponding to the area identification information; and a transmission unit for transmitting, to the terminal, the replaced output information in which the replacement unit has replaced the area identification information included in the output information with the object. [Effects of the Invention]

[0018] The present invention has the effect of making it easier for a user to determine whether an answer to a question is correct, and also to understand the content of the answer. [Brief explanation of the drawings]

[0019] [Figure 1] FIG. 1 is a diagram illustrating a configuration of an information processing system. [Figure 2] FIG. 1 is a block diagram of an information processing device. [Figure 3] FIG. 10 is a diagram illustrating an example of a process in which an information processing device generates a replaced document. [Figure 4] FIG. 10 is a diagram illustrating an example of a process in which an information processing apparatus generates an area document. [Figure 5] 10 is a flowchart showing a flow of processing by the information processing device. DETAILED DESCRIPTION OF THE INVENTION

[0020] [Outline of Information Processing System S] FIG. 1 is a diagram showing the configuration of an information processing system S. The information processing system S is a system used to provide an information processing service. The information processing service is a service that presents answers to questions. The information processing system S has a user terminal 1 and an information processing device 2.

[0021] The user terminal 1 is a terminal used by a user, such as a smartphone, a tablet terminal, a personal computer, or XR (Extended Reality) glasses. For example, a dedicated application program (hereinafter referred to as a "dedicated app") for providing an information processing service is installed on the user terminal 1. By using the dedicated app, the user can obtain answers to questions input by hand or voice.

[0022] The information processing device 2 is a device that manages information processing services, and is, for example, a server. The information processing device 2 manages document data. The document data is data including a document and an object associated with an area ID. The document is a document composed of multiple objects including text and figures, such as a product manual. The object includes at least one of text and figures. The document may include at least a portion of text, or may include at least a portion of an image. The area ID is information for identifying an area in the document that includes an object. Note that the area ID may also be information for identifying an object included in the document. The information processing device 2 also manages an LLM. The LLM is, for example, a known machine learning model such as ChatGPT. The processing executed by the information processing system S will be described below.

[0023] First, when a user inputs a question using the user terminal 1, the user terminal 1 transmits the question input by the user to the information processing device 2 ((1) in FIG. 1). When the information processing device 2 receives the user's question, it generates a prompt corresponding to the question ((2) in FIG. 1). The prompt is information for requesting an answer to the user's question. The prompt includes, for example, the user's question, multiple objects contained in the document, and an instruction to present the area ID corresponding to the object related to the answer.

[0024] The information processing device 2 inputs the generated prompt into the LLM to obtain output information output by the LLM ((3) in FIG. 1). The output information includes an answer sentence indicating an answer to the user's question and an area ID corresponding to an object related to the answer. The information processing device 2 replaces the area ID included in the output information with the object corresponding to the area ID ((4) in FIG. 1).

[0025] Then, the information processing device 2 transmits replaced output information in which the area ID included in the output information is replaced with the object corresponding to the area ID to the user terminal 1 ((5) in FIG. 1). Thereafter, the user terminal 1 displays the replaced output information on a display.

[0026] In this way, the information processing system S can present to the user an object in the document related to the answer along with the answer to the user's question. This allows the user to check the answer while referring to the object. As a result, the information processing system S can make it easier for the user to determine whether the answer to the question is correct. Furthermore, by presenting to the user the answer to the user's question and the object related to the answer, the information processing system S can make it easier for the user to understand the content of the answer.

[0027] Furthermore, when a single document is input directly into an LLM, the greater the amount of information in the document (e.g., the number of character strings or the number of objects), the greater the likelihood that the key points of each object contained in the document will be diluted or missing. If the key points of each object contained in the document are diluted or missing, the LLM will be unable to accurately interpret the key points of each object, which may reduce the accuracy of inferring answers to questions for the user. In response to this, the information processing system S can prevent the key points of each object from being diluted or missing by dividing the document into multiple objects and inputting them into the LLM. This allows the information processing system S to accurately interpret the key points of each object. As a result, the information processing system S can improve the accuracy of inferring answers to questions for the user. The configuration of the information processing device 2 will be described below.

[0028] [Configuration of information processing device 2] FIG. 2 is a block diagram of an information processing device 2. In FIG. 2, arrows indicate main data flows, and data flows other than those shown in FIG. 2 may also exist. In FIG. 2, each block indicates a functional configuration rather than a hardware (device) configuration. Therefore, the blocks shown in FIG. 2 may be implemented in a single device, or may be implemented separately in multiple devices. Data may be exchanged between blocks via any means, such as a data bus, a network, or a portable storage medium.

[0029] The information processing device 2 has a communication unit 21, a storage unit 22, and a control unit 23. The information processing device 2 may be configured by two or more physically separate devices connected by wire or wirelessly. The information processing device 2 may also be configured by a cloud, which is a collection of computer resources.

[0030] The communication unit 21 is a communication interface for connecting to a network, and includes a communication controller for receiving data from an external device.

[0031] The memory unit 22 is a large-capacity storage device such as a ROM (Read Only Memory) that stores the BIOS (Basic Input Output System) of the computer that realizes the information processing device 2, a RAM (Random Access Memory) that serves as the working area of ​​the information processing device 2, an HDD (Hard Disk Drive) or SSD (Solid State Drive) that stores the OS (Operating System), application programs, and various information referenced when the application programs are executed.

[0032] The control unit 23 is a processor such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), or NPU (Neural network Processing Unit) of the information processing device 2, and functions as a document acquisition unit 231, a generation unit 232, a reception unit 233, an output information acquisition unit 234, a replacement unit 235, a transmission unit 236, a selection unit 237, and a judgment unit 238 by executing programs stored in the memory unit 22.

[0033] The document acquisition unit 231 acquires a document that is made up of a plurality of objects including text and figures. The document acquisition unit 231 acquires the document from, for example, a terminal (not shown) used by an administrator of the information processing device 2. The document acquisition unit 231 may acquire the document by reading out a document that is stored in advance in the storage unit 22.

[0034] Furthermore, the document acquisition unit 231 may acquire a document attached by the user when the later-described acceptance unit 233 accepts a question from the user. The document attached by the user may be a document file stored in advance in the user terminal 1, or may be a captured image of the document captured by a camera of the user terminal 1.

[0035] The generation unit 232 functions as an identification information generation unit, and generates an area ID for identifying each area that includes an object in the document acquired by the document acquisition unit 231. The details will be described later, but the area for which the generation unit 232 generates an area ID includes one or more objects.

[0036] For example, first, the generation unit 232 analyzes the document acquired by the document acquisition unit 231 and identifies the type of object (e.g., text or image) included in the document and the area of ​​the object. The generation unit 232 can identify the type of object included in the document and the area of ​​the object, for example, using a known technique. Then, for each area of ​​the identified object, the generation unit 232 generates an area ID corresponding to the area.

[0037] After generating the area ID, the generation unit 232 stores in the storage unit 22 document data including a document ID for identifying the document acquired by the document acquisition unit 231, and a plurality of objects each associated with a plurality of area IDs. The objects stored in the storage unit 22 are, for example, text sentences indicating sentences included in the document, images displaying the sentences, or images displaying diagrams included in the document. The generation unit 232 associates the area ID with the object, for example, by adding the area ID to the file name of the object.

[0038] The generation unit 232 may further store in the storage unit 22 document data further including a summary that summarizes the document acquired by the document acquisition unit 231. The document summary is, for example, information output by the LLM when a prompt requesting a summary of the document acquired by the document acquisition unit 231 is input to the LLM. The generation unit 232 may further store in the storage unit 22 document data further including, for each area in the document acquired by the document acquisition unit 231 that includes an object (for each area for which an area ID has been generated), a summary that summarizes an object included in the area. The object summary is, for example, information output by the LLM when a prompt requesting a summary of the object is input to the LLM.

[0039] The reception unit 233 receives a user's question from the user terminal 1. For example, the reception unit 233 receives the user's question by acquiring a question input by the user using the user terminal 1. The reception unit 233 may further receive, from the user terminal 1, a designation of a document related to the user's question. For example, a list of documents is displayed on the display screen of the dedicated app, and when the user performs an operation to designate a specific document from the list of documents, the reception unit 233 receives the designation of the document by acquiring, from the user terminal 1, a document ID corresponding to the document designated by the user. Furthermore, as described above, the reception unit 233 may acquire a document related to the question along with the user's question.

[0040] The generation unit 232 further functions as a prompt generation unit, and when the reception unit 233 receives a user's question, generates a prompt for requesting an answer to the user's question, the prompt including the user's question, multiple objects included in a document, and an instruction to present an area ID corresponding to an object related to the answer. The generation unit 232 generates, for example, a prompt with multiple objects attached, the prompt including a request statement such as "Please answer the user's question by referring to the multiple attached objects. Also, please present an area ID associated with the object related to the answer to the user's question." The generation unit 232 may generate a prompt that further includes a document acquired by the document acquisition unit 231.

[0041] When multiple documents are stored in the storage unit 22, the generation unit 232 may generate a prompt corresponding to a document related to the user's question among the multiple documents. For example, when the reception unit 233 acquires a document ID along with the user's question, the generation unit 232 identifies a document corresponding to the document ID acquired by the reception unit 233 from the multiple documents stored in the storage unit 22, and generates a prompt corresponding to the identified document.

[0042] Furthermore, when the document acquisition unit 231 acquires a captured image attached by a user, the generation unit 232 may perform a vector search on the multiple documents stored in the storage unit 22, select a document that is the same as the document included in the captured image from the multiple documents, and generate a prompt corresponding to the selected document. For example, first, the selection unit 237 converts each document stored in the storage unit 22 into a vector, and then converts the captured image acquired by the document acquisition unit 231 into a vector. The information processing device 2 can convert documents, captured images, etc. into vectors using a known technique such as Word2Vec, for example.

[0043] After converting the captured image and the document into vectors, the selection unit 237 calculates, for each document stored in the storage unit 22, a similarity between the vector data obtained by converting the document into a vector and the vector data obtained by converting the captured image into a vector. The selection unit 237 selects, from among the documents stored in the storage unit 22, a document corresponding to a similarity indicating the highest degree of similarity as the document included in the captured image. Then, the generation unit 232 generates a prompt corresponding to the selected document.

[0044] Based on the user's question, the generation unit 232 may identify a document related to the user's question from among a plurality of documents stored in the storage unit 22. For example, first, the generation unit 232 converts each summary sentence of a document stored in the storage unit 22 into a vector, and then converts the user's question accepted by the acceptance unit 233 into a vector.

[0045] After converting each of the document abstracts and the user's question into vectors, the generation unit 232 calculates, for each document abstract stored in the storage unit 22, a similarity between the vector data obtained by converting the document abstract into a vector and the vector data obtained by converting the user's question into a vector. The generation unit 232 then identifies, from among the multiple documents stored in the storage unit 22, the document whose abstract corresponds to the similarity indicating the highest degree of similarity, as the document related to the user's question. Thereafter, the generation unit 232 generates a prompt corresponding to the identified document.

[0046] The output information acquisition unit 234 acquires the output information output by the LLM by inputting the prompt generated by the generation unit 232 to the LLM. The output information includes an answer sentence indicating the answer to the user's question and an area ID corresponding to the object related to the answer.

[0047] The replacement unit 235 replaces the area ID included in the output information acquired by the output information acquisition unit 234 with an object corresponding to the area ID. For example, the replacement unit 235 generates replaced output information by replacing the area ID included in the output information acquired by the output information acquisition unit 234 with an object stored in the storage unit 22 in association with the area ID.

[0048] The transmitting unit 236 transmits the replaced output information in which the replacing unit 235 replaced the area ID included in the output information with the object to the user terminal 1. After that, the user terminal 1 displays, on the display, the answer sentence indicating the answer to the user's question and the object as the replaced output information.

[0049] The information processing device 2 transmits, for example, replaced output information in which the area ID included in the output information is replaced with a diagram from among the sentences and diagrams included in the document, to the user terminal 1. The information processing device 2 executes the following five steps to transmit replaced output information in which the area ID included in the output information is replaced with a diagram to the user terminal 1.

[0050] In a first step, when the document acquisition unit 231 acquires a document, the generation unit 232 generates an area ID corresponding to an area that includes one figure, treating the area as a target area for generating an area ID. Specifically, the generation unit 232 generates an area ID corresponding to each area that includes a figure among multiple objects included in the document. For example, the generation unit 232 targets only figures among the multiple types of identified objects, and generates an area ID corresponding to each area of ​​the identified figure.

[0051] In a second step, the generation unit 232 further functions as a document generation unit, and executes a process of generating a replaced document in which, for each area in the document that includes a figure among multiple objects, the figure is replaced with an area ID corresponding to the area. For example, the generation unit 232 extracts each area of ​​the identified figure by cutting out the area from the document, and inputs the area ID corresponding to the area into the extracted area in the document, thereby generating a replaced document in which each figure is replaced with the corresponding area ID. The generation unit 232 associates the extracted area (figure) with the area ID corresponding to the area and stores the result in the storage unit 22.

[0052] 3 is a diagram schematically illustrating an example of a process for generating a replaced document by the information processing device 2. In the example shown in FIG. 3, material D is a document that includes two figures (a first figure included in document area A1 and a second figure included in document area A2).

[0053] In this case, the generation unit 232 first generates "A0001" as the area ID corresponding to document area A1, and generates "A0002" as the area ID corresponding to document area A2. Then, the generation unit 232 generates document D', which is a replaced document in which the first figure included in document area A1 in document D is replaced with "A0001" and the second figure included in document area A2 is replaced with "A0002". In document D' shown in Figure 3, it can be seen that each figure included in document D has been replaced with its respective area ID.

[0054] The generation unit 232 may further include information other than the area ID as information to be replaced from a figure included in a document. Specifically, the generation unit 232 executes a process of generating a replaced document by replacing, for each area in the document that includes a figure, the figure with the area ID corresponding to the area and a summary that summarizes the content of the figure. In this way, the LLM can make inferences based on the summary of the figure included in the replaced document without analyzing the input figure, thereby enabling the information processing device 2 to shorten the time required for the LLM to make inferences.

[0055] In a third step, the generating unit 232 generates a prompt that includes, as a plurality of objects included in the document, the replacement document and a diagram extracted from the document and associated with the corresponding area ID, and includes an instruction to present the area ID corresponding to the diagram related to the answer as an instruction to present the area ID corresponding to the object related to the answer. The "diagram related to the answer" is, for example, a diagram that includes the answer to the user's question.

[0056] The generation unit 232 generates a prompt that includes, for example, a replacement document and a diagram extracted from the document, and includes a request statement such as, "Please answer the user's question by referring to the attached replacement document and diagram. Also, please provide the area ID associated with the diagram referenced to answer the user's question." Note that if the replacement document includes a summary of the diagram, the generation unit 232 may generate a prompt without including the diagram extracted from the document.

[0057] The "figure related to the answer" may be a figure adjacent to a sentence in the document that contains the answer to the user's question. In this case, the generation unit 232 generates a prompt including an instruction to present an area ID adjacent to the sentence related to the answer in the replacement document as an instruction to present an area ID corresponding to the figure related to the answer.

[0058] The generation unit 232 may generate a prompt that includes, for example, a replacement document and a diagram extracted from the document, and includes a request statement such as, "Please answer the user's question by referring to the attached replacement document and diagram. Also, please provide the area ID adjacent to the sentence related to the answer in the replacement document." The generation unit 232 may generate a prompt that includes a request statement such as, "Please answer the user's question by referring to the attached replacement document and diagram. Also, please provide the area ID associated with the diagram related to the answer, or the area ID adjacent to the sentence related to the answer in the replacement document." In this way, the information processing device 2 can identify diagrams that may be related to the answer to the user's question.

[0059] The generation unit 232 generates a prompt including, for example, the generated replacement document and all of the figures extracted from the document. The generation unit 232 may generate a prompt including a part of the generated replacement document and a part of the figures extracted from the document. Specifically, the generation unit 232 performs vector search on the generated replacement document and all of the figures extracted from the document, respectively, to identify a part of the replacement document that may be related to the user's question and select a figure that may be related to the user's question, and generate a prompt that includes the identified part of the replacement document and the selected figure.

[0060] More specifically, first, the selection unit 237 identifies a sentence region including a sentence related to the user's question from the replacement document generated by the generation unit 232. For example, first, the selection unit 237 converts each block of sentence included in the replacement document into a vector, and then converts the user's question received by the reception unit 233 into a vector. A block of sentence is, for example, a sentence that has a certain gap between it and other adjacent sentences, or a paragraph of sentence.

[0061] The selection unit 237 calculates, for each block of sentence, a similarity between vector data obtained by converting the sentence into a vector and vector data obtained by converting the user's question into a vector. The selection unit 237 then identifies a sentence region including a sentence corresponding to a similarity equal to or greater than a sentence threshold among the calculated similarities as part of a replacement document that may be related to the user's question. The sentence threshold is a threshold established for determining whether a sentence may be related to the user's question. In other words, the sentence threshold is a threshold established for excluding sentences that may not be related to the user's question. The sentence region may be a region including only sentences corresponding to a similarity equal to or greater than the sentence threshold, a region including sentences corresponding to a similarity equal to or greater than the sentence threshold and other sentences adjacent to the sentence, or a region of a chapter, section, or paragraph including sentences corresponding to a similarity equal to or greater than the sentence threshold.

[0062] Next, the selection unit 237 selects diagrams that may be related to the user's question from among the multiple diagrams extracted from the document. For example, the selection unit 237 first converts each diagram extracted from the document into a vector. The selection unit 237 can convert the diagram into a vector using, for example, a known technique. For each diagram, the generation unit 232 calculates the similarity between the vector data obtained by converting the diagram into a vector and the vector data obtained by converting the user's question into a vector. The generation unit 232 selects, from one or more diagrams extracted from the document, diagrams corresponding to a similarity greater than or equal to the diagram threshold as part of the diagrams that may be related to the user's question. The diagram threshold is a threshold determined for determining whether a diagram is related to the user's question. In other words, the diagram threshold is a threshold determined for excluding diagrams that may not be related to the user's question.

[0063] The selection unit 237 may select a figure associated with an area ID included in an area adjacent to the identified text area as a figure that may be related to the user's question. The selection unit 237 may also specify, as a text area, an area including text adjacent to an area including an area ID associated with the selected figure in the replacement document.

[0064] When the selection unit 237 identifies a text region and selects a diagram, the generation unit 232 generates a prompt that includes the extracted document and the diagram, extracted by cutting out the text region from the replacement document. In this way, the information processing device 2 can exclude parts of the replacement document and diagrams that may not be related to the user's question, thereby shortening the time required for the LLM to make inferences.

[0065] In a fourth step, when the output information acquisition unit 234 acquires the output information, the replacement unit 235 replaces the area ID included in the output information with a figure corresponding to the area ID. For example, the replacement unit 235 generates the replaced output information by replacing the area ID included in the output information acquired by the output information acquisition unit 234 with a figure stored in the storage unit 22 in association with the area ID.

[0066] In a fifth step, the transmitting unit 236 transmits the replaced output information, in which the replacing unit 235 has replaced the area ID included in the output information with the graphic associated with the area ID, to the user terminal 1. Thereafter, the user terminal 1 displays on the display, as the replaced output information, an answer sentence and a graphic indicating the answer to the user's question.

[0067] In this way, by dividing a document into sentences and figures and inputting them into the LLM, the information processing device 2 can prevent the main points of each sentence and figure from being diluted or missing. As a result, the information processing device 2 can improve the accuracy of inferring answers to questions from users. Furthermore, although it may be difficult for a user to understand the content of an answer from the answer text output by the LLM alone, the information processing device 2 can make the content of the answer easier to understand by presenting the user with figures that may be related to the answer.

[0068] Furthermore, if a document is simply divided into sentences and figures and input into the LLM, the relationship between the positions of the sentences and figures in the document will be lost, which may reduce the accuracy with which the LLM infers answers to user questions. For example, in a product manual, the position of each sentence and each figure is important, so the LLM will not be able to infer answers to user questions by taking into account the position or order of each sentence and figure. In response to this, the information processing device 2 inputs a replaced document in which figures in the document are replaced with area IDs into the LLM, thereby dividing the document into sentences and figures while maintaining the relationship between the positions of the sentences and figures, thereby preventing a decrease in the accuracy with which the LLM infers answers to user questions.

[0069] The information processing device 2 may transmit replaced output information in which the area ID included in the output information has been replaced with an area document to the user terminal 1. The area document is at least a part of a document. Specifically, the information processing device 2 transmits replaced output information in which the area ID included in the output information has been replaced with an area document to the user terminal 1 by executing the following five steps.

[0070] In a first step, when the document acquisition unit 231 acquires a document, the generation unit 232 generates a plurality of area documents by dividing the document in stages. Specifically, the generation unit 232 executes a process of dividing the document into a plurality of areas, and repeatedly dividing each area into a plurality of areas to generate one or more area documents, which are documents for hierarchical areas for each hierarchy.

[0071] 4 is a diagram schematically illustrating an example of a process in which the information processing device 2 generates an area document. FIG. 4(a) shows an area document at the first level. In the example shown in FIG. 4(a), the generation unit 232 generates an area document at the first level for document area A3, where the entire document is one level area. Note that if the document is made up of multiple pages, the generation unit 232 may generate multiple area documents at the first level by dividing the document for each page.

[0072] FIG. 4(b) shows a second-level area document. The generation unit 232 generates multiple second-level area documents by dividing the first-level area document. In the example shown in FIG. 4(b), the generation unit 232 first analyzes the first-level area document and identifies each object included in the area document. The generation unit 232 identifies each object based on, for example, the differences in the types of the multiple objects, the width of the gaps between the multiple objects, etc.

[0073] Then, the generation unit 232 generates, as area documents of the second layer, area documents of document area A4 and document area A5, which are hierarchical areas obtained by dividing the area document of the first layer into two so that the two area documents are equal. The generation unit 232 generates area documents of document area A4 and document area A5 in the second layer by dividing the area document of the first layer into two so that, for example, the number of characters, the number of objects, or the area of ​​area R4 and area R5 are equal.

[0074] Fig. 4(c) shows a third-level area document. The generation unit 232 generates a third-level area document by dividing each second-level area document. In the example shown in Fig. 4(c), the generation unit 232 first analyzes each second-level area document and identifies each object included in each area document.

[0075] Then, the generation unit 232 generates, as a third-level area document, area documents for document area A6, document area A7, document area A8, document area A9, document area A10, document area A11, and document area A12, which are hierarchical areas obtained by dividing each area document in the second level so that each area document contains only one object. One object is a group of objects, such as an object (e.g., a figure) that is different in type from other adjacent objects (e.g., text), or an object that has a certain gap between it and other adjacent objects. One object may also be a paragraph of text.

[0076] The division unit for each area is not limited to the example shown in Fig. 4. For example, the generation unit 232 may divide each of one or more area documents included in each layer into two, and repeat the division of the area document until an area document containing only one object is obtained, thereby generating one or more area documents that are hierarchical area documents for each layer. Furthermore, when the amount of information in the document is small, the generation unit 232 may generate area documents for hierarchical areas by dividing the area document of the first layer so that only one object is included in the second layer.

[0077] In a second step, when the generation unit 232 generates the area document, it generates an area ID by treating the hierarchical area (area document) as the area for which the area ID is to be generated. Specifically, the generation unit 232 generates an area ID corresponding to each hierarchical area. When the generation unit 232 generates the area ID, it stores the generated area document in the storage unit 22 in association with the area ID. The generation unit 232 may further generate a hierarchical ID for identifying the hierarchical level, and store the generated area document in the storage unit 22 in association with the hierarchical ID.

[0078] In a third step, when the receiving unit 233 receives a user's question, the generating unit 232 generates a prompt that treats one or more area documents included in each of a plurality of hierarchies as a plurality of objects included in the document, and includes an instruction to present an area ID corresponding to an area document related to the answer as an instruction to present an area ID corresponding to an object related to the answer. The generating unit 232 generates a prompt that includes the generated area documents for each hierarchy attached thereto, and includes a request statement such as "Please answer the user's question by referring to the attached area documents. Also, please present the area ID associated with the area document referred to in answering the user's question."

[0079] The generation unit 232 may further generate a prompt including information indicating the relationship between the hierarchies and the domain documents. For example, the generation unit 232 generates a prompt that further includes a relationship list indicating the relationship between the hierarchies and the domain documents as information indicating the relationship between the hierarchies and the domain documents. The relationship list is, for example, a list in which, for each combination of a hierarchies and domain documents included in the hierarchies, a hierarchical ID corresponding to the hierarchies is associated with a domain document ID for identifying the domain documents. In this way, the information processing device 2 can cause the LLM to infer an answer to the user's question based on the relationship between the hierarchies and the domain documents.

[0080] When the document includes an image displaying a character string, the generation unit 232 may further generate a prompt including a text sentence in which the character string of the image included in the area document to be included in the prompt is converted into text. For example, the generation unit 232 may extract character strings from the area document by performing a process related to OCR (Optical Character Recognition) on the area document including the image, and generate a prompt including a text sentence in which the extracted character string is converted into text. In this case, the generation unit 232 may generate a prompt including a text sentence corresponding to the area document instead of the area document. In this way, the information processing device 2 can omit the process of the LLM analyzing the image.

[0081] The generation unit 232 may further generate a prompt including a summary that summarizes the domain document to be included in the prompt. The summary of the domain document is, for example, information output by the LLM by inputting a prompt requesting a summary of the domain document to the LLM for each generated domain document. In this case, the generation unit 232 may generate a prompt including the summary of the domain document instead of the domain document. In this way, the information processing device 2 can reduce the amount of information referenced by the LLM.

[0082] The generation unit 232 generates a prompt that includes, for example, all of the generated domain documents. The generation unit 232 may also generate a prompt that includes some of the generated domain documents. Specifically, the generation unit 232 performs a vector search on the generated domain documents to select domain documents that may be related to the user's question from the domain documents, and generates a prompt that includes the selected domain documents.

[0083] More specifically, first, the selection unit 237 selects a domain document related to the user's question from one or more domain documents in each of a plurality of layers (for example, a domain document generated by the generation unit 232, a text sentence of the domain document, or a summary sentence of the domain document). For example, first, the selection unit 237 converts each domain document into a vector, and converts the user's question received by the reception unit 233 into a vector. For each domain document, the selection unit 237 calculates a similarity between vector data obtained by converting the domain document into a vector and vector data obtained by converting the user's question into a vector. Then, the selection unit 237 selects the domain document corresponding to the highest similarity from the plurality of domain documents as the domain document related to the user's question.

[0084] When the selection unit 237 selects an area document related to the user's question, it selects at least one of an upper area document and a lower area document based on the selected area document, which is the selected area document. The upper area document is an area document that is included in a higher level of the level including the selected area document and that includes the selected area document. In the example shown in FIG. 4, when the area document related to the user's question is the area document of the document area A5, the upper area document is an area document of the document area A3 that is included in the first level, which is a level higher than the second level including the area document of the document area A5, and that includes the area document of the document area A5. The lower area document is an area document that is included in a lower level of the level including the selected area document and that includes at least a part of the selected area document. In the example shown in Figure 4, when the area document related to the user's question is the area document of document area A5, the sub-area documents are the area documents of document area A10, the area document of document area A11, and the area document of document area A12, which are included in the third hierarchy, which is a hierarchy lower than the second hierarchy that includes the area document of document area A5, and which include at least a portion of the area document of document area A5.

[0085] The selection unit 237 may select, for example, a higher-level area document corresponding to the level immediately above the level including the selected area document, and may select multiple lower-level area documents corresponding to the level immediately below the level including the selected area document. The selection unit 237 may select the higher-level area document and the lower-level area document that may be related to the user's question.

[0086] For example, in the process of selecting a higher-level area document, the selection unit 237 first determines the layer immediately above the layer containing the selected area document as the target layer, and calculates the similarity between vector data obtained by converting the higher-level area document corresponding to the target layer into vectors and vector data obtained by converting the user's question into vectors.The selection unit 237 then selects the higher-level area document corresponding to the calculated similarity if the calculated similarity is equal to or greater than the upper-level threshold.The upper-level threshold is a threshold set for determining whether the higher-level area document is related to the user's question.

[0087] If a higher-level area document is selected and there is a level above the target level, the selection unit 237 sets the level immediately above the target level as the target level, determines the selected higher-level area document as a new selected area document, and executes a series of processes, including calculating the similarity corresponding to the target level and selecting the higher-level area document corresponding to the similarity. The selection unit 237 repeatedly executes the series of processes until there is no level above the target level or the calculated similarity becomes less than the upper level threshold.

[0088] In addition, as a process for selecting a lower-area document, the selection unit 237 first determines the layer immediately below the layer including the selected area document as a target layer, and for each lower-area document corresponding to the target layer, calculates the similarity between vector data obtained by converting the layer area document into a vector and vector data obtained by converting the user's question into a vector. Then, for each lower-area document corresponding to the target layer, the selection unit 237 selects the lower-area document if the similarity corresponding to the layer area document is equal to or greater than a lower-layer threshold. The lower-layer threshold is a threshold set for determining whether the lower-area document may be related to the user's question.

[0089] If a lower-area document is selected and a layer exists below the target layer, the selection unit 237 sets the layer immediately below the target layer as the target layer, determines the selected lower-area document as a new selected-area document, and executes a series of processes: calculating the similarity corresponding to the target layer and selecting the lower-area document corresponding to the similarity. The selection unit 237 repeatedly executes the series of processes until there are no layers below the target layer or until the calculated similarity becomes less than the lower-layer threshold.

[0090] When the selection unit 237 selects multiple domain documents, the generation unit 232 generates a prompt that includes the multiple domain documents selected by the selection unit 237 as one or more domain documents included in each of multiple hierarchies (for example, the domain document generated by the generation unit 232, the text sentence of the domain document, or the summary sentence of the domain document). In this way, the information processing device 2 can exclude domain documents that may not be related to the user's question, thereby shortening the time required for the LLM to make inference.

[0091] In a fourth step, when the output information acquisition unit 234 acquires the output information, the replacement unit 235 replaces the area ID included in the output information with the area document corresponding to the area ID. For example, the replacement unit 235 generates replaced output information by replacing the area ID included in the output information acquired by the output information acquisition unit 234 with the area document stored in the storage unit 22 in association with the area ID.

[0092] As a fifth step, the sending unit 236 sends the replaced output information in which the replacing unit 235 replaced the area ID included in the output information with the area document associated with the area ID to the user terminal 1. Thereafter, the user terminal 1 displays on the display, as the replaced output information, an answer sentence indicating an answer to the user's question and the area document.

[0093] As mentioned above, the greater the amount of information in a document, the greater the likelihood that the key points of each object contained in the document will be diluted or lost. However, the information processing device 2 can prevent this by inputting multiple area documents into which the document is divided in stages into LLM. As a result, the information processing device 2 can improve the accuracy of inferring answers to user questions.

[0094] Depending on the amount of information in the document, the information processing device 2 may execute either a series of processes for transmitting replacement output information in which sentences and figures included in the document have been replaced with figures to the user terminal 1, or a series of processes for transmitting replacement output information in which sentences and figures have been replaced with area documents to the user terminal 1. For example, first, the determination unit 238 determines whether the amount of information in the document acquired by the document acquisition unit 231 is equal to or greater than an information amount threshold. The information amount threshold is a value set for determining whether the amount of information in a document is large.

[0095] If the determination unit 238 determines that the amount of information in the document is less than a predetermined threshold, the generation unit 232 executes a process of generating an area ID corresponding to each area containing a figure among the multiple objects included in the document. Thereafter, the information processing device 2 executes the remaining four steps of the five steps described above for transmitting replacement output information in which the sentences and figures included in the document have been replaced with figures to the user terminal 1.

[0096] On the other hand, if the determination unit 238 determines that the amount of information in the document is equal to or greater than a predetermined threshold, the generation unit 232 executes a process of generating one or more area documents for each layer. Thereafter, the information processing device 2 executes the remaining four steps of the five steps described above for transmitting the replaced output information substituted with the area document to the user terminal 1. In this way, the information processing device 2 can execute a process according to the amount of information in the document and transmit the replaced output information to the user terminal 1.

[0097] [Processing of information processing device 2] Next, a description will be given of the processing flow of the information processing device 2. Fig. 5 is a flowchart showing the processing flow of the information processing device 2. This processing starts when the document acquisition unit 231 acquires a document (S1). The generation unit 232 generates an area ID corresponding to each area that includes an object in the document acquired by the document acquisition unit 231 (S2).

[0098] The reception unit 233 receives a user's question from the user terminal 1 (S3). When the reception unit 233 receives the user's question, the generation unit 232 generates a prompt for requesting an answer to the user's question, the prompt including the question, a plurality of objects included in the document, and an instruction to present area IDs corresponding to objects related to the answer (S4).

[0099] The output information acquisition unit 234 acquires the output information output by the LLM by inputting the prompt generated by the generation unit 232 to the LLM (S5). The replacement unit 235 replaces the area ID included in the output information acquired by the output information acquisition unit 234 with an object corresponding to the area ID (S6). Then, the transmission unit 236 transmits the replaced output information in which the replacement unit 235 replaced the area ID included in the output information with the object to the user terminal 1 (S7).

[0100] [Effects of this embodiment] As described above, the information processing device 2 first acquires output information output by the LLM by inputting a prompt for requesting an answer to a user's question, the prompt including the question, multiple objects included in the document, and an instruction to present area IDs corresponding to objects related to the answer.The information processing device 2 then transmits replaced output information, in which the area IDs included in the acquired output information are replaced with objects corresponding to the area IDs, to the user terminal 1.

[0101] In this way, the information processing device 2 can present the user with an answer to the user's question as well as an object in the document related to the answer. This allows the user to check the answer while referring to the object. As a result, the information processing device 2 can make it easier for the user to determine whether the answer to the question is correct and to understand the content of the answer.

[0102] Furthermore, this invention will make it possible to contribute to Goal 9 of the United Nations' Sustainable Development Goals (SDGs), which is "Build resilient infrastructure, promote inclusive and sustainable industrialization, and promote innovation and resilience."

[0103] The present invention has been described above using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by functionally or physically distributing or integrating any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments. [Explanation of symbols]

[0104] 1. User terminal 2. Information processing equipment 21 Communications Department 22 Memory section 23 Control Unit 231 Document Acquisition Department 232 Generation part 233 Reception Department 234 Output information acquisition unit 235 Replacement part 236 Transmitter 237 Selection Section 238 Judgment section S Information Processing System

Claims

1. a document acquisition unit that acquires a document composed of a plurality of objects including text and figures; an identification information generating unit that generates, for each area in the document acquired by the document acquisition unit, area identification information for identifying the area including the object; a reception unit that receives questions from users from the terminal; a prompt generation unit that generates a prompt for requesting an answer to the question received by the reception unit, the prompt including the question, the plurality of objects included in the document, each of the plurality of objects being associated with region identification information corresponding to each of the plurality of regions, and an instruction to present the region identification information corresponding to the object related to the answer; an output information acquisition unit that acquires output information output by a large-scale language model by inputting the prompt generated by the prompt generation unit into the large-scale language model; a replacement unit that replaces the area identification information included in the output information acquired by the output information acquisition unit with the object corresponding to the area identification information; a transmitting unit that transmits, to the terminal, replaced output information obtained by the replacing unit replacing the area identification information included in the output information with the object; An information processing device having the above.

2. the information processing device further includes a document generation unit that executes, for each of the areas including the figure among the plurality of objects, a process of generating a replaced document, which is the document in which the figure is replaced with the area identification information corresponding to the area; the prompt generation unit generates the prompt, which sets the replaced document and the drawing extracted from the document and associated with the corresponding region identification information as the plurality of objects included in the document, and includes an instruction to present the region identification information corresponding to the drawing related to the answer as an instruction to present the region identification information corresponding to the object related to the answer. The information processing device according to claim 1 .

3. the document generation unit further executes a process of generating the replaced document by replacing the figure with a summary that summarizes the content of the figure for each of the regions including the figure. The information processing device according to claim 2 .

4. the prompt generation unit generates the prompt including an instruction to present the region identification information adjacent to the sentence related to the answer in the replacement document as an instruction to present the region identification information corresponding to the figure related to the answer.

4. The information processing device according to claim 2 or 3.

5. the information processing device further includes a document generation unit that executes a process of dividing the document into a plurality of areas, and further dividing each area into a plurality of areas, thereby generating one or more area documents that are the documents of hierarchical areas for each hierarchical level; the identification information generation unit generates, for each of the hierarchical regions, the region identification information corresponding to the hierarchical region; the prompt generation unit generates the prompt by setting the one or more area documents included in each of the plurality of hierarchies, the one or more area documents being associated with the area identification information corresponding to each of the plurality of hierarchical areas, as the plurality of objects included in the document, and including an instruction to present the area identification information corresponding to the area document related to the answer as an instruction to present the area identification information corresponding to the object related to the answer; the replacing unit replaces the area identification information included in the output information acquired by the output information acquiring unit with the area document corresponding to the area identification information; the transmitting unit transmits, to the terminal, the replaced output information obtained by the replacing unit replacing the area identification information included in the output information with the area document. The information processing device according to claim 1 .

6. The information processing device further comprises a selection unit that selects the area document related to the question from the one or more area documents in each of the plurality of hierarchies, and further selects at least one of an upper area document included in a hierarchical level above the hierarchical level including the selected area document, which is the selected area document, and a lower area document included in a hierarchical level below the hierarchical level including the selected area document, which is the lower area document including at least a part of the selected area document; the prompt generation unit generates the prompt including the plurality of area documents selected by the selection unit as the one or more area documents included in each of the plurality of hierarchies. The information processing device according to claim 5 .

7. The prompt generation unit further generates the prompt including a summary sentence summarizing the domain document to be included in the prompt.

7. The information processing device according to claim 5 or 6.

8. The document includes an image on which a character string is displayed, The prompt generation unit further generates the prompt including a text sentence obtained by converting a character string of the image included in the area document to be included in the prompt into text.

7. The information processing device according to claim 5 or 6.

9. The prompt generation unit further generates the prompt including information indicating a relationship between the hierarchy and the domain document.

7. The information processing device according to claim 5 or 6.

10. the document generation unit executes a process of generating the one or more area documents for each layer when the amount of information of the document is equal to or greater than a predetermined threshold.

7. The information processing device according to claim 5 or 6.

11. The computer executes obtaining a document consisting of a plurality of objects including text and graphics; generating area identification information for identifying each area in the acquired document that includes the object; receiving a question from a user from the terminal; generating a prompt for requesting an answer to the received question, the prompt including the question, the plurality of objects included in the document, each of the objects being associated with region identification information corresponding to a plurality of the regions, and an instruction to present the region identification information corresponding to the object related to the answer; inputting the generated prompt into a large-scale language model to obtain output information output by the large-scale language model; replacing the area identification information included in the acquired output information with the object corresponding to the area identification information; transmitting, to the terminal, replaced output information in which the area identification information included in the output information is replaced with the object; An information processing method comprising:

12. Computer, a document acquisition unit that acquires a document composed of a plurality of objects including text and figures; an identification information generation unit that generates, for each area including the object in the document acquired by the document acquisition unit, area identification information for identifying the area; a reception unit that receives questions from users from the terminal; a prompt generation unit that generates a prompt for requesting an answer to the question received by the reception unit, the prompt including the question, the plurality of objects included in the document, each of which is associated with region identification information corresponding to a plurality of the regions, and an instruction to present the region identification information corresponding to the object related to the answer; an output information acquisition unit that acquires output information output by a large-scale language model by inputting the prompt generated by the prompt generation unit into the large-scale language model; a replacement unit that replaces the area identification information included in the output information acquired by the output information acquisition unit with the object corresponding to the area identification information; and a transmitting unit that transmits, to the terminal, replaced output information obtained by the replacing unit replacing the area identification information included in the output information with the object; A program to function as a

Citation Information

Patent Citations

  • Information extraction method and training method and device of information extraction model

    CN118172786A

  • Method and system for extracting data from tables within regulatory content

    US11823477B1

  • Layout analysis system, layout analysis method, and program

    WO2024047764A1

  • Method, computer device, and computer program for providing dialogue dedicated to domain by using language model

    JP2023076413A