Search program, search method, and information processing device
Patent Information
- Application Number
- PCT/JP2025/012966
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012966_01102026_PF_FP_ABST
Abstract
Description
Search Program, Search Method, and Information Processing Apparatus
[0001] The present invention relates to a search program, a search method, and an information processing apparatus.
[0002] In recent years, in software development, document reading technology has been demanded due to the need for automating the creation and maintenance of design document groups. Furthermore, attempts have been progressing to cause Large Language Models (LLMs) to understand external documents, including Retrieval-Augmented Generation (RAG).
[0003] As a prior art, there is a technique that accepts a plurality of keywords a user desires to search for, and searches for link destination candidates matching the keyword string from the keyword string, which is the plurality of keywords, based on the hierarchical structure of a document tree. There is also a technique that extracts a target sentence, which is a sentence addressing a topic of interest indicated by user information, from among sentences included in content to be processed, and generates a summary sentence summarizing the content of the content to be processed based on the target sentence.
[0004] There is also a technique that analyzes text segments within a document section, identifies a text segment associated with a pointer indicating sample content, and generates a link between the identified text segment and an associated topic. There is also a technique that searches for information resources and displays search results in a collapsible / expandable format based on a display criterion or hierarchy selected by a user.
[0005] Japanese Unexamined Patent Application Publication No. 2015-162004 Japanese Unexamined Patent Application Publication No. 2021-131769 US Patent Application Publication No. 2016 / 0292153 US Patent No. 6175830
[0006] However, in the prior art, there is a problem that it is difficult to search for useful information from a design document group or the like when solving some task using an LLM.
[0007] In one aspect, an object of the present invention is to improve information search accuracy.
[0008] In one embodiment, a search program is provided that receives instruction information indicating a search target in one or more documents, obtains summary information which summarizes a part of the content corresponding to each of several nodes included in tree structure information which represents the contents of the one or more documents in a hierarchical structure, and searches for the information of the search target indicated by the received instruction information from the one or more documents by referring to the multiple summaries obtained.
[0009] According to one aspect of the present invention, it has the effect of improving the accuracy of information retrieval.
[0010] Figure 1 is an explanatory diagram showing one embodiment of the search method according to the embodiment. Figure 2 is an explanatory diagram showing an example of the system configuration of the information processing system 200. Figure 3 is a block diagram showing an example of the hardware configuration of the related information aggregation device 201. Figure 4 is a block diagram showing an example of the functional configuration of the related information aggregation device 201. Figure 5 is an explanatory diagram showing a specific example of a related information aggregation query. Figure 6 is an explanatory diagram showing a specific example of a task description prompt. Figure 7 is an explanatory diagram showing an example of the creation of tree structure information. Figure 8 is an explanatory diagram showing a specific example of tree structure information. Figure 9 is an explanatory diagram showing an example of the search process for information to be searched. Figure 10 is a flowchart showing an example of the preprocessing procedure of the related information aggregation device 201. Figure 11 is a flowchart (part 1) showing an example of the related information aggregation processing procedure of the related information aggregation device 201. Figure 12 is a flowchart (part 2) showing an example of the related information aggregation processing procedure of the related information aggregation device 201.
[0011] Embodiments of the search program, search method, and information processing device according to the present invention will be described in detail below with reference to the drawings.
[0012] (Embodiment) Figure 1 is an explanatory diagram showing one embodiment of the search method according to the embodiment. In Figure 1, the information processing device 101 is a computer that searches for information to be searched from one or more documents. One or more documents are the source information for the search, and their contents can be represented in a hierarchical structure.
[0013] In software development, test code is sometimes created to verify that the developed system functions as expected. There is a desire to use LLM (Language Level Management) to determine which test items should be tested.
[0014] For example, one could provide the LLM with a set of system design documents and execute a task (test item generation) to create a test specification document that describes the test items to be tested. Another possibility is to execute a task (revision determination) to have the LLM determine whether the test specification document used before the update can still be used after the update, or whether it should be revised, when the set of design documents is updated.
[0015] When using a Design Language Manager (LLM) to solve tasks such as test item generation and revision detection, it is sometimes necessary to input important sections and information from the design documents into the LLM. For example, if the size of the design documents exceeds the upper limit of the context length that can be input into the LLM, it will be necessary to extract important information and provide it to the LLM. Therefore, when solving tasks such as test item generation and revision detection, techniques for searching for useful information for the LLM from the design documents are required.
[0016] Traditionally, there is a technology called LLM agents. An LLM agent is a program in which an LLM acts as an agent and autonomously performs tasks using available tools. For example, an LLM agent could autonomously search for information to use when solving tasks such as test item generation or revision judgment using a tool that provides search functionality. Search functionality could include, for example, keyword search or vector search.
[0017] However, conventional LLM agents sometimes search for information that is only superficially similar to the information being sought, even if the topic is different. For example, if you are looking for a description of how to handle Server A when it is restarted (startup / shutdown category), the keyword "Server A" might influence the search, causing it to search for a description of the reliability requirements for Server A (reliability category).
[0018] To avoid such problems, it is conceivable to filter the information included in the search results based on the hierarchical structure of the design document set. For example, when searching for a description of how to handle server A during restart (startup / shutdown category), if the searched information falls under the reliability category, it can be excluded from the search results. However, conventional LLM agents have difficulty understanding the hierarchical structure of large design document sets, and are unable to determine if the information found using the search tool is incorrect.
[0019] Therefore, in this embodiment, a search method that improves the accuracy of information retrieval by enabling searches based on the hierarchical structure of one or more documents will be described. Here, an example of processing by the information processing device 101 will be described.
[0020] (1) The information processing device 101 receives instruction information 120 that indicates the search target in the document group 110. Here, the document group 110 includes one or more documents. The document group 110 is information whose contents can be represented in a hierarchical structure. The document group 110 may be, for example, a set of design documents for a system under development. The document group 110 may also be a collection of articles of the Constitution or laws.
[0021] The search target is the information to be searched from the document group 110. For example, if the document group 110 is a set of design documents, the search target may be information used to solve specific tasks such as test item generation or revision determination. Also, if the document group 110 is a collection of legal provisions, the search target may be information used to solve specific tasks such as case law analysis.
[0022] (2) The information processing device 101 acquires summary information that summarizes a part of the content of the document group 110, corresponding to each of the multiple nodes included in the tree structure information 130, which represents the content of the document group 110 in a hierarchical structure. Each node included in the tree structure information 130 is associated with summary information that summarizes the information corresponding to each node in the immediately lower hierarchy connected to that node.
[0023] The tree structure information 130 is a tree structure with two or more levels, and each level contains one or more nodes. The top-level node in the tree structure is called the root node, and the lowest-level node is called a leaf node. The connection between adjacent nodes in different levels is called an edge. Of the two nodes connected by an edge, the higher-level node is called the parent node, and the lower-level node is called the child node.
[0024] Specifically, for example, the information processing device 101 obtains summary information 141, 142, and 143, respectively, associated with the second nodes 132, 133, and 134 (child nodes) in the direct subordinate hierarchy connected to the first node 131 (parent node) included in the tree structure information 130. The tree structure information 130 may be created in the information processing device 101, or it may be created in advance and stored in the information processing device 101.
[0025] (3) The information processing device 101 refers to the acquired summary information and searches for the information to be searched for, as instructed by the received instruction information 120, from the document group 110. Specifically, for example, the information processing device 101 refers to the acquired summary information 141, 142, 143 and identifies the second node among the second nodes 132, 133, 134 that is related to the search target.
[0026] Identifying a second node related to the search target is performed, for example, using an LLM (Large-Scale Language Model). Here, let's assume that the second node 134 is identified as the second node related to the search target. In this case, the information processing device 101 searches for information about the search target from the information 150 under the identified second node 134. The information 150 under the second node refers to information at a lower level that branches off from the second node 134.
[0027] Thus, the information processing device 101 makes it possible to narrow down the search range to the relevant parts of the search target by referring to summary information which summarizes a part of the content of the hierarchically structured document group 110, thereby improving the accuracy of information retrieval.
[0028] In the example shown in Figure 1, the information processing device 101 can narrow down the search range to the information 150 under the second node 134 related to the search target by referring to the summary information 141, 142, and 143. This allows the information processing device 101 to improve the accuracy of information retrieval by suppressing cases where information that is only superficially similar to the search target is included in the search results.
[0029] (Example of System Configuration of Information Processing System 200) Next, an example of the system configuration of the information processing system 200, including the information processing device 101 shown in Figure 1, will be described. Here, the example will be given of the case in which the information processing device 101 shown in Figure 1 is applied to the related information aggregation device 201 within the information processing system 200. The related information aggregation device 201 is applied to, for example, the LLM agent.
[0030] Figure 2 is an explanatory diagram showing an example of the system configuration of the information processing system 200. In Figure 2, the information processing system 200 includes an associated information aggregation device 201 and a client device 202. In the information processing system 200, the associated information aggregation device 201 and the client device 202 are connected via a wired or wireless network 210. The network 210 is, for example, the Internet, a LAN (Local Area Network), or a WAN (Wide Area Network).
[0031] Here, the related information aggregation device 201 is a computer that searches for the target information from one or more documents and outputs the search results. The target information is, for example, information related to a specific task and is used when solving a specific task using an LLM (for example, the LLM 420 shown in Figure 4 below). The related information aggregation device 201 is, for example, a server.
[0032] The client device 202 is a computer used by a user of the information processing system 200. The user is, for example, a software developer. The client device 202 may be, for example, a PC (Personal Computer), a tablet PC, etc.
[0033] In this example, the related information aggregation device 201 and the client device 202 are provided as separate components, but this is not the only option. For example, the related information aggregation device 201 may be implemented by the client device 202. Furthermore, the information processing system 200 may include multiple client devices 202.
[0034] (Example of hardware configuration of the related information aggregation device 201) Next, an example of the hardware configuration of the related information aggregation device 201 will be described.
[0035] Figure 3 is a block diagram showing an example of the hardware configuration of the related information aggregation device 201. In Figure 3, the related information aggregation device 201 includes a CPU (Central Processing Unit) 301, memory 302, disk drive 303, disk 304, communication interface 305, portable recording medium interface 306, and portable recording medium 307. Each component is connected by a bus 300.
[0036] Here, the CPU 301 is responsible for the overall control of the related information aggregation device 201. The CPU 301 may have multiple cores. The memory 302 includes, for example, ROM (Read Only Memory) and RAM (Random Access Memory). The program stored in the memory 302 is loaded into the CPU 301, causing the CPU 301 to execute the coded processing.
[0037] The disk drive 303 controls the reading and writing of data to the disk 304 according to the control of the CPU 301. The disk 304 stores the data written under the control of the disk drive 303. The disk 304 is, for example, a magnetic disk, an optical disk, etc.
[0038] The communication interface 305 is connected to the network 210 via a communication line, and through the network 210, it is connected to an external computer (for example, the client device 202 shown in Figure 2). The communication interface 305 manages the interface between the network 210 and the inside of the device, and controls the input and output of data from the external computer. The communication interface 305 is, for example, a modem or a LAN adapter.
[0039] The portable recording medium interface 306 controls the reading and writing of data to the portable recording medium 307 according to the control of the CPU 301. The portable recording medium 307 stores the data written under the control of the portable recording medium interface 306. The portable recording medium 307 is, for example, a CD (Compact Disc)-ROM, a DVD (Digital Versatile Disc), or a USB (Universal Serial Bus) memory.
[0040] Furthermore, the related information aggregation device 201 may have, in addition to the components described above, a GPU (Graphics Processing Unit), GPU memory, an input device, a display, etc. Also, the related information aggregation device 201 does not have to have, for example, the portable recording medium I / F 306 and the portable recording medium 307 among the components described above. The client device 202 shown in Figure 2 can also be realized with the same hardware configuration as the related information aggregation device 201. However, the client device 202 may have, in addition to the components described above, an input device, a display, etc.
[0041] (Example of functional configuration of the related information aggregation device 201) Next, an example of the functional configuration of the related information aggregation device 201 will be described.
[0042] FIG. 4 is a block diagram showing an example of the functional configuration of the related information aggregating apparatus 201. In FIG. 4, the related information aggregating apparatus 201 includes a reception unit 401, a creation unit 402, an acquisition unit 403, a search unit 404, an output unit 405, and a storage unit 410. The reception unit 401 to output unit 405 are functions serving as a control unit 400. Specifically, for example, the functions are realized by causing a CPU 301 to execute a program stored in a storage device such as the memory 302, the disk 304, or the portable recording medium 307 shown in FIG. 3, or by a communication I / F 305. Processing results of each functional unit are stored, for example, in a storage device such as the memory 302 or the disk 304. The storage unit 410 is realized by a storage device such as the memory 302 or the disk 304, for example. Specifically, for example, the storage unit 410 stores an LLM 420.
[0043] The reception unit 401 receives a document group. Here, the document group is a set of documents (document data) described in natural language, and includes one or more documents. The document group is, for example, a group of design documents related to a system under development. A design document is, for example, a specification document that describes functions, requirements, and the like of a system.
[0044] Specifically, for example, the reception unit 401 may receive a received document group by receiving the document group from the client apparatus 202 shown in FIG. 2. The reception unit 401 may also receive a document group via a user's operation input using an input device (not shown) of the own apparatus.
[0045] The reception unit 401 also receives a related information aggregation query. Here, the related information aggregation query corresponds to instruction information that indicates a search target. For the related information aggregation query, for example, an explanatory text indicating what kind of information the search target is may be indicated, or specific information related to the search target may be indicated. The related information aggregation query is created, for example, for each search target that a user searches for from the document group.
[0046] Further, the receiving unit 401 receives a task description prompt. Here, the task description prompt is a prompt for requesting execution of a related information aggregation task. The related information aggregation task is a task of searching for information of a search target specified by a related information aggregation query. The search processing performed by the search unit 404 corresponds to, for example, the related information aggregation task. The search target is used, for example, when solving a specific task (subsequent task) executed after the related information aggregation task.
[0047] The task description prompt may include supplementary information for improving search efficiency or search accuracy when searching for information of a search target. The supplementary information may be, for example, information representing characteristics of a document group. The information representing characteristics of a document group may specify what kind of application (spreadsheet software, word processing software, etc.) is used to create the document group. The information representing characteristics of a document group may be information that the document group is categorized.
[0048] Further, the supplementary information may be information useful for reading and understanding the document group. The information useful for reading and understanding the document group may indicate categories for classifying the document group. Further, the supplementary information may be information representing key points for searching for information of a search target. The information representing key points may indicate what purpose the search target information is used for in a specific task (subsequent task).
[0049] Further, the supplementary information may be information for controlling a thinking process when searching for information of a search target. The information for controlling the thinking process may be intended to improve the efficiency of search processing or cause an operation to prevent an infinite loop. The task description prompt is created, for example, by a user for each document group.
[0050] Specifically, for example, the reception unit 401 may accept the received related information aggregation query and task description prompt by receiving the related information aggregation query and task description prompt from the client device 202. Alternatively, the reception unit 401 may accept the related information aggregation query and task description prompt by user input using the device's input device (not shown).
[0051] Related information aggregation queries and specific examples of related information aggregation queries will be described later using Figures 5 and 6.
[0052] The creation unit 402 creates tree structure information that represents the contents of the document group in a hierarchical structure. Here, the tree structure information includes, for example, a tree graph in which the node representing the document group is the root node, and the nodes representing chunks divided from each document included in the document group are the leaf nodes. A chunk is partial data obtained by dividing a document into predetermined units.
[0053] The predetermined unit is a small, cohesive unit that can be arbitrarily defined. For example, the predetermined unit might be a sentence unit of about 1 to 3 sentences. Alternatively, the predetermined unit might be a paragraph unit. Each leaf node is associated with the content of a chunk. The content of a chunk is the content (original text) of the section of the document group corresponding to that chunk.
[0054] Each node between the root node and the leaf nodes is associated with summary information that summarizes the information corresponding to each node in the immediately lower hierarchy connected to that node. The root node may also be associated with summary information that summarizes the information corresponding to each node in the immediately lower hierarchy connected to the root, or it may not be associated with any summary information.
[0055] Specifically, for example, the creation unit 402 creates tree structure information based on the received document group and hierarchical definition information. Here, the hierarchical definition information is information that defines the hierarchical structure of the document group. The hierarchical definition information is created, for example, according to the characteristics of the document group.
[0056] As an example, suppose the document group is a collection of documents created using spreadsheet software, and the information is categorized. In this case, the hierarchical definition information may define the document group as having a hierarchical structure such as "Document Group (Top Level) - Category - Document - Sheet - Chunk (Lowest Level)". The hierarchical definition information also includes information that allows identification of which document in the document group belongs to which category.
[0057] Furthermore, let's assume that the document group is a collection of documents created using word processing software, and that the information is divided into chapters. In this case, the hierarchical definition information may, for example, define the document group as having a hierarchical structure such as "document group (top level) - document - chapter - chunk (bottom level)". Note that each chapter in a document can be identified, for example, by a table of contents or indentation.
[0058] The hierarchical definition information may be received, for example, by the receiving unit 401 along with the document group. Alternatively, the creation unit 402 may create the hierarchical definition information by performing multi-stage clustering based on the contents of each document group to identify the hierarchical structure of the document group.
[0059] To explain in more detail, for example, first, the creation unit 402 divides each document in the received document group into chunks. Then, the creation unit 402 refers to the hierarchical definition information and creates a tree graph in which the node representing the document group is the root node and the nodes representing each divided chunk are leaf nodes. Next, the creation unit 402 associates the contents of the chunk corresponding to each leaf node in the tree graph.
[0060] Then, the creation unit 402 creates summary information by summarizing the information corresponding to each child node of a node, starting from the lower-level nodes, excluding leaf nodes, and associates the created summary information with the node. A child node is a node in the immediate lower level that is connected to a given node, which is the parent node.
[0061] The information corresponding to a child node is, if the child node is a leaf node, the content of the chunk associated with that child node. On the other hand, if the child node is not a leaf node, the information corresponding to the child node is the summary information associated with that child node. The summary is performed, for example, using LLM. Note that the root node does not need to be associated with summary information.
[0062] This allows the creation unit 402 to create tree structure information that represents the contents of the document group in a hierarchical structure. However, the tree structure information may be obtained from another computer. Specifically, for example, the creation unit 402 may obtain the tree structure information created in the client device 202 from the client device 202.
[0063] Examples and specific examples of creating tree structure information will be described later using Figures 7 and 8.
[0064] The acquisition unit 403 acquires summary information that summarizes a portion of the content of the document group for each of the multiple nodes included in the tree structure information. The search unit 404 then refers to the multiple acquired summary pieces of information to search for the information to be searched from the document group, as instructed by the received related information aggregation query.
[0065] Here, searching for information to be searched means, for example, identifying the section containing the information to be searched from a set of documents, extracting the information from that section, or summarizing the extracted information. The search process by the search unit 404 is performed, for example, using LLM420. LLM420 is, for example, a trained machine learning model generated by performing machine learning such as deep learning using a large amount of text data.
[0066] For example, the search unit 404 searches for information to be searched by inputting the received related information aggregation query and task description prompt to the LLM 420. The acquired summary information is referenced, for example, when the LLM 420 searches for information to be searched. The LLM 420 may be stored, for example, in the storage unit 410.
[0067] Furthermore, LLM420 may be stored on a computer other than the related information aggregation device 201. The other computer may be, for example, an external server that provides a service making LLM420 available. In this case, the related information aggregation device 201 can utilize LLM420 by accessing the other computer.
[0068] Specifically, for example, the acquisition unit 403 acquires information (summary information, or chunk description) associated with each of the multiple nodes (child nodes) in the immediate lower hierarchy connected to the specified node (parent node) when any node included in the tree structure information created by the search unit 404 (LLM 420) is specified.
[0069] For example, the acquisition unit 403 acquires summary information associated with each of the multiple second nodes in the immediate lower hierarchy connected to the first node, in response to the specification of a first node included in the tree structure information. As an example, suppose the root node is specified by the LLM 420. In this case, the acquisition unit 403 acquires summary information associated with each of the multiple nodes (second nodes) in the immediate lower hierarchy connected to the root node.
[0070] Then, the search unit 404 refers to the multiple summary pieces of information obtained and identifies the second node related to the search target from among the multiple second nodes. For example, the search unit 404 uses the LLM 420 to refer to the multiple summary pieces of information obtained and identifies the second node related to the search target from among the multiple second nodes.
[0071] In this case, the search unit 404 may identify a second node related to the search target by considering supplementary information contained in the received task description prompt. For example, by considering the characteristics of the document group and information useful for reading the document group, the search unit 404 can make it easier for the LLM 420 to understand the contents of the document group and improve the accuracy of identifying a second node related to the search target.
[0072] Furthermore, by considering information that represents key points in searching for information on the search target, the search unit 404 can more easily understand how the search target will be used in the LLM 420, thereby improving the accuracy of identifying second nodes related to the search target. In addition, by considering information to control the thought process when searching for information on the search target, the search unit 404 can suppress redundant processing in the LLM 420 and improve the efficiency of the search process. For example, the search unit 404 can suppress redundant tool calls and promote planned tool use by presenting the remaining context length.
[0073] The search unit 404 then searches for information on the target of the search from the information under the identified second node. The information under the identified second node refers to information in the lower levels that branch off from the identified second node. To explain in more detail, for example, the search unit 404 determines whether or not multiple third nodes in the immediate lower level connected to the identified second node are leaf nodes.
[0074] If multiple third nodes are leaf nodes, the search unit 404 specifies a second node to the acquisition unit 403. In response to the specification of a second node included in the tree structure information, the acquisition unit 403 retrieves the contents associated with each of the multiple third nodes in the immediately lower hierarchy connected to the second node.
[0075] The search unit 404 then searches for the target information based on the multiple entries obtained. More specifically, for example, the search unit 404 may extract the target information from the multiple entries obtained. The search unit 404 may also summarize the extracted information.
[0076] In this case, the search unit 404 may search for the information to be searched for by taking into account the supplementary information contained in the task description prompt. For example, by taking into account the characteristics of the document group and information useful for reading the document group, the search unit 404 can make it easier for the LLM 420 to understand the contents of the document group and improve the accuracy of searching for the information to be searched.
[0077] Furthermore, by considering information that represents key points for searching for information on the target, the search unit 404 can more easily understand what purpose the target will be used for in the LLM 420, thereby improving the accuracy of the search for information on the target.
[0078] On the other hand, if the search unit 404 finds that none of the multiple third nodes are leaf nodes, it identifies the identified second node as the first node (parent node) and the multiple third nodes as multiple second nodes (child nodes). The search unit 404 then obtains summary information by specifying the first node (identified second node) to the acquisition unit 403 and repeats the process of identifying the second node related to the search target from among the multiple second nodes.
[0079] An example of the information search process will be described later using Figure 9.
[0080] The output unit 405 outputs the search results. Specifically, for example, the output unit 405 may output the search results in association with a specific task (subsequent task) that is executed after the related information aggregation task. The output format of the output unit 405 may include storage in a storage device such as memory 302 or disk 304, or transmission to another computer (for example, client device 202) via the communication interface 305.
[0081] Furthermore, the output unit 405 may also output information that identifies which category of the design document the search result belongs to and which sheet within which design document it corresponds to. This allows the related information aggregation device 201 to make the search results more user-friendly, for example, in a specific task (subsequent task).
[0082] The functional units (reception units 401 to output units 405) of the related information aggregation device 201 are implemented, for example, by an LLM agent. Furthermore, some of the functional units of the related information aggregation device 201, such as the creation unit 402 and the acquisition unit 403, may be implemented by separate programs that operate in conjunction with the LLM agent.
[0083] Furthermore, the functional units of the related information aggregation device 201 may be implemented by multiple computers within the information processing system 200 (for example, the related information aggregation device 201 and client devices 202). In this case, communication between functional units of different computers is performed, for example, by sending and receiving data between functional units via the network 210.
[0084] (Specific examples of related information aggregation queries and task description prompts) Next, specific examples of related information aggregation queries and task description prompts will be explained using Figures 5 and 6.
[0085] Figure 5 is an explanatory diagram illustrating a specific example of a related information aggregation query. In Figure 5, the related information aggregation query 500 is an example of instruction information that specifies the search target. The related information aggregation query 500 shows test cases from an existing test specification.
[0086] Here, the information related to the test cases in the existing test specification corresponds to the information to be searched. This searched information is used in a specific task (a subsequent task) to determine whether or not the test cases in the existing test specification are necessary when creating a new test specification for the system under development.
[0087] Figure 6 is an explanatory diagram showing a specific example of a task description prompt. In Figure 6, the task description prompt 600 is an example of a prompt that requests the execution of a related information aggregation task to search for information specified in the related information aggregation query 500 (see Figure 5). For example, the related information aggregation query 500 is entered in the {query} of [Test Case] in the task description prompt 600. The task description prompt 600 shows a description of the related information aggregation task 601. According to the description 601, the LLM 420 is instructed to search for information related to the test case shown in the related information aggregation query 500.
[0088] Furthermore, the task description prompt 600 includes supplementary information 602-604 to improve the search efficiency and accuracy when searching for the target information. Supplementary information 602 is an example of information useful for reading the document set and information that represents key points when searching for the target information. According to supplementary information 602, LLM420 becomes easier to understand the contents of the design document set.
[0089] Supplementary information 603 is an example of information (output control prompt) that instructs the output format when outputting the execution results (search results) of the related information aggregation task. According to supplementary information 603, LLM420 knows what format the search results should be output in. Supplementary information 604 is an example of information (thought process control prompt) for controlling the thought process when searching for information to be searched. According to supplementary information 604, LLM420 can understand what to pay attention to when using tools (for example, the child node information list acquisition tool and the content acquisition tool described later) and can streamline the search process.
[0090] (Examples and specific examples of creating tree structure information) Next, we will explain examples and specific examples of creating tree structure information using Figures 7 and 8. Here, we will use "Design Document Group 710" as an example of a group of documents. Design Document Group 710 includes, for example, requirements specifications, functional specifications, and test specifications related to the system under development.
[0091] Figure 7 is an explanatory diagram showing an example of creating tree structure information. In Figure 7, the related information aggregation device 201 divides each design document included in the design document group 710 into chunks using the chunk divider 701. The chunk divider 701 is a module that divides documents into predetermined units. Here, the set of chunks divided from the design document group 710 is referred to as the "chunk-divided design document group 711".
[0092] Next, the related information aggregation device 201 obtains a hierarchical tree graph with the entire design document group 710 as the root using the hierarchical structure information assigner 702. The hierarchical structure information assigner 702 is a module that creates a tree graph using hierarchical definition information. Here, the design document group 710 is assumed to be a collection of documents created using spreadsheet software and pre-categorized.
[0093] Furthermore, the hierarchical definition information defines a hierarchical structure as "document group (top level) - category - document - sheet - chunk (bottom level)". A category is used to classify documents according to their content. The category to which each document included in the design document group 710 belongs is predetermined. However, the related information aggregation device 201 may determine which category each document belongs to using LLM or the like. A sheet is a table composed of cells in columns and rows.
[0094] For example, a node representing a category will have nodes representing documents belonging to that category. Similarly, a node representing a document will have nodes representing the sheets contained within that document. And a node representing a sheet will have nodes representing the chunks contained within that sheet.
[0095] In this case, the related information aggregation device 201 obtains a tree graph in which the node representing the design document group 710 is the root node and the nodes representing each chunk of the chunked design document group 711 are the leaf nodes. The related information aggregation device 201 then associates the contents (original text) of each chunk with each leaf node (the lowest-level node) in the tree graph.
[0096] Next, the related information aggregation device 201 uses the hierarchical summary assigner 703 to associate summary sentences (summary information) that summarize the information corresponding to the child nodes, starting from the lowest-level nodes, excluding the leaf nodes. The hierarchical summary assigner 703 is a module that assigns summary sentences to nodes.
[0097] The information corresponding to a child node is the contents of the chunk associated with the child node, or summary information. Summarization is performed by an LLM such as GPT4o. The summary instruction prompt 720 given to the LLM includes, for example, a list of the contents or summary information of the child node, and includes summary instructions and output control prompts.
[0098] This creates a pre-processed set of design documents 730. The pre-processed set of design documents 730 corresponds to tree structure information that represents the contents of the design document set 710 in a hierarchical structure.
[0099] Figure 8 is an explanatory diagram illustrating a specific example of tree structure information. In Figure 8, the pre-processed design document group 730 is an example of tree structure information that represents the contents of the design document group 710 in a hierarchical structure. The pre-processed design document group 730 has five levels, from level 0 to level 4. Level 0 is the highest level, and level 4 is the lowest level.
[0100] The pre-processed design document group 730 includes a root node N0 representing the entire design document group 710, nodes representing categories (for example, nodes N1-1 to N1-3 representing categories A to C), nodes representing design documents (for example, nodes N2-1 to N2-3 representing design documents C1 to C3), nodes representing sheets (for example, nodes N3-1 and N3-2 representing sheets C2-1 and C2-2), and nodes representing chunks (for example, nodes N4-1 and N4-2 representing chunks C2-2-1 and C2-2-2).
[0101] Each node representing a chunk is associated with the contents of that chunk (for example, contents ck1, ck2). For example, node N4-1 is associated with the contents ck1 of chunk C2-2-1.
[0102] Each node representing a sheet is associated with a summary sentence (for example, summary sentences A3-1 and A3-2) that summarizes the contents of the chunks corresponding to its child nodes. For example, summary sentence A3-2 is a summary that summarizes the contents ck1, ck2, etc. of chunks C2-2-1 and C2-2-2, which correspond to child nodes N4-1 and N4-2, etc.
[0103] Each node representing a design document is associated with a summary document (for example, summaries A2-1 to A2-3) that summarizes the summary documents corresponding to its child nodes. For example, summary document A2-2 is summary information that summarizes summary documents A3-1, A3-2, etc., which correspond to child nodes N3-1, N3-2, etc.
[0104] Each node representing a category is associated with a summary sentence (for example, summaries A1-1 to A1-3) that summarizes the summary sentences corresponding to its child nodes. For example, summary sentence A1-2 is summary information that summarizes summaries A2-1 to A2-3, etc., which correspond to child nodes N2-1 to N2-3, etc.
[0105] The root node N0, which represents the entire set of design documents 710, is associated with summary A0, which is a summary of the summary sentences corresponding to the child nodes (for example, summary sentences A1-1 to A1-3). However, it is not necessary to associate summary A0 with the root node N0.
[0106] (Example of information search process) Next, an example of information search process will be explained using Figure 9. Here, we will assume a case where we identify a section of the design document group 710 shown in Figure 7 that contains information (the information to be searched) used when solving a specific task (subsequent task), extract the information from that section, and summarize the extracted information.
[0107] Figure 9 is an explanatory diagram showing an example of the search process for information to be searched. In Figure 9, the action generator 901 corresponds to, for example, the search unit 404 of the related information aggregation device 201 (see Figure 4). The child node list acquisition query 902 and the content acquisition query 903 correspond to, for example, the acquisition unit 403 of the related information aggregation device 201 (see Figure 4).
[0108] The action generator 901 uses the LLM 420 to search for information to be searched based on the related information aggregation query 910 and the task description prompt 920. The related information aggregation query 910 is instruction information that specifies the search target, such as the related information aggregation query 500 shown in Figure 5. The task description prompt 920 is a prompt to request the execution of a related information aggregation task, such as the task description prompt 600 shown in Figure 6.
[0109] The available tool information 930 is information that identifies the tools that can be used by the LLM420 (for example, a tool for obtaining a list of child node information, a tool for obtaining the contents described). The available tool information 930 also includes information that identifies how to use the tools.
[0110] Specifically, for example, the action generator 901 executes a child node list acquisition query 902 in response to a request from the LLM 420. The child node list acquisition query 902 is a query for the child node information list acquisition tool. The child node information list acquisition tool is a tool that returns a list of information (child node information list 940) of the child nodes of a specified node. The child node information includes, for example, the node name and a summary.
[0111] For example, suppose that in the child node list retrieval query 902, node N0 representing "the entire design document group 710" is specified. In this case, the child node information list 940 will be "Category A: """[Summary of A]""", Category B: """[Summary of B]""",... For example, Category A is the node name of child node N1-1 of node N0. [Summary of A] is the summary A1-1 associated with child node N1-1.
[0112] The list of child node information 940 is stored in the tool execution history 970. The tool execution history 970 stores information obtained using, for example, the LLM generation thought process 960 and various tools. The LLM generation thought process 960 is the result of the thought process considered by the LLM 420 when searching for information on the target. The tool execution history 970 is referenced by the LLM 420. For example, the result of what the LLM 420 should do based on the contents of the tool execution history 970 in order to search for information on the target is stored each time as the LLM generation thought process 960.
[0113] The action generator 901, for example, uses the LLM 420 to refer to the tool execution history 970 and uses the child node information list acquisition tool to narrow down the search scope in the order of category, design document, and sheet. In this case, the action generator 901 may also narrow down the search scope by considering supplementary information contained in the task description prompt 920 (see, for example, Figure 6).
[0114] For example, LLM420 refers to the list of child node information 940 included in the tool execution history 970 to identify the node related to the search target from among the child nodes (for example, nodes N1-1 to N1-3) of node N0, which represents the "entire set of design documents 710". LLM420 may also obtain the summary text A0 associated with the root node N0 and refer to it when identifying the node related to the search target.
[0115] Here, let's assume that node N1-3, which represents "Category C", has been identified as a node related to the search target. In this case, the action generator 901 specifies node N1-3 and executes the child node list acquisition query 902 to obtain the child node information list 940. This child node information list 940 will be in the format "Design document C1.xlsx: """[Summary of C1]""", Design document C2.xlsx: """[Summary of C2]""", ... For example, Design document C1.xlsx is the node name of child node N2-1 of node N1-3. [Summary of C1] is the summary A2-1 associated with child node N2-1.
[0116] Next, LLM420 refers to the tool execution history 970 and identifies the child nodes (for example, nodes N2-1 to N2-3) of node N1-3, which represent "Category C", that are related to the search target. Here, let's assume that node N2-2, which represents "Design Document C2.xlsx", is identified as the node related to the search target.
[0117] In this case, the action generator 901 obtains the child node information list 940 by executing the child node list acquisition query 902, specifying node N2-2. This child node information list 940 is in the format "Sheet C2-1: """[Summary of C2-1]""", Sheet C2-2: """[Summary of C2-2]""",...". For example, Sheet C2-1 is the node name of child node N3-1 of node N2-2. [Summary of C2-1] is the summary A3-1 associated with child node N3-1.
[0118] Next, LLM420 refers to the tool execution history 970 and identifies the node related to the search target among the child nodes (for example, nodes N3-1 and N3-2) of node N2-2, which represents "design document C2.xlsx". Here, it is assumed that node N3-2, which represents "sheet C2-2", is identified as the node related to the search target.
[0119] Here, the child nodes of node N3-2, which represents "sheet C2-2" (for example, nodes N4-1 and N4-2), are leaf nodes. In this case, the action generator 901 obtains the list of contents 950 by specifying node N3-2 and executing the content acquisition query 903.
[0120] The content retrieval query 903 is a query for the content retrieval tool. The content retrieval tool is a tool that returns a list of information (content list 950) of the child nodes of a specified node. The child node information includes the contents of the chunks. The content list 950 is stored in the tool execution history 970.
[0121] For example, suppose that in the content retrieval query 903, node N3-2 representing "sheet C2-2" is specified. In this case, the content list 950 will include, for example, the content ck1 of chunk C2-2-1 and the content ck2 of chunk C2-2-2.
[0122] Then, the action generator 901 uses the LLM 420 to search for information on the target based on the list of contents 950 contained in the tool execution history 970. Specifically, for example, the action generator 901 uses the LLM 420 to refer to the tool execution history 970 and extract information on the target from the list of contents 950, or summarize the extracted information.
[0123] The action generator 901 outputs the search results found by the LLM 420 as the LLM-generated response 980. If the action generator 901 determines, based on the LLM 420, that the tool execution history 970 does not contain enough information for the response, it may return to any node in the pre-processed design document group 730 and repeat the series of processes. The action generator 901 autonomously determines, for example, which node to return to using the LLM 420.
[0124] This allows the related information aggregation device 201 to search for information useful for solving a specific task (subsequent task) from the design document group 710 (see Figure 7).
[0125] (Various Processing Procedures of the Related Information Aggregation Device 201) Next, various processing procedures of the related information aggregation device 201 will be explained using Figures 10 to 12. First, the pre-processing procedure of the related information aggregation device 201 will be explained using Figure 10. The pre-processing of the related information aggregation device 201 is the process of creating tree structure information that represents the contents of the design document group in a hierarchical structure.
[0126] Figure 10 is a flowchart showing an example of the preprocessing procedure of the related information aggregation device 201. In the flowchart of Figure 10, first, the related information aggregation device 201 acquires a group of design documents and hierarchical definition information that defines the hierarchical structure of the group of design documents (step S1001).
[0127] Next, the related information aggregation device 201 divides the acquired design document group into chunks (step S1002). Then, the related information aggregation device 201 refers to the acquired hierarchical definition information and creates a tree graph in which the node representing the document group is the root node and the nodes representing each divided chunk are the leaf nodes (step S1003).
[0128] In the following explanation, the hierarchy of the tree graph may be denoted as "hierarchy 0 to M" (M: a natural number). Hierarchy 0 is the highest level. Hierarchy M is the lowest level.
[0129] Next, the related information aggregation device 201 associates the contents (original text) of each chunk with each leaf node in the tree graph (step S1004). Then, the related information aggregation device 201 sets the j of hierarchy j to "j = M-1" (step S1005) and selects the unselected nodes from among the nodes of hierarchy j (step S1006).
[0130] Next, the related information aggregation device 201 generates a summary sentence by aggregating and summarizing the information associated with each child node of the selected node (step S1007). The information associated with the child node is either the contents of the chunk or the summary sentence. Then, the related information aggregation device 201 associates the generated summary sentence with the selected node (step S1008).
[0131] Next, the related information aggregation device 201 determines whether or not there are any unselected nodes among the nodes of hierarchy j (step S1009). If there are unselected nodes (step S1009: Yes), the related information aggregation device 201 returns to step S1006.
[0132] On the other hand, if there are no unselected nodes (step S1009: No), the related information aggregation device 201 sets the value of j in hierarchy j to "j = j - 1" (step S1010) and determines whether j has become less than 0 (step S1011). If j is 0 or greater (step S1011: No), the related information aggregation device 201 returns to step S1006.
[0133] On the other hand, if j becomes less than 0 (step S1011: Yes), the related information aggregation device 201 terminates the series of processes according to this flowchart.
[0134] As a result, the related information aggregation device 201 can create tree structure information (for example, the pre-processed design document group 730 shown in Figure 8) that represents the contents of the design document group in a hierarchical structure.
[0135] Next, the related information aggregation process of the related information aggregation device 201 will be explained using Figures 11 and 12. The related information aggregation process of the related information aggregation device 201 is, for example, a process that searches for related information related to a specific task (subsequent task) as the search target. Related information is used, for example, when solving a specific task (subsequent task).
[0136] Figures 11 and 12 are flowcharts illustrating an example of the related information aggregation process of the related information aggregation device 201. In the flowchart of Figure 11, first, the related information aggregation device 201 receives a related information aggregation query and a task description prompt (step S1101).
[0137] Next, the related information aggregation device 201 sets j in hierarchy j to "j=0" (step S1102), specifies a node in hierarchy j, and executes a child node list acquisition query to obtain a list of child node information (step S1103). Then, the related information aggregation device 201 stores the acquired list of child node information in the tool execution history (step S1104).
[0138] Next, the related information aggregation device 201 sets the value of j in hierarchy j to "j = j + 1" (step S1105) and determines whether or not it has become "j = M - 1" (step S1106).
[0139] Here, if "j ≠ M-1" (step S1106: No), the related information aggregation device 201 refers to the tool execution history and identifies the node in hierarchy j that is related to the search target indicated in the related information aggregation query (step S1107). In this case, the related information aggregation device 201 may identify the node related to the search target by considering, for example, the supplementary information contained in the task description prompt.
[0140] Next, the related information aggregation device 201 obtains a list of child node information by specifying the node of the identified hierarchy j and executing a child node list acquisition query (step S1108). Then, the related information aggregation device 201 stores the acquired child node information list in the tool execution history (step S1109) and returns to step S1105.
[0141] If "j = M-1" occurs in step S1106 (step S1106: Yes), the related information aggregation device 201 proceeds to step S1201 shown in Figure 12.
[0142] In the flowchart of Figure 12, first, the related information aggregation device 201 refers to the tool execution history and identifies the node related to the search target among the nodes of hierarchy j (S1201). Next, the related information aggregation device 201 specifies the identified node of hierarchy j and executes a content acquisition query to obtain a list of content (step S1202).
[0143] Then, the related information aggregation device 201 stores the acquired list of recorded contents in the tool execution history (step S1203). Next, the related information aggregation device 201 determines whether or not sufficient information for the answer is stored in the tool execution history (step S1204).
[0144] If sufficient information is not stored at this point (step S1204: No), the related information aggregation device 201 returns to step S1102 shown in Figure 11. Note that while this explanation uses the case of returning to step S1102 as an example, the related information aggregation device 201 autonomously determines whether to return to that step, for example, using the LLM420.
[0145] On the other hand, if sufficient information is stored (step S1204: Yes), the related information aggregation device 201 extracts related information from the list of contents included in the tool execution history and generates a summary of the extracted related information (step S1205). Then, the related information aggregation device 201 outputs the generated summary of related information (step S1206) and terminates the series of processes according to this flowchart.
[0146] This allows the related information aggregation device 201 to search the design document set for information (summary of related information) that is useful when solving a specific task (subsequent task).
[0147] As described above, the related information aggregation device 201 according to the embodiment receives instruction information that specifies the search target in the document group, acquires summary information that summarizes a part of the content of the document group, corresponding to each of the multiple nodes included in the tree structure information that represents the content of the document group in a hierarchical structure, and searches for the search target information specified by the received instruction information from the document group by referring to the multiple acquired summary information. Specifically, for example, the related information aggregation device 201 uses the LLM420 to search for the search target information by referring to the multiple acquired summary information.
[0148] As a result, the related information aggregation device 201 can refer to summary information that summarizes a portion of the content of the hierarchically structured document group, and perform searches while narrowing the search range to the parts related to the search target, thereby improving the accuracy of information retrieval.
[0149] Furthermore, the related information aggregation device 201 can acquire summary information associated with each of the multiple second nodes (child nodes) in the direct hierarchy connected to the first node (parent node) included in the tree structure information.The related information aggregation device 201 can then refer to the acquired summary information to identify the second node related to the search target from among the multiple second nodes, and search for information on the search target from the information under the identified second node.Each node included in the tree structure information has summary information associated with it, which summarizes the information corresponding to each of the nodes in the direct hierarchy connected to that node.
[0150] As a result, the related information aggregation device 201 can improve the accuracy of information retrieval by narrowing the search range to information under child nodes related to the search target among the child nodes branching off from the parent node.
[0151] Furthermore, according to the related information aggregation device 201, if multiple third nodes in the immediate lower hierarchy connected to the identified second node are leaf nodes, the contents associated with each of the multiple third nodes can be acquired, and the information to be searched can be searched based on the acquired contents. The tree structure information includes a root node representing a group of documents and leaf nodes representing chunks separated from each document included in the group. The contents of the chunks are associated with the leaf nodes, and each node between the root node and the leaf nodes is associated with summary information that summarizes the information corresponding to each of the multiple nodes in the immediate lower hierarchy connected to that node.
[0152] As a result, the related information aggregation device 201 can understand the hierarchical structure of the document group and narrow down the content (original text) related to the search target, thereby improving the accuracy of information retrieval.
[0153] Furthermore, according to the related information aggregation device 201, if the multiple third nodes are not leaf nodes, the identified second node can be designated as the first node (parent node), and the multiple third nodes can be designated as multiple second nodes (child nodes). The related information aggregation device 201 can then acquire summary information associated with each of the multiple second nodes and repeat the process of identifying the second node related to the search target from among the multiple second nodes.
[0154] As a result, the related information aggregation device 201 can understand the hierarchical structure of the document group and perform searches while gradually narrowing the search range, thereby improving the accuracy of information retrieval.
[0155] Furthermore, the related information aggregation device 201 can create tree structure information based on a group of documents and hierarchical definition information that defines the hierarchical structure of the group of documents.
[0156] This allows the related information aggregation device 201 to create tree structure information that represents the contents of the document group in a hierarchical structure.
[0157] Furthermore, the related information aggregation device 201 can output the search results.
[0158] As a result, the related information aggregation device 201 can extract and output information useful to the LLM420 from lengthy documents such as design documents when solving specific tasks such as test item generation and revision determination.
[0159] Based on these considerations, the related information aggregation device 201 can understand the hierarchical structure of the document group and search for the target information while gradually narrowing the search range. This reduces the number of cases in which the related information aggregation device 201 includes superficially similar content from documents with little relation to the search target in the search results, making it easier to find the correct content. Furthermore, because the related information aggregation device 201 can perform searches based on the hierarchical structure of the document group using the LLM 420, it can efficiently search for information in specific locations within other documents, for example, when it is necessary to refer to a specific location within another document in order to understand a certain description within a document.
[0160] For example, the related information aggregation device 201 can apply the document set to a "system design document set" to search for information useful to the LLM420 when solving specific tasks such as test item generation and revision determination. Furthermore, the related information aggregation device 201 can apply the document set to a "constitution and legal text collection" to search for information useful to the LLM420 when solving specific tasks such as case law analysis.
[0161] The search method described in this embodiment can be implemented by executing a pre-prepared program on a computer such as a personal computer or workstation. This search program is recorded on a computer-readable recording medium such as a hard disk, flexible disk, CD-ROM, DVD, or USB memory, and is executed when read from the recording medium by the computer. This search program may also be distributed via a network such as the Internet.
[0162] Furthermore, the information processing device 101 (related information aggregation device 201) described in this embodiment can also be realized using application-specific ICs such as standard cells and structured ASICs (Application Specific Integrated Circuits), or PLDs (Programmable Logic Devices) such as FPGAs.
[0163] 101 Information Processing Device 110 Document Group 120 Instruction Information 130 Tree Structure Information 131 First Node 132, 133, 134 Second Node 141, 142, 143 Summary Information 150 Subordinate Information 200 Information Processing System 201 Related Information Aggregation Device 202 Client Device 210 Network 300 Bus 301 CPU 302 Memory 303 Disk Drive 304 Disk 305 Communication I / F 306 Portable Recording Medium I / F 307 Portable Recording Medium 400 Control Unit 401 Reception Unit 402 Creation Unit 403 Acquisition Unit 404 Search Unit 405 Output Unit 410 Storage Unit 420 LLM 500, 910 Related Information Aggregation Query 600, 920 Task Description Prompt 601 Description 602, 603, 604 Supplementary Information 701 Chunk Splitter 702 Hierarchical Structure Information Assigner 703 Hierarchical Summary Assigner 710 Design Document Group 711 Chunked Design Document Group 720 Summary Instruction Prompt 730 Preprocessed Design Document Group 901 Action Generator 902 Child Node List Acquisition Query 903 Content Acquisition Query 930 Available Tool Information 940 Child Node Information List 950 Content List 960 LLM Generation Thought 970 Tool Execution History 980 LLM Generation Response Nodes N0, N1-1 to N1-3, N2-1 to N2-3, N3-1, N3-2, N4-1, N4-2
Claims
1. A search program characterized by causing a computer to execute a process that involves receiving instruction information indicating a search target in one or more documents, obtaining summary information which summarizes a part of the content corresponding to each of several nodes included in tree structure information which represents the contents of the one or more documents in a hierarchical structure, and searching for the information of the search target indicated by the received instruction information in the one or more documents by referring to the multiple summaries obtained.
2. The search program according to claim 1, wherein each node included in the tree structure information is associated with summary information that summarizes the information corresponding to each node in the immediately lower hierarchy connected to that node, the acquisition process acquires the summary information associated with each of a plurality of second nodes in the immediately lower hierarchy connected to the first node included in the tree structure information, and the search process refers to the acquired plurality of summary information to identify the second node among the plurality of second nodes that is related to the search target, and searches for the information of the search target from the information under the identified second node.
3. The tree structure information includes a root node representing one or more documents and leaf nodes representing chunks divided from each of the one or more documents, the leaf nodes are associated with the contents of the chunks, each node between the root node and the leaf nodes is associated with summary information that summarizes the information corresponding to each of the multiple nodes in the immediately lower hierarchy connected to that node, and the search process is characterized in that, if the multiple third nodes in the immediately lower hierarchy connected to the identified second node are leaf nodes, the contents associated to each of the multiple third nodes are obtained, and the search program searches for the information to be searched based on the multiple contents obtained.
4. The search program according to claim 3, characterized in that, if the plurality of third nodes are not leaf nodes, the computer is repeatedly made to perform the process of acquiring the summary information and the search process, with the identified second node being the first node and the plurality of third nodes being the plurality of second nodes.
5. The search program according to claim 1, characterized in that it causes the computer to perform a process of creating the tree structure information based on the one or more documents and hierarchical definition information that defines the hierarchical structure of the one or more documents.
6. The search program according to claim 1, characterized in that it causes the computer to execute a process that outputs the search results.
7. The search program according to any one of claims 1 to 6, characterized in that the search process uses LLM (Large Language Models) to search for information on the target of the search.
8. A search method characterized in that a computer receives instruction information indicating a search target in one or more documents, obtains summary information which summarizes a part of the content corresponding to each of a plurality of nodes included in tree structure information which represents the contents of the one or more documents in a hierarchical structure, and searches for the information of the search target indicated by the received instruction information from the one or more documents by referring to the plurality of obtained summary information.
9. An information processing device having a control unit that receives instruction information indicating a search target in one or more documents, obtains summary information which summarizes a part of the content corresponding to each of a plurality of nodes included in tree structure information which represents the contents of the one or more documents in a hierarchical structure, and searches for the information of the search target indicated by the received instruction information from the one or more documents by referring to the plurality of obtained summary information.