Information processing device, method, and program

The information processing device uses a large-scale language model to extract and classify document data, enhancing the interpretability and comparability of analysis results, addressing the limitations of existing tools in providing diverse and user-friendly outputs.

JP7836130B2Active Publication Date: 2026-03-26NEEDS EXPLORER CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing document analysis tools often struggle to provide diverse and easily interpretable analytical results that meet the varied needs of users.

Method used

An information processing device that utilizes a large-scale language model to extract structured data from documents, classifies them into groups based on relevance, and outputs associated information for multiple items, facilitating easy interpretation of analysis results.

Benefits of technology

Enables the generation of varied and easily interpretable analytical results from documents, allowing for efficient comparison and identification of similarities and differences in solutions, problems, and effects across documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007836130000001
    Figure 0007836130000001
  • Figure 0007836130000002
    Figure 0007836130000002
  • Figure 0007836130000003
    Figure 0007836130000003
Patent Text Reader

Abstract

Obtain a variety of easy-to-interpret analysis results for your documents. [Solution] An extraction unit 32 extracts information related to each of a predetermined number of items from each of a plurality of documents in accordance with a predetermined data structure that is unified into a format that allows each of the plurality of documents to be compared, a classification unit 34 classifies the plurality of documents into groups based on the information extracted for one or more items selected from the plurality of items, and an output unit 36 ​​outputs, for each group, the information extracted for the selected item and the information extracted for at least one item other than the selected item, in association with each other.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] When considering business directions, etc., various documents such as patent documents, technical papers, and securities reports may be analyzed.

[0003] As a technique related to document analysis, for example, population information regarding a population including a plurality of patent documents is received, a first patent document is extracted from the plurality of patent documents included in the population, and a first output result output in response to inputting a prompt including the extracted first patent document to a large language model is obtained. A second patent document is extracted from the plurality of patent documents included in the received population, and a second output result output in response to inputting the second patent document to a learning model that has learned information based on the first output result as teacher data is obtained. An information processing apparatus has been proposed (Patent Document 1).

[0004] Also, for example, in a technical document search apparatus that searches for technical documents such as patent documents and technical papers, a search means, a technical map storage means that stores a technical map including information on technical documents and keywords associated with technical elements, and a document similarity determination means that determines the similarity between technical documents are provided. The document similarity determination means determines the similarity between the technical document searched by the search means and the technical documents included in the technical map, and classifies the searched technical document into the technical elements of the technical map based on the similarity determined by the document similarity determination means. This has been proposed (Patent Document 2).

[0005] Furthermore, for example, a trained model storage unit stores a trained model, which is a classification model trained using a predetermined learning algorithm, using training data that includes a set of A term pairs, which are pairs of terms included in the A patent document group, and labels representing status concept relationships, including the relationships between higher and lower concepts, between the A term pairs. When feature data including a set of B term pairs, which are pairs of terms included in the B patent document group, is input to the trained model, the output of the trained model is used to store a group of status concept data representing the relationships between the higher and lower concepts between the B term pairs of the feature data. A data processing device has been proposed that comprises: a positional concept data group creation unit; an output unit that performs output processing to output the positional concept data group to an output device; a patent document group acquisition unit that obtains the B patent document group containing a predetermined keyword by searching the database of patent documents based on a predetermined keyword; a term group acquisition unit that obtains a term group consisting of the terms in the B patent document group up to a predetermined rank in order of frequency of occurrence by data mining the B patent document group; and a feature data creation unit that creates the feature data from the term group (Patent Document 3).

[0006] Furthermore, for example, an abstract creation support system has been proposed that includes an extraction rule dictionary that defines rules for extracting essential parts for each certain classification; a document structure analysis unit that analyzes the chapter and section structure of the digitized original patent specification; an essential part selection unit that selects an extraction rule dictionary according to the classification of the analyzed document structure analysis results and selects essential parts from the document structure analysis results based on the rules of this dictionary; a machine translation unit that translates the document structure analysis results into the native language; an abstract candidate extraction unit that extracts native language abstract candidates from the translated native language translation results for each essential part selected by the essential part selection unit; and an abstract editing and display control unit that displays the extracted native language abstract candidates and the original text essential parts corresponding to these native language abstract candidates, or displays the native language abstract candidates in association with the original text essential parts, and creates a native language abstract including processing for correcting the native language abstract candidates (Patent Document 4).

[0007] For example, an input device has been proposed that, in accordance with the input information received by the input device, obtains a published patent gazette for analysis from a publication database device, performs linguistic analysis on the documents contained in the published patent gazette for analysis, extracts terms related to the problem (provisional problem term and provisional problem category) and terms related to the method of solving the problem (provisional solution term and provisional solution category) for each published patent gazette, obtains the patent classification description for each publication from a classification database device, verifies and examines the validity of the terms related to the problem and the terms related to the method of solving the problem using the patent classification description for each published patent gazette, and determines the terms that represent the gist of the published patent gazette for analysis (problem term, problem category, solution term and solution category) (Patent Document 5). [Prior art documents] [Patent Documents]

[0008] [Patent Document 1] Patent No. 7493195 [Patent Document 2] Japanese Patent Publication No. 2002-163275 [Patent Document 3] Patent No. 7431379 [Patent Document 4] Japanese Patent Publication No. 2005-31813 [Patent Document 5] Japanese Patent Publication No. 2007-148630 [Overview of the Initiative] [Problems that the invention aims to solve]

[0009] Existing document analysis tools can sometimes make it difficult to interpret the analysis results. Furthermore, they may not provide diverse analysis results that meet the various needs of users.

[0010] This disclosure is made in view of the above points and aims to provide an information processing device, method, and program that can obtain a variety of easily interpretable analytical results from a document. [Means for solving the problem]

[0011] An information processing device according to a first aspect of the disclosed technology includes: an extraction unit that inputs a prompt to a large-scale language model instructing it to extract information related to each of a predetermined set of items from each of a set of documents, according to a predetermined data structure unified into a comparable format, and acquires information including structured data which is the output of the large-scale language model; a classification unit that classifies the set of documents into groups based on the degree of relevance of the information including the structured data acquired by the extraction unit for one or more items selected from the set of items; and an output unit that outputs to each of the groups the information acquired by the extraction unit for the selected items and the information acquired by the extraction unit for at least one item other than the selected items, in association with each of the groups.

[0012] Furthermore, the information processing method relating to the second aspect of the disclosed technology involves an extraction unit inputting a prompt to a large-scale language model instructing it to extract information related to each of a predetermined set of items from each of a set of documents, according to a predetermined data structure unified into a comparable format for each of the set of documents, thereby obtaining information including structured data which is the output of the large-scale language model, a classification unit classifying the set of documents into groups based on the degree of relevance of the information including the structured data obtained by the extraction unit for one or more items selected from the set of items, and an output unit outputting each of the groups, associating the information obtained by the extraction unit for the selected items with the information obtained by the extraction unit for at least one item other than the selected items.

[0013] Furthermore, an information processing program relating to a third aspect of the disclosed technology is a program that causes a computer to input prompts to a large-scale language model instructing it to extract information related to each of a predetermined set of items from each of a set of documents, according to a predetermined data structure unified into a comparable format for each of the set of documents, and to acquire information including structured data which is the output of the large-scale language model; a classification unit that classifies the set of documents into groups based on the degree of relevance of the information including the structured data acquired by the extraction unit for items selected from the set of items; and an output unit that outputs to each of the groups the information acquired by the extraction unit for the selected items and the information acquired by the extraction unit for at least one item other than the selected items, in association with each of the groups. [Effects of the Invention]

[0014] According to the information processing device, method, and program relating to this disclosure, a variety of easily interpretable analytical results can be obtained from a document. [Brief explanation of the drawing]

[0015] [Figure 1] This is a block diagram showing the hardware configuration of the information processing device according to the first to third embodiments. [Figure 2] These are functional block diagrams of the information processing apparatus according to the first and third embodiments. [Figure 3] This diagram schematically shows an example of a data structure. [Figure 4] This figure shows a specific example of a prompt. [Figure 5] This figure shows a concrete example of structured data. [Figure 6] This figure shows an example of the analysis results. [Figure 7] This figure shows an example of the analysis results. [Figure 8] This figure shows an example of the analysis results. [Figure 9] This figure shows an example of the analysis results. [Figure 10] It is a diagram showing a specific example of a prompt. [Figure 11] It is a diagram showing a specific example of an output result. [Figure 12] It is a diagram showing an example of an analysis result. [Figure 13] It is a diagram showing an example of an analysis result. [Figure 14] It is a diagram showing an example of an analysis result. [Figure 15] It is a flowchart showing an example of an information processing routine. [Figure 16] It is a diagram showing an example of a knowledge graph. [Figure 17] It is a diagram for explaining a modification example. [Figure 18] It is a diagram for explaining the use of independent component analysis. [Figure 19] It is a diagram for explaining the use of independent component analysis. [Figure 20] It is a diagram showing an example of a proposal. [Figure 21] It is a diagram for explaining the use of independent component analysis. [Figure 22] It is a diagram for explaining the use of independent component analysis. [Figure 23] It is a diagram for explaining the use of independent component analysis. [Figure 24] It is a diagram for explaining the use of independent component analysis. [Figure 25] It is a diagram for explaining the use of independent component analysis. [Figure 26] It is a diagram for explaining the use of independent component analysis. [Figure 27] It is a functional block diagram of an information processing apparatus according to the second embodiment. [Figure 28] It is a diagram showing an example of a prompt. [Figure 29] It is a diagram showing an example of an output of a problem. [Figure 30] It is a diagram showing an example of a proposal. [Figure 31] It is a flowchart showing an example of a presentation process. [Figure 32] This diagram illustrates the effectiveness of presenting differences in the solutions or suggestions adopted for each user attribute. [Figure 33] This diagram illustrates the effectiveness of presenting differences in the solutions or suggestions adopted for each user attribute. [Figure 34] This figure shows an example of the analysis results. [Figure 35] This figure shows an example of the analysis results. [Figure 36] This figure shows an example of the analysis results. [Figure 37] This figure shows a specific example of prompts to enter into Deep Research, etc. [Figure 38] This figure shows an example of output from Deep Research, etc. [Figure 39] This figure shows an example of output from Deep Research, etc. [Modes for carrying out the invention]

[0016] Hereinafter, an example of an embodiment of this disclosure will be described with reference to the drawings. In each drawing, identical or equivalent components and parts are given the same reference numerals. Also, the dimensions and proportions in the drawings are exaggerated for illustrative purposes and may differ from actual proportions.

[0017] <First Embodiment> Figure 1 is a block diagram showing the hardware configuration of the information processing device 10 according to the first embodiment. As shown in Figure 1, the information processing device 10 includes a CPU (Central Processing Unit) 12, a memory 14, a storage device 16, an input device 18, an output device 20, a storage medium reader 22, and a communication I / F (Interface) 24. Each component is connected to the others via a bus 26 so as to be able to communicate with each other.

[0018] The storage device 16 stores information processing programs for executing information processing routines described later. The CPU 12 is a central processing unit that executes various programs and controls each component. Specifically, the CPU 12 reads a program from the storage device 16 and executes the program using memory 14 as a workspace. The CPU 12 controls each component and performs various calculations according to the program stored in the storage device 16.

[0019] Memory 14 is composed of RAM (Random Access Memory) and temporarily stores programs and data as a working area. Storage device 16 is composed of ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), etc., and stores various programs and data, including the operating system.

[0020] The input device 18 is a device for performing various types of input, such as a keyboard or mouse. The output device 20 is a device for outputting various types of information, such as a display or printer. By using a touch panel display as the output device 20, it may also function as the input device 18.

[0021] The storage medium reader 22 reads data stored on various storage media such as CD (Compact Disc)-ROM, DVD (Digital Versatile Disc)-ROM, Blu-ray disc, and USB (Universal Serial Bus) memory, and writes data to the storage media. The communication I / F 24 is an interface for communication with other devices, and standards such as Ethernet (registered trademark), FDDI, or Wi-Fi (registered trademark) are used.

[0022] Next, the functional configuration of the information processing device 10 according to the first embodiment will be described.

[0023] Figure 2 is a block diagram showing an example of the functional configuration of the information processing device 10. As shown in Figure 2, the information processing device 10 includes, as a functional configuration, an extraction unit 32, a classification unit 34, an output unit 36, and a Large Language Model (LLM) 40. In addition, a predetermined storage area of ​​the information processing device 10 may be stored using the LLM 40. The LLM 40 may be located externally and accessible from the information processing device 10. Each functional configuration is realized by the CPU 12 reading the information processing program stored in the storage device 16, expanding it into memory 14, and executing it.

[0024] The extraction unit 32 extracts information related to each of a predetermined set of items from each of the multiple documents, according to a predetermined data structure that unifies the multiple documents into a comparable format. Hereinafter, the predetermined data structure information extracted by the extraction unit 32 will also be referred to as "structured data".

[0025] Specifically, the extraction unit 32 retrieves multiple documents from the document database 50. The documents may be, for example, patent documents, technical papers, securities reports, web articles, product catalogs, etc. The extraction unit 32 inputs each retrieved document and an instruction as a prompt to the LLM 40. The instruction may be, for example, "Extract information regarding item A, item B, ... from the following documents." The instruction may also specify a display format appropriate for each item, or specify an upper limit on the number of characters in the information to be extracted. The extraction unit 32 may also extract the information for each item in JSON format or Markdown format. If the document is a patent document, the extraction unit 32 may extract information related to each of the multiple items on a claim basis. The extraction unit 32 may also include in the instruction that if there is no description corresponding to an item in the document, it should output "No description".

[0026] The multiple items may include at least one of the following: the problem, the solution to the problem, the effects of the solution, and the application examples of the solution. In this case, if the document contains an item corresponding to the solution, the extraction unit 32 inputs a prompt to the LLM 40 instructing it to extract information related to the problem, effects, and application examples based on a summary of the content described in the item corresponding to the solution in the document, and obtains the information related to the problem, effects, and application examples which are the output of the LLM 40. If the document does not contain an item corresponding to the solution, but contains an item corresponding to the problem, the extraction unit 32 inputs a prompt to the LLM 40 instructing it to extract information related to the effects and application examples based on a summary of the content described in the item corresponding to the problem in the document, and obtains the information related to the effects and application examples which are the output of the LLM 40.

[0027] More specifically, when the document is a patent document, the extraction unit 32, as step 1, inputs a prompt to the LLM 40 instructing it to extract at least one of the claim features and the claim abstract from the patent document, and obtains at least one of the claim features and the claim abstract, which are the output of the LLM 40, as information related to the solution. As step 2, the extraction unit 32 inputs a prompt to the LLM 40 instructing it to generate information relating to at least one of the problem, solution, and effect based on at least one of the obtained claim features and the claim abstract, and obtains the output of the LLM 40. During extraction, a prompt may be created and input to output the problem, effect, etc., separated into technical and social matters, and the output of the LLM 40 may be obtained. Then, the extraction unit 32 inputs a prompt to the LLM 40 instructing it to generate a summary that includes at least one of the obtained claim features and the claim abstract, and the generated information, and obtains the summary, which is the output of the LLM 40, thereby extracting information related to the solution and at least one of the problem and effect. Steps 1 and 2 may also be prompted to perform the next step after each step is completed.

[0028] Furthermore, as step 1, the extraction unit 32 inputs a prompt to the LLM 40 instructing it to extract at least one of the claim features and the claim abstract from the patent document, and obtains at least one of the claim features and the claim abstract, which are the output of the LLM 40, as information related to the solution. As step 2, the extraction unit 32 inputs a prompt to the LLM 40 instructing it to extract information related to application examples described in the patent document based on at least one of the obtained claim features and the claim abstract, and obtains information related to application examples, which are the output of the LLM 40. Alternatively, the extraction unit 32 may input a prompt to the LLM 40 instructing it to generate application examples not described in the patent document based on at least one of the obtained claim features and the claim abstract, and obtains application examples, which are the output of the LLM 40. Alternatively, the extraction unit 32 may input a prompt to the LLM 40 instructing it to extract information related to application examples described in the patent document that are not based on the claims described in the patent document, and obtains information related to application examples, which are the output of the LLM 40.

[0029] The extraction unit 32 may achieve the extraction of the above-described information by stepwise inputting prompts to the LLM 40. For example, the extraction unit 32 inputs multiple patent documents and an instruction statement to the LLM 40 as prompts, instructing it to extract features of the claims described in the target patent document that are not present in the claims of other patent documents, and extracts the features of the claims described in the target patent document. Next, the extraction unit 32 inputs the target patent document and an instruction statement to the LLM 40 as prompts, instructing it to generate technical problems corresponding to the extracted claim features, and generates technical problems related to the target patent document. Next, the extraction unit 32 inputs the target patent document, the extracted technical problems, and an instruction statement to the LLM 40 as prompts, instructing it to generate social problems corresponding to the extracted technical problems, and generates social problems related to the target patent document. Finally, the extraction unit 32 inputs the generated technical problems and social problems, and an instruction statement to the LLM 40 as prompts, instructing it to create summaries thereof, and generates summaries of the problems.

[0030] Furthermore, for example, the extraction unit 32 extracts the features of each claim from the patent document in the same manner as described above. The extraction unit 32 then inputs the extracted features of each claim and an instruction statement that instructs the LLM 40 to list as many or as necessary number of elemental technologies of each claim as prompts, and extracts the elemental technologies of each claim. Next, the extraction unit 32 inputs the target patent document and an instruction statement that instructs the LLM 40 to extract the technical effects as prompts, and extracts the technical effects. The extraction unit 32 then inputs the target patent document, the extracted elemental technologies, and an instruction statement that instructs the LLM 40 to extract the social effects as prompts, and extracts the social effects. Finally, the extraction unit 32 inputs the extracted technical effects and social effects and an instruction statement that instructs the creation of summaries thereof as prompts, and generates summaries of the effects.

[0031] Furthermore, for example, the extraction unit 32 inputs a patent document and an instruction statement instructing the extraction of application examples described in that patent document as prompts to the LLM 40, and extracts the application examples. As a measure against hallucination in large-scale language models, an instruction statement instructing the extraction of the original text of a patent document containing an application example described in the extracted patent document may also be input to the LLM 40 as a prompt, and the original text of the patent document containing the application example may be extracted. In addition, the extraction unit 32 inputs a patent document, the extracted application examples, and an instruction statement instructing the estimation of application examples not described in the patent document as prompts to the LLM 40, and generates other application examples. Furthermore, the extraction unit 32 inputs the extracted and generated application examples and an instruction statement instructing the assignment of categories to these application examples as prompts to the LLM 40, and assigns categories to the application examples.

[0032] Furthermore, information extraction is not limited to stepwise prompt input to LLM40 as described above, but may also be performed with a single prompt input. In addition, information such as problems, effects, and application examples is not limited to extraction for each claim, but may also be extracted for each patent document by stepwise or single prompt input to LLM40.

[0033] Figure 3 schematically shows an example of structured data extracted by the extraction unit 32 when the document is a patent document. In Figure 3, the categories of claims, problems, solutions, application examples, elemental technologies, problem summaries, solution summaries, and application examples correspond to items, respectively. In the example in Figure 3, information on these items is extracted in a predetermined data format, either as a patent unit, as a unit of claims output by LLM40, or as individual claims, thereby forming a unified data structure. The "analysis values" portion in Figure 3 will be described in detail later.

[0034] Figure 4 shows a more specific example of prompts that the extraction unit 32 inputs to the LLM 40 when the document is a patent document. Figure 5 shows an example of information including structured data obtained by the prompts in Figure 4. In Figures 4 and 5, "technical elements," "novel perspectives that are not found elsewhere," and "field of invention" are examples of at least one of the features of the claims and the abstract of the claims. In this way, by specifying the invention and then extracting information on items such as problems and effects, it is possible to extract information for each item with high accuracy even from a large amount of information.

[0035] Furthermore, the extraction unit 32 may extract additional information, including date and time information related to the document, in addition to the above items. For example, if the document is a patent document, the extraction unit 32 may extract additional information such as the application date, application number, publication number, applicant, whether the patent right is maintained or abandoned, the number of times it has been cited as a reference in a rejection reason, and the number of times it has been cited as a reference in other patents, etc. Specifically, the extraction unit 32 may extract additional information from a document in a predetermined format, using headings and keywords as clues and following predetermined rules, or it may use LLM 40 for extraction. The extraction unit 32 may also aggregate the extracted information (for example, information on references) to calculate additional information such as the number of times it has been cited.

[0036] The classification unit 34 classifies multiple documents into groups based on information including structured data obtained by the extraction unit 32 for one or more items selected from a plurality of items. The information including structured data may be the structured data itself, or it may be information with the above additional information added to the structured data. Specifically, the classification unit 34 vectorizes the information including structured data obtained by the extraction unit 32, either data for each of the multiple items or data by combining items. General methods such as using an embedding model may be applied for vectorization. The classification unit 34 classifies multiple documents into groups based on the similarity between the vectors of information related to the selected items or the characteristics of the vectors.

[0037] More specifically, the classification unit 34 calculates the cosine similarity between vectorized information for each of several items in a unit of information extracted from multiple documents or from each of multiple documents (hereinafter referred to as the "analysis unit"), using this as a numerical value that serves as the criterion for classifying the information (hereinafter referred to as the "classification value"). Figure 3 shows an example in which the classification value is calculated for each of the items of claims, problems, solutions, and application examples, on a patent unit or claim unit basis. Note that the patent unit or claim unit is just one example of an analysis unit. The classification unit 34 may also perform principal component analysis of the vectorized information and calculate the principal component score (PCA value) for each item as the classification value for each analysis unit. The classification unit 34 may also perform independent component analysis (hereinafter also referred to as "ICA") of the vectorized information. When ICA is performed and the data is classified into n categories, the skewness of n independent components is calculated for each analysis unit, such as a sentence or word. The number of the independent component to which the sentence skewness (hereinafter referred to as "sentence skewness") belongs is designated as the group number (hereinafter referred to as "group number"), and the group number and sentence skewness of each item may be calculated as classification values. Alternatively, the classification unit 34 may calculate classification values ​​based on information that has not been vectorized.

[0038] The classification unit 34 uses the classification values ​​calculated for the selected items to group the information of the analysis units into groups. For example, the classification unit 34 groups information of analysis units whose cosine similarity is greater than or equal to a predetermined value, or it groups the information of analysis units based on the PCA value or group number calculated by ICA, or the skewness of the text, etc.

[0039] Furthermore, the classification unit 34 may classify information of analysis units into groups not only based on classification values, but also based on keywords included in the information. Specifically, the classification unit 34 may classify a group of documents extracted by performing a search on information including structured data, specifying items and keywords, as a single group. For example, if you want to find a patent from multiple patent documents that contains "heat" in the "effect" section, and the search target is not structured data, then all documents containing "heat" will be extracted from the entire document, and you would have to visually search through the extracted documents to find the one that contains "heat" in the "effect" section. On the other hand, in this embodiment, since the search target is structured data, by using only the "effect" field, which is one of the multiple items included in the structured data, and performing a search for "heat," it is possible to easily extract patents that contain "heat" in the "effect" section.

[0040] The similarity between the vectors or the characteristics of the vectors mentioned above, as well as the search results specifying items and keywords, are examples of the degree of relevance of information, including structured data, as disclosed herein.

[0041] The output unit 36 ​​outputs the analysis results by associating the information extracted by the extraction unit 32 for the selected item with the information extracted by the extraction unit 32 for at least one item other than the selected item, for each group of information of the analysis unit. Specifically, the output unit 36 ​​outputs the user's desired analysis results by inputting the results of the classification unit 34 grouping the structured data extracted by the extraction unit 32, and an instruction statement indicating what kind of analysis results to output, as prompts to the LLM 40. The classification unit 34 may also create and output the analysis results according to predetermined creation rules.

[0042] Figure 6 shows an example of analysis results for the structured data shown in Figure 3, where information for each patent document is grouped based on the classification value of the item "Summary of Problem". In the example in Figure 6, "Group" is the group identification number, and "Summary of Problem" is the selected item, which is the information extracted by the extraction unit 32 for that item. "Category of Application Example" and "Summary of Solution" are examples of items other than the selected item, which are the information extracted by the extraction unit 32 for these items. "Publication Number" is an example of additional information, which is the publication number assigned to the published patent gazette that is the relevant patent document. In the example in Figure 6, for "Summary of Solution", the summaries of solutions extracted from each patent document are used as the headings for each column, and the column corresponding to the summary of solution for each patent document is set to "1". For example, by specifying the applicant and technical field (in the example in Figure 6, conference systems) for a patent document and outputting analysis results as shown in Figure 6, it is possible to easily understand what kind of problems are addressed and what kind of solutions and application examples are proposed.

[0043] Furthermore, the output unit 36 ​​may output analysis results in which date and time information is added to the information of the analysis units classified by group. An example of this is shown in Figure 7. In the example in Figure 7, the number of patent documents filed each year is added to the analysis results. For example, as shown in Figure 7, by outputting analysis results that show the number of patent documents filed each year for the information of the patent documents grouped by the problem summary, it is possible to grasp the trend of applications for a problem, support the consideration of problems for which solutions need to be explored in depth, and grasp the chronological flow and continuity of development of a particular company.

[0044] Figure 8 shows an example of analysis results when the unit of analysis is the claim unit. In the example in Figure 8, the items include "Claim Features," "Technical Problems," "Social Problems," "Technical Effects," "Social Effects," "Application Example 1 (described in the patent)," and "Application Example 2 (not described in the patent)." Additional information includes "Publication Number" and "Filing Year." "No." is a sequential number assigned to all claims in all target patent documents. In the analysis results shown in Figure 8, by searching for predetermined keywords in a specific item, groups of claims containing those keywords can be easily identified. In the analysis results shown in Figure 8, for example, by searching the item "Application Example 2" with keywords such as "car" and "driving," it is possible to search for potential technologies that could be used in cars or driving without using knowledge graphs, vector similarities, etc., which will be described later.

[0045] Furthermore, the output unit 36 ​​may, for each group, associate the details of the solutions (elemental technologies) extracted on a claim-by-claim basis, arrange the claim-by-claim information in chronological order (e.g., filing year), and output the analysis results with highlighted information indicating the details of solutions that do not appear in the solutions of earlier patent documents. Figure 9 shows an example of the analysis results in this case. In the example in Figure 9, the underlined portion in the "Elemental Technology" column represents the highlighted portion. That is, among the elemental technologies extracted from the claims of the patent document filed in 2020, the elemental technologies indicated by the underlined portion are not disclosed in the patent document filed in 2018, and among the elemental technologies extracted from the claims of the patent document filed in 2022, the elemental technologies indicated by the underlined portion are not disclosed in the patent documents filed in 2018 and 2020. By outputting such analysis results, the user can grasp the development trends of the relevant technology.

[0046] The classification unit 34 and the output unit 36 ​​may also use the LLM 40 to perform group classification and analysis together. In this case, the LLM 40 is input with prompts containing information including structured data obtained by the extraction unit 32 and instruction sentences that instruct classification and analysis, and the output results of the LLM 40 are obtained. Figure 10 shows an example of the instruction sentences in the prompts in this case, and Figure 11 shows an example of the output results.

[0047] Furthermore, the output unit 36 ​​outputs a scatter plot in which, based on the results of principal component analysis or independent component analysis for the selected items, points corresponding to each of the multiple analysis units of information are arranged in the principal component space or independent component space in different display modes for each group to which the information corresponding to the points belongs.

[0048] Figure 12 shows an example of a scatter plot in the principal component space with principal component A and principal component B as axes, when claim-level information is classified into groups by PCA value. In the example in Figure 12, each point (circle) corresponds to the information of each claim-level, and the difference in the pattern within the circle represents the difference in groups. The numerical value next to each point is the group number, and the string next to each point is the summary of the problem. Note that for the sake of plotting, some group numbers and problem summaries have been omitted. The output unit 36 ​​may enlarge the range specified in the scatter plot. Figure 13 shows an enlarged view when the frame indicated by the dashed line in the scatter plot of Figure 12 is specified. By outputting such a scatter plot as an analysis result, the grouping status of the information of the analysis units and the relationships between groups can be easily grasped.

[0049] Furthermore, the output unit 36 ​​may output a scatter plot in which points corresponding to each of the documents included in the first and second document groups are arranged in different display modes depending on whether the document corresponding to a point belongs to the first or second document group, in a principal component space that is based on two or more principal components among the nth principal components obtained from the principal component analysis results obtained from the first document group having the first attribute, where there is a predetermined difference between the first attribute and the second attribute, and the points corresponding to each of the documents included in the first and second document groups are arranged in different display modes depending on whether the document corresponding to a point belongs to the first or second document group. The first and second attributes may be, for example, the applicant in the patent document. Figure 14 shows a scatter plot of claim-level information extracted from patent documents with companies A and B as applicants, in a principal component space that is based on the second principal component of the problem and the third principal component of the solution, where differences in characteristics between companies A and B tend to be apparent. By outputting such a scatter plot as an analysis result, meaningful comparisons of differences in characteristics between the two companies can be made.

[0050] Furthermore, the output unit 36 ​​presents to the user, as candidates for the name of the first group, one or more words included in documents belonging to the first group whose frequency of occurrence in documents belonging to groups other than the first group is less than or equal to a predetermined value, and assigns the word selected by the user from the presented words as the name of the first group. Specifically, the output unit 36 ​​creates a histogram of the frequency of occurrence of words included in group A, and also creates a histogram of the frequency of occurrence of words included in groups other than group A. The output unit 36 ​​selects the word with the highest frequency of occurrence from the histogram of group A and determines whether the frequency of occurrence of the selected word in the histogram of groups other than group A is less than or equal to a predetermined value. If it exceeds the predetermined value, the output unit 36 ​​selects the next highest-frequency word from the histogram of group A and repeats the same determination until a word whose frequency of occurrence in the histogram of groups other than group A is found to be less than or equal to the predetermined value. The output unit 36 ​​assigns the word found in this way as the name of group A (group name). The group name may be included in each analysis result.

[0051] Next, the operation of the information processing device 10 according to the first embodiment will be described.

[0052] Figure 15 is a flowchart showing the flow of an information processing routine executed by the CPU 12 of the information processing device 10. The CPU 12 reads an information processing program from the storage device 16, loads it into memory 14, and executes it. In this way, the CPU 12 functions as one of the functional components of the information processing device 10, and the information processing routine shown in Figure 15 is executed.

[0053] First, in step S10, the extraction unit 32 retrieves multiple documents from the document DB 50. Next, in step S12, the extraction unit 32 inputs the retrieved documents and instruction statements indicating the items and data format of the information to be extracted to the LLM 40 as prompts, thereby extracting information item by item from each document according to a predetermined data structure that unifies each of the multiple documents into a comparable format. If the documents are patent documents, the extraction unit 32 may extract information on a claim basis.

[0054] Next, in step S14, the classification unit 34 vectorizes the information of each item into an analysis unit from which information has been extracted from multiple documents, and calculates classification values ​​such as cosine similarity between vectors, PCA value, group number calculated by ICA, and sentence skewness (hereinafter referred to as "ICA value"). Next, in step S16, the classification unit 34 accepts the selection of items from the user and classifies the information of the analysis unit into groups using the classification values ​​calculated for the selected items.

[0055] Next, in step S18, the output unit 36 ​​inputs the results of grouping the structured data extracted in step S12 in step S16, along with an instruction statement specifying what kind of analysis results to output, to the LLM 40. This creates and outputs analysis results, for example, as shown in Figures 6 to 14. The information processing routine then terminates.

[0056] As described above, the information processing device according to the first embodiment extracts information related to each of a predetermined set of items from each of a set of documents, according to a predetermined data structure that unifies each of the documents into a comparable format, classifies the documents into groups based on the information extracted for the selected items, and outputs each group in association with the information extracted for the selected items and the information extracted for at least one item other than the selected items. This makes it possible to obtain a variety of easily interpretable analysis results for the documents.

[0057] For example, by extracting information from patent documents as structured data that includes items such as problems, solutions, effects, and application examples, it becomes easier to compare examples between documents or claims where different solutions are applied to the same problem, or where different effects or application examples exist for the same problem. Therefore, it becomes easier to, for example, search for markets for materials developed in-house.

[0058] Furthermore, if you wish to compare specific conditions such as problems, solutions, and effects (hereinafter referred to as "specific conditions of problems, etc."), you may create and structure the problems, solutions, and effects using prompts that do not omit the specific conditions of problems, etc. when summarizing LLM, and then determine their similarity. In this case, it is possible to find examples where very similar problems and effects are solved with different solutions, i.e., cases where the problems and effects are similar but the solutions are dissimilar, and also examples where similar solutions are used to solve different problems and effects, i.e., cases where the solutions are similar but the problems and effects are dissimilar.

[0059] Furthermore, when extracting information from patent documents as structured data consisting of sets of items such as problems, solutions, effects, and application examples, there may be documents where, for example, the "[Problem to be Solved by the Invention]" column is missing, and the problem cannot be found without looking at the entire text. In such cases, comparing the descriptions in the "[Problem to be Solved by the Invention]" column may be meaningless. In this case, in order to determine the similarity of problems, a comparison of unstructured full texts would be made, which would worsen the accuracy of the similarity calculation. Also, even if the "[Problem to be Solved by the Invention]" column is present and a comparison is possible, there is a possibility that meaningless parts of the problem (for example, parts like paragraph number [XXXX]) may be included in the judgment of similarity or features. Moreover, if the documents to be compared are long, manual analysis and judgment of which parts are similar may be required, making automation and labor saving difficult in practice. Even in such cases, according to the first embodiment, by using LLM to extract or generate problems, solutions, effects, application examples, etc., and then extracting the information as unified structured data, it is possible to easily compare problems between patent documents or between claims.

[0060] Furthermore, for example, if there is variation in the length of the sentences described in each item in a document, or in the order of each item within the structured data, it becomes difficult to determine similarity or non-similarity. In the first embodiment, by vectorizing the information of each item in the unified structured data, variations in sentence length can be absorbed, making it easier to determine similarity or non-similarity.

[0061] The LLM used in the above embodiment may be a general-purpose model, or RAG (Retrieval-Augmented Generation) may be applied. In the case of a general-purpose LLM, the model is trained using a general database, so the LLM's responses will also be general. On the other hand, by applying RAG, the likelihood of obtaining analysis results that are specific to the user's needs, such as the user company's industry and developed products, increases. Furthermore, adopting RAG also has the effect of improving security and preventing LLM hallucination.

[0062] To efficiently utilize RAG, structured data can be used as a knowledge graph. Figure 16 shows an example of a knowledge graph. A knowledge graph represents entities and their relationships in a graph structure. Therefore, it can handle complex questions that cannot be answered by simple searches. In addition, any relationship and its related entities can be added to a knowledge graph. For example, in the example in Figure 16, the "Data" node contains descriptions of each patent. As shown by the dashed line in Figure 16, by adding the user's desired subjects or parameters, such as materials used or temperature conditions, to a data node and creating a graph, it is possible to provide accurate responses by searching and extracting features. However, even when using a knowledge graph, if the text entered into the entities is long or contains symbols such as paragraph numbers, it can result in a knowledge graph that contains redundancy. Therefore, by utilizing structured data in LLM that has relatively short key points or features such as problems, solutions, and effects, it is possible to perform highly accurate searches based on keywords and key sentences, as well as similarity judgments and feature extraction at the combination level, using any combination of entities and their relationships. Furthermore, because it uses relatively short, structured data rather than analyzing the similarity of entire texts, it makes it easy to understand where similarities and unique features lie. This allows for a significant degree of automation of manual judgment and analysis, contributing to labor savings. Additionally, knowledge graphs can handle real-time information and reflect the latest data, enabling the generation of responses based on the most up-to-date information whenever necessary.

[0063] A modified example of the first embodiment, applying a combination of knowledge graph and RAG, will now be described. For example, suppose company A, which manufactures a certain product, has a problem X regarding the properties of the material used in that product, and company A cannot find a solution to this problem. Also, suppose a certain research institute has developed a material that can solve problem X, but cannot find an application for it. In this case, as shown in Figure 17, the solution is unclear in company A's knowledge graph, and the application examples are unclear in the research institute's knowledge graph. Therefore, in this modified example, by vectorizing the problem or effect alone, or the problem and effect (hereinafter referred to as "problem, etc.") (dashed line in Figure 17) in the knowledge graphs of each company obtained by RAG, similar problems, etc. can be searched for, and needs as application examples and seeds as solutions can be matched. Then, in this modified example, knowledge graphs with similar problems, etc. are vectorized and similarities are searched for, and they are compiled into a knowledge graph that includes solutions, problems, effects, and application examples. This gives a knowledge graph of what would happen if company A and research institute B co-created. In this case, potential matches between Company A's needs and Research Institute B's capabilities were quickly identified. Therefore, it is possible to support the search for co-creation partners and the consideration of solutions and application examples in the event of co-creation. Note that in Figure 17, "Company A" is not necessarily limited to a company. Also, "Research Institute B" does not necessarily have to be limited to a research institute; it could be a company, an individual, a university, etc. Furthermore, while the above mainly describes the case where the similarity of problems and effects is the target, the target of the similarity investigation may be selected according to the application, such as effects and application examples. Even without using a knowledge graph, in order to find co-creation partners, for example, one could vectorize only the problems and effects of structured patent data and examine their similarity.

[0064] Furthermore, in the above embodiment, Figures 12 to 14 illustrate an example of analysis results based on principal component analysis, but the system may also output analysis results based on independent component analysis. For example, as shown in Figure 18, the distribution based on the filing date of claims (feature n) where the skewness calculated by ICA for the nth independent component is greater than or equal to a predetermined value is shown for a specific item of claim-level information (e.g., summary of the problem) extracted from the patent documents of companies A and B. From such analysis results, let's say we focus on feature 4, which appears only in company B's documents. Figure 19 shows the skewness of the fourth independent component. Then, for example, by inputting the skewness calculated by ICA for feature 4, as shown in Figure 19, and an instruction to create a proposal into the LLM as prompts, a proposal concerning feature 4 of focus may be created. This makes it possible to create, for example, a proposal for company A proposing complementary technologies, a proposal for company B proposing differentiation that leverages the strengths of feature 4, and a proposal for other companies proposing to utilize company B's advanced strengths as an example in development in other fields. Figure 20 shows an example of a proposal prepared by LLM using the first embodiment.

[0065] Furthermore, we will conduct independent component analysis with multiple companies to extract their distinctive fields and set axes for each company to consider new co-creation business ideas. We will then develop new co-creation business ideas by incorporating the distinctive fields and elements of leading companies, thereby devising new co-creation business ideas. For example, as shown in Figure 21, for specific items of claim-level information extracted from patent documents of Company A and Research Institute B (e.g., summary of the problem), we will extract patent documents (feature n) with a skewness above a predetermined value for each group number for each company and Research Institute B. We will then plot the skewness of Company A's patent documents on the vertical axis and the skewness of Research Institute B's patent documents on the horizontal axis, and consider new business matching based on the positioning for each group number. An example of the relationship between document skewness and group number is shown. Once business matching ideas have emerged, we will extract features and their application technologies from the ICA values ​​of information extracted from patent documents of leading companies specializing in the application of Research Institute B's basic research, and develop new business ideas that incorporate the characteristics of leading companies. Furthermore, by extracting the characteristics of multiple rival companies and using a different axis, comparisons can be made. The above process allows for the rapid generation of many co-creation business ideas. Previously, even when co-creation ideas were generated, it was difficult to imagine future business models, and the ideas tended to be extensions of conventional ideas generated within a limited number of members. However, with the above process, it is easy to create multiple proposals that incorporate business models and ideas selected based on the results of trial and error carried out by leading companies. In addition, by comparing with competitor companies using independent component analysis, or by comparing and examining competitor companies instead of leading companies after generating co-creation business ideas, it is possible to identify business gaps and analyze and examine business domains in which one's own strengths can be utilized.

[0066] Figure 22 also shows an example of grouping a company's patents based on group numbers calculated by ICA. Visually, three of the 83 patent documents were identified as being related to regenerative medicine. As shown in Group 0 of Figure 22, the same classification was achieved using group numbers calculated by ICA. Additionally, Group 1 represents features related to the difficulty of optimizing robot parameters.

[0067] Figure 23 shows an example of a comparison between Company A and Company B using ICA. The example in Figure 23 shows the results of extracting patent summaries from patents searched by the automated assistant using LLM, and then performing ICA on them, from patents filed by Company A and Company B. Figure 24 shows an example of the results of a text skewness distribution analysis using ICA for Company B's automated assistant patent. Figure 25 shows an example of a business domain visible from the characteristics of Company B's patent, output by inputting information as a prompt into LLM as shown in Figure 22 or Figure 23, and Figure 26 shows an example of a business domain visible from the characteristics of Company A's patent.

[0068] By using the results of independent component analysis, it is possible to identify distinctive technologies and challenges even within independent categories. Furthermore, distinctive technologies and challenges can be found within similar categories, transcending linguistic boundaries, such as in foreign language documents, photographs, and illustrations. In addition, in the first embodiment, since ICA is performed item by item based on structured data, the extracted features and functions become easier to understand compared to when ICA features are extracted on a document-by-document basis.

[0069] <Second Embodiment> Next, a second embodiment will be described.

[0070] The hardware configuration of the information processing device 210 according to the second embodiment is the same as that of the information processing device 10 according to the first embodiment shown in Figure 1, so a description will be omitted.

[0071] Next, the functional configuration of the information processing device 210 according to the second embodiment will be described. Note that, for the functional configuration of the information processing device 210 according to the second embodiment, the same reference numerals are used for components similar to those in the information processing device 10 according to the first embodiment, and detailed explanations are omitted.

[0072] Figure 27 is a block diagram showing an example of the functional configuration of the information processing device 210. As shown in Figure 27, the information processing device 210 includes, as a functional configuration, an extraction unit 32, a classification unit 34, an output unit 236, and an LLM 40. A knowledge database 242 is stored in a predetermined storage area of ​​the information processing device 210. The knowledge database 242 stores information including structured data extracted from the document database 50 by the extraction unit 32 as knowledge information. Specifically, information including structured data, such as issues and proposals for those issues, acquired by the extraction unit 32, is stored in the knowledge database 242 as knowledge information.

[0073] The output unit 236 acquires the task related to the situation information collected by the user. The situation information may be, for example, image data such as photographs taken by the user, video data, illustrations, text data, audio data, etc. Specifically, as shown in Figure 28, the output unit 236 inputs a prompt to the LLM40 that includes image data, which is an example of situation information, and an instruction sentence that instructs the output of a task related to the situation information, and acquires the task, which is the output of the LLM40.

[0074] Furthermore, the output unit 236 retrieves and outputs suggestions for a problem based on the similarity between the acquired problem and the problems contained in the knowledge information stored in the knowledge database 242. Specifically, the output unit 236 vectorizes each of the sentences indicating the problem acquired from the situation information, and vectorizes each of the problems contained in the knowledge information, and extracts suggestions corresponding to problems for which the cosine similarity between the two vectors is equal to or greater than a threshold. Knowledge information may be stored in the knowledge database 242 in a vectorized state beforehand.

[0075] Here, we assume that multiple documents related to a certain company are stored in document DB50, and information including structured data extracted from those documents is stored as knowledge information in knowledge DB242. The knowledge information includes the content proposed by the company and the results of negotiations regarding those proposals. In this case, output unit 236 extracts the content that the company can propose and the results of those negotiations for the acquired issue. Specifically, output unit 236 obtains the content that the company can propose and the results of those negotiations for the issue indicated by the situation information, based on the similarity between the acquired issue and the issue included in the knowledge information.

[0076] For example, suppose a housing manufacturer engaged in businesses such as renovation, real estate, and urban development has internal documents such as product catalogs, product specifications, repair history, and parts replacement history, as well as documents such as patent literature, sales pitches for related business negotiations, repair proposals, and negotiation results such as the success or failure of proposals, stored in document DB 50. Structured data including issues extracted from these documents, proposed solutions for those issues, and negotiation results is stored in knowledge DB 242. An employee of this housing manufacturer sends a photograph taken in the city or elsewhere to the information processing device 210. The output unit 236 receives the transmitted photograph and inputs a prompt as shown in Figure 28 to LLM 40, and obtains an issue from LLM 40, for example, as shown in Figure 29. The output unit 236 then obtains a proposal, for example, as shown in Figure 30, based on the similarity between the obtained issue and the issues included in the knowledge information. When obtaining a proposal, the output unit 236 may obtain a proposal whose issue similarity is above a threshold value and which has a corresponding negotiation result indicating that the business negotiation was concluded. Furthermore, the output unit may acquire proposals with a similarity to the problem equal to or greater than a threshold value, regardless of whether the negotiation was successful or not, along with the negotiation results related to those proposals.

[0077] Furthermore, when multiple proposals are output, the output unit 236 retrieves information on the proposal adopted by the user and stores the adopted proposal in association with the user's attributes. Then, when instructed by the user, the output unit 236 outputs which proposals were adopted for each of the user's attributes in a comparable manner. User attributes may include, for example, age, occupation, department, position, etc. This allows for the provision of information such as similar success stories from structured data, including the challenges and solutions to successful orders, as well as the sales pitches used, for example, for an old house owned by a salaried worker in his 50s. It also provides information on challenges and solutions to unsuccessful orders, as well as the sales pitches used, thereby improving the efficiency of sales and proposals. This is not limited to the examples above; for instance, in car or motorcycle repair, it can provide not only the problems and solutions themselves, but also methods of persuasion, which can be applied in various ways. For example, it can be used to determine what kind of repairs are possible and what kind of explanations are most acceptable to different types of customers. Or, in nursing care and medical treatment, it can provide information on what kind of treatments and therapies are possible for different types of patients, what kind of explanations will convince patients and their families, and what kind of advice will lead to effective treatment and rehabilitation.

[0078] Next, the operation of the information processing device 210 according to the second embodiment will be described.

[0079] In the second embodiment, the CPU 12 reads an information processing program from the storage device 16, loads it into memory 14, and executes it. The CPU 12 functions as one of the functional configurations of the information processing device 210, and in addition to the information processing routine shown in Figure 14, the presentation process shown in Figure 31 is executed. The presentation process shown in Figure 31 will now be explained.

[0080] In step S210, the output unit 236 receives situation information such as a photograph from the user. Next, in step S212, the output unit 236 inputs a prompt to the LLM40 that includes the situation information and an instruction statement that instructs the output of a task corresponding to the situation information, and obtains the task which is the output of the LLM40.

[0081] Next, in step S214, the output unit 236 obtains suggestions for the issues based on the similarity between the acquired issues and the issues included in the knowledge information stored in the knowledge DB 242, and presents them to the user as suggested information.

[0082] Next, in step S216, the output unit 236 acquires information on the proposals adopted by the user and stores the adopted proposals in association with the user's attributes. Then, in step S218, if instructed by the user, the output unit 236 outputs which solutions or proposals were adopted for each user attribute in a comparable manner, and the presentation process ends.

[0083] As explained above, the information processing device according to the second embodiment stores structured data extracted from documents as knowledge information. The information processing device then uses LLM to acquire issues based on situational information such as photographs received from the user, and presents suggestions that refer to the knowledge information. By referring to the knowledge information, it is possible to obtain suggestions that are not limited to general answers but are tailored to specific companies, thereby acquiring useful information that can lead to business opportunities.

[0084] Furthermore, the information processing device according to the second embodiment displays, in a comparable manner, which solution or proposal was adopted for each user attribute. This information can be used as material to promote internal communication. For example, as shown in Figures 32 and 33, it is expected to show responses that correspond to differences in user attributes, such as differences in the challenges and solutions adopted between company executives and on-site staff, or differences in the challenges and solutions adopted between experienced engineers and junior employees. This can be used to improve internal and inter-organizational communication and to foster mutual understanding of each other's opinions.

[0085] <Third Embodiment> Next, a third embodiment will be described.

[0086] The hardware configuration of the information processing device 310 according to the third embodiment is the same as that of the information processing device 10 according to the first embodiment shown in Figure 1, so a description will be omitted.

[0087] Next, the functional configuration of the information processing device 310 according to the third embodiment will be described. Note that, for the functional configuration of the information processing device 310 according to the third embodiment, the same reference numerals are used for components similar to those in the information processing device 10 according to the first embodiment, and detailed explanations are omitted.

[0088] As shown in Figure 2, the information processing device 310 includes, functionally, an extraction unit 32, a classification unit 34, and an output unit 336. Furthermore, the LLM 40 is stored in a predetermined storage area of ​​the information processing device 310.

[0089] The output unit 336 obtains and outputs analysis results from the structured data extracted by the extraction unit 32, including multifaceted analysis and extraction of insights.

[0090] Specifically, the output unit 336 inputs a prompt to the LLM 40 that includes structured data extracted by the extraction unit 32 and an instruction that instructs the LLM 40 to compare information related to the two entities, including multifaceted analysis and insight extraction, for each of the two entities being compared. The LLM 40 then obtains and outputs the comparison results of the information related to the two entities, which are the output of the LLM 40.

[0091] For example, if the document is a patent document, the output unit 336 vectorizes the information of each patent document contained in the structured data extracted from patent documents in which companies A and B are applicants, and performs independent component analysis to classify each piece of information into categories. The output unit 336 inputs a prompt to the LLM 40 that includes an instruction to output a graph comparing the structured data for companies A and B with the number of applications by category, and obtains the comparison result, which is the output of the LLM 40. Figure 34 shows an example of the comparison result. In this example, it is easy to compare the categories in which companies A and B have strengths.

[0092] Furthermore, the output unit 336 creates prompts that include the acquired comparison results, instructions specifying the content of the analysis such as application trends, medium-term plans, division of roles, and strategies, and instructions specifying the analysis to be performed by a model that performs processing including multifaceted analysis and extraction of insights, and inputs these prompts into LLM40. The model that performs processing including multifaceted analysis and extraction of insights may be, for example, OpenAI's Deep Research, Google Gemini, etc. (hereinafter referred to as "Deep Research, etc."). Figure 35 shows the output of LLM40 in this case.

[0093] Furthermore, the output unit 336 may input a prompt to the LLM 40 that includes structured data categorized by a specific item of the structured data (for example, the "effect" item) and an instruction to compare the number of applications by company, thereby obtaining a comparison result as shown in Figure 36. By using structured data, for example, Deep Research will not be overwhelmed by inputting a large amount of patent documents, and only the essential points will be extracted. Since the structured data is created at the same level in terms of item order, item length, and syntax, it will have a structure that is easy for systems such as Deep Research to understand for comparison, and appropriate comparison results can be obtained.

[0094] The operation of the information processing device 310 according to the third embodiment differs from that of the first embodiment in that step S18 of the information processing routine shown in Figure 15 also includes a process for outputting analysis results, including the multifaceted analysis and extraction of insights described above.

[0095] As described above, the information processing device according to the third embodiment outputs analysis results, including multifaceted analysis and extraction of insights, using structured data extracted from documents. By using structured data, for example, by inputting a large amount of patent documents, Deep Research and other systems can input more information into LLM without exceeding the input volume limit, making it easier to perform comparisons between two parties using independent component analysis, for example. Furthermore, because only the essential points are extracted, and the structured data is created at the same level in terms of item order, item length, and syntax, it has a structure that is easy for Deep Research and other systems to understand for comparison, resulting in highly accurate analysis results. It is also useful when creating comparison tables to compare information by issue, effect, etc.

[0096] Figure 37 shows a specific example of prompts containing instructions and structured data to be input into Deep Research, etc., when the document is a patent document. Figures 38 and 39 show examples of output from Deep Research, etc., obtained using the prompts in Figure 37. Figure 39 shows the detailed analysis results output by Deep Research, etc., for cluster B from the output results of Figure 38. As shown above, structured data is well compatible with Deep Research, etc., and because structured and compressed data can be input, it is possible to feed Deep Research, etc., with many more patents than the original patent text, and obtain highly accurate analysis results.

[0097] Furthermore, the program processing flow described in each of the above embodiments is merely an example, and unnecessary steps may be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0098] Furthermore, the information processing that the CPU reads and executes in each of the above embodiments may be executed by various processors other than the CPU. Examples of such processors include PLDs (Programmable Logic Devices) such as FPGAs (Field-Programmable Gate Arrays) whose circuit configuration can be changed after manufacturing, and dedicated electrical circuits that are processors with circuit configurations specifically designed to execute specific processing, such as ASICs (Application Specific Integrated Circuits). In addition, the information processing may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (for example, multiple FPGAs, and a combination of a CPU and an FPGA). More specifically, the hardware structure of these various processors is an electrical circuit that combines circuit elements such as semiconductor elements.

[0099] Furthermore, while the above embodiments describe a configuration in which the information processing program is pre-stored (installed) in a storage device, the invention is not limited to this configuration. The program may be provided in a form recorded on a recording medium such as a CD-ROM, DVD-ROM (Digital Versatile Disc Read Only Memory), or USB (Universal Serial Bus) memory. Alternatively, the program may be provided in a form that can be downloaded from an external device via a network. [Explanation of Symbols]

[0100] 10, 210, 310 Information Processing Devices 12 CPU 14 memory 16 Storage device 18 Input device 20 Output device 22 Storage medium reader 24 Communication I / F 26 bus 32 Extraction part 34 Classification Department 36, 236, 336 Output section 40 LLM 242 Knowledge DB 50 Document Databases

Claims

1. An extraction unit inputs a prompt to a large-scale language model instructing it to extract information related to each of a predetermined set of items from each of a set of documents, according to a predetermined data structure that unifies each of the documents into a comparable format, and obtains information including structured data, which is the output of the large-scale language model and is information of the predetermined data structure. A classification unit that classifies the plurality of documents into groups based on the degree of relevance of the information, including the structured data, obtained by the extraction unit for one or more items selected from the plurality of items, Each of the above groups includes an output unit that outputs information obtained by the extraction unit for the selected one or more items and information extracted by the extraction unit for at least one item other than the selected one or more items, in association with each other. The aforementioned items include at least one of the following: a problem, a solution to the problem, the effect of the solution, and an example of the application of the solution. The extraction unit extracts from each of the plurality of documents, If the document contains an item corresponding to the solution, prompts are input to the large language model instructing it to extract information related to each of the problems, effects, and application examples based on a summary of the content described in the item corresponding to the solution in the document, thereby obtaining the information related to each of the problems, effects, and application examples which are the output of the large language model. If the document does not contain an item corresponding to the solution, but does contain an item corresponding to the problem, a prompt is input to the large-scale language model instructing it to extract information related to the effects and application examples based on a summary of the content described in the item corresponding to the problem in the document, and the information related to the effects and application examples, which is the output of the large-scale language model, is obtained. Information processing device.

2. In the case where the aforementioned document is a patent document, As step 1, the extraction unit inputs a prompt to the large-scale language model instructing it to extract at least one of the features of the claims and the abstract of the claims in the patent document, and obtains at least one of the features of the claims and the abstract of the claims, which are outputs of the large-scale language model, as information related to the solution. Step 2 involves inputting a prompt to the large language model instructing it to generate matters relating to at least one of the problem, the solution, and the effect, based on at least one of the acquired features of the claim and the summary of the claim, thereby obtaining the matters which are the output of the large language model. A prompt is input to the large language model instructing it to generate the summary, which includes at least one of the acquired features of the claim and the summary of the claim, and at least one of the aforementioned matters, thereby obtaining the summary, which is the output of the large language model. Based on the summary obtained, information relating to the solution and at least one of the problem and the effect is obtained. The information processing apparatus according to claim 1.

3. In the case where the aforementioned document is a patent document, As step 1, the extraction unit inputs a prompt to the large-scale language model instructing it to extract at least one of the features of the claims and the abstract of the claims in the patent document, and obtains at least one of the features of the claims and the abstract of the claims, which are outputs of the large-scale language model, as information related to the solution. Step 2 involves inputting a prompt into the large-scale language model instructing it to extract information related to the application described in the patent document based on at least one of the acquired claim features and the claim summary, thereby obtaining information related to the application, which is the output of the large-scale language model; or inputting a prompt into the large-scale language model instructing it to generate application examples not described in the patent document based on at least one of the acquired claim features and the claim summary, thereby obtaining the application, which is the output of the large-scale language model; or inputting a prompt into the large-scale language model instructing it to extract information related to the application described in the patent document that is not based on the claims in the patent document, thereby obtaining information related to the application, which is the output of the large-scale language model. The information processing apparatus according to claim 1.

4. The information processing apparatus according to claim 1, wherein the classification unit vectorizes the information including the structured data extracted by the extraction unit, either the data for each of the multiple items or the data by combining the items, and classifies the multiple documents into groups based on the similarity between the vectors of information related to the selected items or the characteristics of the vectors.

5. The classification unit performs principal component analysis or independent component analysis on each of the multiple items in each of the multiple documents, The output unit outputs a scatter plot in which points corresponding to each of the multiple documents are arranged in the principal component space or independent component space in a different display manner for each group to which the documents corresponding to the points belong, based on the results of principal component analysis or independent component analysis for the selected items. The information processing apparatus according to claim 4.

6. When the principal component analysis is performed by the classification unit, The output unit outputs a scatter plot in which points corresponding to documents included in the first and second document groups are arranged in different display modes depending on whether the document corresponding to a point belongs to the first or second document group, in a principal component space that is based on two or more principal components that have a predetermined difference between the first attribute and the second attribute, among the nth principal components obtained from the principal component analysis results of the information obtained by the extraction unit as the output of the large-scale language model from each of the first document group having the first attribute and the second document group having the second attribute, respectively. The information processing apparatus according to claim 5.

7. The information processing apparatus according to any one of claims 1 to 6, wherein the output unit presents to the user, as candidates for the name of the first group, one or more words from among the words contained in documents belonging to the first group whose frequency of appearance in documents belonging to groups other than the first group is less than or equal to a predetermined value, and assigns the word selected by the user from the presented words as the name of the first group.

8. The extraction unit extracts additional information, including date and time information related to the document, from each of the plurality of documents. The output unit further outputs the additional information in association with it. The information processing apparatus according to any one of claims 1 to 6.

9. The extraction unit inputs a prompt to the large-scale language model instructing it to extract information related to the issue, details of the solution to the issue, and date and time information from each of the plurality of documents, and obtains the output of the large-scale language model, which is information related to the issue, details of the solution to the issue, and date and time information. The classification unit classifies the plurality of documents into groups based on information related to the problem. death, The output unit, for each group, associates the details of the solution extracted from the documents with each of the multiple documents, arranges the multiple documents in chronological order, and outputs information highlighting the details of the solution that do not appear in the details of the solution of the aforementioned patent document for the given date and time. The information processing apparatus according to claim 8.

10. In the case where the aforementioned document is a patent document, The extraction unit extracts information related to each of the plurality of items on a claim basis, The output unit outputs the information extracted by the extraction unit in the units of the claims. The information processing apparatus according to any one of claims 1 to 6.

11. The structured data, including the problem and the proposal for the problem, is stored in the storage unit as knowledge information. The output unit inputs a prompt to the large-scale language model, which includes situational information collected by the user and an instruction sentence instructing the extraction of issues from the situational information, to obtain the issues which are the output of the large-scale language model, and based on the similarity between the obtained issues and the issues included in the knowledge information, it obtains and outputs suggestions for the obtained issues. The information processing apparatus according to any one of claims 1 to 6.

12. The information processing apparatus according to claim 11, wherein the aforementioned situation information includes image data.

13. The information processing apparatus according to claim 11, wherein, when the output unit outputs a plurality of proposals, it obtains the proposal adopted by the user and outputs the adopted proposal in a manner that can be compared for each attribute of the user.

14. The information processing apparatus according to any one of claims 1 to 6, wherein the output unit inputs a prompt to the large-scale language model that includes information including the structured data extracted by the extraction unit and an instruction sentence that instructs an analysis including multifaceted analysis and extraction of insights, and obtains and outputs the analysis result which is the output of the large-scale language model.

15. The information processing apparatus according to claim 14, wherein the output unit classifies the information, including the structured data acquired by the extraction unit, into categories by independent component analysis and outputs the analysis results for each category.

16. The extraction unit inputs a prompt to the large-scale language model instructing it to extract information related to each of a predetermined set of items from each of a set of documents, according to a predetermined data structure that unifies each of the set of documents into a comparable format, and obtains information including structured data, which is the output of the large-scale language model and is information of the predetermined data structure. The classification unit classifies the multiple documents into groups based on the degree of relevance of the information, including the structured data obtained by the extraction unit, for one or more items selected from the multiple items. The output unit outputs to each of the groups the information obtained by the extraction unit for the selected item and the information extracted by the extraction unit for at least one item other than the selected item, in association with each group. The aforementioned items include at least one of the following: a problem, a solution to the problem, the effect of the solution, and an example of the application of the solution. The extraction unit extracts from each of the plurality of documents, If the document contains an item corresponding to the solution, prompts are input to the large language model instructing it to extract information related to each of the problems, effects, and application examples based on a summary of the content described in the item corresponding to the solution in the document, thereby obtaining the information related to each of the problems, effects, and application examples which are the output of the large language model. If the document does not contain an item corresponding to the solution, but does contain an item corresponding to the problem, a prompt is input to the large-scale language model instructing it to extract information related to the effects and application examples based on a summary of the content described in the item corresponding to the problem in the document, and the information related to the effects and application examples, which is the output of the large-scale language model, is obtained. Information processing methods.

17. Computers, This command instructs the system to extract information related to each of several predetermined items from each of several documents, according to a predetermined data structure that unifies each of the documents into a comparable format. An extraction unit inputs a prompt to a large-scale language model and obtains information including structured data, which is the output of the large-scale language model and is information of a predetermined data structure. A classification unit that classifies the plurality of documents into groups based on information including the structured data obtained by the extraction unit for one or more items selected from the plurality of items, and An information processing program for each of the aforementioned groups to function as an output unit that outputs information obtained by the extraction unit for one or more selected items in association with information obtained by the extraction unit for at least one item other than the selected items, The aforementioned items include at least one of the following: a problem, a solution to the problem, the effect of the solution, and an example of the application of the solution. The extraction unit extracts from each of the plurality of documents, If the document contains an item corresponding to the solution, prompts are input to the large language model instructing it to extract information related to each of the problems, effects, and application examples based on a summary of the content described in the item corresponding to the solution in the document, thereby obtaining the information related to each of the problems, effects, and application examples which are the output of the large language model. If the document does not contain an item corresponding to the solution, but does contain an item corresponding to the problem, a prompt is input to the large-scale language model instructing it to extract information related to the effects and application examples based on a summary of the content described in the item corresponding to the problem in the document, and the information related to the effects and application examples, which is the output of the large-scale language model, is obtained. Information processing program.

Citation Information

Patent Citations

  • Technical document retrieving device

    JP2002163275A

  • Abstract preparation supporting system, program, abstract preparation supporting method, patent document retrieving system, and patent document rerieving method

    JP2005031813A

  • Patent analyzing device, patent analyzing system, patent analyzing method and program

    JP2007148630A

  • DATA PROCESSING APPARATUS AND DATA PROCESSING METHOD

    JP7431379B1

  • Program, method, information processing device, and system

    JP7493195B1