Document visualization method, document visualization device, document visualization program, computer-readable storage medium storing the document visualization program, and document generation method using the document visualization method.
The document visualization method leverages a language model to quantify and cluster electronic documents, addressing the limitations of conventional AI methods by providing clear cluster identification and improved usability through structured text block analysis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2026-03-13
AI Technical Summary
Conventional methods for extracting and visualizing relationships between electronic documents using AI are either opaque (black box) or lack expressiveness and processing speed, and traditional statistical methods are inconvenient to use.
A document visualization method that utilizes a language model to quantify and cluster electronic documents into multidimensional data, combining morphological and language model-based approaches for explicit identification of cluster information and improved usability, with interactive visualization options.
Enables clear identification of cluster characteristics and enhanced usability through macroscopic and microscopic analyses, facilitating user-friendly visualizations and accurate classification of text blocks based on document structure.
Smart Images

Figure 2026046633000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a document visualization method, a document visualization apparatus, a document visualization program, a computer-readable storage medium storing the document visualization program, and a document generation method using the document visualization method.
Background Art
[0002] For example, Patent Document 1 discloses a computer-implemented method for generating a text string by processing an image. Specifically, the method disclosed in Patent Document 1 extracts a line image from a filled form and is configured to digitize the line image into a text string. This method generates a bounding box around the text and executes various processes based on the position of the bounding box.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Technologies for extracting and outputting text from various electronic documents, as in Patent Document 1, are well-known.
[0005] On the other hand, the inventors of the present application considered not only extracting text but also inputting the text extracted from an electronic document into a language model such as a generative AI. This enables two-dimensional or three-dimensional mapping of each of a plurality of electronic documents and grouping of the mapped electronic documents into a plurality of clusters. Based on this, it was considered that the relationships between electronic documents can be visualized through each grouped cluster.
[0006] However, when conventional AI is used to map electronic documents, the processes involved, including the so-called embedding process, remain a black box. Therefore, analysts cannot understand the information (features) that led to each electronic document being classified into a different cluster.
[0007] In contrast, while traditional methods that statistically analyze morphological elements to map electronic documents are easy to understand, they also have drawbacks in terms of expressiveness, processing speed, and other aspects, making them inconvenient to use.
[0008] This disclosure is based on the aforementioned points and aims to achieve both the explicit identification of information characterizing each cluster and the improvement of the usability of the clustering process when clustering mapped data based on multiple electronic documents. [Means for solving the problem]
[0009] A first aspect of this disclosure relates to a document visualization method that uses a computer having a processing unit to read multiple electronic documents, each containing multiple texts, and visualize the relationships between the multiple electronic documents.
[0010] According to the first embodiment, the document visualization method involves the arithmetic unit reading the plurality of electronic documents as a collection of multiple text blocks, each containing one or more texts; the arithmetic unit quantifying the plurality of electronic documents into multidimensional data via a language model that takes the multiple text blocks corresponding to each electronic document as input; aggregating the multidimensional data obtained for each text block for each electronic document and clustering it into a plurality of first clusters; the arithmetic unit obtaining feature quantities for each of the plurality of electronic documents, statistically quantifying each morpheme contained in each electronic document, using the multiple text blocks corresponding to each electronic document as input; the arithmetic unit associating the feature quantities for each first cluster based on the electronic documents constituting each of the plurality of first clusters; and the arithmetic unit visualizing the plurality of clusters as a two-dimensional or three-dimensional first map, and controlling the display mode of the first map according to the value of the feature quantities.
[0011] According to the first embodiment described above, by reading multiple electronic documents as a collection of text blocks, it becomes possible to generate multidimensional data for each text block. Compared to a configuration in which multidimensional data is generated on a per-electronic document basis, the data size input to the language model can be reduced.
[0012] Furthermore, according to the first embodiment, the calculation unit performs quantification of multidimensional data for clustering via a language model, while quantifying the information (features) that characterize each first cluster via statistical processing of morphemes. Transformation to multidimensional data via a language model offers superior expressive power and processing speed compared to transformation via morphemes. On the other hand, quantification via morphemes provides a clearer meaning, such as the number of occurrences or frequency of occurrence of each morpheme.
[0013] Therefore, it is possible to leverage the advantages of both language model-based and morphological methods. This allows for both explicit identification of information characterizing each first cluster of electronic documents and improved usability of clustering processes.
[0014] Furthermore, since each feature is assigned to each morpheme, the value of the feature for each morpheme remains unchanged even if the number of clusters in the first cluster changes. Therefore, it becomes possible to smoothly and quickly associate the features with each changed first cluster. This also contributes to improved usability, such as increased processing speed.
[0015] Furthermore, according to a second aspect of this disclosure, the calculation unit may cluster each of the text blocks corresponding to the multidimensional data into a plurality of second clusters composed of one or more text blocks, associate the feature quantities with each second cluster based on the text blocks that constitute each of the plurality of second clusters, visualize the plurality of second clusters as a two-dimensional or three-dimensional second map, and control the display mode of the second map according to the level of the feature quantities.
[0016] According to the second embodiment described above, in addition to clustering and visualization of electronic documents, clustering and visualization of text blocks are also performed. This makes it possible to visualize the similarities and differences between electronic documents on a text block basis.
[0017] By enabling the combined use of a more macroscopic analysis based on the first map and a more microscopic analysis based on the second map, it becomes possible to achieve more user-friendly visualizations.
[0018] Furthermore, according to a third aspect of this disclosure, the calculation unit may be connected to a reception unit that receives user input, and the calculation unit may select two or more of the plurality of electronic documents based on the user input to the reception unit, and visualize the second map for each text block of the two or more selected electronic documents.
[0019] According to the third embodiment described above, an interactive interface with the analyst is realized. This enables visualization that is more user-friendly. Furthermore, since the second map is a map for each text block, it is thought to be less visually appealing than the first map due to the larger number of plots. By configuring the system to visualize the second map for the desired electronic document, visualization that is more user-friendly can be achieved in terms of the visual appeal of the second map as well.
[0020] Furthermore, according to a fourth aspect of this disclosure, the calculation unit may obtain the feature quantities based on morphological analysis using the plurality of text blocks as input, and the calculation unit may obtain an index based on the frequency or number of occurrences of each morpheme as the feature quantities.
[0021] According to the fourth embodiment described above, it becomes possible to extract morphemes that characterize each electronic document. This enables more user-friendly visualization.
[0022] Furthermore, according to a fifth aspect of this disclosure, the calculation unit may be connected to a reception unit that receives user input, and when the reception unit receives a predetermined input, the calculation unit may superimpose and display a first word cloud representing the high or low of the index for each morpheme for each of the first clusters onto the first map.
[0023] According to the fifth embodiment described above, the analyst can visualize the morphemes characterizing each first cluster through their display in the first word cloud. This further improves the usability of the first map.
[0024] Also, according to the sixth aspect of the present disclosure, the plurality of text blocks include author information characterizing the author of each of the electronic documents corresponding to the plurality of text blocks, and when the reception unit receives a predetermined operation input, the calculation unit represents the frequency or the number of occurrences of the author information for each of the first clusters as a second word cloud together with the first word cloud representing the frequency or the number of occurrences of each of the morphemes regarding information other than the author information, and superimposes and displays them on the first map.
[0025] According to the sixth aspect, an analyst can visually recognize the author information characterizing each of the first clusters through the second word cloud. Thus, the usability of the first map can be further improved.
[0026] Also, according to the seventh aspect of the present disclosure, the author information may include information indicating the organization to which the author of each of the electronic documents belongs.
[0027] According to the seventh aspect, an analyst can visually recognize the organization to which the author characterizing each of the first clusters belongs through the second word cloud. Thus, the usability of the first map can be further improved.
[0028] Furthermore, according to an eighth aspect of this disclosure, each of the plurality of electronic documents is composed of document data containing the plurality of texts, and the calculation unit generates and outputs structured data based on the electronic document, associating text attributes that characterize at least one of the display position and display form of each of the plurality of texts on the electronic document, text corresponding to the text attribute, and identification information for identifying the text on the electronic document, and the calculation unit inputs a prompt to the generation AI that includes accepting the structured data as input, the input format of the structured data, and outputting second data representing the breakdown of the texts belonging to each of the plurality of text blocks, which is divided into the plurality of text blocks by referring to the first data, using the identification information associated with the text, and the calculation unit inputs the first data to the generation AI, thereby obtaining the second data via the generation AI, and the calculation unit selects the text by referring to the first data and combines the selected text in an order based on the identification information, thereby outputting at least one of the plurality of text blocks as a document.
[0029] Traditionally, it was widely known that text extraction from electronic documents was done on a page-by-page basis.
[0030] In response to this, the inventors of this application attempted to output text in blocks, each consisting of multiple texts such as the introduction, body, and conclusion of an electronic document, instead of outputting on a page-by-page basis, by inputting text extracted from the electronic document into a generating AI. Block-based output offers superior usability compared to conventional technologies.
[0031] However, simply inputting text is insufficient for ensuring the accuracy of the classification into blocks by the generating AI. Furthermore, even if output in block units were possible, the data size output by the generating AI would be enormous, which would also be insufficient for ensuring processing speed.
[0032] To solve the aforementioned problems, it is not enough to simply output text in blocks; it is necessary to achieve both high accuracy in classifying the text for each block and high classification speed.
[0033] In contrast, according to the eighth embodiment described above, the generating AI performs division into multiple text blocks by referring to text attributes that characterize at least one of the display position and display form. The display position of each text reflects the text structure of the electronic document, such as line breaks between sections and chapters. On the other hand, the display form of each text also reflects the text structure of the electronic document, such as the font size in section titles and chapter titles.
[0034] Therefore, by dividing document data into multiple text blocks based on such text structure, a more accurate classification can be performed that reflects the actual section and chapter structure in electronic documents. This contributes to improving the accuracy of text classification into each text block.
[0035] Furthermore, by dividing an electronic document into multiple text blocks, it becomes possible to process the electronic document with other document analysis AIs in text block units rather than inputting it page by page. This makes even electronic documents with large data sizes suitable for various processing methods using generative AI (e.g., LLM).
[0036] Furthermore, according to the eighth aspect, the generating AI outputs second data consisting of identification information associated with each text block, instead of outputting the text body of each text block.
[0037] This allows for a more compact data size output from the generating AI, improving the speed of text classification into text blocks. This is particularly effective when processing a large number of electronic documents with the generating AI.
[0038] Furthermore, according to a ninth aspect of the present disclosure, the plurality of text blocks may be constructed to reflect the section or chapter structure of the electronic document, and the calculation unit may input to the generating AI a prompt to divide the document data by section or by chapter.
[0039] According to the ninth embodiment described above, by explicitly indicating in the prompt that the section or chapter structure of the electronic document should be reflected, the generating AI can be made to actively estimate the section or chapter structure. This further improves the usability of the text block.
[0040] Furthermore, according to a tenth aspect of this disclosure, the calculation unit may determine whether the document structure of the electronic document is known or not, and if the document structure is known, the calculation unit may input the first data into a rule-based estimation model to obtain the second data through the estimation model, and if the document structure is unknown, the calculation unit may input the first data into the generation AI to obtain the second data through the generation AI.
[0041] According to the tenth embodiment described above, the calculation unit acquires second data using two methods depending on the document structure of the electronic document. By using a rule-based estimation model, it is possible to achieve more accurate classification for documents with clearly defined electronic document formats, such as patent documents. On the other hand, for documents with unclear formats, it is possible to achieve more versatile classification by using a method like that of the first embodiment described above.
[0042] Furthermore, an eleventh aspect of this disclosure relates to a document visualization device that uses a computer having a processing unit to read multiple electronic documents, each containing multiple texts, and visualizes the relationships between the multiple electronic documents.
[0043] According to the 11th embodiment, the document visualization device includes means for reading the plurality of electronic documents as a collection of a plurality of text blocks, each containing one or more texts; means for quantifying the plurality of electronic documents into multidimensional data via a language model that takes the plurality of text blocks corresponding to each electronic document as input, and for aggregating the multidimensional data obtained for each text block for each electronic document and clustering it into a plurality of first clusters; means for obtaining, for each of the plurality of electronic documents, a feature quantity obtained by statistically quantifying each morpheme contained in each electronic document, using the plurality of text blocks corresponding to each electronic document as input; means for associating the feature quantities for each first cluster based on the electronic documents that constitute each of the plurality of first clusters; and means for visualizing the plurality of first clusters as a two-dimensional or three-dimensional first map, and controlling the display mode of the first map according to the value of the feature quantities.
[0044] According to the 11th embodiment described above, it is possible to achieve both explicit identification of information characterizing each cluster of electronic documents and improved usability of the clustering process.
[0045] Furthermore, a twelfth aspect of this disclosure relates to a document visualization program that, when executed by a computer having a processing unit, reads multiple electronic documents, each containing multiple texts, and visualizes the relationships between the multiple electronic documents.
[0046] According to the 12th embodiment, the document visualization program causes the computer to execute the following processes: the calculation unit reads the plurality of electronic documents as a collection of a plurality of text blocks, each containing one or more texts; the calculation unit quantifies the plurality of electronic documents into multidimensional data via a language model that takes the plurality of text blocks corresponding to each electronic document as input, aggregates the multidimensional data obtained for each text block for each electronic document, and clusters it into a plurality of first clusters; the calculation unit obtains feature quantities for each of the plurality of electronic documents, statistically quantifying each morpheme contained in each electronic document, using the plurality of text blocks corresponding to each electronic document as input; the calculation unit associates the feature quantities with each first cluster based on the electronic documents that constitute each of the plurality of first clusters; and the calculation unit visualizes the plurality of first clusters as a two-dimensional or three-dimensional first map and controls the display mode of the first map according to the value of the feature quantities.
[0047] According to the 12th embodiment described above, when clustering multiple electronic documents, it is possible to achieve both explicit identification of information characterizing each cluster and improved usability of the clustering process.
[0048] Furthermore, a thirteenth aspect of this disclosure relates to a computer-readable storage medium that stores the document visualization program.
[0049] According to the 13th embodiment described above, when clustering multiple electronic documents, it is possible to achieve both explicit identification of information characterizing each cluster and improved usability of the clustering process.
[0050] Furthermore, a fourteenth aspect of this disclosure relates to a document generation method using the document visualization method according to the first aspect.
[0051] Furthermore, according to the 14th embodiment, if an electronic document describing a newly patentable technical idea is defined as a new technical document, and multiple patent application documents similar to the new technical document are defined as multiple prior art documents, then the document generation method includes: the calculation unit generating the multiple text blocks for the new technical document; the calculation unit extracting classification information from the multiple text blocks generated based on the new technical document, indicating at least one of the technical field, background, prior art, technical problems, features, and summary of the technical idea; the calculation unit obtaining multiple prior art documents based on search extension generation using the classification information as input; the calculation unit generating the multiple text blocks for each of the multiple prior art documents; and the calculation unit obtaining the features and clustering into multiple first clusters for the document group comprising the new technical document and the multiple prior art documents. The calculation unit performs clustering of the document collection into the plurality of second clusters, associates the feature quantities with each of the plurality of first clusters and the plurality of second clusters, visualizes the plurality of first clusters on the first map, extracts one or more prior art documents from the plurality of prior art documents that belong to the same first cluster as the new technical book, in order of proximity on the first map, visualizes the plurality of second clusters on the second map, extracts the points of agreement and difference between the new technical book and the extracted prior art documents for each text block according to the level of the feature quantities, and inputs the points of agreement and difference to the second generating AI, causing the second generating AI to determine whether or not the patentability of the technical idea is affirmed, and visualizes the result of that determination.
[0052] According to the 14th embodiment described above, by using the first cluster and the second cluster in combination, it is possible to output similarities and differences at the text block level of an electronic document, such as sections and chapters. By making judgments at the text block level rather than at the electronic document level, similarities and differences can be determined with greater accuracy.
[0053] Furthermore, according to a 15th aspect of this disclosure, in the document generation method, the calculation unit may acquire progress information showing the review process of the prior art, the calculation unit may input the progress information to the second generation AI, and cause the AI to perform a judgment based on the progress information, using the points of agreement and the points of difference as input.
[0054] According to the 15th embodiment described above, by referring to the examination history of prior art, it is possible to make a more accurate determination as to whether or not patentability is affirmed.
[0055] Furthermore, according to a sixteenth aspect of this disclosure, in the document generation method, the calculation unit may obtain the examination standards for patent applications in the target country, the calculation unit may input standard information indicating the examination standards to the second generation AI, and cause the AI to perform a judgment based on the points of agreement and the points of difference based on the standard information.
[0056] According to the 16th embodiment described above, by referring to the examination standards of the country where the application is being filed, the determination of whether or not patentability is affirmed can be made with greater accuracy.
[0057] Furthermore, according to a 17th aspect of this disclosure, in the document generation method, if the patentability of the novel technical document is affirmed, the calculation unit may cause the third generation AI to output an invention summary document describing the technical idea based on the novel technical document, the extracted prior art, the similarities, and the differences.
[0058] According to the 17th embodiment described above, by using visualization based on the first and second clusters in combination, as mentioned above, it is possible to accurately determine the similarities and differences, and as a result, it becomes possible to generate a more appropriate invention summary. By exchanging this invention summary between the company and its agent, it can be used to assist in the preparation of patent claims, specifications, etc. This makes it possible to proceed with patent applications efficiently.
[0059] Furthermore, according to the 18th aspect of this disclosure, in the document creation method, if the patentability of the new technical document is affirmed, the calculation unit may cause the third generating AI to output at least a portion of the application for a patent application describing the technical idea, based on the new technical document, the extracted prior art, the similarities, and the differences.
[0060] According to the 18th embodiment described above, by using visualization based on the first and second clusters in combination, it is possible to accurately determine the similarities and differences, as described above, and as a result, it becomes possible to generate a more appropriate application. This makes it possible to proceed with patent applications efficiently.
[0061] Furthermore, according to a 19th aspect of this disclosure, in the document creation method, the calculation unit may update the invention summary based on the contents of the application actually submitted to the office after the patent application based on the new technical document.
[0062] According to the 19th embodiment described above, by having the generating AI update the invention summary, discrepancies between the invention summary and the actual application are eliminated. This makes it possible to manage the contents of each patent application more efficiently.
[0063] Furthermore, according to a 20th aspect of this disclosure, the calculation unit may output the contents of at least one of the invention summary and the application via an interactive fourth generating AI.
[0064] According to the 20th embodiment, the fourth generating AI can output the contents of each document. This makes it possible to manage the contents of each patent application more efficiently. [Effects of the Invention]
[0065] As explained above, this disclosure makes it possible to achieve both explicit identification of information characterizing each cluster and improved usability of the clustering process when clustering mapped data based on multiple electronic documents. [Brief explanation of the drawing]
[0066] [Figure 1] Figure 1 is a diagram illustrating the hardware configuration of a document visualization device. [Figure 2] Figure 2 is an example of the software configuration of a document visualization device. [Figure 3] Figure 3 is a flowchart illustrating the first part of the document visualization method. [Figure 4A] Figure 4A is a block diagram illustrating the input and output of a document visualization method. [Figure 4B] Figure 4B is a block diagram illustrating the input and output of a document visualization method. [Figure 4C] Figure 4C is a block diagram illustrating the input and output of a document visualization method. [Figure 5] Figure 5 is a diagram illustrating document data. [Figure 6] Figure 6 is a flowchart illustrating the steps of the document retrieval process. [Figure 7] Figure 7 is a flowchart illustrating the steps of the electronic document structuring process. [Figure 8] Figure 8 is a diagram illustrating structured data. [Figure 9] Figure 9 is a flowchart illustrating the steps of the index acquisition process. [Figure 10A] Figure 10A illustrates a system prompt that is input to the first generation AI. [Figure 10B] Figure 10B is a table illustrating each input field in the system prompt. [Figure 11] Figure 11 illustrates a user prompt that is input to the first generation AI. [Figure 12] Figure 12 is an example of a display screen showing the output results from the first generated AI. [Figure 13] Figure 13 is a flowchart illustrating the steps of the index matching process. [Figure 14A]Figure 14A illustrates the output of a text block. [Figure 14B] Figure 14B is an example illustrating the output of a text block for the document data shown in Figure 5. [Figure 15] Figure 15 is a flowchart illustrating the latter part of the document visualization method. [Figure 16] Figure 16 is a flowchart illustrating the steps of the clustering process. [Figure 17] Figure 17 is a flowchart illustrating the steps of the feature acquisition process. [Figure 18] Figure 18 is a diagram illustrating the operations related to feature extraction. [Figure 19] Figure 19 is a flowchart illustrating the steps of the feature assignment process. [Figure 20] Figure 20 is a flowchart illustrating the steps of the map display process. [Figure 21] Figure 21 is a scatter plot illustrating the display modes of the first visualization map corresponding to each electronic document. [Figure 22] Figure 22 is a scatter plot illustrating an example of the display mode of the first visualization map. [Figure 23] Figure 23 is a scatter plot illustrating an example of the display mode of the first visualization map. [Figure 24] Figure 24 illustrates the display modes according to the features in the first visualization map. [Figure 25] Figure 25 illustrates an example of extraction from the first visualization map. [Figure 26] Figure 26 is a scatter plot illustrating the display configuration of the second visualization map corresponding to each text block. [Figure 27] Figure 27 illustrates an example of how the second visualization map can be displayed. [Figure 28] Figure 28 illustrates the display modes in the second visualization map according to the features. [Figure 29] Figure 29 is a scatter plot illustrating an example of how the second visualization map can be displayed in relation to the patent document. [Figure 30] Figure 30 is a network diagram showing another example of the display mode of the second visualization map for the patent document. [Figure 31] Figure 31 is a flowchart illustrating the first part of the document creation process. [Figure 32] Figure 32 is a flowchart illustrating the latter part of the document creation process. [Figure 33] Figure 33 is a diagram illustrating the procedure for obtaining prior art. [Figure 34] Figure 34 is a diagram illustrating how to obtain progress information and reference information. [Figure 35] Figure 35 is a diagram illustrating the determination of patentability. [Figure 36] Figure 36 is a diagram illustrating the generation of an invention summary. [Figure 37] Figure 37 is a diagram illustrating the generation of a patent application form. [Figure 38] Figure 38 is a diagram illustrating the updating of the invention summary. [Figure 39] Figure 39 is a diagram illustrating the output of the invention. [Figure 40] Figure 40 is an example of output data from a generating AI. [Modes for carrying out the invention]
[0067] The embodiments of this disclosure will be described below with reference to the drawings. Note that the following description is illustrative.
[0068] <1.Device configuration> Figure 1 is a diagram illustrating the hardware configuration of the document visualization device and document generation device (specifically, the computer 1 that constitutes the document visualization device and the document generation device, respectively) related to this disclosure, and Figure 2 is a diagram illustrating the software configuration thereof.
[0069] As illustrated in Figure 1, computer 1 comprises a Central Processing Unit (CPU) 3 that controls the entire computer 1, a Read Only Memory (ROM) 5 that stores boot programs and the like, a Random Access Memory (RAM) 7a that functions as main memory, and a Solid State Drive (SSD) 7b as secondary storage. Note that a Hard Disk Drive (HDD) or the like can be used instead of the SSD 7b as secondary storage.
[0070] Of these elements, the CPU 3 executes various programs. The CPU 3 constitutes the arithmetic unit in this embodiment. The RAM 7a and SSD 7b temporarily or continuously store the programs executed by the CPU 3. The RAM 7a and SSD 7b constitute the storage unit 7 in this embodiment.
[0071] Computer 1 also includes a display 9, graphics memory (Video RAM: VRAM) 11 for storing image data displayed on the display 9, and a keyboard 13a and mouse 13b as a human-machine interface. The keyboard 13a and mouse 13b each accept at least one of input and / or operation (hereinafter collectively referred to as "operation input") from the user (analyst).
[0072] The keyboard 13a and mouse 13b are configured to receive user (analyst) input and constitute the reception unit (input unit) 13 in this embodiment. The reception unit 13 is electrically connected to the CPU 3 by wireless or wired connection. The display 9 can display a screen based on the calculation results of the CPU 3, which will be described later, and constitutes the display unit in this embodiment.
[0073] Furthermore, the computer 1 according to this embodiment can send and receive data with external devices via a communication unit 15 configured as a communication interface. Specifically, the computer 1 is connected to a server machine 101 and an external server 102 via the communication unit 15. The computer 1 and server machine 101 form a closed network within the analyst's affiliated organization to prevent them from being used for machine learning by other companies or organizations.
[0074] In this context, the term "affiliated institution" refers to a concept that encompasses both public or private institutions to which the analyst belongs, such as universities and administrative corporations, and companies to which the analyst belongs, such as manufacturers.
[0075] The server machine 101 is composed of a computer equipped with a storage device 7c, which is, for example, an SSD or HDD, a CPU (not shown), and a GPU (Graphics Processing Unit) 19. The server machine 101 may also be read as "external computer" or "second computer." The storage device 7c, together with the RAM 7a and SSD 7b, constitutes the storage unit 7 in this embodiment.
[0076] Furthermore, server machine 101 is equipped with a first generation AI 201 utilizing GPU 19, a second generation AI 208, a third generation AI 209, and a fourth generation AI 210 (see Figure 4A and Figures 35-39 below).
[0077] The first generation AI 201 is an interactive generation AI for analyzing electronic documents Dc, which has been pre-trained on a large number of electronic documents Dc. The first generation AI 201 is an example of a "generation AI" in this embodiment. Other generation AIs will be described later.
[0078] Specifically, the first generation AI201 is composed of a Large Language Model (LLM) pre-trained on a large number of electronic documents (Dc). This LLM is constructed, for example, by a transferer composed of a neural network.
[0079] External server 102 is a public institution in any country, or a server machine within the aforementioned affiliated institution. This external server 102 forms an open or closed network with computer 1 and server machine 101.
[0080] Note that both server machine 101 and external server 102 are not essential. At least some of the functions performed by server machine 101 and external server 102 may be implemented by another server machine, or by computer 1 and its CPU 3. Furthermore, as described later, multiple computers 1, and by extension multiple CPUs 3, may be made to execute the following processes separately or simultaneously.
[0081] As illustrated in Figure 2, the program memory of SSD7b stores a document visualization program 21, a document generation program 27, an operating system (OS) (not shown in the document), and an application program (also not shown).
[0082] Here, the document visualization program 21 is a program configured to cause the computer 1 to execute each process that constitutes the document visualization method according to this embodiment. The document visualization program 21 is pre-stored in a computer-readable storage medium 17. This storage medium 17 is a tangible storage medium, such as a disk medium.
[0083] Specifically, the document visualization program 21 according to this embodiment consists of a text block generation program 23 and an electronic document visualization program 25.
[0084] Similarly, the document generation program 27 is a program configured to cause the computer 1 to execute each process constituting the document generation method according to this embodiment. The document generation program 27 is pre-stored in a computer-readable storage medium 17. This storage medium 17 is a tangible storage medium, such as a disk medium.
[0085] The classification of each program illustrated in Figure 2 is merely a convenient classification established to clarify the functions constituting this disclosure. For example, it is not necessary to consider the text block generation program 23 and the electronic document visualization program 25 as independent programs. The same applies to the individual programs that make up the text block generation program 23 and the electronic document visualization program 25, respectively.
[0086] For example, the text block generation program 23 according to this embodiment consists of an electronic document acquisition program 231, an electronic document structuring program 232, an index acquisition program 233, and an index matching program 234. By executing this text block generation program 23 on the computer 1, the text block generation method (text block generation process) according to this embodiment is implemented.
[0087] Furthermore, the electronic document visualization program 25 according to this embodiment consists of a text block acquisition program 251, a clustering program 252, a feature acquisition program 253, a feature assignment program 254, and a map display program 255. By executing this electronic document visualization program 25 on the computer 1, the electronic document visualization method (electronic document visualization process) according to this embodiment is implemented.
[0088] In the program memory of SSD7b, each program constituting the document visualization program 21 and the document generation program 27 is started in response to commands input from the reception unit 13, etc. At that time, each program is loaded from SSD7b into RAM7a and executed by CPU3.
[0089] On the other hand, the data memory of SSD7b pre-stores document data 31, structured data 33, index data 35, prompt data 37, text block data 39, embedded data 41, aggregated data 42, low-dimensional data 43, clustering data 45, feature data 47, and specific morphological data 49, or stores them sequentially, temporarily, or continuously as each process progresses. Details of this data will be described later. Structured data 33 is an example of "first data" in this embodiment. Index data 35 is an example of "second data" in this embodiment.
[0090] Furthermore, the data illustrated in Figure 2 may be stored in the storage device 7c instead of the data memory of the SSD 7b, or it may be read from the storage device 7c.
[0091] In addition, the various data generated by executing the aforementioned programs are stored in the data memory of SSD7b or in RAM7a, which serves as main memory, as needed.
[0092] The following describes the document analysis method.
[0093] <2. Outline of Document Visualization Methods> Figure 3 is a flowchart illustrating the first half of the document visualization method. Figure 15 is a flowchart illustrating the second half of the document visualization method. Figures 4A to 4C are block diagrams illustrating the input and output in the document visualization method. Figure 5 is a diagram for explaining document data 31.
[0094] The document visualization method according to this embodiment uses a computer 1 to read multiple electronic documents Dc and visualize the relationships between the multiple electronic documents Dc. Each electronic document Dc includes document data 31 and supplementary data (metadata) 32.
[0095] In this embodiment, document data 31 is digital data containing multiple texts Ta. Each of the multiple texts Ta is digital data corresponding to a string of characters written in the electronic document Dc.
[0096] The document data 31 may further include digital data corresponding to figures in the electronic document Dc, in addition to multiple texts Ta corresponding to the body of the electronic document Dc.
[0097] Furthermore, as in this embodiment, the text Ta referred to herein may include digital data (hereinafter also referred to as table text Dt) that represents a string of characters described in a table within the electronic document Dc.
[0098] For example, in this embodiment, the table text Dt includes not only the item names in the table and the text Ta that shows the numerical values and words corresponding to each item, but also the text Ta that shows the table number and title, as shown in "Table 1" in Figure 5.
[0099] The supplementary data 32 is digital data indicating the name of the editing software used for text editing, the date and time of creation, etc. It is not mandatory for the electronic document Dc to include supplementary data 32. If an electronic document Dc without supplementary data is analyzed, in the following description, the electronic document Dc may be treated as identical to document data 31. Furthermore, if an electronic document Dc containing supplementary data 32 is analyzed, the CPU 3 may refer to this supplementary data 32 when determining the document structure of the electronic document Dc.
[0100] Furthermore, the file formats for electronic documents (DC) include PDF (Portable Document Format) and HTML (HyperText Markup Language). In the case of HTML format, at least a portion of the supplementary data 32 will be written in its source code.
[0101] As shown in Figures 3 to 15, the document visualization method is implemented by having computer 1 execute a text block generation process (steps S1 to S4) and an electronic document visualization process (steps S101 to S105) that is executed following the text block generation process.
[0102] As shown in Figure 3, the text block generation process (text block generation method) is carried out by having computer 1 execute the following in order: the electronic document acquisition process (step S1), the electronic document structuring process (step S2), the index acquisition process (step S3), and the index matching process (step S4).
[0103] As shown in Figure 15, the electronic document visualization process (electronic document visualization method) is carried out by having computer 1 execute a text block acquisition process (step S101), a clustering process (step S102), a feature acquisition process (step S103), a feature assignment process (step S104), and a map display process (step S105).
[0104] When CPU3 executes the electronic document acquisition program 231, etc., computer 1 configures a document visualization device that includes a text block generation device and an electronic document visualization device.
[0105] In other words, as shown in Figures 4A to 4C, the computer 1 functions as a text block generation device comprising an electronic document acquisition means 301 that executes an electronic document acquisition process, an electronic document structuring means 302 that executes an electronic document structuring process, an index acquisition means 303 that executes an index acquisition process, and an index matching means 304 that executes an index matching process.
[0106] Furthermore, computer 1 will function as an electronic document visualization device comprising a text block acquisition means 305 for executing a text block acquisition process, a clustering means 306 for executing a clustering process, a feature acquisition means 307 for executing a feature acquisition process, a feature assignment means 308 for executing a feature assignment process, and a map display means 309 for executing a map display process.
[0107] The following describes each process that constitutes the document visualization method in order. For details on data input and output in each process, and the relationship between the input / output data and each means, please refer to Figures 4A to 4C as appropriate.
[0108] <3. Details on how to generate text blocks> (3-1. Electronic Document Acquisition Process) Figure 6 is a flowchart illustrating the steps of the electronic document acquisition process. When the control process proceeds to step S1 in Figure 3, the CPU 3 executes each step sequentially from step S11 in Figure 6. Each step in Figure 6 is executed by the electronic document acquisition means 301, which is one of the functional elements configured by the CPU 3 (see Figure 4A).
[0109] First, in step S11, the CPU 3 obtains the electronic document Dc to be analyzed from the storage unit 7, such as the SSD 7b or storage device 7c, or from a database (not shown in the figure).
[0110] In the following step S12, the CPU3 determines the file format of the acquired electronic document Dc. The file format of the electronic document Dc may be determined automatically by the CPU3 based on the extension of the electronic document Dc, or the file format specified by the analyst via the reception unit 13 may be used for the determination by the CPU3.
[0111] In the subsequent step S13, the CPU3 stores the acquired electronic document Dc and its file format in the storage unit 7, and terminates the electronic document acquisition process. The CPU3 then proceeds through the control process from step S1 to step S2 in Figure 3, and starts the electronic document structuring process.
[0112] (3-2. Electronic Document Structuring Process) Figure 7 is a flowchart illustrating the steps of the electronic document structuring process. Figure 8 is a diagram illustrating the structured data 33. When the control process proceeds to step S2, the CPU 3 executes each step from step S21 in Figure 7. Each step in Figure 7 is executed by the electronic document structuring means 302, which is a functional element of the computer 1 (see Figure 4A).
[0113] In this embodiment, the electronic document structuring process is configured such that the CPU 3 generates and outputs structured data 33 based on the electronic document Dc acquired in the electronic document acquisition process.
[0114] Here, as shown in Figure 8, structured data 33 is a dataset that associates text attributes Tb that characterize at least one of the display position and display form of each of multiple texts Ta, the text Ta corresponding to that text attribute Tb, and the line number Tc of that text Ta on the electronic document Dc. Structured data 33 is, for example, text data in CSV or TSV format. Structured data 33 is not limited to CSV or TSV format; it can be any text data that uses separators that can distinguish between text attributes Tb, text Ta, and line number Tc. For example, in the specific example in Figure 11, separators such as [] and <> are used.
[0115] Furthermore, as mentioned above, the document data 31 of the electronic document Dc contains multiple texts Ta. For example, a break between texts Ta may be determined when each text Ta is separated by a line break.
[0116] Specifically, the text attribute Tb according to this embodiment is configured to characterize both the display position and display form of each of the multiple texts Ta.
[0117] In detail, as illustrated in Figure 8, the text attribute Tb according to this embodiment consists of both position information Tb1 that characterizes the display position of each text Ta and font information Tb2 that characterizes the display form of each text Ta.
[0118] More specifically, location information Tb1 includes the X and Y coordinates (X,Y) of each text Ta on each page of the electronic document Dc. Here, the X and Y coordinates may also be the X and Y coordinates of the first character of each text Ta, respectively. Location information Tb1 also includes the page number on which each text Ta is displayed. <page>It may include.
[0119] For more details, font information Tb2 refers to the font size of each text Ta on each page of the electronic document Dc. <size>And, font color<Cоlоr> And, whether or not it is displayed in bold (whether or not it is displayed in bold).<Bоld> And, whether or not italicized text is used (whether or not it is displayed in italics). <italics>It includes at least one of the following.
[0120] Furthermore, as illustrated in Figure 8, the text Ta corresponding to the text attribute Tb, as mentioned above, represents the content of the string "Text" indicated by each text Ta, separated by line breaks.
[0121] Furthermore, the line number Tc corresponding to each text Ta indicates the line number [Index] of that text Ta on each page of the electronic document Dc. This line number Tc may be reset for each page, or it may be counted so that it is a continuous number across all pages. The line number Tc is identification information for identifying text Ta on the electronic document Dc. The line number Tc, as identification information, is assigned to each text Ta on the electronic document Dc.
[0122] It is not mandatory to use the line number Tc as identification information. The identification information may be the Y coordinate of each text Ta. In other words, at least a portion of the text attribute Tb in this disclosure may also serve as identification information.
[0123] Returning to Figure 7, in step S21 of the same figure, CPU3 first determines whether the file format of the electronic document Dc is HTML format, based on the processing performed in step S12 of Figure 6.
[0124] If the file format is PDF (step S21: NO), CPU3 proceeds to step S23 of the control process. If the process proceeds to step S23, CPU3 generates structured data 33 based on at least one of the character codes embedded in the electronic document Dc and the OCR processing performed on the electronic document Dc.
[0125] In this case, if a character code is associated with each text Ta in the electronic document Dc, CPU3 determines each text Ta by extracting that character code. On the other hand, if no character code is embedded in the electronic document Dc and only glyph information exists, CPU3 estimates and obtains each text Ta based on OCR (Optical Character Recognition) processing of the electronic document Dc. CPU3 also obtains text attributes Tb and line numbers based on the OCR processing. CPU3 generates and outputs structured data 33 by associating each of the obtained pieces of information.
[0126] On the other hand, if the file format is HTML (step S12: YES), CPU3 proceeds the control process to step S22. If the process proceeds to step S22, CPU3 generates structured data 33 based on the source code of the electronic document Dc.
[0127] The source code of the electronic document Dc contains the character encoding settings of the electronic document Dc (information indicating which character encoding is set), each text Ta, and supplementary information such as tags and font settings set for each text Ta. Based on this information, CPU3 acquires various information that constitutes structured data 33. Of the information that constitutes structured data 33, information that is not thought to be included in the source code, such as location information Tb1, is acquired based on OCR processing, similar to PDF electronic documents Dc. CPU3 generates and outputs structured data 33 by associating each of the acquired pieces of information.
[0128] In step S24, which follows steps S22 and S23 respectively, the CPU 3 stores the output structured data 33 in the storage unit 7.
[0129] Once step S24 is complete, CPU3 terminates the electronic document structuring process. CPU3 then proceeds through the control process from step S2 to step S3 in Figure 3 and starts the index acquisition process.
[0130] (3-3. Index Acquisition Process) Figure 9 is a flowchart illustrating the steps of the index acquisition process. Figure 10A is an example of a system prompt 203 input to the first generation AI 201. Figure 10B is a diagram illustrating each input field in the system prompt 203. Figure 11 is an example of a user prompt 204 input to the first generation AI 201. Figure 12 is an example of a display screen 205 showing the output results from the first generation AI 201.
[0131] Figure 10A can also be considered a concrete example of prompt data 37 input to the first generation AI 201 as a system prompt 203. Figure 11 can also be considered a concrete example of structured data 33 input to the first generation AI 201 as a user prompt 204. Figure 12 can also be considered a concrete example of index data 35 displayed on the display 9.
[0132] Figure 11 is an excerpt from the paper "Toshiki Kondo, Takeo Kodaira, Hiromasa Kenmochi, "Development of Design Support Technology for Efficient Discovery of Structural Knowledge in Automobile Bodies (1) Proposal of Nonlinear Sparse Modeling Using Evolutionary Factor Extraction and Factor Selection Probability", Proceedings of the 31st Conference of the Design Engineering and Systems Division (2021), 2202", published by the inventors of the present invention.
[0133] When the control process proceeds to step S3, the CPU 3 executes each step from step S31 in Figure 9. Each step in Figure 9 is executed by the index acquisition means 303, which is one of the functional elements configured by the CPU 3 (see Figure 4A).
[0134] -Basic Concepts of the Index Acquisition Process- The index acquisition process according to this embodiment includes the CPU 3 inputting a prompt (system prompt 203) to the first generation AI 201 (prompt input processing in step S33).
[0135] Here, the system prompt 203 in this embodiment is an instruction given to the first generating AI 201 in the background. This system prompt 203 includes at least three instructions.
[0136] The input of the system prompt 203 to the first generated AI 201 may be performed by the CPU 3 when the reception unit 13 receives a predetermined operation by the analyst (for example, a specific key input or a specific mouse operation), triggered by that reception.
[0137] Furthermore, the content of the system prompt 203 may be pre-configured, for example, as text data in Markdown format (prompt data 37). In that case, the content of the system prompt 203 will be appropriately set by the CPU 3 to reflect the content of prompt data 37.
[0138] The first of the three instructions (the first instruction) is, as shown in Figure 10A, that "the first generating AI 201 accepts structured data 33 as input." The structured data 33 is input, for example, to the command prompt of the first generating AI 201 as a so-called user prompt 204.
[0139] Note that the parenthetical phrase "The first generated AI 201..." indicating the first instruction merely conceptually represents the content of the instruction, and it is not necessary to configure the system prompt 203 exactly as it is written. Similarly, for the second and third instructions, which will be described later, it is not necessary to configure the system prompt 203 exactly as it is written in each instruction.
[0140] For example, in the example shown in Figure 10A, as shown by the leader line C1 in the figure, the first instruction is "Information from the document will be provided on N It says "pages.". As you can see, in the example in Figure 10A, the phrase "Information from the document" is used instead of the phrase "structured data".
[0141] For more details, in the example shown in Figure 10A, you can specify the number of pages to be analyzed from the electronic document (document) Dc by entering any natural number in the input field labeled "N" (see also Figure 10B).
[0142] Furthermore, the first instruction may include information indicating the type of electronic document Dc to be entered, such as "information from a research paper will be entered" or "information from a newsletter will be entered."
[0143] The second of the three instructions (the second instruction) is the "input format of the structured data 33," as shown by the leader line C2 in Figure 10A. The "input format of the structured data 33" consists of descriptions of each data that makes up the structured data 33.
[0144] In this embodiment, the description of each data item is configured to reflect the CSV or TSV format of the structured data 33 and to describe the data content written in that format.
[0145] In the example in Figure 10A, the input format for structured data 33 in electronic document Dc can be specified by entering any input format in the input field labeled "F" (italicized and underlined) (see also Figure 10B).
[0146] As an example, in the input field "F" mentioned above, the input format for structured data 33 is entered in the following order, similar to the example in Figure 8: [Index] indicating the input format for line number Tc, (x coordinate, y coordinate) indicating the input format for position information Tb1, <Maximum font size in the document> indicating the input format for font information Tb2, and "Text" indicating the input format for text Ta.
[0147] In this case, as shown in the specific example in Figure 11, the structured data 33 will be entered according to the input format. For example, the first row in Figure 11 is: row number =
[12] , (x coordinate, y coordinate) = (400, 109), font size = <14> The data will be entered in the following order: text = "We have developed the interactive design support technology in order to efficiently obtain finding of lightweight-car-"
[0148] Thus, in this embodiment, the input format of the structured data 33 includes an input format for the text attribute Tb, an input format for the line number Tc as identification information, and an input format for the text Ta (see also Figure 10B).
[0149] In other words, as illustrated in the "[Index]" section of Figure 10A mentioned above, the "input format of structured data 33" includes the fact that "structured data 33 has the line number Tc of the text Ta in the electronic document Dc as identification information."
[0150] Furthermore, as illustrated by "(x coordinate, y coordinate)" and "<Maximum font size in the document>" in Figure 10A mentioned above, the "input format of structured data 33" includes the fact that "structured data 33 has, as text attribute Tb, position information Tb1 and font information Tb2 for each of the multiple texts Ta in the electronic document Dc."
[0151] Furthermore, as illustrated in the "text" section of Figure 10A mentioned above, the "input format of structured data 33" includes the requirement that "structured data 33 has text Ta corresponding to text attribute Tb." In the example of Figure 10A, "text Ta corresponding to text attribute Tb" refers to text Ta to which the same line number Tc as text attribute Tb is assigned.
[0152] The third of the three instructions (the third instruction) is, as shown by the leader lines C3 and C4 in Figure 10A, that "the first generating AI 201 outputs, for document data 31 which is divided into multiple text blocks Bl by referring to structured data 33, as index data 35 which represents the breakdown of text Ta that will belong to each of the multiple text blocks Bl, using identification information (e.g., line number Tc) associated with the text Ta."
[0153] For details, in the example in Figure 10A, as shown by the leader line C3 in the figure, the first part of the third instruction is "Extract blocks of text by B It says ".". This instruction tells the first generation AI201 to divide the document data 31, which was input as structured data 33, into multiple text blocks Bl.
[0154] In this context, a text block Bl is a "block of documents" composed of one or more texts Ta. A "block of documents" here may be, for example, a "group of documents" composed of one or more texts Ta according to a predetermined rule. Furthermore, a "group" here may be a "group" of sections or chapters that constitute an electronic document Dc.
[0155] To approximate "grouping" at the section or chapter level, in the example in Figure 10A, you can specify the division criteria for document data 31 by entering an arbitrary division criterion in the input field labeled "B" (italicized and underlined). The division criterion consists of one or more of the following: "chapter," "section," "part," and "paragraph" (see also Figure 10B).
[0156] Specifically, in the example shown in Figure 10A, by entering "Section" or "Chapter" in the input field "B", the system prompt 203 can be configured to "divide the document data 31 into sections or chapters". By using such a system prompt 203, multiple text blocks Bl will be constructed to reflect the section or chapter structure of the electronic document Dc.
[0157] Furthermore, in the example shown in Figure 10A, as indicated by the leader line C4 in the figure, the latter part of the third instruction is "Output extracted text blocks by their Identifiable information I. The instructions state: "This instruction causes the first generation AI201 to output index data 35 that represents the breakdown of text Ta belonging to each of the multiple text blocks Bl, using identification information (e.g., line number Tc) associated with the text Ta."
[0158] In this case, the type of identification information to be output as index data 35 is specified by entering the "row number," "column number," "Y coordinate," etc., in the input field labeled "I" which is italicized and underlined (see also Figure 10B). The specific examples shown in Figure 12 and Figure 14A described later show the case where the row number Tc is entered in the input field labeled "I."
[0159] Furthermore, as illustrated by the leader lines C11 and C12 in Figure 12, which will be described later, the index data 35 is not the body text of the text Ta that constitutes each text block Bl, but rather a dataset that represents the breakdown of each text block Bl by index (line number Tc).
[0160] Furthermore, index data 35 indicates the range of text Ta that constitutes each text block Bl. The range of text Ta is represented by the line number Tc of the text Ta that constitutes each text block Bl.
[0161] In this embodiment, the index data 35 does not include the text body Ta, but only the line number Tc. Excluding the text body Ta from the index data 35 is advantageous in reducing the data size of the index data 35.
[0162] In addition, in the example shown in Figure 10A, as illustrated by the leader line C5 in the figure, another instruction (the fourth instruction) related to the output of the index data 35 is described as "Simplify and output the attributes that enable continuous expressions using C." This instruction specifies to the first generated AI2021 the notation for consecutive row numbers Tc when outputting the index data 35.
[0163] In this case, the notation for consecutive line numbers Tc is specified by entering "-", "~", "_", etc. in the italicized and underlined "C" input field (see also Figure 10B). The specific examples shown in Figures 12 and 14A show the case where "-" is entered in the "C" input field. The notation for adjacent Y coordinates can also be specified when using identification information other than line number Tc. The notation for consecutive or adjacent identification information in general can be entered in the "C" input field.
[0164] In addition, in the example shown in Figure 10A, as indicated by the leader line C6 in the figure, there is an additional instruction (the fifth instruction) that reads, "If it's possible to identify M, output those as well." This instruction causes the first generation AI 201 to output attribute information that can be read from the electronic document Dc, along with the index data 35 that shows the classification results of multiple text blocks Bl.
[0165] In this case, the type of attribute information to be read from the electronic document Dc is specified by entering the type of attribute information in the input field marked "M" in italics and underlined (see also Figure 10B). The attribute information to be read from the electronic document Dc consists of one or more of the following from the electronic document Dc: "Chapter Title," "Section Title," "References," "Reference Information," and "Author Information (one or more of the author's name, affiliation, and email address)."
[0166] Specifically, in the example shown in Figure 10A, by entering "Section Title" or "Chapter Title" in the "M" input field, the system prompt 203 can include the instruction to "estimate the section title or chapter title for each of the multiple text blocks Bl".
[0167] In this case, CPU3 will have the first generating AI201 estimate the section title or chapter title, and output the estimation result together with the index data 35, or as data independent of the index data 35.
[0168] In addition, in the example in Figure 10A, as shown by the leader line C7 in the figure, the instruction "Exclude E." is written as an additional instruction (sixth instruction). This instruction specifies to the first generation AI201 information to be excluded from the text block Bl and index data 35.
[0169] In this case, by entering the type of information in the input field labeled "E" (which is italicized and underlined), the information entered in the input field is excluded from the output of the index data 35 (see also Figure 10B). The types of information to be excluded consist, for example, of the "page number," "cited references," "references," "author information (one or more of the author's name, affiliation, and email address)," "title of the paper (electronic document Dc)," "title of the figure," and one or more of the specific supplementary data 32 of the electronic document Dc.
[0170] For example, consider the case where "the paper title" is entered in input field "E". In this case, CPU3 instructs the first generating AI201 to remove "the paper title" from the text block Bl. Generating AI201 will then remove the line number Tc corresponding to the paper title from the index data 35.
[0171] In addition, in the example shown in Figure 10A, as indicated by the leader line C8 in the figure, the following additional instruction (the seventh instruction) is given: "When the information contains S, please output it without breaking the structure, using strings such as separators." This instruction instructs the first generation AI201 to output information such as table text Dt, either as information embedded within the index data 35 or as information independent of the index data 35.
[0172] In this case, by entering the type of information to be output independently in the italicized and underlined "S" input field, the first generation AI 201 can output information belonging to the type entered in the input field (see also Figure 10B). The type of information to be input consists of, for example, one or more of the "table data (table text Dt)", "reference information", and "author information" in the electronic document Dc. The information that can be entered in the "S" input field is information that includes multiple texts Ta and can generate a text block Bl, such as table text Dt.
[0173] The CPU 3 then instructs the first generation AI 201 to exclude the information entered in the "S" input field from the output using the identification information (row number Tc) in the index data 35. The CPU 3 also instructs the generation AI 201 to output the information entered in the "S" input field in an output format different from the row number Tc.
[0174] For example, consider the case where "table text Dt" is entered in the input field "E". In this case, the line number Tc of the text Ta that makes up table text Dt will be excluded from the index data 35. In other words, the system prompt 203 in this case includes the instruction to "exclude the text Ta or text block Bl corresponding to table text Dt from the output using line number Tc in the index data 35".
[0175] Furthermore, as illustrated by the leader line C13 in Figure 12, which will be described later, the system prompt 203 in this case includes the instruction to "include in the index data 35 text Ta, which represents the table text Dt in Markdown format, instead of using the line number Tc, for the text block Bl corresponding to the table text Dt."
[0176] Furthermore, the index acquisition process according to this embodiment includes the CPU 3 inputting structured data 33 to the first generating AI 201, thereby acquiring index data 35 via the first generating AI 201 (the index acquisition process in steps S34 and S35).
[0177] By inputting the aforementioned system prompt 203 into the first generation AI 201, the first generation AI 201 is ready to take structured data 33 as input and output index data 35.
[0178] Therefore, by inputting structured data 33 in accordance with the input format of the system prompt 203 illustrated in Figure 11, the first generation AI 201 will output index data 35 in accordance with the output format of the system prompt 203, as illustrated in Figure 12.
[0179] As illustrated by leader lines C11 and C12 in Figure 12, the index data 35 includes information indicating each text block Bl. Furthermore, as illustrated by leader line C13 in the same figure, the index data 35 also includes information indicating table text Dt, written in Markdown format.
[0180] For example, the leader line C11 in Figure 12 is attached to index data 35, which indicates that the first text block Bl is composed of text Ta from line 12 to line 21 of the electronic document Dc.
[0181] On the other hand, the leader line C12 in Figure 12 is attached to index data 35, which indicates that the second text block Bl is composed of text Ta from lines 25 to 39 of the electronic document Dc and text Ta from lines 47 to 55.
[0182] -Specific example of the index acquisition process- First, in step S31 of Figure 9, the CPU 3 reads the structured data 33 generated by the electronic document structuring process.
[0183] In the following step S32, the CPU3 determines whether the document structure of electronic document Dc is known. If the determination is NO, the CPU3 proceeds to step S33. On the other hand, if the determination in step S32 is YES, the CPU3 proceeds to step S36.
[0184] For example, the CPU 3 according to this embodiment determines whether the electronic document Dc is a patent document, and if it is determined that the electronic document Dc is a patent document, it determines that the document structure is known. This determination may be made automatically by the CPU 3, or it may be made based on the input content of the analyst via the reception unit 13.
[0185] More generally, this can be applied not only to patent documents but also to any publications, academic journals, commercial magazines, papers (including academic journals and collected papers), etc. In that case, the terms "patent document" and "publication name" in the following explanation should be replaced with terms such as "academic journal" and "journal name," respectively.
[0186] To illustrate further, in the former configuration example (when the determination is made automatically), the determination in step S32 may be made based on whether or not the structured data 33 contains information indicating that the electronic document Dc is a patent document. "Information indicating that it is a patent document" means that the structured data 33 contains a string of characters indicating the publication name of the patent document.
[0187] In this case, the memory unit 7 has a string of characters indicating the publication name of the patent document stored in advance, and the CPU 3 compares the stored content with the text Ta included in the structured data 33 to determine whether or not the electronic document Dc is a patent document.
[0188] For example, if the text Ta constituting the structured data 33 contains text Ta indicating the publication name of a patent document, the CPU 3 will determine that "the electronic document Dc is a patent document." To prevent confusion with cited documents etc. described in the patent document, the determination may be made by combining the text Ta indicating the publication name and the text attribute Tb of that text Ta, or by combining the line number Tc of that text Ta with, or by using the text attribute Tb instead of, or in addition to, the text attribute Tb.
[0189] First, if the document structure of the electronic document Dc is unknown (step S32: NO), the CPU 3 inputs structured data 33 to the first generation AI 201, thereby obtaining index data 35 via the first generation AI 201. Specifically, the CPU 3 sequentially executes the prompt input process (step S33) and the index acquisition process (steps S34 and S35) described above.
[0190] In other words, in step S33, the CPU 3 inputs the system prompt 203, as exemplified in Figure 10A, etc., to the first generating AI 201 of the server machine 101. This prepares the first generating AI 201. Note that the processing in step S33 may be performed in advance by the computer 1 prior to the electronic document structuring process S2, the electronic document acquisition process S1, etc.
[0191] In the following step S34, the CPU 3 inputs the structured data 33, as illustrated in Figures 8 and 11, to the first generation AI 201. The first generation AI 201 generates the index data 35, as illustrated in Figure 12, according to the instructions provided in the system prompt 203.
[0192] Subsequently, in the following step S35, the CPU 3 obtains the index data 35 generated by the first generation AI 201 via the communication unit 15.
[0193] On the other hand, if the document structure of the electronic document Dc is known (step S32: YES), the CPU 3 inputs structured data 33 into a rule-based estimation model 206, as illustrated in steps S36 and S37 described later, and obtains index data 35 through the estimation model 206. The estimation model 206 is a pre-built rule-based model (illustrated only in Figure 4B). The estimation model 206 is implemented, for example, on the server machine 101. This estimation model 206 takes structured data 33 as input and outputs index data 35.
[0194] Furthermore, the estimated model 206 is not limited to rule-based models. The estimated model 206 may, for example, be a deep learning model pre-trained on a large number of electronic documents Dc.
[0195] When using a rule-based model, the estimation model 206 may determine the boundaries of text blocks Bl based, for example, on changes in the font size of text Ta appearing in the electronic document Dc (for example, the relationship between the font size of each text Ta and a predetermined threshold), or it may determine the boundaries of text blocks Bl based on whether or not the text Ta appearing in the electronic document Dc corresponds to a pre-set phrase.
[0196] Let's specifically explain the case where the latter structure (judgment based on pre-defined wording) is adopted. For example, in the case of patent documents in the United States, the titles of each section are clearly defined, such as "ABSTRACT," "BACKGROUND OF THE INBENTION," "SUMMORY OF THE DESCRIPTION," and "DETAILED DESCRIPTION OF THE DESCRIPTION."
[0197] In this case, for example, the rules that the estimation model 206 refers to can be pre-configured so that the text Ta representing "BACKGROUND OF THE INBENTION" to the text Ta immediately preceding "SUMMORY OF THE DESCRIPTION" is treated as a single text block Bl. By pre-configuring it in this way, the estimation model 206 can generate index data 35 for generating text blocks Bl divided into sections, based on the input structured data 33 and the pre-configured rules.
[0198] On the other hand, in the case of patent publication and patent registration in Japan, the same processing can be applied to each section, such as "Title of Invention," "Technical Field," "Background Art," "Summary of Invention," "Modes for Carrying Out the Invention," and "Claims."
[0199] In other words, in step S36, the CPU 3 inputs structured data 33, as exemplified in Figure 11, etc., into the estimation model 206. The estimation model 206 generates index data 35, as exemplified in Figure 12, etc., according to pre-set rules.
[0200] Therefore, in the following step S37, the CPU 3 obtains the index data 35 generated by the estimation model 206 via the communication unit 15.
[0201] Subsequently, in step S38, which follows steps S35 and S37 respectively, the CPU 3 stores the acquired index data 35 in the storage unit 7.
[0202] Once step S38 is complete, CPU3 terminates the index acquisition process. CPU3 then proceeds through the control process from step S3 to step S4 in Figure 3 and starts the index matching process.
[0203] (3-4. Index matching process) Figure 13 is a flowchart illustrating the steps of the index matching process. Figure 14A is a diagram illustrating the output of the text block Bl. When the control process proceeds to step S4, the CPU 3 executes each step from step S41 in Figure 13. Each step in Figure 13 is executed by the index matching means 304, which is one of the functional elements configured by the CPU 3 (see Figure 4A).
[0204] The index matching process according to this embodiment is configured such that the CPU 3 refers to the index data 35 to select text Ta from document data 31 or structured data 33, and combines the selected text Ta in an order based on the line number Tc as identification information, thereby creating and outputting at least one of a plurality of text blocks Bl as a document.
[0205] As an example, in this embodiment, the CPU 3 performs the document creation and outputs the text by referring to the index data 35 and combining the selected text Ta in the order of the line numbers Tc used as identification information. If Y coordinates are used instead of line numbers Tc as identification information, the CPU 3 will combine the text Ta associated with each Y coordinate in the order of the Y coordinates.
[0206] Specifically, in step S41 of Figure 13, the CPU 3 reads structured data 33 and index data 35 from the storage unit 7.
[0207] In the following step S42, the CPU 3 obtains the line number Tc contained in the index data 35 for each text block Bl, and selects the text Ta corresponding to the obtained line number Tc from the structured data 33.
[0208] In the following step S43, the CPU3 combines the text Ta selected in step S42 in line number Tc order, thereby outputting a text block Bl using the text Ta instead of line numbers Tc.
[0209] This makes it possible to generate and output a text block Bl, which is composed of one or more texts Ta, as illustrated in Figure 14A. The CPU 3 generates text block data 39 that represents the output text block Bl and stores it in the storage unit 7. This text block data 39 is associated with an ID for identifying the corresponding electronic document Dc.
[0210] Furthermore, as illustrated by the leader line C8 in Figure 10A, by configuring the system to exclude "information from academic societies," the line number
[0364] corresponding to the information from academic societies (specifically, the text "Japan Society of Mechanical Engineers" Ta) is excluded from the index data 35 in Figure 14A. As a result, "information from academic societies" can be excluded from each text block Bl, as shown in the lower part of Figure 14A.
[0211] Subsequently, in the following step S44, the CPU 3 stores the output text block Bl in the storage unit 7. Once this process is complete, the CPU 3 terminates the index matching process. The computer 1 terminates the text block generation method configured in steps S1 to S4 and starts a new electronic document visualization method according to the pre-configurations. At this time, the CPU 3 advances the control process from step S4 in Figure 3 to step S101 in Figure 15.
[0212] When performing the electronic document visualization method, CPU3 applies the text block generation method shown in Figure 3 to each of the multiple electronic documents Dc that are to be visualized. As a result, text block data 39 is generated for each of the multiple electronic documents Dc.
[0213] Figure 14B illustrates the output result of text block Bl for the document data 31 in Figure 5. By entering "Chapter" in the input field "B" in Figure 10A, the generating AI 201 classifies the document data 31 by chapter and generates text block Bl for each chapter.
[0214] In the example shown in Figure 14B, the first generated AI 201 will classify the document data 31 into a first text block Bl1 indicating the title of the document data 31, a second text block Bl2 indicating the "Abstract" of the document data 31, a third text block Bl3 indicating the "1. Introduction" of the document data 31, and a fourth text block Bl4 indicating the "2. Main Body" of the document data 31.
[0215] <4. Details of the Electronic Document Visualization Method> (4-1. Text block acquisition process) When the control process proceeds to step S101 in Figure 15, the CPU 3 executes the text block acquisition process. This process is performed by the text block acquisition means 305, which is one of the functional elements configured by the CPU 3 (see Figure 4C).
[0216] This text block acquisition process includes the CPU3 reading the plurality of electronic documents as a collection of multiple text blocks Bl, each containing one or more texts.
[0217] Specifically, CPU3 reads text block data 39 corresponding to each of the multiple electronic documents Dc that are to be visualized.
[0218] As explained with reference to Figure 15, each text block data 39 is a collection of text data in which multiple text blocks Bl are arranged in line number and page number order for each electronic document Dc.
[0219] For example, if the number of electronic documents Dc to be visualized is N (where N is a natural number), then CPU3 reads N different text block data 39.
[0220] Once all the text block data 39 has been read, the CPU 3 terminates the processing related to step S101 in Figure 15. Subsequently, the CPU 3 proceeds with the control process, either executing steps S102 and S103 in parallel or executing steps S102 and S103 sequentially.
[0221] (4-2. Clustering Process) Figure 16 is a flowchart illustrating the steps of the clustering process. When the control process proceeds to step S102, the CPU 3 executes each step from step S121 in Figure 16. Each step in Figure 16 is executed by the clustering means 306, which is a functional element of the computer 1 (see Figure 4C).
[0222] The clustering process according to this embodiment includes the CPU 3 quantifying multiple electronic documents Dc into embedded data 41 via a language model 207, aggregating the embedded data 41 obtained for each text block Bl for each electronic document Dc, and then clustering them into multiple clusters.
[0223] Here, the language model 207 is an AI that takes each text block Bl corresponding to each electronic document Dc as input and outputs embedded data 41 as multidimensional data. The AI is input with text tokenized by a tokenizer.
[0224] More specifically, language model 207 is an embedding model that converts variable-length text into fixed-length vector data (multidimensional data). Language model 207 receives each text block Bl, which constitutes the text block data 39 corresponding to each electronic document Dc, as input, as variable-length text. Based on this, language model 207 generates embedded data 41 for each text block Bl for each electronic document Dc.
[0225] Returning to Figure 16, first, in step S121 of the same figure, the CPU 3 inputs text block data 39 corresponding to the electronic document Dc into the language model 207, and through the language model 207, acquires embedded data 41 as vector data for each text block Bl. The acquired embedded data 41 is stored in the storage unit 7 for each electronic document Dc. This embedded data 41 is acquired for each electronic document Dc and for each text block Bl.
[0226] In the subsequent step S122, the CPU 3 aggregates the embedded data 41 acquired in step S122 for each electronic document Dc. The aggregated embedded data 41 is stored in the storage unit 7 as aggregated data 42. The embedded data 41 is vector data assigned to each text block Bl. The aggregated data 42 is vector data assigned to the entire electronic document Dc (document data 31).
[0227] For example, if the number of electronic documents Dc to be visualized is N (where N is a natural number), the CPU 3 generates N types of aggregated data 42 and stores them in the storage unit 7.
[0228] Specifically, the CPU 3 according to this embodiment is configured to generate aggregated data 42 for each document data 31 by averaging each component of each embedded data 41, or by obtaining the maximum and minimum values of each component of each embedded data 41.
[0229] In the subsequent step S123, the CPU3 outputs, for each electronic document Dc, at least the aggregated data 42 from among the multiple embedded data 41 and aggregated data 42, in a 2D or 3D vector data format. Any method can be used for this dimensionality reduction, such as Principal Component Analysis, t-SNE (t-distributed Stochastic Neighbor Embedding), or UMAP (Uniform Manifold Approximation and Projection).
[0230] This reduced-dimensional vector data (hereinafter also referred to as "low-dimensional data 43") is output for all electronic documents Dc that are targeted for visualization.
[0231] For example, if the number of electronic documents Dc to be visualized is N (where N is a natural number), the CPU 3 generates a data set (low-dimensional data set) consisting of N low-dimensional data 43 and stores it in the storage unit 7.
[0232] Furthermore, in the specific example described later, CPU3 converts both the multiple embedded data 41 and aggregated data 42 into two-dimensional low-dimensional data 43 using the t-SNE method. Each component of the low-dimensional data 43 can be considered as the coordinates of a plot Pr on a two-dimensional plane (see Figure 21).
[0233] In this case, the CPU 3 will generate and output, for each electronic document Dc, low-dimensional data 43 based on aggregated data 42 and low-dimensional data 43 based on each embedded data 41.
[0234] For example, if the number of electronic documents Dc to be visualized is N (where N is a natural number), and the number of text blocks Bl in each electronic document Dc is M (where M is a natural number), then the CPU 3 generates a low-dimensional data set consisting of multiple (N) low-dimensional data 43 based on each aggregated data 42, and a low-dimensional data set consisting of multiple (N × M) low-dimensional data 43 based on each embedded data 41, and stores them in the storage unit 7.
[0235] More generally, the number of text blocks Bl may differ for each electronic document Dc. In that case, the total number of data in the low-dimensional data set can be obtained by adding up the number of text blocks Bl for each of the N electronic documents Dc.
[0236] In the subsequent step S124, the CPU3 clusters the low-dimensional data set generated in step S123 into multiple clusters. Each cluster will contain one or more low-dimensional data sets 43. The number of clusters may be specified by the analyst via the reception unit 13 connected to the CPU3, or it may be set in advance by reading a CSV file or the like.
[0237] CPU3 may cluster the low-dimensional data set using hard clustering or soft clustering. In the former case, CPU3 may perform hard clustering using the k-means method or the Ward method. In the latter case, CPU3 may perform soft clustering using the c-means method.
[0238] In the specific example described later, CPU3 is configured to cluster both the low-dimensional data set based on aggregated data 42 and the low-dimensional data set based on each embedded data 41 using the Ward method.
[0239] The former case corresponds to CPU3 clustering the aggregated data 42 as multidimensional data into multiple clusters (hereinafter also referred to as "first cluster Cl1") for each electronic document Dc through its dimensionality reduction. The latter case corresponds to CPU3 clustering each embedded data 41 as multidimensional data into multiple clusters (hereinafter also referred to as "second cluster Cl2") for each text block Bl through its dimensionality reduction.
[0240] In the following step S125, the CPU 3 generates clustering data 45, which is data representing the cluster to which each electronic document Dc belongs among multiple clusters, and stores it in the storage unit 7.
[0241] When clustering is performed on a text block Bl basis, that is, on a low-dimensional data set based on each embedded data 41, the CPU 3 may convert information indicating the cluster to which each text block Bl belongs into data and include it in the clustering data 45.
[0242] Once the generation and storage of the clustering data 45 is complete, the CPU 3 terminates the processing related to step S102 in Figure 15. The CPU 3 then proceeds with the control process to start the feature acquisition process related to step S103 in the same figure, either in parallel with or before / after step S102.
[0243] (4-3. Feature Acquisition Process) Figure 17 is a flowchart illustrating the steps of the feature acquisition process. Figure 18 is a diagram illustrating the calculations related to features. When the feature acquisition process starts, CPU 3 executes each step from step S131 in Figure 17. Each step in Figure 17 is executed by the feature acquisition means 307, which is a functional element of computer 1 (see Figure 4C).
[0244] The feature acquisition process according to this embodiment includes the CPU 3 acquiring feature quantities for each of the multiple electronic documents Dc, which are statistically quantified features of each morpheme (word) contained in each electronic document Dc, using multiple text blocks Bl corresponding to each electronic document Dc as input.
[0245] More specifically, CPU3 obtains features based on morphological analysis using multiple text blocks Bl as input. CPU3 obtains indices based on the frequency or count of occurrence of each morpheme as features.
[0246] First, in step S131 of Figure 17, the CPU 3 divides each text block Bl that makes up the text block data 39 into multiple morphemes for each electronic document Dc. By performing this division into multiple morphemes, each text block data 39 is subjected to preprocessing.
[0247] In the subsequent step S132, the CPU3 applies filtering to each morpheme that was divided in step S131. For example, as shown in "1." in Figure 18, the CPU3 may select the type of word (morpheme) to be used for feature calculation from among particles, verbs, nouns, adjectives, etc. The selection of the morpheme type may be performed by the CPU3 based on pre-configured settings.
[0248] In the subsequent step S133, the CPU3 calculates an index for each morpheme belonging to the type selected in step S132, which increases or decreases according to the frequency or number of occurrences in the electronic document Dc containing that morpheme, and uses the calculation result as a feature.
[0249] Furthermore, as mentioned above regarding the clustering process, if the CPU3 clusters the multidimensional data into multiple clusters (second cluster Cl2) for each text block Bl, the aforementioned feature quantities may be calculated for each text block Bl that contains each morpheme.
[0250] For example, as shown in "2." in Figure 18, CPU3 quantifies the feature quantities associated with each morpheme using Bow (Bag of Word), which increases as the frequency (number of occurrences) of each morpheme increases, as mentioned above.
[0251] Furthermore, CPU3 weights morphemes that appear relatively frequently within a particular text to increase their feature size, and weights morphemes that are used across multiple texts to decrease their feature size. This weighting can be achieved using so-called TF-IDF (Term Frequency - Inverse Document Frequency). By using TF-IDF, the feature sizes of "the foregoing" and punctuation marks in the patent document can be reduced. The feature sizes after applying TF-IDF are examples of "indicators" in this embodiment. Hereafter, these will also be simply referred to as "feature sizes."
[0252] In the following step S134, the CPU 3 generates feature data 47 that shows the feature quantities obtained for each morpheme constituting each electronic document Dc, and stores it in the storage unit 7.
[0253] Once the generation and storage of the feature data 47 is complete, the CPU 3 terminates the processing related to step S103 in Figure 15. If the processing in step S102 has also been completed, the CPU 3 proceeds to step S104 of the control process.
[0254] (4-4. Feature Assignment Process) FIG. 19 is a flowchart illustrating the procedure of the feature amount assignment process. When the feature amount acquisition process is started, the CPU 3 executes each step from step S141 in FIG. 19. Each step in FIG. 19 is executed by the feature amount assignment means 308 among the functional elements in the computer 1 (see FIG. 4C).
[0255] The feature amount assignment process according to the present embodiment includes the CPU 3 associating feature amounts for each first cluster Cl1 based on the electronic documents Dc that respectively constitute the plurality of first clusters Cl1. Further, the feature amount assignment process according to the present embodiment includes the CPU 3 associating feature amounts for each second cluster Cl2 based on the text blocks Bl that respectively constitute the plurality of second clusters Cl2.
[0256] Through the clustering process, clustering data 45 indicating the breakdown of the electronic documents Dc that constitute each first cluster Cl1 and the breakdown of the text blocks Bl that constitute each second cluster Cl2 is output. Also, through the feature amount acquisition process, feature amount data 47 indicating the feature amounts of each morpheme that constitutes each electronic document Dc is output.
[0257] Therefore, by combining the clustering data 45 and the feature amount data 47, it becomes possible to quantify the feature amounts of the morphemes included in each cluster for each of the first cluster Cl1 and the second cluster Cl2.
[0258] Specifically, in step S141 of FIG. 19, the CPU 3 reads the clustering data 45 and the feature amount data 47 from the storage unit 7.
[0259] In the subsequent step S142, as shown in "3." of FIG. 18, the CPU 3 calculates the average value (average feature amount Na) of the feature amounts Ns for each cluster and for each morpheme. For the average feature amount, refer to the reference numeral Na in the same figure. In the subsequent step S143, as shown in "4." of FIG. 18, the CPU 3 averages each average feature amount Na calculated in step S142 over all clusters for each morpheme. As a result, the average value (overall average) of each morpheme over a plurality of clusters is calculated.
[0260] In the subsequent step S144, the CPU 3 calculates the standard deviation (overall standard deviation) of each average feature amount Na with respect to the overall average for each morpheme based on the average feature amount Na and the overall average.
[0261] In the subsequent step S145, as shown in "4." of FIG. 18, the CPU 3 compares, for each cluster and for each morpheme, the average feature amount Na with a comparison target value calculated based on the overall average and the overall standard deviation. Assuming that a pre-set threshold value is X, the comparison target value is calculated for each morpheme based on the following formula.
[0262] Comparison target value = overall average + X · overall standard deviation The CPU 3 extracts, for each first cluster Cl1, morphemes (specific morphemes) whose average feature amount exceeds the comparison target value. For the specific morphemes, refer to the morphemes corresponding to the numerical values shown in bold in "5." of FIG. 18.
[0263] In the subsequent step S146, the CPU 3 generates specific morpheme data 49 indicating the specific morphemes extracted for each cluster and stores it in the storage unit 7.
[0264] When the generation and storage of the specific morpheme data 49 are completed, the CPU 3 ends the process related to step S104 in FIG. 15 and advances the control process to step S105.
[0265] Note that these processes are common to both the first cluster Cl1 corresponding to the clustering for each electronic document Dc and the second cluster Cl2 corresponding to the clustering for each text block Bl.
[0266] (4-5. Map display process) Figure 20 is a flowchart illustrating the steps of the map display process. Figures 21 to 23 are scatter plots illustrating the display modes of the first visualization map M1 corresponding to each electronic document Dc. Figure 24 illustrates the display modes of the first visualization map M1 according to the feature quantity Ns (the feature quantity Ns is illustrated in Figure 18). Figure 25 illustrates an example of extraction from the first visualization map M1. The first visualization maps M1 illustrated in Figures 21 to 25 each show application examples to several hundred academic papers published over a five-year period from 2018 to 2022.
[0267] Figure 26 is a scatter plot illustrating the display modes of the second visualization map M2 corresponding to each text block Bl. Figure 27 is an example of the display modes of the second visualization map M2.
[0268] Note that the first visualization map M1 is an example of the "first map" in this embodiment. Similarly, the second visualization map M2 is an example of the "second map" in this embodiment.
[0269] When the map display process starts, the CPU 3 executes each step from step S151 in Figure 20. Each step in Figure 20 is performed by the map display means 309, which is a functional element of the computer 1 (see Figure 4C).
[0270] The map display process according to this embodiment includes the CPU 3 visualizing a plurality of first clusters Cl1 as a two-dimensional or three-dimensional map, and controlling the display mode of the map according to the value of the feature quantity Ns (particularly the high or low value of the feature quantity Ns) as illustrated in Figure 18.
[0271] In this embodiment, a two-dimensional first visualization map M1 is used as a two-dimensional or three-dimensional map containing information about multiple first clusters Cl1 (see Figures 23-24 in particular).
[0272] Specifically, in step 151 of Figure 20, the CPU 3 reads text block data 39 and low-dimensional data 43 for each of the multiple electronic documents Dc.
[0273] In the following step S152, the CPU3 reads clustering data 45 and specific morphological data 49 for the entirety of multiple electronic documents Dc.
[0274] In the following step S153, the CPU 3 visualizes the first visualization map M1. This visualization may be performed by displaying the first visualization map M1 on the display unit, the display 9. The CPU 3 then plots on the first visualization map M1 a set of low-dimensional data 43 (low-dimensional data set) corresponding to the aggregated data 42 of each of the multiple electronic documents Dc. Each plot Pr on the first visualization map M1 corresponds to the low-dimensional data 43 of each electronic document Dc.
[0275] Here, CPU3 changes the display mode of the plot Pr corresponding to each electronic document Dc (for example, at least one of the shape and color of each plot Pr) based on the bibliographic information of the electronic document Dc.
[0276] In the case of the first visualization map M1 in Figure 21, CPU3 is configured to make each plot Pr different in color based on the publication year of each electronic document Dc, as shown in the legend Ug located on the right side of the first visualization map M1.
[0277] In the subsequent step S154, the CPU 3 changes the display mode of the first visualization map M1 based on the operation input to the reception unit 13. The change in the display mode of the first visualization map M1 is performed based on interactive operation input to the computer 1.
[0278] For example, in the case of FIG. 21, when the mouse cursor is superimposed on each plot Pr on the first visualization map M1, the CPU 3 extracts bibliographic information (publication year, paper ID, title, and author information) of the electronic document Dc corresponding to the plot Pr from the text block Bl of the electronic document Dc. The CPU 3 may generate a first window W1 that describes part or all of the bibliographic information and display it in a pop-up manner. The same applies when using the rule-based estimation model 206.
[0279] Furthermore, the plurality of publication years listed in the above-mentioned legend Ug can be switched between light and dark through a click operation of the mouse 13b. As a result, as shown in FIG. 22, it is possible to switch the on / off of the display of the plot Pr corresponding to each of the plurality of electronic documents Dc on the first visualization map M1. In the case of the legend, the publication year in white characters corresponds to the display off, and the publication year in black characters corresponds to the display on.
[0280] Also, above the paper surface of the first visualization map M1, a first interface (InterFace: IF) 401 for switching the display of the boundary line Lb indicating each first cluster Cl1 on or off is arranged. The first IF 401 is a user interface that receives a user's operation.
[0281] For example, as shown in FIGS. 21 to 23, when the item "no cluster information" is selected in the first IF 401 through a mouse operation or the like, the CPU 3 does not display the boundary line on the first visualization map M1.
[0282] On the other hand, as shown in FIG. 23, when any one of "large", "medium", and "small" (in the legend, "medium") other than "no cluster information" is selected in the first IF 401, the CPU 3 superimposes and displays the boundary line Lb indicating each first cluster Cl1 on the first visualization map M1.
[0283] Here, among the selection items that make up the first IF401, the three items "Large," "Medium," and "Small" each represent the number of clusters in the first cluster Cl1 (the total number of clusters in the first cluster Cl1). As shown in Figure 23, when "Medium" is selected, the number of clusters is, for example, "20." When "Large" is selected, the number of clusters changes to a larger number than when "Medium" is selected, and when "Small" is selected, the number of clusters changes to a smaller number than when "Medium" is selected.
[0284] When the number of clusters is changed, CPU3 re-executes the clustering process (step S102) in Figure 15, specifically the processes shown in steps S124 and S125 in Figure 16, and all of the feature allocation process (step S104) corresponding to "3." onwards in Figure 18. The feature acquisition process (step S103) does not need to be executed again.
[0285] Once these processes are complete, CPU3 overlays the boundary line Lb corresponding to the changed number of clusters onto the first visualization map M1 on display 9.
[0286] Furthermore, in this embodiment, when the reception unit 13 receives a predetermined operation input, the CPU 3 superimposes and displays a second window W2 on the first visualization map M1 for each first cluster Cl1, which includes a first word cloud Cr1 representing the high or low levels of the index (feature). In this embodiment, the first word cloud Cr1 represents the frequency or number of occurrences of each morpheme related to information other than author information.
[0287] The term "predetermined operation input" as used herein includes an operation to select at least one of the boundary line Lb and plot Pr that constitute one of the multiple first clusters Cl1 (for example, a click operation on the boundary line Lb and / or plot Pr).
[0288] Furthermore, the phrase "high or low feature quantity Ns for each morpheme" includes "high or low values of other indicators that increase or decrease in accordance with feature quantity Ns." In the example diagram, the first word cloud Cr1 represents the high or low average feature quantity calculated for each first cluster Cl1, as described above.
[0289] For example, as shown in Figure 24, the first word cloud Cr1 according to this embodiment is a word cloud composed of frequently occurring words (topics) in the electronic document Dc. The first word cloud Cr1 according to this embodiment shows a set of specific morphemes extracted for each first cluster Cl1.
[0290] Furthermore, in this embodiment, when the reception unit 13 receives the aforementioned "predetermined operation input," the CPU 3 superimposes and displays a second word cloud Cr2, which represents the frequency or number of occurrences of author information for each first cluster Cl1, onto the first visualization map M1 together with the first word cloud Cr1.
[0291] For example, as shown in Figure 24, the second word cloud Cr2 according to this embodiment is placed on the same second window W2 as the first word cloud Cr1. This second word cloud Cr2 is a word cloud composed of words from the author information that indicate the author's affiliated institution.
[0292] Specifically, in the example shown in Figure 24, the first word cloud Cr1, the third word cloud Cr3, and the second word cloud Cr2 are arranged in order from left to right on the second window W2. Here, the third word cloud Cr3 is a word cloud composed of the keyword groups attached to each paper.
[0293] Furthermore, an input bar (second IF402) is located at the top of the first visualization map M1 and at the bottom of the first IF401 to extract specific plots Pr. By inputting information contained in the first window W1 into this second IF402, only plots Pr corresponding to that information can be displayed on the first visualization map M1. The second IF402 is a user interface that accepts user input.
[0294] Specifically, in the example shown in Figure 25, the word "XX Corporation," which indicates a specific affiliated organization, is entered in the second IF402. CPU3 searches for electronic documents Dc that contain that word in their bibliographic information, and displays only the low-dimensional data 43 corresponding to the search result, and consequently, only the plot Pr corresponding to that low-dimensional data 43.
[0295] Furthermore, the map display process according to this embodiment includes the CPU 3 visualizing a plurality of second clusters Cl2 as a two-dimensional or three-dimensional map, and controlling the display mode of the map according to the level of the feature quantity Ns.
[0296] In this embodiment, a two-dimensional second visualization map M2 is used as a two-dimensional or three-dimensional map containing information about multiple second clusters Cl2 (see Figures 26-27 in particular).
[0297] Switching from the first visualization map M1 to the second visualization map M2 may be performed by the CPU3, for example, based on an operation input on the first visualization map M1.
[0298] For example, as illustrated in one of the two Figures 26, the CPU 3 selects two or more electronic documents Dc based on the operation input via the reception unit 13. In the example shown, the CPU 3 receives an operation to select two or more electronic documents Dc from the first visualization map M1.
[0299] Then, as illustrated in the other part of Figure 26, CPU3 visualizes the second visualization map M2 for each text block Bl for two or more selected electronic documents.
[0300] Specifically, the CPU 3 displays a second visualization map M2 on the display unit 9, and plots a set of low-dimensional data 43 (low-dimensional data set) corresponding to the embedded data 41 of each of the multiple electronic documents Dc on the second visualization map M2. Each plot Pr2 on the second visualization map M2 corresponds to the low-dimensional data 43 of each text block Bl.
[0301] Furthermore, in Figure 26, the five plots Pr2 connected by the dashed line L1 each correspond to a text block Bl that constitutes a single electronic document Dc. Each plot Pr2 may be given a name equivalent to the title of the corresponding text block Bl, such as "Chapter 1," "Chapter 2," "Chapter 3," "Chapter 4," and "Chapter 5." In the case of rule-based estimation, the section titles included in the rules used for estimation can be used.
[0302] Furthermore, in Figure 26, the three plots Pr2 connected by the solid line Ll2 each correspond to a text block Bl that makes up a different electronic document Dc. Each plot Pr2 is labeled with a name equivalent to the title of its corresponding text block Bl, such as "Chapter 1," "Chapter 2," and "Chapter 3." In the example shown, the plot Pr2 corresponding to "Chapter 3" is obscured by the plot Pr2 corresponding to "Chapter 4," which is connected by the dashed line Ll1.
[0303] As shown in both figures of Figure 26, the two electronic documents Dc selected based on the first visualization map M1 are close together on that first visualization map M1. This suggests that the document data 31 constituting each electronic document Dc are similar to each other.
[0304] Therefore, as shown in Figure 26, by displaying plot Pr2 for each text block Bl, the analyst can gain insights such as, "The content of Chapter 1 of the two electronic documents Dc is similar to each other. The content of Chapter 3 of the two electronic documents Dc is relatively very different."
[0305] Furthermore, CPU3 can also display boundary lines Lb representing each second cluster Cl2 on the second visualization map M2. The display of boundary lines Lb can be switched on or off. Since the clustering targets differ from those of the first visualization map M1, the number of clusters in the second visualization map M2 and the shape of each boundary line Lb will differ from those in the first visualization map M1.
[0306] Furthermore, a legend Ug2 related to the second visualization map M2 is displayed on the right side of Figure 27. This legend Ug2 may, for example, display the paper ID of each electronic document Dc selected from the first visualization map M1.
[0307] Furthermore, in this embodiment, when the reception unit 13 receives a predetermined operation input, the CPU 3 superimposes and displays a second window W2 on the second visualization map M2 for each second cluster Cl2, which includes a first word cloud Cr1 representing the high or low levels of the index (feature). The configuration of the various information placed in the second window W2, such as the first word cloud Cr1, the second word cloud Cr2, and the third word cloud Cr3, is the same as the configuration of the second window W2 with respect to the first visualization map M1.
[0308] Here, the phrase "high or low feature quantity Ns for each morpheme" includes "high or low values of other indicators that increase or decrease in accordance with feature quantity Ns." In the example diagram, the first word cloud Cr1 represents the high or low average feature quantity calculated for each of the second clusters Cl2, as described above.
[0309] For example, as shown in Figure 28, the first word cloud Cr1 for the second cluster Cl2 is a word cloud composed of frequently occurring words (topics) in each text block Bl. In this embodiment, the first word cloud Cr1 represents a set of specific morphemes extracted for each of the second clusters Cl2.
[0310] Furthermore, Figure 29 shows an example of displaying the second visualization maps M101 and M102 side by side between Japanese patent registrations and patent publications. As mentioned above, by establishing rules according to the format of patent documents and dividing each electronic document Dc into multiple text blocks Bl based on those rules, it is possible to compare and examine the differences section by section.
[0311] In the example in Figure 29, the number of electronic documents Dc corresponding to patent registrations and patent publications is 6 each. Sections that are relatively long, such as "Modes for Carrying Out the Invention," are divided into multiple text blocks BL, so more plots than the number of electronic documents Dc can be observed. The "Claims" section is also divided into multiple text blocks BL.
[0312] It is not mandatory to make the first visualization map M1 and the second visualization map M2 scatter plots. Instead of plotting each morpheme Pr and Pr2, a directed graph structure (DAG structure) connecting the morphemes can be generated as exemplified in Figure 30, and the network corresponding to that structure can be designated as the "first visualization map" or the "second visualization map".
[0313] Figure 30 shows an example of a directed graph structure (DAG structure), visualizing a Bayesian network. The numerical values attached to each edge indicate the strength of the causal relationship between nodes. It can be seen that the patent registration is represented by a more complex network than the patent document. This reflects that the connections between morphemes became more complex due to the revisions.
[0314] <5. Significance of Document Visualization Methods> As described above, according to the embodiment described in relation to the text block acquisition process in Figure 15, it is possible to generate embedded data 41 for each text block Bl by reading multiple electronic documents Dc as a collection of text blocks Bl. Compared to a configuration in which embedded data 41 is generated for each electronic document Dc, the data size input to the language model 207 can be reduced.
[0315] Furthermore, as illustrated in Figures 4C, 16, 17, and 18, CPU3 performs quantification of the embedded data 41 for clustering via the language model 207, while quantifying the information (feature quantity Ns) that characterizes each first cluster Cl1 via statistical processing of morphemes. The conversion to embedded data 41 via the language model 207 is superior in expressive power and processing speed compared to the conversion via morphemes. On the other hand, quantification via morphemes makes the meaning of each morpheme, such as its number of occurrences or frequency, clearer.
[0316] Therefore, it is possible to leverage the advantages of both the language model 207-based method and the morphological method. This allows for both explicit identification of information characterizing each first cluster Cl1 of electronic document Dc and improved usability of the clustering process.
[0317] Furthermore, since each feature vector Ns is assigned to each morpheme, the value of the feature vector Ns for each morpheme remains unchanged even if the number of clusters in the first cluster Cl1 changes. Therefore, it becomes possible to smoothly and quickly associate the feature vector Ns with each changed first cluster Cl1. This also contributes to improved usability, such as increased processing speed.
[0318] Furthermore, as illustrated in Figures 26 and 27, in addition to clustering and visualization by electronic document Dc, clustering and visualization by text block Bl will also be performed. This makes it possible to visualize the similarities and differences between electronic documents Dc at the text block Bl level.
[0319] By enabling the combined use of a more macroscopic analysis based on the first visualization map M1 and a more microscopic analysis based on the second visualization map M2, it becomes possible to achieve more user-friendly visualizations.
[0320] Furthermore, as illustrated in Figure 26, based on the operation input to the reception unit 13, two or more of the multiple electronic documents Dc are selected, and the second visualization map M2 is visualized for each text block Bl of the two or more selected electronic documents Dc.
[0321] This enables an interactive interface with the analyst, resulting in a more user-friendly visualization. Furthermore, because the second visualization map M2 is a map for each text block Bl, it has a larger number of plots, which is likely to make it less visually appealing than the first visualization map M1. By configuring the system to visualize the second visualization map M2 for the desired electronic document Dc, a more user-friendly visualization can be achieved in terms of the visibility of the second visualization map M2.
[0322] Furthermore, as explained using Figure 18, by using the feature vector Ns after applying TF-IDF, it becomes possible to extract morphemes that characterize each electronic document Dc more appropriately. This enables the realization of a more user-friendly visualization.
[0323] Furthermore, as illustrated in the first window W1 of Figure 24, the analyst can visually identify the morphemes characterizing each first cluster Cl1 through their display in the first word cloud Cr1. This further improves the usability of the first visualization map M1.
[0324] Furthermore, as illustrated in the first window W1 of Figure 24, the analyst can visualize author information (e.g., affiliated institution) characterizing each first cluster Cl1 through the second word cloud Cr2. This further improves the usability of the first visualization map M1.
[0325] Traditionally, it was widely known that text extraction from electronic documents (DC) was performed on a page-by-page basis.
[0326] The inventors of this application attempted to output text in units of text blocks Bl, each consisting of multiple texts such as the introduction, body, and conclusion of an electronic document, instead of outputting on a page-by-page basis, by inputting text Ta extracted from an electronic document Dc into a generation AI (first generation AI 201). Outputting in blocks offers superior usability compared to conventional technologies.
[0327] However, simply inputting the text "Ta" is insufficient for ensuring the classification accuracy of the first generating AI 201 when classifying it into each text block "Bl". Furthermore, even if output in block units were possible, the data size output by the generating AI would be enormous, which would also be insufficient for ensuring processing speed.
[0328] Therefore, it is not enough to simply output text Ta for each text block Bl; it is necessary to achieve both high classification accuracy for text Ta in each text block Bl and high classification speed.
[0329] In contrast, according to the above embodiment, the first generating AI 201 performs division into multiple text blocks Bl by referring to text attribute Tb that characterizes at least one of the display position and display form of each text Ta, as illustrated in the structured data 33 in Figure 8, the leader line C2 in Figure 10A, and the text attribute Tb in Figure 11. The display position of each text Ta reflects the text structure of the electronic document Dc, such as line breaks between sections and chapters. On the other hand, the display form of each text Ta also reflects the text structure of the electronic document Dc, such as the font size in section titles and chapter titles.
[0330] Therefore, by dividing the document data 31 into multiple text blocks Bl based on such text structure, a more accurate classification can be performed that reflects the actual section and chapter structure in the electronic document Dc. This contributes to improving the accuracy of classifying text Ta into each text block Bl.
[0331] Furthermore, as illustrated in step S52 of Figure 15, by dividing the electronic document Dc into multiple text blocks Bl, when processing the electronic document Dc with another AI for document analysis (second generation AI 208), it becomes possible to input it in units of text blocks Bl rather than inputting it in units of pages into the second generation AI 208. This makes it possible to make even electronic documents Dc with a large data size suitable for various processing (e.g., LLM) using the first generation AI 201.
[0332] Furthermore, as illustrated by the leader lines C11 and C12 in Figure 12, and the index data 35 in Figure 14A, the first generation AI 201 outputs index data 35 representing the line number Tc of each text Ta that constitutes each text block Bl, instead of outputting the text body of each text block Bl.
[0333] This makes it possible to reduce the data size output from the first generation AI201 and improve the classification speed of text Ta into text blocks Bl. This is particularly effective when processing a large number of electronic documents Dc with the first generation AI201.
[0334] Furthermore, as explained with reference to the leader line C3 in Figure 10A, by explicitly indicating in the system prompt 203 that the section or chapter structure of the electronic document Dc should be reflected, the first generating AI 201 can be made to actively estimate the section or chapter structure. This further improves the classification accuracy of the first generating AI 201 into text blocks Bl.
[0335] Furthermore, as explained with reference to step S32 in Figure 9, the CPU 3 acquires index data 35 using two methods depending on the document structure of the electronic document Dc. By using the rule-based estimation model 206, it is possible to achieve more accurate classification for documents such as patent documents, where the format of the electronic document Dc is clearly defined. On the other hand, for documents where the format is not clearly defined, it is possible to achieve more versatile classification by using the first generation AI 201.
[0336] <6. Outline of Document Generation Method> Next, I will explain the document visualization method.
[0337] The document generation method according to this embodiment is a method for generating an electronic document Dc using the document visualization method described above. This document generation method is performed by having computer 1 execute the document generation program 27.
[0338] Here, the document generation program 27 is a program configured to cause the computer 1 to execute each process that constitutes the document generation method according to this embodiment. The document generation program 27 is pre-stored in a computer-readable storage medium 17. This storage medium 17 is a tangible storage medium made up of a disk medium or the like.
[0339] The following provides a detailed explanation of the document visualization method.
[0340] <7. Details of the document generation method> Figure 31 is a flowchart illustrating the first half of the document generation method. Figure 32 is a flowchart illustrating the second half of the document generation method. Hereafter, the electronic document Dc containing the newly patentable technical idea Th will be referred to as the new technical document Dn (see Figure 33).
[0341] First, in step S201 of Figure 31, CPU 3 generates multiple text blocks Bl for the new technical document Dn. This generation may be performed by having computer 1 execute the text block generation program 23. In this case, CPU 3 may generate the text blocks Bl via the first generation AI 201, or via the rule-based estimation model 206.
[0342] In the subsequent step S202, the CPU3 extracts classification information In1 from multiple text blocks Bl generated based on the new technical document Dn, indicating at least one of the following: the technical field, background, prior art, technical challenges, features, and summary of the technical idea Th (see Figure 33). Using text blocks Bl allows for more accurate extraction of classification information In1.
[0343] Specifically, classification information In1 can be associated with each text block Bl. Hereafter, classification information In1 and text block Bl will be described as the same, but the term "text block Bl" can be replaced with "classification information In1" as appropriate.
[0344] In the subsequent step S203, the CPU3 obtains multiple patent application documents (hereinafter referred to as "prior documents") Da that are similar to the new technical document Dn, based on Retrieval Augmented Generation (RAG) using classification information In1 as input (see Figure 33). Specifically, the CPU3 obtains multiple prior documents Da by executing a query to an external server 102 such as the Japan Patent Office via a generation AI 211 capable of executing RAG.
[0345] In the subsequent step S204, the CPU3 obtains progress information In2, which shows the examination history of each prior art document Da, and standard information Im3, which shows the examination standards for patent applications in the target country (e.g., Japan). Specifically, the CPU3 obtains the progress information In2 and standard information In3 by executing a query to an external server 102 such as the Japan Patent Office via a generating AI 212 capable of executing RAG (see Figure 34). This information may also be obtained manually by the analyst without using the generating AI.
[0346] In the subsequent step S205, the CPU3 generates multiple text blocks Bl for each prior art document Da. At this time, the CPU3 may generate multiple text blocks Bl for each prior art document Da using the same method as for the new technical book Dn.
[0347] In the subsequent step S206, the CPU3 performs the following operations for the group of documents (multiple electronic documents Dc) consisting of the new technical document Dc and each prior art document Da: obtaining feature quantities Ns, clustering into multiple first clusters Cl1, and clustering into multiple second clusters Cl2. This process may also be performed by having computer 1 execute a part of the electronic document visualization program 25.
[0348] In the following step S207, the CPU3 associates the feature quantities Ns with each of the multiple first clusters Cl1 and the multiple second clusters Cl2 (details are described above). This process may also be performed by having the computer 1 execute a part of the electronic document visualization program 25.
[0349] In the subsequent step S208, the CPU3 visualizes multiple first clusters Cl1 on the first visualization map M1, and then extracts one or more prior art documents Da from among multiple prior art documents Da that belong to the same first cluster Cl1 as the new technical book Dn, in order of proximity on the first visualization map M1. This extraction may be performed based on the Euclidean distance on the first visualization map M1.
[0350] In the following step S209, the CPU3 visualizes multiple second clusters Cl2 on the second visualization map M2, and then extracts the points of agreement and difference between the new technical book Dn and the extracted prior art Da, for each text block Bl, according to the level of the feature quantity Ns.
[0351] In the following step S210, the CPU 3 inputs the points of agreement and disagreement to the second generating AI 208, causing the second generating AI 208 to determine whether or not the patentability of the technical idea Th of the new technical book Dn is affirmed (see Figure 35). As in the following step S211, the CPU 3 may visualize the result of this determination. This visualization can be done by displaying the determination result on the display 9.
[0352] Specifically, CPU3 can also input progress information In2 and reference information In3 to the second generation AI208, in addition to the points of agreement and differences. In that case, CPU3 instructs the second generation AI208 via its system prompt to make a judgment that takes progress information In2 and reference information In3 into consideration. By inputting progress information In2 and reference information In3 to the second generation AI208, CPU3 causes it to perform a judgment based on progress information In2 and reference information In3, using the points of agreement and differences as input.
[0353] Specifically, the second generation AI208 is composed of an LLM pre-trained on a large number of electronic documents Dc. This LLM is constructed, for example, by a transferer composed of a neural network.
[0354] Furthermore, the determination of whether or not patentability is affirmed can be made based on the remarkable effects and benefits of the technical idea Th described in the new technical document Dn, and the inhibiting factors in the combination of extracted prior art documents Da. The determination based on the effects and inhibiting factors can be made more accurately by inputting progress information In2 and reference information In3.
[0355] Furthermore, for at least one of the effects and inhibiting factors, the user may, via prompts or the like, have the second generation AI208 read a file they have created themselves (for example, a digital file in CSV format) through the CPU3.
[0356] Furthermore, in the subsequent step S212, the CPU3 determines whether the patentability of the new technical document Dn is affirmed or not. If it is affirmed, the control step proceeds to step S213; otherwise, the process exemplified in Figures 31 and 32 is terminated. Note that either step S214 or step S214 may be omitted.
[0357] In step S213, the CPU3 causes the third generation AI209 to output an invention summary Dx that describes the technical idea Th based on the new technical document Dn, extracted prior art Da, similarities, and differences (see Figure 36). An invention summary Dx is a document that textualizes the features of the technical idea Th, along with surrounding information such as background and problems. When generating the invention summary Dx, the third generation AI209 may be given classification information In1 (not shown). The invention summary Dx may include at least a part of the classification information In1.
[0358] The invention summary Dx may include, for example, the remarkable effects of the technical idea Th, and the inhibiting factors in the combination of extracted prior art documents Da, as illustrated in the output data Inx of Figure 40. When analyzing the similarities and differences, a portion of the output data Inx ("Comparison with Prior Art Document 1") may be visualized.
[0359] The invention summary Dx may include, for example, summaries of each prior art document Da. Furthermore, the display of various texts included in the invention summary Dx, such as the technical idea Th and summaries of each prior art document Da, may be varied according to the feature quantity Ns.
[0360] For example, the font size, font color, etc., of the morpheme corresponding to the specific feature Na may be different from those of other morphemes in the invention summary Dx.
[0361] Specifically, the third-generation AI209 is composed of an LLM pre-trained on a large number of electronic documents Dc. This LLM is constructed, for example, by a transferer composed of a neural network.
[0362] In this embodiment, document data 31 is digital data containing multiple texts Ta. Each of the multiple texts Ta is digital data corresponding to a string of characters written in the electronic document Dc.
[0363] In the subsequent step S214, the CPU3 causes the third generating AI209 to output at least a portion of the patent application Dy describing the technical idea Th, based on the new technical document Dc, the extracted prior art Da, the similarities and differences (see Figure 37). At this time, the CPU3 may optionally input an invention summary Dx, the remarkable effects of the technical idea Th, and the inhibiting factors between the extracted prior art Da combinations. By inputting this information, the third generating AI209 can output a more appropriate application Dy.
[0364] In Japan, "at least a part of the application form Dy" may refer to one or more of the following sections: "Title of Invention," "Technical Field," "Background Art," "Summary of Invention," "Modes for Carrying Out the Invention," and "Claims." It may also refer to a portion of each section, such as a portion of "Summary of Invention" and "Claims."
[0365] Furthermore, as illustrated in Figure 38, CPU3 can also update the Invention Summary Dx based on the description in the Application Dy by inputting the Application Dy and Invention Summary Dx actually submitted to the Patent Office after the application has been filed into the Third Generation AI 209 or another Generation AI (see Figure 38). The Application Dy actually submitted to the Patent Office may be the Application Dy generated by the Third Generation AI 209, another Application Dy created based on that Application Dy, or an Application Dy created independently of that Application Dy. Another Application Dy refers to, for example, the Application Dy generated by the Third Generation AI 209 that has undergone processes such as additions, deletions, and modifications.
[0366] Furthermore, as illustrated in Figure 39, after filing a patent application using the application form Dy, the CPU 3 can also output the contents of the application form Dy via the interactive fourth generation AI (fourth generation AI) 210. Alternatively, the system may be configured to take an invention summary Dx as input instead of the application form Dy and output the contents of the invention summary Dx. The "content" output here includes one or more of the following: "abstract of the invention," "applicant's intellectual property manager," "assigned patent attorney," "summary of features," and "details of competitors."
[0367] Specifically, the fourth generation AI210 is composed of an LLM pre-trained on a large number of electronic documents Dc. This LLM is constructed, for example, by a transferer composed of a neural network.
[0368] As described above, the document creation method, as explained using steps S208 and S209 in Figures 31 and 32, can output points of agreement and disagreement in terms of text block Bl units of the electronic document Dc, such as sections and chapters, by using the first cluster Cl1 and the second cluster in combination. By making the judgment at the text block Bl unit level rather than at the electronic document Dc unit level, points of agreement and disagreement can be determined with greater accuracy.
[0369] Furthermore, as explained in relation to step S210 in Figure 32, by referring to the examination history (procedure information In2) of prior art Da, the determination of whether or not patentability is affirmed can be made with greater accuracy.
[0370] Furthermore, as explained in relation to step S210 in Figure 32, referring to the examination standards of the country where the application is being filed allows for a more accurate determination of whether or not patentability is affirmed.
[0371] Furthermore, as explained using steps S208 and S209 in Figures 31 and 32, by using visualization based on the first cluster Cl1 and the second cluster Cl2 in combination, as described above, it is possible to accurately determine the similarities and differences, and as a result, it becomes possible to generate a more appropriate invention summary Dx. By exchanging this invention summary Dx between the company and its agent, it can be used to assist in the creation of patent claims, specifications, etc. This makes it possible to proceed with patent applications efficiently.
[0372] Furthermore, as explained using steps S208 and S209 in Figures 31 and 32, by using visualization based on the first cluster Cl1 and the second cluster Cl2 in combination, as described above, it is possible to accurately determine the points of agreement and differences, and as a result, it becomes possible to generate a more appropriate application Dy. This makes it possible to proceed with patent applications efficiently.
[0373] Furthermore, as explained using Figure 34, updating the invention summary Dx with the third generation AI209 eliminates discrepancies between the invention summary Dx and the actual application Dy. This makes it possible to manage the contents of each patent application more efficiently.
[0374] Furthermore, as illustrated in Figure 35, the contents of each document can be output to the fourth generation AI 210. This makes it possible to manage the contents of each patent application more efficiently.
[0375] <8. Other Embodiments> In the above embodiment, a generative AI for the Embedding model was used for the language model 207, but the use of a generative AI is not essential. The language model 207 may be a machine learning model that has learned distributed representations of text.
[0376] Furthermore, while the above-described embodiment shows an example in which a document visualization device and a document generation device are configured by a single computer 1, this disclosure is not limited to that example. The document visualization method, document visualization device and document visualization program 21, and the document generation method, document generation device and document generation program 27 related to this disclosure may be executed using multiple computers 1, for example, by having a first computer execute some of the processing while a second computer executes the remaining processing. In addition, the computer 1 in this disclosure also includes parallel computers such as supercomputers and PC clusters. Each computer 1 may be equipped with multiple CPUs 3, and it is not necessary to have all processing executed by the same CPU 3.
[0377] Furthermore, the screen on which the visualized information can be displayed is not limited to the display screen on computer 1's display 9. Scatter plots and the like may be displayed on a screen prepared separately from computer 1. In other words, the "display unit" in this disclosure only needs to be connected to the CPU 3, and it is not necessary for it to be part of computer 1.
[0378] Furthermore, it is not mandatory to implement the first generation AI 201, second generation AI 208, third generation AI 209, and second generation AI 210 outside of computer 1. The functions performed by these generation AIs may be executed by computer 1. Moreover, it is not mandatory to classify the generation AIs into four categories in the first place. A single generation AI may perform multiple types of processing. [Explanation of symbols]
[0379] 1. Computer (document analysis device) 3 CPU (arithmetic unit) 7a RAM (storage unit) 7b SSD (storage unit) 7c Storage device (storage unit) 9. Display (Display Unit) 13 Reception Department 13a Keyboard 13b Mouse 17 Storage medium 19 GPU 21 Document Visualization Program 23 Text block generation program 25 Electronic Document Visualization Program 27 Document generation program 31 Document Data 33. Structured Data (Data 1) 35 Index data (second data) 37 Prompt data 41. Embedded data (multidimensional data) 201 1st generation AI (generation AI) 203 System prompt (prompt) 207 Language Models 208 Second Generation AI (Second Generation AI) 209 Third Generative AI (Third Generative AI) 210 Fourth Generative AI (Fourth Generative AI) Cl1 1st cluster Cl2 Second Cluster Cr1 First Word Cloud CR2 Second Word Cloud DC Electronic Documents Dn New Technical Book Da Previous Literature Dx Invention Summary Application form for Dy patent application Bl Text Block M1 First Visualization Map (First Map) M2 Second Visualization Map (Second Map) Ns features Na average feature Ta Text Tb Text Attributes Tc line number In1 classification information In2 Progress Information In3 Standard Information Inx output data< / italics> < / size> < / page>
Claims
1. A document visualization method that uses a computer having a processing unit to read multiple electronic documents, each containing multiple texts, and visualizes the relationships between the multiple electronic documents, The calculation unit reads the plurality of electronic documents as a collection of multiple text blocks, each containing one or more texts, The calculation unit quantifies the plurality of electronic documents into multidimensional data via a language model that takes the plurality of text blocks corresponding to each electronic document as input, aggregates the multidimensional data obtained for each text block for each electronic document, and then clusters it into a plurality of first clusters. The calculation unit obtains a feature quantity for each of the plurality of electronic documents, which is a statistically quantified representation of each morpheme contained in each electronic document, using the plurality of text blocks corresponding to each electronic document as input. The calculation unit associates the feature quantities for each of the multiple first clusters based on the electronic documents that constitute each of the first clusters, The calculation unit visualizes the plurality of first clusters as a two-dimensional or three-dimensional first map and controls the display mode of the first map according to the value of the feature quantity. A document visualization method characterized by the following:
2. In the document visualization method described in claim 1, The calculation unit clusters each of the text blocks corresponding to the multidimensional data into a plurality of second clusters, each consisting of one or more text blocks. The calculation unit associates the feature quantities for each second cluster based on the text blocks that constitute each of the plurality of second clusters, The calculation unit visualizes the plurality of second clusters as a two-dimensional or three-dimensional second map and controls the display mode of the second map according to the level of the feature quantities. A document visualization method characterized by the following:
3. In the document visualization method described in claim 2, The calculation unit is connected to a reception unit that receives user input, The calculation unit selects two or more of the multiple electronic documents based on the operation input to the reception unit, and visualizes the second map for each text block of the two or more selected electronic documents. A document visualization method characterized by the following:
4. In the document visualization method described in claim 1, The calculation unit obtains the feature quantities based on morphological analysis using the plurality of text blocks as input. The calculation unit obtains an index based on the frequency or number of occurrences of each morpheme as the feature quantity. A document visualization method characterized by the following:
5. In the document visualization method described in claim 4, The calculation unit is connected to a reception unit that receives user input, When the receiving unit receives a predetermined operation input, the calculation unit superimposes a first word cloud representing the high or low of the index for each morpheme onto the first map for each of the first clusters. A document visualization method characterized by the following:
6. In the document visualization method described in claim 5, Each of the plurality of text blocks includes author information that characterizes the author of each of the electronic documents corresponding to the plurality of text blocks. When the receiving unit receives a predetermined operation input, the calculation unit superimposes on the first map a second word cloud representing the frequency or number of occurrences of the author information for each of the first clusters, together with the first word cloud representing the frequency or number of occurrences of each morpheme relating to information other than the author information. A document visualization method characterized by the following:
7. In the document visualization method described in claim 6, The aforementioned author information includes information indicating the affiliation of the author of each of the aforementioned electronic documents. A document visualization method characterized by the following:
8. In the document visualization method described in claim 1, Each of the aforementioned multiple electronic documents is composed of document data containing the aforementioned multiple texts, The calculation unit generates and outputs first data based on the electronic document, associating text attributes that characterize at least one of the display position and display form of each of the plurality of texts on the electronic document, text corresponding to the text attributes, and identification information for identifying the text on the electronic document. The calculation unit then generates the AI, Accepting the aforementioned first data as input, The input format of the aforementioned first data, With respect to the document data which is divided into a plurality of text blocks by referring to the first data, the breakdown of the text belonging to each of the plurality of text blocks is output as second data represented by the identification information associated with the text, Enter a prompt that includes: The calculation unit inputs the first data to the generating AI, thereby acquiring the second data via the generating AI. The calculation unit selects the text by referring to the first data, and combines the selected text in an order based on the identification information, thereby forming and outputting at least one of the plurality of text blocks as a sentence. A document visualization method characterized by the following:
9. In the document visualization method described in claim 8, The aforementioned multiple text blocks are constructed to reflect the section or chapter structure of the electronic document. The calculation unit inputs to the generating AI the instruction to divide the document data into sections or chapters, including this instruction in the prompt. A document visualization method characterized by the following:
10. In the document visualization method described in claim 8, The calculation unit determines whether the document structure of the electronic document is known or not. The aforementioned arithmetic unit, If the aforementioned sentence structure is known, the first data is input into a rule-based estimation model, and the second data is obtained through the estimation model. If the aforementioned document structure is unknown, the first data is input to the generating AI, and the second data is obtained via the generating AI. A document visualization method characterized by the following:
11. A document visualization device that uses a computer having a processing unit to read multiple electronic documents, each containing multiple texts, and visualizes the relationships between the multiple electronic documents, Means for reading the aforementioned plurality of electronic documents as a collection of multiple text blocks, each containing one or more texts, A means for quantifying the aforementioned multiple electronic documents into multidimensional data via a language model that takes the aforementioned multiple text blocks corresponding to each electronic document as input, and for aggregating the multidimensional data obtained for each text block for each electronic document, and then clustering them into multiple first clusters, For each of the aforementioned multiple electronic documents, a means for obtaining a feature quantity obtained by statistically quantifying each morpheme contained in each electronic document, using the aforementioned multiple text blocks corresponding to each electronic document as input, A means for associating the feature quantities for each of the multiple first clusters based on the electronic documents that constitute each of the first clusters, The system includes means for visualizing the plurality of first clusters as a two-dimensional or three-dimensional first map, and for controlling the display mode of the first map according to the value of the feature quantities. A document visualization device characterized by the following features.
12. A document visualization program that, when executed on a computer having a processing unit, reads multiple electronic documents, each containing multiple texts, and visualizes the relationships between the multiple electronic documents, To the aforementioned computer, The processing unit reads the plurality of electronic documents as a collection of multiple text blocks, each containing one or more texts, The calculation unit quantifies the plurality of electronic documents into multidimensional data via a language model that takes the plurality of text blocks corresponding to each electronic document as input, aggregates the multidimensional data obtained for each text block for each electronic document, and then clusters it into a plurality of first clusters. The calculation unit performs a process of obtaining, for each of the plurality of electronic documents, a feature quantity obtained by statistically quantifying each morpheme contained in each electronic document, using the plurality of text blocks corresponding to each electronic document as input. The calculation unit performs a process of associating the feature quantities for each of the multiple first clusters based on the electronic documents that constitute each of the first clusters, The calculation unit performs a process of visualizing the plurality of first clusters as a two-dimensional or three-dimensional first map, and controlling the display mode of the first map according to the value of the feature quantities. A document visualization program characterized by the following features.
13. It stores the document visualization program described in claim 12. A computer-readable storage medium characterized by the following features.
14. A document generation method using the document visualization method described in claim 1, If an electronic document describing a newly patentable technical idea is considered a new technical document, and multiple patent application documents similar to this new technical document are considered multiple prior art documents, The calculation unit performs the generation of the plurality of text blocks for the new technical document, The calculation unit extracts classification information from the plurality of text blocks generated based on the new technical document, indicating at least one of the following: the technical field, background, prior art, technical challenges, features, and summary of the technical idea. The calculation unit obtains multiple prior art documents based on the search extension generation using the classification information as input. The calculation unit performs the generation of the plurality of text blocks for each of the plurality of prior art documents, The calculation unit performs the acquisition of feature quantities and clustering into the plurality of first clusters for the collection of documents comprising the new technical book and the plurality of prior art documents. The calculation unit performs clustering of the document collection into the plurality of second clusters, The calculation unit associates the feature quantities with respect to each of the plurality of first clusters and the plurality of second clusters, The calculation unit visualizes the plurality of first clusters on the first map, and then extracts one or more prior art documents from the plurality of prior art documents that belong to the same first cluster as the new technical book, in order of proximity on the first map. The calculation unit visualizes the second cluster on the second map, and then, according to the level of the feature quantities, extracts the points of agreement and differences between the new technical book and the extracted prior art documents for each text block. The calculation unit inputs the points of agreement and the points of difference to the second generating AI, causing the second generating AI to determine whether or not the patentability of the technical idea is affirmed, and visualizes the result of that determination. A document generation method characterized by the following:
15. In the document generation method described in claim 14, The calculation unit acquires progress information showing the review process of the prior art, The calculation unit inputs the progress information to the second generating AI and causes it to perform a judgment based on the progress information, using the points of agreement and the points of difference as input. A document generation method characterized by the following:
16. In the document generation method described in claim 14, The calculation unit obtains the examination standards for patent applications in the target country, The calculation unit inputs the criteria information indicating the evaluation criteria into the second generating AI, and causes it to perform a judgment based on the points of agreement and the points of difference based on the criteria information. A document generation method characterized by the following:
17. In the document generation method described in claim 14, If the calculation unit determines that the patentability of the new technical document is affirmed, it causes the third generating AI to output an invention summary document describing the technical idea based on the new technical document, the extracted prior art, the points of agreement, and the points of difference. A document generation method characterized by the following:
18. In the document generation method described in claim 17, If the calculation unit determines that the patentability of the new technical document is affirmed, it causes the third generating AI to output at least a portion of the patent application form describing the technical idea, based on the new technical document, the extracted prior art, the points of agreement, and the points of difference. A document generation method characterized by the following:
19. In the document creation method described in claim 17. The calculation unit updates the invention summary based on the contents of the application actually submitted to the patent office after the patent application based on the new technical document. A document generation method characterized by the following:
20. In the document creation method described in claim 18, The calculation unit causes the contents of at least one of the invention summary and the application to be output via an interactive fourth generation AI. A document generation method characterized by the following:
Citation Information
Patent Citations
Method and system for identifying field label on form
JP2024009774A