Document search apparatus, document search method, and document search program
The document search device enhances document search capabilities by extracting and labeling characteristic words, creating a co-occurrence network to facilitate the discovery of relevant use cases and materials.
Patent Information
- Application Number
- JP2023223405
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-10
AI Technical Summary
Existing document search systems, such as those described in Patent Document 1, fail to effectively search for information related to characteristic words in documents, limiting the ability to find highly relevant use cases and materials.
A document search device that extracts characteristic words from documents, assigns labels to these words based on co-occurrence relationships, and outputs information in a predetermined order using a co-occurrence network, allowing for the search of relevant use cases and materials.
Enables the search for highly relevant use cases and materials by analyzing the co-occurrence of characteristic words, expanding the applicability of document search systems.
Smart Images

Figure 2025105100000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a document search device, a document search method, and a document search program.
Background Art
[0002] For example, Patent Document 1 discloses a document classification device including: an input unit that inputs a document to be classified; a document similarity calculation unit that calculates the similarity between the input document and a set of documents with pre-assigned classifications by extracting keywords; a document extraction unit that extracts a specified number of documents most similar to the input document from the set of documents with pre-assigned classifications; a score calculation unit that calculates the classification score of the specified number of documents extracted based on the number of documents in the same classification considering the similarity; and a classification set extraction unit that extracts classifications for which the calculated score is greater than a specified value.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, the technique described in Patent Document 1 only determines the importance of documents such as patent documents, and cannot search for information related to a characteristic word from the characteristic words in the documents such as patent documents. That is, in the prior art, it is impossible to search for highly relevant use cases and materials from the characteristic words representing arbitrary materials and use cases obtained from documents such as patent documents.
[0005] In view of the above points, the disclosed technique is made, and an object thereof is to provide a document search device, a document search method, and a document search program capable of searching for information related to a characteristic word from the characteristic words in a document.
Means for Solving the Problems
[0006] In order to achieve the above object, a document search device according to one aspect of the present disclosure includes an acquisition unit that acquires a document to be searched, an extraction unit that identifies valid sentences from the document acquired by the acquisition unit and extracts characteristic words included in the identified valid sentences, an assignment unit that assigns a label corresponding to a predetermined search item to the characteristic words extracted by the extraction unit, a creation unit that creates a co-occurrence network representing the co-occurrence relationship of the characteristic words to which the label has been assigned by the assignment unit, and an output unit that arranges and outputs information on the search items related to the label in a predetermined order based on the co-occurrence network created by the creation unit.
[0007] Further, it further includes a reception unit that receives an input of a reference word related to the label, and a search unit that searches the co-occurrence network starting from the reference word to identify information on the search items related to the reference word and obtains a first score indicating the degree of association between the reference word and the identified information on the search items. The output unit may output the identified information on the search items in descending order of the first score.
[0008] Further, the reception unit further receives an input of a sentence including the information on the search items output by the output unit, searches a plurality of pieces of document information related to the sentence from a document information database, and further includes a search unit that obtains a second score indicating the degree of association between the sentence and each of the plurality of pieces of document information searched. The output unit may output the plurality of pieces of document information searched by the search unit together with the second score.
[0009] Further, it further includes a classification unit that classifies important document information among the plurality of pieces of document information output by the output unit using a classification model for classifying important document information based on predetermined classification conditions. The output unit may output the important document information classified by the classification unit in an identifiable manner.
[0010] Further, the search item is a use case, and further includes a search unit that searches for information on compounds related to the information of the use case output by the output unit from a compound information database. The output unit may output the information on the compound searched by the search unit in place of the information on the use case or together with the information on the use case.
[0011] Further, the search item may be any one of a material, a property, and a use case, and the label may also be any one of a material, a property, and a use case.
[0012] Further, the document is a patent document, the patent document includes claims in which the invention is described and an abstract in which the summary of the invention is described, and the effective sentence may include at least one of the sentences described in the claims and the sentences described in the abstract.
[0013] Further, the extraction unit may remove at least one of specific characters, numerical values, and symbols from the sentences described in the claims or the sentences described in the abstract, and extract the characteristic words from the sentences from which at least one of the specific characters, numerical values, and symbols has been removed.
[0014] A document search method according to an aspect of the present disclosure includes: a document search device obtaining a document to be searched, specifying an effective sentence from the obtained document, extracting characteristic words included in the specified effective sentence, assigning a label corresponding to a predetermined search item to the extracted characteristic words, creating a co-occurrence network representing a co-occurrence relationship of the characteristic words to which the label is assigned, and arranging and outputting information on the search item related to the label in a predetermined order based on the created co-occurrence network.
[0015] A document search program according to one aspect of the present disclosure acquires documents to be searched, identifies valid sentences from the acquired documents, extracts characteristic words included in the identified valid sentences, assigns a label corresponding to a predetermined search item to the extracted characteristic words, creates a co-occurrence network representing the co-occurrence relationship of the characteristic words to which the label is assigned, and causes a computer to execute a process of arranging and outputting information on the search items related to the label in a predetermined order.
Effect of the Invention
[0016] According to the present disclosure, information related to a characteristic word can be searched from the characteristic words in a document.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0018] Hereinafter, with reference to the drawings, an example of a mode for carrying out the technology of the present disclosure will be described in detail. Note that components and processes having the same functions of operation, action, and function may be given the same reference numerals throughout all the drawings, and redundant descriptions may be omitted as appropriate. Each drawing only schematically shows the technology of the present disclosure to such an extent that it can be sufficiently understood. Therefore, the technology of the present disclosure is not limited only to the illustrated examples. In addition, in the present embodiment, descriptions of configurations not directly related to the technology of the present disclosure and well-known configurations may be omitted.
[0019] FIG. 1 is a block diagram showing an example of the configuration of a document search system 100 according to the present embodiment.
[0020] As shown in FIG. 1, the document search system 100 according to the present embodiment includes a document search device 10, a user terminal 30, a document information DB (database) 40, and a compound information DB 50. As the document search device 10, for example, a general-purpose computer device such as a personal computer or a server computer is used. The document search device 10 is communicably connected to the user terminal 30 via a network N. The document search device 10 is communicably connected to each of the document information DB 40 and the compound information DB 50 via the network N.
[0021] The document search device 10 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, an input / output interface (I / O) 14, a storage unit 15, a display unit 16, an operation unit 17, and a communication unit 18.
[0022] The CPU 11, ROM 12, RAM 13, and I / O 14 are each connected via a bus. To the I / O 14, each functional unit including a storage unit 15, a display unit 16, an operation unit 17, and a communication unit 18 is connected. These functional units are enabled to communicate with each other with the CPU 11 via the I / O 14.
[0023] The control unit is constituted by the CPU 11, ROM 12, RAM 13, and I / O 14. The control unit may be configured as a sub-control unit that controls some operations of the document search device 10, or may be configured as a part of the main control unit that controls the entire operation of the document search device 10. For some or all of the blocks of the control unit, for example, an integrated circuit such as an LSI (Large Scale Integration) or an IC chip set is used. Individual circuits may be used for the above-mentioned respective blocks, or circuits in which some or all are integrated may be used. The above-mentioned respective blocks may be provided integrally, or some of the blocks may be provided separately. Also, in each of the above-mentioned blocks, a part thereof may be provided separately. For the integration of the control unit, not limited to LSI, a dedicated circuit or a general-purpose processor may be used.
[0024] As the storage unit 15, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), a flash memory, etc. are used. In the storage unit 15, the document search program 15A according to the present embodiment is stored. Note that this document search program 15A may be stored in the ROM 12.
[0025] The document search program 15A may be pre-installed in the document search device 10, for example. The document search program 15A may be realized by storing it in a non-volatile storage medium, distributing it via the network N, and appropriately installing it in the document search device 10. Examples of non-volatile storage media include CD-ROM (Compact Disc Read Only Memory), magneto-optical disk, HDD, DVD-ROM (Digital Versatile Disc Read Only Memory), flash memory, memory card, etc.
[0026] For the display unit 16, for example, a liquid crystal display (LCD), an organic EL (Electro Luminescence) display, etc. are used. The display unit 16 may integrally have a touch panel. For the operation unit 17, for example, devices for operation input such as a keyboard and a mouse are provided. The display unit 16 and the operation unit 17 receive various instructions from the user of the document search device 10. The display unit 16 displays various information such as the result of the process executed according to the instruction received from the user and the notification for the process.
[0027] As an example, the communication unit 18 is connected to the network N such as the Internet, LAN (Local Area Network), WAN (Wide Area Network), etc., and is capable of communicating with each of the user terminal 30, the document information DB 40, and the compound information DB 50 via the network N.
[0028] The user terminal 30 is operated by the user. Functionally, the user terminal 30 includes a control unit 31 and a display unit 32 as shown in FIG. 1.
[0029] The control unit 31 controls the operation of the user terminal 30. The display unit 32 displays various information according to the control by the control unit 31.
[0030] The document information DB40 is a database that registers document information such as patent documents, academic papers, and news articles. For the document information DB40, databases such as J-Platpat provided by the Japan Patent Office are applied, for example. The compound information DB50 is a database that registers compound information. For the compound information DB50, databases such as the Japanese Chemical Substance Dictionary (Nikkaji) and PubChem are applied, for example.
[0031] The document search device 10 according to the present embodiment extracts characteristic words from a document to be searched, assigns a label corresponding to a predetermined search item to the extracted characteristic words, and arranges and outputs information on the search items related to the label in a predetermined order based on a co-occurrence network representing the co-occurrence relationship of the labeled characteristic words. Thereby, information related to the characteristic words can be searched from the characteristic words in the document. That is, by using the co-occurrence network, it is possible to search for highly relevant use cases and materials from the characteristic words representing a certain material or use case obtained from the document. Thereby, it is possible to expand use cases highly relevant to a certain material or discover materials highly relevant to a certain use case. Note that hereinafter, a patent document will be cited as an example of a document for explanation, but the present invention is not limited to patent documents.
[0032] Specifically, the CPU 11 of the document search device 10 according to the present embodiment functions as each unit shown in FIG. 2 by writing the document search program 15A stored in the storage unit 15 into the RAM 13 and executing it.
[0033] FIG. 2 is a block diagram showing an example of the functional configuration of the document search device 10 according to the present embodiment.
[0034] As shown in FIG. 2, the CPU 11 of the document search device 10 according to the present embodiment functions as an acquisition unit 11A, an extraction unit 11B, an assignment unit 11C, a creation unit 11D, a reception unit 11E, a search unit 11F, an output unit 11G, a search section 11H, and a classification unit 11J.
[0035] The storage unit 15 stores a document search program 15A and an LLM (Large Language Models) 15B. Note that the LLM 15B may be stored in an external storage device. This LLM 15B is a trained model generated by machine learning. The trained model is not limited to the LLM 15B, and other natural language processing models may be used for the trained model. Instead of the LLM 15B, a generative AI (Artificial Intelligence) such as ChatGPT may be used for the trained model.
[0036] The acquisition unit 11A acquires a patent document to be searched from the user terminal 30. Here, a patent document is also referred to as a patent specification, and is a document in which the patent applicant has sufficiently described the invention to the extent that a person skilled in the technical field related to the invention can implement the invention. Note that the patent document includes a utility model specification. A patent document is generally composed of a specification, claims, abstract, and drawings. The claims include one or more claims in which the invention is described. The abstract describes the summary of the invention. The specification describes the embodiments of the invention. The drawings show the diagrams necessary for explaining the invention.
[0037] The extraction unit 11B identifies valid sentences from the patent document acquired by the acquisition unit 11A, and extracts characteristic words included in the identified valid sentences. For example, Python is used for the extraction of characteristic words. Python is an interpreter-type high-level general-purpose programming language. As an example, the extraction unit 11B uses the LLM 15B to extract characteristic words included in the identified valid sentences. In this case, the LLM 15B is a model generated by machine learning the characteristic words of valid sentences included in a group of pre-prepared patent documents. The group of patent documents is, for example, a collection of patent documents related to a specific technical field for a certain development theme. The LLM 15B takes a valid sentence as input and outputs the characteristic words of the valid sentence. A valid sentence is a sentence effective for the extraction of characteristic words. A valid sentence includes, for example, at least one of the sentences described in the claims and the sentences described in the abstract.
[0038] Note that it is desirable for the extraction unit 11B to remove at least one of specific characters, numerical values, and symbols from the text described in the claims or the text described in the abstract. As an example, remove the string within the inked parentheses (e.g., the abstract, etc.), the string within the parentheses, numerical values, alphabets (e.g., (101A), etc.), or unnecessary parts such as "the foregoing", etc.
[0039] When the effective text before noise removal is, for example, "The cooling floor member (100A) is a cooling floor member (100A) that cools the battery cell placed above, ···", the effective text after noise removal is "The cooling floor member is a cooling floor member that cools the battery cell placed above, ···". By inputting the effective text after noise removal into the LLM15B, it becomes possible to accurately extract the feature amount of the effective text.
[0040] FIG. 3 is a diagram showing an example of characteristic words extracted from the effective text of a patent document. Visualization of the patent document is performed by identifying the effective text from the patent document and extracting the characteristic words included in the identified effective text. Note that the extracted characteristic words may be clustered. By clustering the extracted characteristic words, it becomes possible to grasp the approximate distribution of the obtained patent documents.
[0041] The assignment unit 11C assigns a label corresponding to a predetermined search item to the characteristic words extracted by the extraction unit 11B. Here, the search item is, for example, any one of material, property, and use case, and the label is, for example, any one of material, property, and use case. The combination of the search item and the label is not particularly limited and may be appropriately set according to the purpose of the user. Specifically, as an example, the assignment unit 11C uses the LLM15B to assign a label to the characteristic words. In this case, the LLM15B is a model generated by machine learning by associating the characteristic words with the labels. The LLM15B takes the characteristic words as input and outputs the label corresponding to the characteristic words.
[0042] FIG. 4 is a diagram showing an example of a label assigned to a characteristic word. As shown in FIG. 4, if the characteristic word is related to "material", a label of "material" is assigned, and if the characteristic word is related to "property", a label of "property" is assigned. If the characteristic word is related to "use case", a label of "use case" is assigned.
[0043] As an example, as shown in FIG. 5, the creation unit 11D creates a co-occurrence network 55 representing the co-occurrence relationship of the characteristic words to which labels are assigned by the assignment unit 11C. The co-occurrence network 55 may be created, for example, from a co-occurrence relationship based on the appearance frequency of the characteristic words.
[0044] FIG. 5 is a diagram showing an example of a co-occurrence network 55 representing the co-occurrence relationship of characteristic words. The co-occurrence network 55 shown in FIG. 5 is created using a general index representing the co-occurrence relationship. As an example of a general index, the Jaccard coefficient can be used, but it is not limited to the Jaccard coefficient. The co-occurrence network 55 includes information in a co-occurrence relationship with a characteristic word to which any label of, for example, "material", "property", and "use case" is assigned.
[0045] The output unit 11G outputs, in a predetermined order, information on search items related to the label based on the co-occurrence network 55 created by the creation unit 11D. The predetermined order may be, for example, in descending order or ascending order.
[0046] Specifically, the reception unit 11E receives an input of a reference word related to the label from the user. For example, when searching for any of "material", "property", and "use case", it receives an input of a reference word related to any of "material", "property", and "use case" assigned as a label.
[0047] The exploration unit 11F explores the co-occurrence network 55 starting from the received reference word, identifies information on search items related to the reference word, and obtains a first score indicating the degree of association between the reference word and the identified information on the search items. In this case, the output unit 11G outputs the identified information on the search items in descending order of the first score. The output unit 11G causes the information on the search items and the first score to be displayed, for example, on the display unit 32 of the user terminal 30.
[0048] FIG. 6 is a diagram showing an example of the output result 70 of the information on the search items and the first score obtained from the co-occurrence network 55.
[0049] As shown in FIG. 6, when the user inputs a reference word related to, for example, "material" into the input field 60 displayed on the display unit 32 of the user terminal 30, the exploration unit 11F explores the co-occurrence network 55 starting from the received reference word, identifies information on "use case", which is an example of a search item related to the reference word, and obtains a first score indicating the degree of association between the reference word and the identified information on the search item. Then, the output unit 11G causes the output result 70 of the information on the search item and the first score to be displayed on the display unit 32 of the user terminal 30. In the output result 70 of FIG. 6, the search item is "use case", and as the information on "use case", for example, "culture medium", "edible", "meat",... are displayed on the display unit 32 of the user terminal 30 in descending order of the first score.
[0050] Here, when the search item is "use case", it may be configured to cooperate with the compound information database 50 and execute a related compound search process. In this case, the search unit 11H uses the LLM 15B to search the compound information database 50 for information on "compounds" related to the information on "use case" output by the output unit 11G. The LLM 15B takes the information on "use case" as input and outputs information on related "compounds". The output unit 11G outputs the information on "compounds" retrieved by the search unit 11H in place of the information on "use case", or together with the information on "use case". For example, as shown in FIG. 6, instead of the output result 70 including the information on "use case", the output result 71 including the information on "compounds" is displayed on the display unit 32 of the user terminal 30.
[0051] Alternatively, it may be configured to cooperate with the literature information database 40 and execute a related literature search process including the information on the search item. In this case, the reception unit 11E receives an input of a text including the information on the search item output by the output unit 11G. The search unit 11H uses the LLM 15B to search the literature information database 40 for a plurality of pieces of literature information related to the text received by the reception unit 11E, and obtains a second score indicating the degree of relevance between the text and each of the plurality of pieces of retrieved literature information. The LLM 15B takes the received text as input and outputs a plurality of pieces of related literature information. The output unit 11G outputs the plurality of pieces of related literature information retrieved by the search unit 11H together with the second score.
[0052] FIG. 7 is a diagram showing an example of an output result 80 of a plurality of pieces of related literature information and the second score obtained from the literature information database 40. In FIG. 7, as an example, the case where patent documents are the search target is shown.
[0053] As shown in FIG. 7, when the user inputs a text such as "a compound having oo characteristics in the xx system" or "a use case of xx used in oo" into the input field 60 displayed on the display unit 32 of the user terminal 30, the search unit 11H searches, as an example, for a plurality of patent documents related to the text received from the literature information DB 40 using the approximate nearest neighbor search method, and acquires the plurality of related patent documents and the second score. Then, the output unit 11G causes the display unit 32 of the user terminal 30 to display the output result 80 of the plurality of related patent documents and the second score. In the output result 80 of FIG. 7, the plurality of related patent documents and the second score acquired from the literature information DB 40 are displayed on the display unit 32 of the user terminal 30. However, in the case of patent documents, it is also possible to associate them with the applicant and company name.
[0054] Further, important document classification processing may be executed on the plurality of related document information acquired from the literature information DB 40.
[0055] The classification unit 11J classifies important document information among the plurality of related document information output by the output unit 11G using the LLM 15B. The LLM 15B is an example of a classification model and classifies important document information based on predetermined classification conditions. Here, as a machine learning method, for example, known machine learning and deep learning methods such as BERT (Bidirectional Encoder Representations from Transformers) are used. BERT is a type of natural language processing model. In BERT, the relationship between words is learned by the Masked Language Model, and the relationship between sentences is learned by the Next Sentence Prediction.
[0056] As an example, the output unit 11G outputs the important document information classified by the classification unit 11J in an identifiable manner, as shown in FIG. 8.
[0057] FIG. 8 is a diagram showing an example of important document information classified from among a plurality of related document information. In FIG. 8, as an example, the case where patent documents are the classification target is shown.
[0058] As shown in FIG. 8, the output unit 11G causes the display unit 32 of the user terminal 30 to display the important patent documents classified by the classification unit 11J in a distinguishable manner with respect to the output result 81 of the related patent documents acquired from the document information DB 40. In the example of FIG. 8, black circles are added to the important patent documents to make them distinguishable. However, for example, any form may be used as long as it is distinguishable from other related patent documents, such as adding a colored marker or underlining.
[0059] Further, an important part specifying process for specifying an important part (a part to be read) among specific parts (for example, "mode for carrying out the invention") described in the important patent documents may be executed.
[0060] The classification unit 11J may extract, from the specification, a sentence having a feature amount similar to the feature amount of the effective sentence included in the important patent document by using the LLM 15B. The LLM 15B inputs a sentence and outputs a feature amount (for example, a feature vector) of the sentence. Specifically, a feature amount is extracted from the sentence described in the specification, and the similarity to the feature amount of the effective sentence described in the claims or the abstract is calculated. For calculating the similarity, for example, a known method such as cosine similarity is used.
[0061] FIG. 9 is a diagram showing an example of a similar sentence extracted from the specification.
[0062] As shown in FIG. 9, when the sentence described in claim 1 is input to the LLM 15B as an effective sentence, a feature amount is extracted for the sentence of claim 1. It is desirable to use, as the effective sentence, a sentence obtained by removing noise such as the above from claim 1. Similarly, by inputting the sentence described in the specification to the LLM 15B, a feature amount for the sentence in the specification is extracted. The similarity between the sentences is calculated from these feature amounts, and a sentence having a similarity equal to or greater than a predetermined value in the specification is specified. In the example of FIG. 9, the underlined sentence in paragraph 0020 is specified as a sentence having a similarity equal to or greater than the predetermined value.
[0063] Next, with reference to FIG. 10, the operation of the document search device 10 according to the present embodiment will be described.
[0064] Figure 10 is a flowchart showing an example of the processing flow by the document search program 15A according to the present embodiment.
[0065] First, when the document search device 10 is instructed to execute the document search process, the CPU 11 starts the document search program 15A and executes the following steps.
[0066] In step S101 of FIG. 10, the CPU 11 acquires a patent document to be the target of document search from the user terminal 30.
[0067] In step S102, as an example, as shown in FIG. 3 above, the CPU 11 uses the LLM 15B to identify valid sentences from the patent document acquired in step S101 and extracts characteristic words included in the identified valid sentences. Here, the valid sentences include, for example, at least one of the sentences described in the claims and the sentences described in the abstract.
[0068] In step S103, as an example, as shown in FIG. 4 above, the CPU 11 uses the LLM 15B to assign a label corresponding to a predetermined search item to the characteristic words extracted in step S102. Here, the search item is, for example, any one of material, property, and use case, and the label is, for example, any one of material, property, and use case.
[0069] In step S104, as an example, as shown in FIG. 5 above, the CPU 11 creates a co-occurrence network 55 representing the co-occurrence relationship of the characteristic words labeled in step S103.
[0070] In step S105, as an example, the CPU 11 receives an input of a reference word related to the label from the user via the input field 60 shown in FIG. 6 above.
[0071] In step S106, the CPU 11 searches the co-occurrence network 55 starting from the reference word received in step S105, identifies information on search items related to the reference word, and obtains a first score indicating the degree of association between the reference word and the identified information on the search items.
[0072] In step S107, the CPU 11 causes the display unit 32 of the user terminal 30 to display the information on the search items identified in step S106 in descending order of the first score, and ends a series of processes by the document search program 15A.
[0073] Note that although patent documents have been exemplified and described above, it is not limited to patent documents as described above. For example, it is also applicable to the following documents (1) to (3). These documents (1) to (3) are all documents in a form that can be extracted as valid articles. Also, the diversity of terms in the documents can be covered by learning with BERT as described above.
[0074] (1) Classification of net news article genres. Example of valid article: text within HTML (text within the body tag), Example of genre: presence or absence of association with in-house technology / self-department theme, etc. (2) Classification of academic papers. Example of valid article: text within the abstract, Example of genre: presence or absence of association with in-house technology / self-department theme, etc. (3) Classification of questionnaires. Example of valid article: text written in the free description column such as impressions, Example of genre: determination of whether it is a positive document or a negative document (applied to questionnaire result analysis).
[0075] As described above, according to this embodiment, information related to a characteristic word can be searched from the characteristic words in a document such as a patent document. Thereby, it is possible to search for a highly relevant use case, property, or material from the characteristic words representing any material, property, or use case obtained from a document such as a patent document.
[0076] In addition, in each of the above embodiments, the processing executed by the CPU by loading software (program) may be executed by various processors other than the CPU. Examples of the processor in this case include a PLD (Programmable Logic Device) whose circuit configuration can be changed after manufacture, such as an FPGA (Field-Programmable Gate Array), and a dedicated electric circuit, which is a processor having a circuit configuration dedicated to executing specific processing, such as an ASIC (Application Specific Integrated Circuit).
[0077] Also, the operation of the processor in each of the above embodiments may be achieved not only by one processor but also by a plurality of physically separated processors cooperating with each other. Also, the order of each operation of the processor is not limited to the order described in each of the above embodiments, and may be changed as appropriate.
[0078] As described above, the document search device according to the embodiment has been exemplified and explained. The embodiment may be in the form of a program for causing a computer to execute the functions of each part included in each document search device. The embodiment may be in the form of a non-transitory storage medium readable by a computer storing these programs.
[0079] In addition, the configuration of the document search device described in the above embodiment is an example, and may be changed according to the situation within the scope not departing from the gist.
[0080] Also, the flow of the program processing described in the above embodiment is an example, and unnecessary steps may be deleted, new steps may be added, or the processing order may be changed within the scope not departing from the gist.
[0081] In the above-described embodiment, the case where the processing according to the embodiment is realized by software configuration using a computer by executing a program has been described, but the present invention is not limited thereto. The embodiment may be realized by, for example, a hardware configuration or a combination of a hardware configuration and a software configuration.
[0082] Regarding the above embodiments, the following additional remarks are disclosed. (Supplementary Note 1) An acquisition unit that acquires a document to be searched; An extraction unit that identifies valid sentences from the document acquired by the acquisition unit and extracts characteristic words included in the identified valid sentences; An assignment unit that assigns a label corresponding to a predetermined search item to the characteristic words extracted by the extraction unit; A creation unit that creates a co-occurrence network representing the co-occurrence relationship of the characteristic words to which the label has been assigned by the assignment unit; An output unit that arranges and outputs information on the search items related to the label in a predetermined order based on the co-occurrence network created by the creation unit; A document search device comprising the above. (Supplementary Note 2) A reception unit that receives an input of a reference word related to the label; A search unit that searches the co-occurrence network starting from the reference word to identify information on the search items related to the reference word and obtains a first score indicating the degree of association between the reference word and the identified information on the search items; The document search device further comprising: The output unit outputs the identified information on the search items in descending order of the first score. The document search device according to Supplementary Note 1. (Supplementary Note 3) The reception unit further receives an input of a sentence including the information on the search items output by the output unit, The document search device further comprising a search unit that searches a plurality of pieces of document information related to the sentence from a document information database and obtains a second score indicating the degree of association between the sentence and each of the plurality of pieces of searched document information. The output unit outputs the plurality of pieces of document information retrieved by the retrieval unit together with the second score. The document search device according to Supplementary Note 2. (Supplementary Note 4) The apparatus further includes a classification unit that classifies important document information among the plurality of pieces of document information output by the output unit, using a classification model for classifying important document information based on predetermined classification conditions. The output unit outputs the important document information classified by the classification unit in an identifiable manner. The document search device according to Supplementary Note 3. (Supplementary Note 5) The search item is a use case. The apparatus further includes a retrieval unit that retrieves information on a compound related to the information on the use case output by the output unit from a compound information database. The output unit outputs the information on the compound retrieved by the retrieval unit instead of, or together with, the information on the use case. The document search device according to any one of Supplementary Notes 2 to 4. (Supplementary Note 6) The search item is any one of a material, a property, and a use case. The label is any one of a material, a property, and a use case. The document search device according to any one of Supplementary Notes 1 to 5. (Supplementary Note 7) The document is a patent document. The patent document includes a claim scope including claims in which the invention is described, and an abstract in which an abstract of the invention is described. The effective sentence includes at least one of the sentences described in the claims and the sentences described in the abstract. The document search device according to any one of Supplementary Notes 1 to 6. (Supplementary Note 8) The extraction unit removes at least one of specific characters, numerical values, and symbols from the sentence described in the claim or the sentence described in the abstract, and extracts the characteristic words from the sentence from which at least one of the specific characters, numerical values, and symbols has been removed. The document search device according to Supplementary Note 7. (Supplementary Note 9) A document search device acquires a document to be searched, identifies valid sentences from the acquired document, extracts characteristic words included in the identified valid sentences, assigns a label corresponding to a predetermined search item to the extracted characteristic words, creates a co-occurrence network representing the co-occurrence relationship of the characteristic words to which the label has been assigned, based on the created co-occurrence network, arranges and outputs information on the search items related to the label in a predetermined order. A document search method. (Supplementary Note 10) acquires a document to be searched, identifies valid sentences from the acquired document, extracts characteristic words included in the identified valid sentences, assigns a label corresponding to a predetermined search item to the extracted characteristic words, creates a co-occurrence network representing the co-occurrence relationship of the characteristic words to which the label has been assigned, a process of arranging and outputting information on the search items related to the label in a predetermined order based on the created co-occurrence network, A document search program for causing a computer to execute.
Explanation of Signs
[0083] 10 Document search device 11 CPU 11A Acquisition unit 11B Extraction unit 11C Assignment unit 11D Creation unit 11E Reception unit 11F Search unit 11G Output unit 11H Search Unit 11J Classification Unit 12 ROM 13 RAM 14 I / O 15 Storage Unit 15A Document Search Program 15B LLM 16 Display Unit 17 Operation Unit 18 Communication Unit 30 User Terminal 31 Control Unit 32 Display Unit 40 Document Information DB 50 Compound Information DB 100 Document Search System
Claims
1. An acquisition unit that acquires a document to be searched; An extraction unit that identifies valid sentences from the document acquired by the acquisition unit and extracts characteristic words included in the identified valid sentences; An assignment unit that assigns a label corresponding to a predetermined search item to the characteristic words extracted by the extraction unit; A creation unit that creates a co-occurrence network representing the co-occurrence relationship of the characteristic words to which the label has been assigned by the assignment unit; An output unit that arranges and outputs information on the search items related to the label in a predetermined order based on the co-occurrence network created by the creation unit; A document search device comprising the above.
2. A reception unit that receives an input of a reference word related to the label; A search unit that searches the co-occurrence network starting from the reference word to identify information on the search items related to the reference word and obtains a first score indicating the degree of association between the reference word and the identified information on the search items; Further comprising: The output unit outputs the identified information on the search items in descending order of the first score. The document search device according to Claim 1.
3. The reception unit further receives an input of a sentence including the information on the search items output by the output unit, Further comprising a search unit that searches a literature information database for a plurality of pieces of literature information related to the sentence and obtains a second score indicating the degree of association between the sentence and each of the plurality of pieces of literature information retrieved; The output unit outputs the plurality of pieces of literature information retrieved by the search unit together with the second score. The document search device according to Claim 2.
4. Further comprising a classification unit that classifies important literature information among the plurality of pieces of literature information output by the output unit using a classification model for classifying important literature information based on predetermined classification conditions; The output unit outputs the important literature information classified by the classification unit in an identifiable manner. The document search device according to Claim 3.
5. The search item is a use case, Further comprising a search unit that searches a compound information database for information on compounds related to the information on the use case output by the output unit; The output unit outputs the information on the compounds retrieved by the search unit in place of, or together with, the information on the use case. The document search device according to Claim 2.
6. The search item is any one of a material, a property, and a use case. The label is any one of material, property, and use case. The document search device according to claim 1.
7. The document is a patent document. The patent document includes a claims scope containing claims in which the invention is described, and an abstract in which an abstract of the invention is described. The effective text includes at least one of the text described in the claims and the text described in the abstract. The document search device according to claim 1.
8. The extraction unit removes at least one of specific characters, numerical values, and symbols from the text described in the claims or the text described in the abstract, and extracts the characteristic words from the text from which at least one of the specific characters, numerical values, and symbols has been removed. The document search device according to claim 7.
9. A document search device acquires a document to be searched, identifies effective text from the acquired document, extracts characteristic words included in the identified effective text, assigns a label corresponding to a predetermined search item to the extracted characteristic words, creates a co-occurrence network representing the co-occurrence relationship of the characteristic words to which the label is assigned, and based on the created co-occurrence network, arranges and outputs information on the search items related to the label in a predetermined order. A document search method.
10. acquires a document to be searched, identifies effective text from the acquired document, extracts characteristic words included in the identified effective text, assigns a label corresponding to a predetermined search item to the extracted characteristic words, creates a co-occurrence network representing the co-occurrence relationship of the characteristic words to which the label is assigned, and a process of arranging and outputting information on the search items related to the label in a predetermined order based on the created co-occurrence network A document search program for causing a computer to execute.
Citation Information
Patent Citations
Document classification device and program
JP2007323454A