Knowledge extraction apparatus and knowledge extraction method
The knowledge extraction device automates the process of ontology definition and knowledge extraction, reducing manual labor and efficiently handling ontology changes, thereby enhancing operational efficiency in extracting knowledge from documents.
Patent Information
- Application Number
- JP2023206812
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-19
AI Technical Summary
Existing knowledge extraction technologies require significant manual effort and manpower to prepare domain-specific natural language processing engines and ontologies, making it inefficient and time-consuming, especially when ontology changes are necessary.
A knowledge extraction device and method that includes an ontology definition unit, a case creation unit, a sentence input unit, a prompt creation unit, a knowledge extraction control unit, a language model, and a knowledge verification unit, which automate the process of ontology definition, case creation, and knowledge extraction, reducing the need for manual intervention.
The proposed solution significantly reduces the manual labor required for knowledge extraction and ontology changes, enabling more efficient processing and validation of extracted knowledge, thus improving overall operational efficiency.
Smart Images

Figure 2025091546000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a knowledge extraction device and a knowledge extraction method, and is suitably applicable to, for example, a knowledge extraction device related to a technique for extracting knowledge from documents such as patent documents and papers.
Background Art
[0002] Documents such as patent documents and papers describe knowledge based on the latest research results. By using the knowledge described in these documents, it is possible to grasp research trends and perform analyses based on the latest data. For example, for the purpose of developing a new material, by utilizing knowledge such as experimental data described in patent documents and papers, it is possible to create a statistical model for predicting the properties of the new material.
[0003] Documents such as patent documents and papers are often unstructured data composed of natural language, figures, tables, etc. Therefore, in order to utilize them for data analysis, etc., it is necessary to extract knowledge from the documents and convert the extracted knowledge into structured data such as table data. The above-described extraction work is often performed manually, and for example, there is a problem that it is very time-consuming because it involves reading and understanding documents such as patent documents and papers and extracting necessary knowledge. In addition, patent documents, papers, etc. are published in large quantities every day, and it is not realistic to extract such documents manually.
[0004] Under such circumstances, Patent Document 1 discloses a technique for automatically extracting useful knowledge from documents using natural language processing. In the technique disclosed in Patent Document 1, necessary information is extracted from documents based on a domain-specific natural language processing engine and a domain-specific ontology. Here, ontology refers to something that defines concepts (classes) to be extracted using natural language processing and their relationships (relations).
Prior Art Documents
Patent Documents
[0005] Patent Document 1 International Publication No. 2021 / 156684 Summary of the Invention Problems to be Solved by the Invention
[0006] In the technology disclosed in Patent Document 1, it is necessary to prepare in advance a natural language processing engine corresponding to the ontology specific to the domain. However, depending on the domain, it is not realistic to prepare a complete ontology in advance, and in many cases, it becomes necessary to change the ontology after the start of operation. If the ontology is changed, it becomes necessary to change the natural language processing engine accordingly. For the natural language engine, methods such as rule-based and machine learning are used. When such an ontology change is made, it becomes necessary to take measures such as adding and reviewing rules, creating and re-learning training data. This correspondence usually needs to be done manually by a data scientist or the like, and there is a problem that the man-hours required for the correspondence are large.
[0007] The present invention has been made in consideration of the above points, and it is intended to propose a knowledge extraction device and a knowledge extraction method that do not require a large amount of manpower and can reduce the man-hours required for dealing with ontology changes. Means for Solving the Problems
[0008] In order to solve such problems, in the present invention, an ontology definition unit that receives, from the outside, a definition of an ontology of knowledge to be extracted and outputs the definition of the ontology as ontology definition data, a case creation unit that receives, from the outside, a combination of the ontology definition data, a case sentence, and extracted knowledge extracted from the case sentence as a case, and outputs the combination as case data, a sentence input unit that receives, from the outside, a case sentence to be a target of knowledge extraction and outputs the case sentence as target sentence data, a prompt creation unit that inputs the ontology definition data, the case data, the target sentence data, and a predetermined prompt template, and outputs a knowledge extraction prompt including the ontology definition data, the case data, the target sentence data, and the predetermined prompt template, a knowledge extraction control unit that inputs the knowledge extraction prompt and knowledge extraction control setting data in which extraction conditions for the knowledge extraction are defined, and outputs a knowledge extraction command indicating that the knowledge extraction should be executed, a language model that inputs the knowledge extraction command and outputs extraction knowledge regarding the extracted knowledge, and a knowledge verification unit that inputs the extraction knowledge and verifies the validity of the extraction knowledge are provided.
[0009] Also, in the present invention, an ontology definition step in which an ontology definition unit receives, from the outside, a definition of an ontology of knowledge to be extracted and outputs the definition of the ontology as ontology definition data; a case creation step in which a case creation unit receives, from the outside, a combination of a case sentence and extracted knowledge extracted from the case sentence as a case, using the ontology definition data as an input, and outputs the combination as case data; a sentence input step in which a sentence input unit receives, from the outside, a case sentence to be the target of knowledge extraction and outputs the case sentence as target sentence data; a prompt creation step in which a prompt creation unit inputs the ontology definition data, the case data, the target sentence data, and a predetermined prompt template, and outputs a knowledge extraction prompt including the ontology definition data, the case data, the target sentence data, and the predetermined prompt template; a knowledge extraction control step in which a knowledge extraction control unit inputs the knowledge extraction prompt and knowledge extraction control setting data in which extraction conditions for the knowledge extraction are defined, and outputs a knowledge extraction command indicating that the knowledge extraction is to be executed; and a knowledge verification step in which a knowledge verification unit uses a language model, inputs the knowledge extraction command, outputs extraction knowledge regarding the extracted knowledge, and verifies the validity of the extraction knowledge.
Effect of the Invention
[0010] According to the present invention, a large amount of manual labor is not required, and the man-hours required for dealing with changes in the ontology can be reduced.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Embodiments for Carrying Out the Invention
[0012] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. (1) First Embodiment FIG. 1 is a system configuration diagram showing an example of the basic configuration of a knowledge extraction device 100 according to the first embodiment. The knowledge extraction device 100 includes an input unit 101, an output unit 102, an arithmetic processing unit 103, and a storage unit 104.
[0013] The input unit 101 is various input devices such as a keyboard, a mouse, and a touch panel. The input unit 101 is used when a user inputs some data to the knowledge extraction device 100.
[0014] The output unit 102 is an output device such as a display device. The output unit 102 displays a screen for interactive processing with the arithmetic processing unit 103.
[0015] The arithmetic processing unit 103 is, for example, a CPU (Central Processing Unit). The arithmetic processing unit 103 executes information processing in the knowledge extraction device 100. The arithmetic processing unit 103 includes an ontology definition unit 105, a case creation unit 106, a text input unit 107, a prompt creation unit 108, a knowledge extraction control unit 109, a large language model 110, an extracted knowledge verification unit 111, and a knowledge storage unit 112.
[0016] The storage unit 104 is a storage unit such as a hard disk drive (HDD) or a solid state drive (SSD: Solid State Drive). The storage unit 104 stores a knowledge base 113 and the like, which will be described later.
[0017] Next, the processing (knowledge extraction method) executed by the knowledge extraction device 100 according to the first embodiment, and the input / output data of the processing will be described. FIG. 2 shows an example of the data flow associated with the processing performed by the knowledge extraction device 100 according to the first embodiment.
[0018] The ontology definition unit 105 receives, from the outside, the definition of the ontology of the knowledge to be extracted, and outputs the definition of the ontology as ontology definition data. Here, the ontology refers to a definition of concepts (classes) to be extracted using natural language processing and their relationships (relations, properties, instances). An example of the screen displayed on the output unit 102 for the ontology definition unit 105 to receive the ontology definition from the outside will be described later.
[0019] The case creation unit 106 takes as input the ontology definition data output by the ontology definition unit 105, accepts from the outside as cases combinations of case sentences and extraction knowledge extracted from the case sentences, and outputs such combinations as case data. That is, this case data includes case sentences and extraction knowledge. On the other hand, the extraction knowledge is knowledge extracted from case sentences. The extraction knowledge is expressed, for example, in the form of an extraction knowledge graph. Although the extraction knowledge graph should originally conform to the ontology defined by the ontology definition data, there may be cases where it does not. Therefore, in this embodiment, the validity of the extraction knowledge is verified based on the extraction knowledge graph as described later.
[0020] The text input unit 107 accepts from the outside a case sentence that is the target for extracting knowledge (hereinafter referred to as "knowledge extraction") from the case data, and outputs the case sentence as target text data. The acceptance of text from the outside may be to directly accept a character string, or to specify a file in which the target text is described, the directory in which the file is stored, etc. and accept the file.
[0021] The prompt creation unit 108 takes as input the ontology definition data, case data, target text data, and a predetermined prompt template, and outputs a knowledge extraction prompt including the ontology definition data, case data, target text data, and the predetermined prompt template. Here, the prompt template describes knowledge extraction instructions and the like. Details of the prompt creation unit 108 will be described later.
[0022] The knowledge extraction control unit 109 takes as input the above-described knowledge extraction prompt and the knowledge extraction control setting data in which the extraction conditions for knowledge extraction are defined, and outputs an instruction (hereinafter referred to as "knowledge extraction instruction") indicating that knowledge extraction should be executed. Here, the knowledge extraction control setting data includes extraction conditions such as, for example, the type of the large language model 110 to be used, its parameters, and the number of attempts for knowledge extraction. The knowledge extraction instruction is formatted according to the knowledge extraction prompt and the large language model that uses the knowledge extraction control setting data. Details of the knowledge extraction prompt will be described later.
[0023] The large language model 110 is an example of a language model, and takes as input the knowledge extraction instruction and outputs the extracted knowledge (hereinafter referred to as "extracted knowledge"). The large language model 110 is a machine learning model trained to output appropriate text in response to input when some instruction or information prompting an output is input as text. The extracted knowledge graph is a knowledge graph extracted from the target case text, and originally should conform to the ontology defined by the ontology definition data.
[0024] The extracted knowledge verification unit 111 is an example of a knowledge verification unit, and takes the extracted knowledge as input and verifies the validity of the extracted knowledge. More specifically, the extracted knowledge verification unit 111 takes the extracted knowledge as input, creates, for example, an extracted knowledge graph related to the extracted knowledge, and verifies the validity of the extracted knowledge. Further, the extracted knowledge verification unit 111 takes the ontology definition data as input, and based on the extracted knowledge graph, determines whether the input extracted knowledge is valid in light of the definition content by the ontology definition data. The extracted knowledge verification unit 111 outputs the verification result. This verification result includes at least an identifier for determining whether the extracted knowledge is valid. Also, when the extracted knowledge is not valid, the extracted knowledge verification unit 111 may include the reason in the output verification result. In the present embodiment, to determine whether the extracted knowledge is valid, predetermined rules, statistical methods, etc. can be utilized. For example, it can be verified according to whether it conforms to the ontology definition data. Details of the ontology definition data will be described later.
[0025] When the extraction knowledge verification unit 111 determines that the input weekly extraction knowledge is valid, the knowledge accumulation unit 112 performs accumulation control of the extraction knowledge on the storage unit 104 having a knowledge base 113 in which the extraction knowledge is accumulated. The knowledge base 113 is, for example, a graph database. When the extraction knowledge verification unit 111 determines that the input extraction knowledge is valid, the extraction knowledge verification unit 111 accumulates the extraction knowledge in the knowledge base 113.
[0026] FIG. 3 is a diagram showing a configuration example of the prompt creation unit 108. The prompt creation unit 108 includes an ontology conversion unit 701, a case conversion unit 702, and a target text conversion unit 703 as an example of a conversion unit that converts data so that the large language model 110 can easily process it. Further, the prompt creation unit 108 includes an information integration unit 704.
[0027] The ontology conversion unit 701 takes ontology definition data as input and outputs converted ontology definition data. The converted ontology definition data is data obtained by converting the ontology definition data so that the large language model 110 can easily process it, and is described, for example, in the Turtle format of the Resource Description Framework.
[0028] The case conversion unit 702 takes case data as input and outputs converted case data. The converted case data is composed of a converted case text and converted case extraction knowledge. The converted case text is a text obtained by performing normalization processing such as removing unnecessary character strings from the case text of the case data. Also, the converted case extraction knowledge graph is data obtained by converting the case extraction knowledge so that the large language model 110 can easily process it, and is described, for example, in the Turtle format of the Resource Description Framework.
[0029] The target text conversion unit 703 takes target text data as input and outputs converted target text data. The converted target text data is a text obtained by performing normalization processing such as removing unnecessary character strings from the target text data.
[0030] The information integration unit 704 integrates the ontology definition data after conversion, the case data after conversion, the target text data after conversion, and the prompt template to create a knowledge extraction prompt and outputs this knowledge extraction prompt. Details of the knowledge extraction prompt will be described later.
[0031] FIG. 4 is a diagram showing an example of the class definition screen 200. The class definition screen 200 is a screen displayed on the output unit 102 and has a plurality of class definition input fields 298 for inputting the definition of each class (hereinafter referred to as "class definition") and an add button 299.
[0032] Each class definition input field 298 has a name input field 201, a parent class input field 202, and a definition input field 203 for inputting a definition. The name input field 201 is an input field for inputting the name of the class to be registered. The parent class input field 202 is an input field for inputting the parent class to be placed above the class input (registered) in the name input field 201. The definition input field 203 is an input field for inputting the definition of each class.
[0033] The add button 299 is a button to be operated to register each input content of each class definition input field 298 in the knowledge base 113.
[0034] FIGS. 5 and 6 are diagrams showing an example of a screen displayed on the output unit 102 for the ontology definition unit 105 to receive an ontology definition from the outside. FIG. 5 shows an example of a screen for defining the classes of the ontology. The user can input the name of the class, the parent class, the definition, etc. through this screen. FIG. 6 shows an example of a screen for defining the relationship between the classes of the ontology. The user can define the name of the relationship, the superordinate concept to inherit, the definition, the domain, the range, etc. through this screen.
[0035] FIG. 5 is a diagram showing an example of the relation definition screen 300. The relation definition screen 300 is a screen displayed on the output unit 102, and has a plurality of relation definition input fields 398 and an add button 399 for inputting definitions regarding the relevance among a plurality of classes (hereinafter referred to as "relation definition").
[0036] Each relation definition input field 398 has a name input field 301, a superordinate concept input field 302, a definition input field 303 for inputting a definition, a domain input field 304, a range input field 305, and a multiplicity input field 306.
[0037] The name input field 301 is an input field for the name of each relation definition. The superordinate concept input field 302 is an input field for inputting the superordinate concept of the relation definition of the name input in the name input field 301. The definition input field 303 is an input field for inputting a definition.
[0038] The domain input field 304 is an input field for inputting a domain such as, for example, "quality", "chemical entity", "quality". The range input field 305 is an input field for inputting a range such as, for example, a character string, "unit", "float", "chemical entity". The multiplicity input field 306 is an input field for multiplicity indicating the degree to which duplication of definition is allowed.
[0039] The add button 399 is a button to be operated for registering each input content of each relation definition input field 398 in the knowledge base 113.
[0040] FIG. 6 shows an example of a case creation screen that the case creation unit 106 displays on the output unit 102 to receive a case from the outside. The case creation screen includes a case text input unit 601, a case extraction knowledge input unit 602, a knowledge addition button 603, and a case addition button 604.
[0041] The user can register both input contents in the knowledge base 113 as a set by inputting an example text into the example text input unit 601, inputting example extraction knowledge into the example extraction knowledge input unit 602, and pressing the example addition button 604. Note that only the extraction knowledge can also be registered in the knowledge base 113 by inputting the extraction knowledge into the example extraction knowledge input unit 602 and pressing the knowledge addition button 603.
[0042] The input of the example extraction knowledge may be, for example, in the form of a triple composed of three elements: a set of instances related to each other and their relation. Also, at that time, the user may be allowed to select the instance classes and relations defined in the ontology definition data using a pull-down menu or the like.
[0043] The knowledge extraction device 100 according to the first embodiment has the above configuration, and next, an example of its operation will be described. FIG. 7 is a flowchart showing an example of the procedure of the knowledge extraction process using the knowledge extraction device 100.
[0044] First, an overview of the knowledge extraction method using the knowledge extraction device 100 will be described. The knowledge extraction method includes an ontology definition step in which the ontology definition unit 105 receives, from the outside, a definition of an ontology of knowledge to be extracted and outputs the ontology definition as ontology definition data; a case creation step in which the case creation unit 106 receives, from the outside, a combination of a case sentence and extracted knowledge extracted from the case sentence as a case using the ontology definition data as an input and outputs the combination as case data; a sentence input step in which the sentence input unit 107 receives, from the outside, a case sentence that is a target of knowledge extraction and outputs the case sentence as target sentence data; a prompt creation step in which the prompt creation unit 108 outputs a knowledge extraction prompt including the ontology definition data, the case data, the target sentence data, and a predetermined prompt template using the ontology definition data, the case data, the target sentence data, and the predetermined prompt template as inputs; and a knowledge extraction control step in which the knowledge extraction control unit 109 outputs a knowledge extraction command to execute knowledge extraction using the knowledge extraction prompt and knowledge extraction control setting data in which extraction conditions for knowledge extraction are defined as inputs. The extraction knowledge verification unit 111 uses the large language model 110 to output extraction knowledge regarding the extracted knowledge using the knowledge extraction command as an input, creates an extraction knowledge graph regarding the extraction knowledge, and verifies the validity of the extraction knowledge. More specifically, the knowledge extraction device 100 operates as follows.
[0045] The ontology definition unit 105 receives an ontology definition from the outside and outputs ontology definition data (step S301). The case creation unit 106 receives, as an input, this ontology definition data, receives a case including a case sentence and extracted knowledge from the outside, and outputs it as case data (step S302).
[0046] The sentence input unit 107 receives a target sentence from the outside and outputs it as target sentence data (step S303). The prompt creation unit 108 outputs a knowledge extraction prompt described later from this ontology definition data, case data, and target sentence data (step S304). Details of the knowledge extraction prompt will be described later (see FIG. 8).
[0047] Based on this knowledge extraction prompt, the knowledge extraction control unit 109 issues a knowledge extraction command to the large language model 110 to obtain the extracted knowledge (step S305). The extracted knowledge verification unit 111 verifies the validity of the extracted knowledge, for example, checks whether the extracted knowledge is in an appropriate format (step S306), and outputs the verification result. Details of the verification regarding the validity of the extracted knowledge will be described later.
[0048] Based on the result of the verification regarding the validity of the extracted knowledge, the extracted knowledge verification unit 111 determines whether the extracted knowledge is valid (step S307). If valid extracted knowledge is obtained, the knowledge accumulation unit 112 saves the extracted knowledge in the knowledge base 113 and ends the process (step S308).
[0049] On the other hand, if valid extracted knowledge is not obtained, the knowledge extraction control unit 109 determines whether the number of attempts of knowledge extraction is equal to or greater than the number of attempts defined in the knowledge extraction control setting data (step S309).
[0050] If the number of attempts of knowledge extraction is not equal to or greater than the number of attempts, the knowledge extraction control unit 109 issues a knowledge extraction command to the large language model 110 and obtains the extracted knowledge graph again as described above (step S305). At this time, the parameters of the large language model 110 may be changed or the verification result may be added to the knowledge extraction command.
[0051] On the other hand, if the number of attempts of knowledge extraction is equal to or greater than the number of attempts, the knowledge extraction control unit 109 presents to the outside that the knowledge extraction was not properly performed and ends the process (step S310).
[0052] FIG. 8 is a diagram showing an example of the knowledge extraction prompt 400. The knowledge extraction prompt 400 describes the knowledge extraction command 401, the above-converted ontology 402, the above-converted example text 403, the above-converted example extracted knowledge 404, the converted text data 405, etc. as text data.
[0053] The knowledge extraction command 401 corresponds to a prompt template and indicates an instruction to extract knowledge. In the knowledge extraction prompt 400, from the beginning of the text data, the knowledge extraction command 401, the converted ontology 402, the converted example sentence 403, the converted example-extracted knowledge 404, and the converted text data 405 are described in this order.
[0054] Figure 9 is a diagram showing an example of ontology definition data. In the illustrated example, classes such as "chemical entity", "quality", "unit", "mass", "unit", "gram", and "mol" are defined. Also, it is defined that the class "mass" is a subclass that inherits from the class "quality", and the classes "gram" and "mol" are subclasses that inherit from the class "unit", respectively.
[0055] Furthermore, it can be seen that the class "chemical entity" is connected by one or more strings (step String) and the property "written as", and is associated with zero or more classes of "quality" and the property "has quality", the class "quality" is associated with a single numerical value (float) and the property "has value", and is associated with zero or more and one or less classes of "unit" and the property "has unit".
[0056] Figure 10 is a diagram showing an example of the result of verifying the extracted knowledge graph using the ontology definition data shown in Figure 9. In the illustrated example, an example of the result of verification is shown in light of the definition content based on the ontology definition data shown in Figure 9 regarding the extracted knowledge graph extracted from the sentence "···nitrogen gas flow was filled with 100 g of organic solvent, N, N-dimethylpropionamide (DMPA),···" as an example of the target text data.
[0057] In this embodiment, the ontology definition data includes at least one class that is an element constituting the extracted knowledge, and a property representing an attribute that the class should have. The extracted knowledge verification unit 111 verifies the validity of the extracted knowledge according to whether the class is consistent with the property.
[0058] In the illustrated example, "Mass-01" is extracted as an instance of the class "mass", the numerical class "100" is associated with the property "has value", and two classes, "gram" and "mol", are associated with the property "has unit".
[0059] For example, according to the ontology definition data shown in FIG. 9, "Mass-01", which is an instance of the class "mass", should be associated with zero or one class "unit" and the property "has unit". However, in the extracted knowledge graph shown in FIG. 10, it is associated with two classes "unit" ("g-01:gram", "mol-01:mol"), resulting in an inconsistency with the ontology definition data.
[0060] In this embodiment, the ontology definition data includes at least one class that is an element constituting the extracted knowledge, and an instance corresponding to the class. The extracted knowledge verification unit 111 verifies the validity of the extracted knowledge according to whether the class corresponds to the instance.
[0061] For example, according to the ontology definition data shown in FIG. 9, "Mass-01", which is an instance of the class "mass", should be associated with an instance of "chemical entity". However, in the extracted knowledge graph shown in FIG. 10, it is not associated with an instance of "chemical entity", resulting in an inconsistency with the ontology definition data.
[0062] The extraction knowledge verification unit 111 determines that the extracted knowledge is valid if there is no inconsistency, for example, based on the consistency with the ontology definition data as described above, while determining that the extracted knowledge is not valid if there is an inconsistency. As described above, the extraction knowledge verification unit 111 outputs the verification result to the output unit 102.
[0063] The knowledge extraction device 100 according to the present embodiment includes an ontology definition unit 105 that receives, from the outside, the definition of the ontology of the knowledge to be extracted and outputs the ontology definition as ontology definition data, a case creation unit 106 that receives, from the outside, a combination of the ontology definition data as input, a case sentence, and the extracted knowledge extracted from the case sentence as a case, and outputs the combination as case data, a sentence input unit 107 that receives, from the outside, the case sentence to be the target of knowledge extraction and outputs the case sentence as target sentence data, a prompt creation unit 108 that inputs the ontology definition data, the case data, the target sentence data, and a predetermined prompt template, and outputs a knowledge extraction prompt 400 including the ontology definition data, the case data, the target sentence data, and the predetermined prompt template, a knowledge extraction control unit 109 that inputs the knowledge extraction prompt 400 and knowledge extraction control setting data in which the extraction conditions for knowledge extraction are defined, and outputs a knowledge extraction command indicating that knowledge extraction should be executed, a large language model 110 that inputs the knowledge extraction command and outputs the extracted knowledge regarding the extracted knowledge, and an extraction knowledge verification unit 111 that inputs the extracted knowledge and verifies the validity of the extracted knowledge.
[0064] By doing so, knowledge extraction can be performed by creating an ontology definition and a small number of cases on the user side without requiring a large amount of manpower, so that the man-hours required for responding to changes in the ontology can be reduced.
[0065] In the present embodiment, the extraction knowledge verification unit 111 creates an extraction knowledge graph regarding the extracted knowledge and verifies the validity of the extracted knowledge. By doing so, the validity of the extracted knowledge can be objectively verified by the extraction knowledge graph.
[0066] In this embodiment, the extracted knowledge verification unit 111 determines whether the input extracted knowledge is valid in light of the defined content in the ontology definition data based on the extracted knowledge graph. By doing so, it is possible to objectively verify whether the extracted knowledge is valid based on the ontology defined in the ontology definition data.
[0067] In this embodiment, the extracted knowledge verification unit 111 outputs the verification result of the validity of the extracted knowledge. By doing so, it is possible to objectively determine the validity of the extracted knowledge based on the output verification result.
[0068] The knowledge extraction device 100 according to this embodiment includes a storage unit 104 having a knowledge base 113 in which the extracted knowledge is stored when it is determined by the extracted knowledge verification unit 111 that the input extracted knowledge is valid. By doing so, it is possible to store the extracted knowledge determined to be valid.
[0069] In this embodiment, the prompt creation unit 108 includes a conversion unit (ontology conversion unit 701, case conversion unit 702, and target text conversion unit 703) that converts data so that it is easy for the large language model 110 to process, and an information integration unit 704 that integrates ontology definition data, case data, target text data, and a prompt template to create a knowledge extraction prompt 400. By doing so, it is possible to optimize the processing using the large language model 110.
[0070] In this embodiment, the ontology definition data includes at least one class that is an element constituting the extracted knowledge and a property representing the attributes that the class should have, and the extracted knowledge verification unit 111 verifies the validity of the extracted knowledge according to whether the class is consistent with the property. By doing so, if the class and the property are set in the ontology definition data in advance, it is possible to accurately determine the validity of the extracted knowledge regarding whether the class corresponds to the property based on the ontology definition data.
[0071] In this embodiment, the ontology definition data includes at least one class that is an element constituting the extracted knowledge, and an instance corresponding to the class. The extracted knowledge verification unit 111 verifies the validity of the extracted knowledge according to whether the class corresponds to the instance. By doing so, if the class and the instance are set in the ontology definition data in advance, it is possible to accurately determine the validity of the extracted knowledge regarding whether the class corresponds to the instance based on the ontology definition data.
[0072] (2) Second Embodiment The knowledge extraction device according to the second embodiment is substantially the same as the knowledge extraction device 100 according to the first embodiment except for a part. Therefore, the description of the same configuration and operation as those of the knowledge extraction device 100 according to the first embodiment is omitted. FIG. 11 is a system configuration diagram showing an example of the basic configuration of the knowledge extraction device 100a according to the second embodiment. FIG. 12 shows an example of the data flow in the second embodiment. FIG. 13 is a flowchart showing an example of the operation of the knowledge extraction device 100a according to the second embodiment.
[0073] In the second embodiment, the arithmetic processing unit 103 includes an extracted knowledge confirmation unit 1101 and an extracted knowledge correction unit 1102 in addition to the configuration in the first embodiment. Also in the second embodiment, the extracted knowledge verification unit 111 takes the ontology definition data and the extracted knowledge as inputs and outputs the verification result to the output unit 102. In the second embodiment, the verification result includes at least an identifier for determining whether the extracted knowledge is valid, and if it is not valid, the reason therefor.
[0074] As such reasons, for example, in the extracted knowledge graph shown in FIG. 10, it can be cited that "two or more 'units' are associated with 'Mass-01'" or "there is no chemical entity associated with 'Mass-01'".
[0075] The extraction knowledge verification unit 1101 takes the extraction knowledge graph and the verification result as inputs, outputs the extraction knowledge to be confirmed (hereinafter referred to as "extraction knowledge to be confirmed"), and presents it externally. This extraction knowledge verification unit 1101 takes the extraction knowledge and the above verification result as inputs, and outputs the extraction knowledge to be confirmed for which confirmation and correction are requested. The extraction knowledge to be confirmed is the knowledge in the extraction knowledge that is determined to be inappropriate by the extraction knowledge verification unit 111.
[0076] The extraction knowledge correction unit 1102 accepts the correction knowledge from the outside for the extraction knowledge to be confirmed, corrects the extraction knowledge, and outputs it as the corrected extraction knowledge. Specifically, this extraction knowledge correction unit 1102 takes the extraction knowledge as an input, accepts the correction knowledge from the outside, corrects the extraction knowledge based on it, and outputs the extraction knowledge graph based on the corrected extraction knowledge (hereinafter referred to as "corrected extraction knowledge") to the output unit 102 (step S310a). That is, in the second embodiment, instead of presenting to the outside that the knowledge extraction was not appropriately performed as in the first embodiment (step S310 in FIG. 7), an opportunity for correction as described above is provided. Here, the corrected extraction knowledge indicates the correction result for the extraction knowledge to be confirmed.
[0077] In this embodiment, when the extraction knowledge verification unit 111 determines that the corrected extraction knowledge is appropriate, the corrected extraction knowledge is registered in the knowledge base 113 (step S308).
[0078] FIG. 14 is a diagram showing an example of a knowledge correction screen. The illustrated example is an example of the screen displayed by the extraction knowledge verification unit 1101 and the extraction knowledge correction unit 1102. The knowledge correction screen is displayed to present the confirmation knowledge and accept the correction knowledge from the outside.
[0079] The illustrated knowledge correction screen has a target location field 901, a confirmation item field 902, and a completion button 903. The target location field 901 displays the target location for which the user is requested to confirm. The confirmation item field 902 includes confirmation items 902a for requesting confirmation regarding the target location displayed in the target location field 901 and an input field 902b for inputting correction items. In the input field 902b, classes (or instances) that should originally be set based on the ontology definition data are displayed in a pull-down menu form, and any of them can be selected.
[0080] The completion button 903 is a button for registering the class selected in the input field 902b and the like in the knowledge base 113.
[0081] To explain with a specific example, in the knowledge correction screen, for example, the numerical value "100" in the target text related to the extraction knowledge to be confirmed is highlighted as a potentially invalid one and is displayed as a confirmation item 902a for the user based on the verification result. The user checks the confirmation item 902a in the confirmation item field 902 and corrects, for example, the unit of the numerical value "100" in the input field 902b. Then, when the completion button 903 is pressed, the corrected information is received by the extraction knowledge correction unit 1102 as corrected extraction knowledge. The knowledge accumulation unit 112 saves the corrected extraction knowledge in the knowledge base 113.
[0082] Next, the flow of the knowledge extraction process according to the second embodiment will be described with reference to FIG. 13. Among the knowledge extraction processes shown in FIG. 13, the description of the procedures similar to the knowledge extraction process shown in FIG. 7 that has already been described will be omitted.
[0083] If, as a result of verification by the extraction knowledge verification unit 111, it is not determined that the extraction knowledge is valid, and the number of attempts of knowledge extraction is equal to or more than the number of attempts defined in the knowledge extraction control setting data, the extraction knowledge confirmation unit 1101 presents the knowledge to be confirmed, and based on the correction result from the outside, the extraction knowledge correction unit 1102 corrects the extraction knowledge graph and displays it on the output unit 102 (step S1307).
[0084] The knowledge extraction device 100a according to this embodiment includes an extraction knowledge confirmation unit 1101 that takes the extracted knowledge and the above verification result as inputs and outputs extraction knowledge to be confirmed for which confirmation and correction are requested, and an extraction knowledge correction unit 1102 that accepts correction knowledge from the outside for the extraction knowledge to be confirmed, corrects the extraction knowledge, and outputs it as corrected extraction knowledge. By doing so, since the knowledge to be confirmed is output, the correction of the extraction knowledge can be performed more easily.
[0085] Note that the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, each element described in parallel in this embodiment may be in a mode in which at least one of each element is connected in series to another element.
Industrial Applicability
[0086] The present invention can be applied to, for example, a knowledge extraction device related to a technique for extracting knowledge from documents such as patent documents and papers.
Explanation of Signs
[0087] 100, 100a... Knowledge extraction device, 102... Output unit, 104... Storage unit, 105... Ontology definition unit, 106... Case creation unit, 107... Document input unit, 108... Prompt creation unit, 109... Knowledge extraction control unit, 110... Large language model, 111... Extracted knowledge verification unit, 112... Knowledge accumulation unit, 113... Knowledge base, 400... Knowledge extraction prompt, 701... Ontology conversion unit, 702... Case conversion unit, 703... Target document conversion unit, 704... Information integration unit
Claims
1. An ontology definition unit that receives, from an external source, a definition of an ontology of knowledge to be extracted and outputs the definition of the ontology as ontology definition data; An example creation unit that receives, from an external source, a combination of the ontology definition data, example sentences, and extraction knowledge extracted from the example sentences as an example, and outputs the combination as example data; A sentence input unit that receives, from an external source, an example sentence to be the target of knowledge extraction and outputs the example sentence as target sentence data; A prompt creation unit that inputs the ontology definition data, the example data, the target sentence data, and a predetermined prompt template, and outputs a knowledge extraction prompt including the ontology definition data, the example data, the target sentence data, and the predetermined prompt template; A knowledge extraction control unit that inputs the knowledge extraction prompt and knowledge extraction control setting data in which extraction conditions for the knowledge extraction are defined, and outputs a knowledge extraction command indicating that the knowledge extraction should be executed; A language model that inputs the knowledge extraction command and outputs extraction knowledge regarding the extracted knowledge; A knowledge verification unit that inputs the extraction knowledge and verifies the validity of the extraction knowledge; A knowledge extraction device, characterized by comprising the above.
2. The knowledge verification unit: Creates an extraction knowledge graph regarding the extraction knowledge and verifies the validity of the extraction knowledge The knowledge extraction device according to claim 1, characterized by the above.
3. The knowledge verification unit: Based on the extraction knowledge graph, determines whether the input extraction knowledge is valid in light of the definition content by the ontology definition data The knowledge extraction device according to claim 2, characterized by the above.
4. The knowledge verification unit: Outputs a verification result of the validity of the extraction knowledge The knowledge extraction device according to claim 1, characterized in that
5. When the knowledge verification unit determines that the input extracted knowledge is valid, it includes a storage unit for storing the extracted knowledge The knowledge extraction device according to claim 3, characterized in that
6. The prompt creation unit A conversion unit that converts data so that the language model can easily process it, and An information integration unit that integrates the ontology definition data, the case data, the target text data, and a predetermined prompt template to create the knowledge extraction prompt The knowledge extraction device according to claim 3, characterized by comprising
7. The ontology definition data At least one class that is an element constituting the extracted knowledge, and A property representing an attribute that the class should have, and includes The knowledge verification unit Verifies the validity of the extracted knowledge according to whether the class is consistent with the property The knowledge extraction device according to claim 3, characterized in that
8. The ontology definition data At least one class that is an element constituting the extracted knowledge, and An instance corresponding to the class, and includes The knowledge verification unit Verifies the validity of the extracted knowledge according to whether the class corresponds to the instance The knowledge extraction device according to claim 3, characterized in that
9. A knowledge confirmation unit that outputs confirmation target knowledge for requesting confirmation and correction with the extracted knowledge and the verification result as inputs A knowledge modification unit that receives external modification knowledge for the knowledge to be confirmed, modifies the extracted knowledge, and outputs it as modified extracted knowledge; The knowledge extraction device according to claim 4, characterized by comprising the above.
10. An ontology definition step in which an ontology definition unit receives an external definition of an ontology of knowledge to be extracted and outputs the definition of the ontology as ontology definition data; An example creation step in which an example creation unit receives, as an input, the ontology definition data, and a combination of an example sentence and extracted knowledge extracted from the example sentence as an example from the outside, and outputs the combination as example data; A sentence input step in which a sentence input unit receives an example sentence to be the target of knowledge extraction from the outside and outputs the example sentence as target sentence data; A prompt creation step in which a prompt creation unit inputs the ontology definition data, the example data, the target sentence data, and a predetermined prompt template, and outputs a knowledge extraction prompt including the ontology definition data, the example data, the target sentence data, and the predetermined prompt template; A knowledge extraction control step in which a knowledge extraction control unit inputs the knowledge extraction prompt and knowledge extraction control setting data in which extraction conditions for the knowledge extraction are defined, and outputs a knowledge extraction command indicating that the knowledge extraction should be executed; A knowledge verification step in which a knowledge verification unit uses a language model, inputs the knowledge extraction command, outputs extraction knowledge regarding the extracted knowledge, and verifies the validity of the extraction knowledge; A knowledge extraction method, characterized by comprising the above.
Citation Information
Patent Citations
Extracting information from unstructured documents using natural language processing and conversion of unstructured documents into structured documents
WO2021156684A1