Knowledge Graph Classification Storage Construction Method and System

By establishing tables and interpreting text to obtain knowledge graph attributes, the problem of high and inaccurate construction of seed vocabulary and dictionary in the prior art is solved, and the efficient, accurate expansion and verification of knowledge graphs are achieved.

CN116304113BActive Publication Date: 2025-07-25ZHONGYUAN ENGINEERING COLLEGE
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310377351.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2025-07-25
Estimated Expiration
2043-04-07

AI Technical Summary

Technical Problem

When the existing technology expands the entity attributes of the knowledge graph, it is necessary to obtain seed vocabulary or construct a synonym dictionary, resulting in large labor and material expenditure and inaccurate expansion, and limited application areas.

Method used

By establishing the first and second tables, obtaining the properties of concepts and entities based on the interpretation text, building a third knowledge graph, and expanding the properties through logical verification, avoiding the construction of seed vocabulary or dictionary.

Benefits of technology

The accurate expansion of the knowledge graph is achieved, reducing manpower and material consumption, the expansion process is reliable and not restricted by application fields, and the extended knowledge graph is easy to verify.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116304113B_ABST
    Figure CN116304113B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for classifying and storing the construction of a knowledge graph, belonging to the technical field of data processing, including the following steps: Step S1: Determine the scenario for establishing the knowledge graph, establish a first table based on the scenario, and establish a first knowledge graph based on the first table; Step S2: Endow the domain with multiple basic attributes, link the corresponding first attribute database based on the scenario, and obtain the first attributes corresponding to the concepts; Step S3: Establish a second table based on the concepts and the entities included therein, link the corresponding second attribute database based on the concepts, obtain the second attributes corresponding to the entities based on the second attribute database, and expand the first knowledge graph into a second knowledge graph; Step S4: Generate a third explanatory text based on the second explanatory texts of different entities, obtain the third attributes corresponding to each entity, so as to expand the second knowledge graph into a third knowledge graph. The present invention not only realizes the accurate expansion of the knowledge graph, but also does not need to construct a corresponding dictionary.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method and system for classifying and storing the construction of a knowledge graph. Background Art

[0002] The knowledge graph was officially proposed by Google in 2012. It combines the theories and methods of disciplines such as applied mathematics, graphics, information visualization technology, and information science with methods such as bibliometric citation analysis and co-occurrence analysis, and uses a visual graph to vividly display the core structure, development history, frontier fields, and overall knowledge architecture of a discipline to achieve the purpose of multi-disciplinary integration.

[0003] In the prior art, there are mainly two ways to construct a knowledge graph: The first way is that technicians obtain the knowledge of the field where the knowledge graph is to be constructed from domain experts or authoritative books, and integrate the obtained knowledge with the help of experts. However, constructing a knowledge graph through artificial learning is a long process that consumes a great deal of manpower and material resources; The second way is to build a framework for fully automatic or semi-automatic construction of a knowledge graph, and then manually review it after construction to improve the construction efficiency of the knowledge graph; For example, Chinese Patent Application "CN112487212A" discloses a method and device for constructing a domain knowledge graph. This method first obtains the seed vocabulary of the target domain, uses the seed vocabulary of the target domain for vocabulary expansion until the expanded vocabulary meets the preset conditions to obtain the relevant vocabulary of the target domain; then extracts the original data corresponding to the relevant vocabulary from the existing database; constructs a knowledge graph based on the original data to generate the knowledge graph of the target domain, so that the construction of the knowledge graph does not need to rely on expert knowledge in a specific domain, effectively improving the construction efficiency of the knowledge graph; Another example is Chinese Patent Application "CN109189946B" which discloses a method for converting device fault statement descriptions into knowledge graph expressions. First, a component dictionary, a state word dictionary, a degree word dictionary, a synonym dictionary, and a relationship dictionary are constructed, and then a sentence pattern template table for knowledge graph conversion is constructed. The historical fault phenomenon descriptions in the device maintenance knowledge base for retrieval are segmented and synonym-replaced, and the sentence patterns of the fault phenomenon descriptions are found to match the sentence pattern template format, semantic fragments, and extended semantic fragments in the sentence pattern template table, so as to convert the device fault statement descriptions into one or more knowledge graphs with minimized semantic structures.

[0004] However, when using the above methods to expand the attributes of knowledge graph entities, it is necessary to first obtain seed vocabulary or construct a synonym dictionary. However, if the selection of seed vocabulary is unreasonable or the construction of the synonym dictionary is unreasonable, it will affect the attribute expansion of knowledge graph entities. Summary of the Invention

[0005] To solve the above problems, the present invention provides a method and system for classifying and storing a knowledge graph, so as to realize the expansion of the knowledge graph without constructing seed words or a dictionary.

[0006] To achieve the above invention object, the present invention proposes a method for classifying and storing a knowledge graph, including:

[0007] Step S1: Determine the scenario for establishing the knowledge graph, establish a first table based on the scenario, the first table includes fields and concepts, where one field includes multiple concepts, generate a scenario node, a field node and a concept node, connect the scenario node, the field node and the concept node to obtain a first knowledge graph, and define the connection line between the field node and the concept node as the first extension line;

[0008] Step S2: Endow the field with multiple basic attributes, use an attribute connection line to connect the field with its corresponding basic attributes, link the corresponding first attribute database based on the scenario, the first attribute database includes the first explanatory text for each concept, split each first explanatory text into an explanatory object and an explanatory statement, match the corresponding concept based on the explanatory object, obtain the first attribute corresponding to the concept based on the explanatory statement, use an attribute connection line to connect the concept with its corresponding first attribute, and store the first explanatory text corresponding to the attribute connection line;

[0009] Step S3: Establish a second table based on the concept and the entities it includes, link the corresponding second attribute database based on the concept, the second attribute database includes the second explanatory text for the entities, obtain the second attribute corresponding to the entity based on the second explanatory text, establish an entity node, use an attribute connection line to connect the entity with its corresponding second attribute, expand the first knowledge graph into a second knowledge graph, and define the connection line between the concept node and the entity node as the second extension line;

[0010] Step S4: Generate a third explanatory text based on the second explanatory texts of different entities, parse the third explanatory text, obtain the third attribute corresponding to each entity, use an attribute connection line to connect the entity with its corresponding third attribute, so as to expand the second knowledge graph into a third knowledge graph, and at the same time store the third explanatory text corresponding to the attribute connection line.

[0011] Further, in the step S4, obtaining the third attribute corresponding to each entity includes the following steps:

[0012] Step S41: Extract the entities in the first order and the second order from the second table, and define them as the first entity and the second entity respectively. Compare the second attributes included in the first entity and the second entity. If the second attributes included in the first entity and the second entity are different, then continue to execute Step S42;

[0013] Step S42: Locate the second attribute in which the first entity and the second entity are different. Based on the attribute connection line between the entity and the second attribute, extract the second explanatory text corresponding to the second entity, replace the explanatory object therein with the first entity, obtain the third explanatory text for explaining the first entity, and judge the logic of the third explanatory text. If the logic of the third explanatory text is correct, then assign the second attribute included in the third explanatory text to the first entity;

[0014] Step S43: Continue to extract the entity in the third order from the second table, and define it as the third entity. Based on Step S41 and Step S42, assign the second attribute in which the third entity and the first entity are different to the first entity. After the comparison is completed, repeat this step until the comparison between the first entity and the remaining entities in the second table is completed.

[0015] Further, in Step S42, judging whether the logic of the third explanatory text is correct includes the following steps:

[0016] Step S421: Set the association strength value based on the association strength between the field and each concept, and between the concept and each entity. Store the association strength value corresponding to the first extension line and the second extension line. Extract the concept corresponding to the first entity, obtain the first attribute corresponding to the concept, extract the explanatory statement from the third explanatory text. If there is the first attribute corresponding to the explanatory statement in the concept, then obtain the association strength value between the concept and the first entity. If the association strength value is greater than the first threshold, then judge that the logic of the third explanatory text is correct. If there is no first attribute corresponding to the explanatory statement, then continue to execute Step S422;

[0017] Step S422: Extract the basic attributes included in the domain. If there are basic attributes corresponding to the explanatory statement, continue to obtain the association strength value between the domain and the concept, and calculate the association score w between the domain and the first entity based on the first formula. The first formula is: w = x1×ω1 + x2×ω2, where x1 is the association strength value between the domain and the concept, x2 is the association strength value between the concept and the first entity, and ω1, ω2 are preset adjustment coefficients. If the association score is greater than the first threshold, it is determined that the logic of the third explanatory text is correct.

[0018] Further, the attributes of the domain and the concept are extended based on the following steps:

[0019] Obtain the second attribute included in the entity, set the extension score of the second attribute, and calculate the generalization score y of the target object based on the second formula, where the target object is the domain or the concept. The second formula is: y = δ×x1×x2, where δ represents the extension score of the second attribute. If the generalization score is greater than the second threshold, the second attribute of the entity is assigned to the target object.

[0020] Further, the knowledge graph is corrected based on the following steps:

[0021] Obtain the target nodes containing the same attributes. The target nodes include the scenario node, the domain node, the concept node, and the entity node. After connecting the target nodes containing the same attributes and the corresponding attributes to each other using auxiliary connection lines, a closed path is obtained. In the closed path, if the target nodes at both ends of the auxiliary connection line are in the same branch, the auxiliary connection line is defined as an internal connection line, otherwise it is defined as an external connection line. Calculate the construction value of the closed path based on the second formula. The second formula is: ε = n×β1 + m×β2 + k×β3, where n, m, k are the numbers of internal connection lines, external connection lines, and attribute connection lines in the closed path respectively, and β1, β2, β3 are the first value of the internal connection line, the second value of the external connection line, and the third value of the attribute connection line respectively. If the construction value is greater than the third threshold, the attributes included in the closed path are deleted to streamline the knowledge graph.

[0022] The present invention also provides a knowledge graph classification storage and construction system, which is used to implement the above-mentioned knowledge graph classification storage and construction method. The system mainly includes:

[0023] A table generation module, which is used to establish a first table and a second table. The first table includes domains and concepts, where one domain includes multiple concepts, and the second table includes concepts and entities;

[0024] The first generation module generates scenario nodes, domain nodes, and concept nodes. The scenario nodes, the domain nodes, and the concept nodes are connected to obtain a first knowledge graph. The connection line between the domain node and the concept node is defined as the first extended line. The first generation module is further based on the first attribute database corresponding to the scenario link. The first attribute database includes the first explanatory text for each concept. Based on the first explanatory text, the first attribute corresponding to the concept is obtained. The concept is connected to its corresponding first attribute using an attribute connection line, and the first explanatory text is stored corresponding to the attribute connection line.

[0025] The second generation module is based on the second attribute database corresponding to the concept link. The second attribute database includes the second explanatory text for the entity. Based on the second explanatory text, the second attribute corresponding to the entity is obtained. An entity node is established. The entity is connected to its corresponding second attribute using an attribute connection line, and the first knowledge graph is extended to a second knowledge graph. The connection line between the concept node and the entity node is defined as the second extended line.

[0026] The third generation module generates a third explanatory text based on the second explanatory text of different entities, parses the third explanatory text, obtains the third attribute corresponding to each entity, and uses an attribute connection line to connect the entity to its corresponding third attribute to extend the second knowledge graph to a third knowledge graph. At the same time, the third explanatory text is stored corresponding to the attribute connection line.

[0027] The parsing module splits each of the first explanatory texts into an explanatory object and an explanatory statement, matches the corresponding concept based on the explanatory object, obtains the first attribute corresponding to the concept based on the explanatory statement, and is based on the second attribute database corresponding to the concept link. The second attribute database includes the second explanatory text for the entity. The parsing module also obtains the second attribute corresponding to the entity based on the second explanatory text. The parsing module also parses the third explanatory text to obtain the third attribute corresponding to each entity.

[0028] Compared with the prior art, the beneficial effects of the present invention are at least as follows:

[0029] The present invention first establishes a first table and a second table to construct an initial architecture of a knowledge graph, and then searches for corresponding explanatory texts from the corresponding databases based on concepts and entities. By disassembling the explanatory texts, the attributes of the concepts and entities are obtained. On this basis, the attributes of the concepts and entities are further expanded by constructing a third explanatory text, so as to further improve the knowledge graph. And during the expansion process, the logic of the third explanatory text is verified to determine whether the attribute expansion of the entity is correct, so as to ensure the accuracy of the expanded knowledge graph. Finally, the present invention also stores the explanatory texts corresponding to the attribute connection lines, so as to facilitate manual verification after the system automatically constructs the knowledge graph. The present invention not only realizes the accurate expansion of the knowledge graph, but also does not need to construct a corresponding dictionary. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a flowchart of the steps of the method for constructing and classifying and storing a knowledge graph of the present invention;

[0031] Figure 2 is a schematic diagram of the principle of constructing a knowledge graph of the present invention;

[0032] Figure 3 is a schematic diagram of the structure of the system for constructing and classifying and storing a knowledge graph of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of the present application, the first xx script may be referred to as the second xx script, and similarly, the second xx script may be referred to as the first xx script.

[0035] As Figure 1 shown, a method for constructing and classifying and storing a knowledge graph includes:

[0036] Step S1: Determine the scenario for establishing a knowledge graph, establish a first table based on the scenario. The first table includes fields and concepts, where one field includes multiple concepts, generate scenario nodes, field nodes and concept nodes, connect the scenario nodes, field nodes and concept nodes to obtain a first knowledge graph, and define the connection line between the field node and the concept node as the first extension line;

[0037] Step S2: Endow multiple basic attributes to the domains, connect the domains with their corresponding basic attributes using attribute connection lines, link the corresponding first attribute database based on the scenarios. The first attribute database includes the first explanatory texts for each concept. Split each first explanatory text into an explanatory object and an explanatory statement, match the corresponding concept based on the explanatory object, obtain the first attribute corresponding to the concept based on the explanatory statement, connect the concept with its corresponding first attribute using an attribute connection line, and store the first explanatory text corresponding to the attribute connection line;

[0038] In this embodiment, an example is given by constructing a biological scenario. For example, Figure 2 As shown, after determining to construct a biological scenario, first establish the scenario node of the knowledge graph, that is, the biological node in the figure. Then, construct the first table based on the biological scenario. In this embodiment, the first table includes two domains: animals and plants. Accordingly, generate animal nodes and plant nodes. Further, animals include two concepts: felines and canines. Correspondingly, generate feline nodes and canine nodes in the knowledge graph; Since the knowledge included in the first table is relatively basic, it can be quickly established by manual means or by using web crawler technology to capture information; After the first table is established, endow attributes to each domain and concept to improve the knowledge graph.

[0039] Since the domain has a very broad scope, its basic attributes can be obtained by consulting materials. For example, the basic attribute of animals is reproduction, which can be quickly obtained by consulting basic materials; When determining the attributes included in a concept, consult the corresponding first attribute database through domain terms. The first attribute database includes the text explanations for each concept in this domain. For example, extract the text "There is a spine in felines" from the first attribute database. Then, split the first explanatory text into an explanatory object and an explanatory statement. In this text, the explanatory object is felines, and the explanatory statement is "There is a spine". Therefore, extract the first attribute "There is a spine". The specific splitting and extraction methods can refer to existing semantic analysis models and will not be elaborated here; Match the felines in the knowledge graph based on the explanatory object and generate a first attribute node. Connect it with the feline node by using an attribute connection line, thereby endowing the feline node with this attribute; And since the first explanatory text is stored corresponding to the attribute connection line, subsequent personnel can conveniently find the text that generates this attribute when verifying the knowledge graph, and verify the concepts and attributes connected at both ends of it, which not only improves the verification efficiency but also improves the verification accuracy.

[0040] Step S3: Establish a second table based on the concept and the entities it includes, link the corresponding second attribute database based on the concept. The second attribute database includes the second explanatory text for the entities. Obtain the second attributes corresponding to the entities based on the second explanatory text, establish entity nodes, use attribute connection lines to connect the entities with their corresponding second attributes, expand the first knowledge graph into a second knowledge graph, and define the connection line between the concept node and the entity node as the second expansion line;

[0041] The method of obtaining entities through concepts can also be based on manual or crawler acquisition. For example, in this embodiment, two entities, namely tiger and cheetah, are obtained under the concept of Felidae. Then, the corresponding second attribute database is consulted through the term "Felidae". The second attribute database includes the text explanations for each entity under this concept. The method of extracting and analyzing the second explanatory text is the same as that of the first explanatory text, which will not be elaborated here. Through this step, the concept can be further expanded, thereby expanding the first knowledge graph into a second knowledge graph.

[0042] Step S4: Generate a third explanatory text based on the second explanatory texts of different entities, parse the third explanatory text, obtain the third attributes corresponding to each entity, and use attribute connection lines to connect the entities with their corresponding third attributes to expand the second knowledge graph into a third knowledge graph, and at the same time store the third explanatory text and the attribute connection lines in a corresponding manner.

[0043] In this embodiment, by extracting data from the constructed knowledge graph, a third explanatory text is constructed, thereby obtaining the third attributes of the entities, so as to achieve the purpose of expanding the knowledge graph. And since the data is extracted from the original data of the knowledge graph, there is no need to determine seed words or construct a dictionary.

[0044] It should be particularly noted that by obtaining data from the knowledge graph itself, the knowledge graph can be expanded without constructing seed words or a dictionary.

[0045] The general way of expanding the attributes of the knowledge graph in the prior art is as follows: First, a knowledge graph is constructed based on a text library, then the concepts and entities included in the knowledge graph are obtained, a corresponding synonym dictionary is constructed, the concepts and entities are expanded through the synonym dictionary, and then based on the expanded synonyms, relevant descriptive texts are retrieved from the database, and the attributes are extracted from them and assigned to the entities; for example, obtaining the synonym "diarrhea" of "loose stools", obtaining the attributes of "diarrhea" and assigning them to "loose stools", so as to expand the attributes of "loose stools"; however, constructing a synonym dictionary takes a certain amount of time, and the way of expanding synonyms often obtains the existing attributes of the entities, resulting in the failure of expansion; in addition, not all knowledge graphs in all fields can be expanded using synonyms, so the application field of this method is limited.

[0046] Based on the above analysis, the present invention proposes the following steps to expand the knowledge graph:

[0047] Step S41: Extract the entities in the first order and the second order from the second table, and define them as the first entity and the second entity respectively. Compare the second attributes included in the first entity and the second entity. If the second attributes included in the first entity and the second entity are different, then continue to execute step S42;

[0048] For example, when constructing the second table, the entity in the first order is a shark, and the entity in the second order is a frog. The shark includes the attributes "carnivorous" and "aquatic", and the frog includes the attributes "carnivorous", "aquatic" and "terrestrial". Then, after comparison, there are different second attributes between the first entity and the second entity, that is, "terrestrial". At this time, step S42 is continued. In addition, if the entities in the first order and the second order have exactly the same attributes, then continue to select the entity in the third order for comparison.

[0049] Step S42: Locate the second attribute that is different between the first entity and the second entity. Based on the attribute connection line between the entity and the second attribute, extract the second explanatory text corresponding to the second entity, replace the explanatory object in it with the first entity, obtain the third explanatory text for explaining the first entity, and judge the logic of the third explanatory text. If the logic of the third explanatory text is correct, then assign the second attribute included in the third explanatory text to the first entity;

[0050] Since the second explanatory text is stored corresponding to the attribute connection line, after locating the different attributes of the first entity and the second entity, the storage location of the second explanatory text can be directly located, and the source text of the second entity attribute can be quickly extracted from it. For example, the second explanatory text here is "A frog can live on land". In particular, the semantic analysis model extracts "can live on land" as terrestrial; after obtaining the second explanatory text, replace the explanatory object of the second explanatory text with the first entity to obtain the third explanatory text. For example, in this embodiment, the third explanatory text obtained after replacement is "A shark can live on land"; then judge its logic. If the logic is correct, then assign the attribute of "terrestrial" to the shark. If the logic is incorrect, then discard the attribute. In practical applications, since the text does not describe all the attributes of the entity, the text description is usually the prominent attributes of the entity, but this attribute is usually ignored in other entities; for example, it is usually described that the reason why an orangutan can stand is that it has a spine, but a shark also has a spine, but the text usually describes its carnivorousness and bloodthirstiness. Therefore, through this step, the ignored "spine" attribute can be assigned to the shark.

[0051] Step S43: Continue to extract the entity in the third order from the second table, define it as the third entity, and based on Step S41 and Step S42, assign the second attribute that is different from the first entity to the first entity. After the comparison is completed, repeat this step until the comparison between the first entity and the remaining entities in the second table is completed.

[0052] After the comparison between the first entity and the second entity is completed, continue to extract the entity in the third order from the second table. For example, if the entity in the third order in the second table is an orangutan, then continue to compare whether the attributes contained in the shark and the orangutan are the same. After the comparison is completed, continue to compare the entities in the fourth order and later in the second table, and repeat this process until the shark is compared with the entity in the last order in the second table.

[0053] Since the present invention first constructs the first table and the second table, relatively perfect concepts and entities can be obtained before constructing the knowledge graph, so there is no need to use the synonym method for expansion; in addition, through the above steps, the attributes of each entity in the knowledge graph can be quickly expanded, and there will be no duplicate attributes after the entity is expanded. Since the expansion of the attributes is carried out on the basis of the original knowledge graph, it will not be restricted by the application field.

[0054] When making a judgment on the third explanatory text, the conventional idea is to perform semantic retrieval within the Internet based on the third explanatory text, and capture the correctness judgment of the third explanatory text from the Internet. For example, based on "sharks can live on land" for semantic retrieval, if relevant descriptions are captured and the answer result is incorrect, then the logic of the third explanatory text can be judged; however, the information existing in the Internet is not always correct, so certain means are still needed to judge the result after obtaining the information. For example, the correct result is determined based on the voting method, which will lead to a complex and long logical judgment process; in addition, if no relevant information is retrieved, it will also lead to the inability to judge the correctness of the third explanatory text. Therefore, the present invention proposes the following steps to judge the logic of the third explanatory text.

[0055] Step S421: Set the association strength value based on the association strength between the domain and each concept, and between the concept and each entity, store the association strength value corresponding to the first extension line and the second extension line, extract the concept corresponding to the first entity, obtain the first attribute corresponding to the concept, extract the explanatory statement from the third explanatory text. If there is a first attribute corresponding to the explanatory statement in the concept, obtain the association strength value between the concept and the first entity. If the association strength value is greater than the first threshold, it is judged that the logic of the third explanatory text is correct. If there is no first attribute corresponding to the explanatory statement, continue to execute Step S422;

[0056] Refer to Figure 2, for example, under the concept of feline, there are entities tiger and cheetah. The cheetah includes the attributes of "agile" and "strong". The third explanatory statement for obtaining the "agile" attribute is "The cheetah is very agile". After replacing it with the entity "tiger", the third explanatory statement "The tiger is very agile" is obtained. When judging the logic of this third explanatory statement, first obtain the attributes included in the upper concept "feline" of the tiger. If the feline also includes the attribute of agility, then obtain the association strength value between the tiger and the feline. If the association strength value is greater than the first threshold, it indicates that the association between the concept and the entity is very strong, and the entity subordinate to the concept is likely to have the same attributes as the concept. At this time, judge that the third explanatory statement is correct and assign the "agile" attribute to the "tiger" entity.

[0057] Step S422: Extract the basic attributes included in the domain. If there are basic attributes corresponding to the explanatory statement, continue to obtain the association strength value between the domain and the concept, and calculate the association score w between the domain and the first entity based on the first formula. The first formula is: w = x1×ω1 + x2×ω2, where x1 is the association strength value between the domain and the concept, x2 is the association strength value between the concept and the first entity, and ω1, ω2 are preset adjustment coefficients. If the association score is greater than the first threshold, judge that the logic of the third explanatory text is correct.

[0058] Specifically, the value ranges of x1 and x2 are between 0 and 1, and the larger the value, the stronger the association between the two. For example, the association between feline and tiger is relatively strong; on this basis, if it is necessary to assign the attribute of "swimming" to the "tiger", and the concept of "feline" does not include the attribute of "swimming", then continue to obtain the basic attributes included in the domain. If the attributes of animals include "swimming", then continue to obtain the association strength value between "animal" and "feline". Since both animal and feline are relatively broad concepts, the association strength value between them is relatively small; calculate the association score between "tiger" and "animal" through the first formula. By setting preset adjustment coefficients, the association strength value between the concept and the first entity can be adjusted to prevent it from being too large and affecting the final relationship score; if the association score is less than the first threshold, it indicates that the association between "tiger" and "animal" is very small, and it is very likely that an error will occur when assigning the attribute of "swimming" to the "tiger". Therefore, judge that the logic of the third explanatory text is incorrect and abandon the assignment of the attribute. If the domain does not contain the attribute to be assigned to the entity, then use the semantic retrieval method for judgment.

[0059] Through the above steps, the logic of the third explanatory text can be quickly judged before using semantic retrieval, and by setting the relationship score, the judgment result has high accuracy, so as to quickly complete the expansion of the knowledge graph.

[0060] Specifically, the present invention also expands the attributes of the domain and concept based on the following steps.

[0061] Expand the attributes of the domain and concept based on the following steps to obtain the fourth knowledge graph:

[0062] Obtain the second attribute included in the entity, set the expansion score of the second attribute, and calculate the generalization score y of the target object based on the second formula, where the target object is a domain or concept, and the second formula is: y = δ × x1 × x2, where δ represents the expansion score of the second attribute. If the generalization score is greater than the second threshold, the second attribute of the entity is assigned to the target object.

[0063] In this embodiment, each attribute has a different expansion score. For example, there is an attribute "omnivorous", and animals are divided into carnivorous, herbivorous and omnivorous, so the expansion score of this attribute is relatively high. Another example is that there is also an attribute "yellow", and there are fewer animals with this attribute, so the expansion score of this attribute is relatively low. After obtaining the expansion score of the attribute, calculate the generalization score of this attribute for the upper concept or domain of the entity through the second formula. For example, if the attribute "omnivorous" is to be assigned to the concept "Felidae", the second formula is y = δ × x1. If the calculation result is greater than the second threshold, the attribute "omnivorous" is assigned to Felidae. If the attribute "omnivorous" is to be assigned to the domain "animal", the second formula is y = δ × x1 × x2. If the calculation result is less than the second threshold, the omnivorous attribute is not assigned to animals. Therefore, the meaning of the second formula is that the farther the target object to be expanded is from the entity, the smaller the probability that it may have the same attribute.

[0064] After obtaining the fourth knowledge graph, when there are multiple entities in the knowledge graph, or there are the same attributes among the domain, concept and entity, the generated graph model will have a relatively high network complexity. The original intention of constructing the knowledge graph graph model is to construct the relationships between various objects, as well as the unique attributes of each object, and display them in the form of a structure diagram. Therefore, an overly complex network structure will, on the one hand, reduce the expression clarity of the knowledge graph, and on the other hand, it will also cause a large number of basic attributes to exist in the data structure of the knowledge graph, thereby increasing the data volume of the knowledge graph. Therefore, the graph model of the knowledge graph is streamlined based on the following steps.

[0065] Obtain target nodes with the same attributes. The target nodes include scenario nodes, domain nodes, concept nodes, and entity nodes. After using auxiliary connection lines to connect the target nodes with the same attributes and their corresponding attributes to obtain a closed path, in the closed path, if the target nodes at both ends of the auxiliary connection line are under the same branch, then define this auxiliary connection line as an internal connection line; otherwise, define it as an external connection line. Calculate the construction value of the closed path based on the second formula. The second formula is: ε = n×β1 + m×β2 + k×β3, where n, m, and k are the numbers of internal connection lines, external connection lines, and attribute connection lines in the closed path respectively, and β1, β2, and β3 are the first value of the internal connection line, the second value of the external connection line, and the third value of the attribute connection line respectively. If the construction value is greater than the third threshold, delete the attributes included in the closed path to streamline the knowledge graph.

[0066] The above steps are explained below. As Figure 2 shown, if the cat family, dog family, tiger, cheetah, and wolf all have a certain attribute, then use auxiliary connection lines to connect the above five nodes in sequence, and then connect each node with the attribute respectively to obtain attribute connection lines, thereby forming a closed path, as Figure 2 shown by the dotted line in; then calculate the closed path value through the second formula. For example, in the above-connected closed path, there are 9 auxiliary connection lines. Among them, since the dog family and the cat family are both in the animal domain, the auxiliary connection line between them is an internal connection line, while the wolf and the cheetah belong to the concepts of the dog family and the cat family respectively, so the auxiliary connection line between them is an external connection line, and in this embodiment, the second value of the external connection line is greater than the first value of the internal connection line, and the first value is greater than the third value; then calculate the score of the closed path through the second formula. If the score is greater than the fourth threshold, it indicates that the attribute is too basic and broad, and its significance in constructing the knowledge graph is small, so it is removed.

[0067] The present invention first establishes a first table and a second table to construct the initial architecture of the knowledge graph, and then searches for corresponding explanatory texts from the corresponding databases based on concepts and entities. By disassembling the explanatory texts, the attributes of the concepts and entities are obtained; on this basis, the attributes of the concepts and entities are further expanded by constructing a third explanatory text to further improve the knowledge graph. And during the expansion process, by verifying the logic of the third explanatory text, it is determined whether the attribute expansion of the entity is correct, so as to ensure the accuracy of the expanded knowledge graph; finally, the present invention also stores the explanatory text and the attribute connection line in a corresponding manner, so that it is convenient for manual verification after the system automatically constructs the knowledge graph; in summary, the present invention not only realizes the accurate expansion of the knowledge graph, but also makes the expanded knowledge graph easy to verify.

[0068] As Figure 3As shown, the present invention also provides a knowledge graph classification storage construction system, which is used to implement the above-mentioned knowledge graph classification storage construction method. The system mainly includes:

[0069] A table generation module for establishing a first table and a second table. The first table includes fields and concepts, where one field includes multiple concepts, and the second table includes concepts and entities.

[0070] A first generation module that generates scenario nodes, field nodes, and concept nodes. The scenario nodes, field nodes, and concept nodes are connected to obtain a first knowledge graph. Define the connection line between the field node and the concept node as the first extension line. The first generation module is also based on the first attribute database corresponding to the scenario link. The first attribute database includes the first explanatory text for each concept. Based on the first explanatory text, obtain the first attribute corresponding to the concept, use an attribute connection line to connect the concept with its corresponding first attribute, and store the first explanatory text corresponding to the attribute connection line.

[0071] A second generation module, based on the second attribute database corresponding to the concept link. The second attribute database includes the second explanatory text for the entity. Based on the second explanatory text, obtain the second attribute corresponding to the entity, establish an entity node, use an attribute connection line to connect the entity with its corresponding second attribute, and expand the first knowledge graph into a second knowledge graph. Define the connection line between the concept node and the entity node as the second extension line.

[0072] A third generation module that generates a third explanatory text based on the second explanatory text of different entities, parses the third explanatory text, obtains the third attribute corresponding to each entity, uses an attribute connection line to connect the entity with its corresponding third attribute, so as to expand the second knowledge graph into a third knowledge graph, and at the same time store the third explanatory text corresponding to the attribute connection line.

[0073] An analysis module that splits each first explanatory text into an explanatory object and an explanatory statement, matches the corresponding concept based on the explanatory object, obtains the first attribute corresponding to the concept based on the explanatory statement, based on the second attribute database corresponding to the concept link. The second attribute database includes the second explanatory text for the entity. The analysis module also obtains the second attribute corresponding to the entity based on the second explanatory text. The analysis module also parses the third explanatory text to obtain the third attribute corresponding to each entity.

[0074] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0075] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0076] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent of the present invention should be subject to the appended claims.

[0077] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for constructing a knowledge graph for classified storage, characterized in that Including: Step S1: Determine the scenario for building the knowledge graph, build a first table based on the scenario, the first table includes domains and concepts, where one domain includes multiple concepts, generate a scenario node, domain nodes and concept nodes, connect the scenario node, the domain nodes and the concept nodes to obtain a first knowledge graph, and define the connection line between the domain node and the concept node as the first extension line; Step S2: Endow the domains with multiple basic attributes, use attribute connection lines to connect the domains with their corresponding basic attributes, link the corresponding first attribute database based on the scenario, the first attribute database includes first explanation texts for each concept, split each first explanation text into an explanation object and an explanation statement, match the corresponding concept based on the explanation object, obtain the first attribute corresponding to the concept based on the explanation statement, use an attribute connection line to connect the concept with its corresponding first attribute, and store the first explanation text corresponding to the attribute connection line; Step S3: Build a second table based on the concepts and the entities they include, link the corresponding second attribute database based on the concepts, the second attribute database includes second explanation texts for the entities, obtain the second attributes corresponding to the entities based on the second explanation texts, build entity nodes, use attribute connection lines to connect the entities with their corresponding second attributes, expand the first knowledge graph into a second knowledge graph, and define the connection line between the concept node and the entity node as the second extension line; Step S4: Generate a third explanation text based on the second explanation texts of different entities, parse the third explanation text, obtain the third attributes corresponding to each entity, use attribute connection lines to connect the entities with their corresponding third attributes to expand the second knowledge graph into a third knowledge graph, and at the same time store the third explanation text corresponding to the attribute connection lines.

2. The method for constructing the classification storage of the knowledge graph according to claim 1, characterized in that In step S4, obtaining the third attributes corresponding to each entity includes the following steps: Step S41: Extract the entities in the first order and the second order from the second table, and define them as the first entity and the second entity respectively. Compare the second attributes included in the first entity and the second entity. If the second attributes included in the first entity and the second entity are different, then continue to execute step S42; Step S42: Locate the second attribute in which the first entity is different from the second entity. Based on the attribute connection line between the entity and the second attribute, extract the second explanation text corresponding to the second entity, replace the explanation object therein with the first entity to obtain the third explanation text for explaining the first entity, judge the logic of the third explanation text. If the third explanation text is logically correct, endow the first entity with the second attribute included in the third explanation text; Step S43: Continue to extract the entity in the third order from the second table, defined as the third entity. Based on Step S41 and Step S42, assign the second attribute that the third entity is different from the first entity to the first entity. After the comparison is completed, repeat this step until the comparison between the first entity and the remaining entities in the second table is completed.

3. The knowledge graph classification storage construction method according to claim 2, wherein In Step S42, judging whether the logic of the third explanatory text is correct includes the following steps: Step S421: Set the association strength value based on the association strength between the domain and each concept, and between the concept and each entity. Store the association strength value corresponding to the first extension line and the second extension line. Extract the concept corresponding to the first entity, obtain the first attribute corresponding to the concept, extract the explanatory statement from the third explanatory text. If there is the first attribute corresponding to the explanatory statement in the concept, obtain the association strength value between the concept and the first entity. If the association strength value is greater than the first threshold, it is judged that the logic of the third explanatory text is correct. If there is no first attribute corresponding to the explanatory statement, continue to execute Step S422; Step S422: Extract the basic attributes included in the domain. If there is a basic attribute corresponding to the explanatory statement, continue to obtain the association strength value between the domain and the concept. Calculate the association score w between the domain and the first entity based on the first formula. The first formula is: w = x1×ω1 + x2×ω2, where x1 is the association strength value between the domain and the concept, x2 is the association strength value between the concept and the first entity, and ω1, ω2 are preset adjustment coefficients. If the association score is greater than the first threshold, it is judged that the logic of the third explanatory text is correct.

4. The knowledge graph classification storage construction method according to claim 3, characterized in that Expand the attributes of the domain and the concept based on the following steps: Obtain the second attribute included in the entity, set the expansion score of the second attribute, and calculate the generalization score y of the target object based on the second formula. The target object is the domain or the concept. The second formula is: y = δ×x1×x2, where δ represents the expansion score of the second attribute. If the generalization score is greater than the second threshold, assign the second attribute of the entity to the target object.

5. The knowledge graph classification storage construction method according to claim 3 or 4, characterized in that, Revise the knowledge graph based on the following steps: Obtain target nodes with the same attributes, where the target nodes include the scenario nodes, the domain nodes, the concept nodes, and the entity nodes. After connecting the target nodes with the same attributes and their corresponding attributes using auxiliary connection lines, a closed path is obtained. In the closed path, if the target nodes at both ends of the auxiliary connection line are under the same branch, then define this auxiliary connection line as an internal connection line; otherwise, define it as an external connection line. Calculate the construction value of the closed path based on the second formula, where the second formula is: ε = n×β1 + m×β2 + k×β3, where n, m, and k are the numbers of internal connection lines, external connection lines, and attribute connection lines in the closed path respectively, and β1, β2, and β3 are the first value of the internal connection line, the second value of the external connection line, and the third value of the attribute connection line respectively. If the construction value is greater than the third threshold, delete the attributes included in the closed path to streamline the knowledge graph.

6. A knowledge graph classification storage construction system for implementing the knowledge graph classification storage construction method according to any one of claims 1-5, characterized in that, Including: A table generation module for establishing a first table and a second table. The first table includes domains and concepts, where one domain includes multiple concepts, and the second table includes concepts and entities; A first generation module that generates scenario nodes, domain nodes, and concept nodes. The scenario nodes, the domain nodes, and the concept nodes are connected to obtain a first knowledge graph. Define the connection line between the domain node and the concept node as the first extension line. The first generation module also bases on the first attribute database corresponding to the scenario link. The first attribute database includes the first explanatory text for each concept. Based on the first explanatory text, obtain the first attributes corresponding to the concepts, use attribute connection lines to connect the concepts with their corresponding first attributes, and store the first explanatory text corresponding to the attribute connection lines; A second generation module that bases on the second attribute database corresponding to the concept link. The second attribute database includes the second explanatory text for entities. Based on the second explanatory text, obtain the second attributes corresponding to the entities, establish entity nodes, use attribute connection lines to connect the entities with their corresponding second attributes, and expand the first knowledge graph into a second knowledge graph. Define the connection line between the concept node and the entity node as the second extension line; A third generation module that generates a third explanatory text based on the second explanatory text of different entities, parses the third explanatory text, obtains the third attributes corresponding to each entity, uses attribute connection lines to connect the entities with their corresponding third attributes to expand the second knowledge graph into a third knowledge graph, and at the same time stores the third explanatory text corresponding to the attribute connection lines; The parsing module splits each of the first explanatory texts into an explanatory object and an explanatory statement, matches a corresponding concept based on the explanatory object, obtains a first attribute corresponding to the concept based on the explanatory statement, links to a corresponding second attribute database based on the concept, the second attribute database includes second explanatory texts of entities, the parsing module also obtains second attributes corresponding to the entities based on the second explanatory texts, and the parsing module also parses the third explanatory text to obtain third attributes corresponding to each entity.

Citation Information

Patent Citations

  • A method for converting equipment fault statements into knowledge graph representations

    CN109189946B

  • Domain knowledge graph construction method and device

    CN112487212A

  • Knowledge graph construction method and device, electronic equipment and storage medium

    CN111324609A

  • Knowledge graph construction method and device based on typhoid theory

    CN114996477A