Knowledge exchange method and system based on knowledge graph

By performing deviation interference outlier analysis and redundant information interference calculation on triple structure data in the knowledge graph, the problem of poor consistency of entity information is solved, and high-quality construction of knowledge graphs and trusted knowledge exchange are realized.

CN120278252BActive Publication Date: 2025-08-15BEIJING ZHONGWEI SHENGDING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510764642.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-08-15
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The prior art fails to fully consider the error interference caused by redundant information characteristics in the entity information during the knowledge graph construction process, resulting in poor consistency of entity information, affecting the quality of knowledge graph construction and the credibility of knowledge exchange.

Method used

By extracting entity information, relationship and attribute information of the knowledge database, vectorization and clustering analysis of triple structure data, calculating deviation interference outliers, obtaining abnormal correlation feature values and redundant information interference, and correcting entity information and building a knowledge graph.

Benefits of technology

It improves the quality of knowledge graph construction and the credibility of knowledge exchange, ensures the consistency of entity information, and improves the reliability of knowledge exchange.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278252B_ABST
    Figure CN120278252B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of knowledge graph analysis and exchange, and specifically to a knowledge exchange method and system based on a knowledge graph, the method comprising: extracting entity information, relationships between entities, and attribute information of entities from textual materials in a knowledge database to obtain triple structure data; analyzing the difference in similarity between each triple structure data and other triple structure data bag-of-words vectors to obtain the deviation interference anomaly value of each triple structure data; obtaining the abnormal correlation feature value between two entities in each triple structure data, and calculating the redundant information interference degree of the target entity to obtain the entity to be corrected for information extraction, and correcting the entity to be corrected, constructing a knowledge graph based on the corrected entity information, building a knowledge exchange system, and realizing knowledge exchange based on the knowledge graph. The present application can improve the quality of knowledge graph construction and the credibility of knowledge exchange based on the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of knowledge graph analysis and exchange technology, and specifically to a knowledge exchange method and system based on knowledge graph. Background Art

[0002] With the increasing application of natural language processing technology in various professional fields, the demand for knowledge exchange between multiple professional fields is growing, in order to meet the communication and collaboration in multiple professional fields, and ultimately realize knowledge exchange and information sharing between professional fields. However, since the knowledge exchange process involves the complex interaction of information and professional knowledge in multiple professional fields, the complexity of knowledge exchange between different professional fields is relatively high, and it will have a strong hindering effect on knowledge exchange between different professional fields. At this stage, by using the knowledge graph construction method to process and integrate information and professional knowledge in multiple professional fields, and building a knowledge exchange system through the knowledge graph, the efficiency of knowledge exchange between different professional fields can be effectively improved, and the rigor of knowledge exchange can be enhanced.

[0003] Existing technologies use knowledge graph construction to process and integrate information and expertise from multiple professional fields, effectively improving the efficiency of knowledge exchange between different professional fields. However, during the knowledge graph construction process, due to the diversity of professional information sources, the extracted entity information is prone to contain a large amount of redundant information fragments. Existing technologies do not fully consider the error interference caused by redundant information features within the entity information to verify and correct the extracted entity information, which affects the consistency of the entity information, thereby reducing the quality of knowledge graph construction and the credibility of knowledge exchange based on knowledge graphs. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of this application is to provide a knowledge exchange method and system based on knowledge graph. The technical solutions adopted are as follows:

[0005] The present invention provides a method for exchanging knowledge based on a knowledge graph, including the following steps:

[0006] Extract entity information, relationship between entities and attribute information of entities from text data in knowledge database to obtain triple structure data;

[0007] Each triple structure data is vectorized to extract each bag-of-words vector, and then cluster analysis is performed. Based on the average level of bag-of-words vector similarity between each triple structure data and other triple structure data in the text cluster, as well as the difference in bag-of-words vector similarity, the deviation interference outlier value of each triple structure data is obtained;

[0008] The triple structure data is divided based on the deviation interference outlier value to obtain a triple test set, the attribute word frequency vectors of the two entities of each triple structure data in the triple test set are extracted, the correlation between the two attribute word frequency vectors is analyzed, and the deviation interference outlier value of each triple structure data is combined to obtain the abnormal correlation feature value between the two entities of each triple structure data;

[0009] Determine the target entity of each triple structure data according to the change of the attribute word frequency vectors of the two entities in each triple structure data in the triple test set, and obtain the redundant information interference degree of the target entity of each triple structure data in the triple test set according to the difference degree of change of the attribute word frequency vector between the target entity and the other entity, combined with the abnormal correlation feature value;

[0010] Based on the redundant information interference degree, the target entities of all triple structure data in the triple test set are divided to obtain the entities to be corrected after information extraction. The entities to be corrected are corrected according to the similarity between each entity to be corrected and other entities. The knowledge graph is constructed based on the corrected entity information and a knowledge exchange system is built to realize knowledge exchange based on the knowledge graph.

[0011] Preferably, the calculation method of the deviation interference outlier value of each triple structure data is:

[0012] Where, is the deviation interference outlier of the i-th triple structure data in the k-th text cluster, is the data mean in the similarity deviation sequence of the i-th triple structure data in the k-th text cluster, is the number of data in the similarity deviation sequence of the i-th triple structure data in the k-th text cluster, and They are the jth and j-1th data in the similarity deviation sequence of the i-th triple structure data in the k-th text cluster respectively;

[0013] Among them, the mean of the bag-of-words vector similarity between each triple structure data in each text cluster and all other triple structure data is calculated as each triple structure data in each text cluster; the absolute value of the difference in word frequency similarity between each triple structure data in each text cluster and other triple structure data is calculated, and all absolute values are arranged from small to large to form a similarity deviation sequence for each triple structure data in each text cluster.

[0014] Preferably, the method for obtaining the triple test set is: performing threshold segmentation on the deviation interference anomaly values of the triple structure data in all text clustering clusters, obtaining the anomaly segmentation threshold, and forming the triple structure data corresponding to the deviation interference anomaly degree greater than the anomaly segmentation threshold in all text clustering clusters into a triple test set.

[0015] Preferably, the method for extracting the attribute word frequency vectors of the two entities of each triple structure data in the triple test set is: extracting the two entities of each triple structure data in the triple test set, inputting the attribute information of the two entities into the bag-of-words model respectively, and obtaining the attribute word frequency vectors of the two entities of each triple structure data in the triple test set.

[0016] Preferably, the calculation method of the abnormal correlation feature value between two entities of each triple structure data is:

[0017] Where, is the abnormal correlation feature value between the two entities of the s-th triple structure data, is the deviation interference outlier of the s-th triple structure data, is the mutual information value between the two entity attribute word frequency vectors of the sth triple structure data, To avoid constants with denominators equal to 0.

[0018] Preferably, the method for determining the target entity of each triple structure data is: respectively calculating the permutation entropy of the attribute word frequency vectors of the two entities in each triple structure data in the triple test set, and taking the entity with the largest permutation entropy as the target entity of each triple structure data in the triple test set.

[0019] Preferably, the method for obtaining the redundant information interference degree of each triple structure data target entity in the triple test set is: calculating the difference in the permutation entropy between each triple structure data target entity and another entity, and taking the normalized result of the product of the difference and the abnormal association eigenvalue as the redundant information interference degree of each triple structure data target entity in the triple test set.

[0020] Preferably, the method for obtaining the entity to be corrected is: performing threshold segmentation on the redundant information interference degree of the target entity of all triple structure data in the triple test set to obtain a redundant segmentation threshold, and taking the target entity corresponding to the redundant information interference degree greater than the redundant segmentation threshold as the entity to be corrected for information extraction.

[0021] Preferably, the method for correcting the entity to be corrected is: clustering the attribute word frequency vectors of all entities to be corrected to obtain entity sets, calculating the similarity between the entity to be corrected and other entities in its entity set with respect to the attribute word frequency vectors, and replacing the entity to be corrected with the entity corresponding to the maximum similarity to complete the correction processing of the entity to be corrected.

[0022] An embodiment of the present application also provides a knowledge exchange system based on a knowledge graph, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, the steps of any one of the above-mentioned knowledge exchange methods based on the knowledge graph are implemented.

[0023] As can be seen from the above, the knowledge exchange method and system based on knowledge graph provided by this application have at least the following beneficial effects:

[0024] This application takes into account that the existing technology does not fully analyze the error interference caused by redundant information features in entity information to verify and correct the entity information extracted, which will affect the consistency of entity information, thereby reducing the quality of knowledge graph construction, and further affecting the credibility of knowledge exchange based on knowledge graph; therefore, this application uses a text clustering algorithm to analyze the abnormal deviation of similarity between triple structure data, and accurately measures the abnormal feature of deviation interference based on the abnormal deviation of similarity, so as to more accurately reflect the error interference effect of redundant information features on triple structure data, so as to more accurately verify and correct entity information in the future;

[0025] Furthermore, this application accurately measures the abnormal correlation features between different entities based on the deviation interference abnormal features and combined with the mutual information features between different entity attributes, which is used for subsequent accurate extraction of the entities to be corrected in information extraction, and improves the consistency of the representation of the same entity in information extraction;

[0026] At the same time, this application is based on the abnormal correlation characteristics between different entities and combined with the random uncertain differences between entities to more accurately extract the redundant information interference degree of the target entity, and more accurately identify the entity to be corrected in the information extraction through the redundant information interference degree of the target entity, and then correct the entity to be corrected in the information extraction and construct a knowledge graph, thereby improving the quality of knowledge graph construction, and finally building a knowledge exchange system through the knowledge graph to improve the credibility of knowledge exchange based on the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0028] Figure 1 A flowchart of the steps of the knowledge exchange method based on knowledge graph provided in this application. DETAILED DESCRIPTION

[0029] In order to further illustrate the technical means and effects adopted by this application to achieve the predetermined invention purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation methods, structures, features and effects of the knowledge exchange method and system based on the knowledge graph proposed in this application. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0030] Unless otherwise specified and limited, terms such as "comprises", "includes" or any other variants thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article or device comprising a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the article or device comprising the element. In addition, the term "and\or" used herein includes any and all combinations of one or more related listed items. All technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs.

[0031] The specific solutions of the knowledge exchange method and system based on knowledge graph provided by this application are described in detail below with reference to the accompanying drawings.

[0032] See also Figure 1 , which shows a flowchart of the steps of a knowledge exchange method based on a knowledge graph provided by an embodiment of the present application, including the following steps:

[0033] Step 1: Extract entity information, relationships between entities, and attribute information of entities from textual materials in the knowledge database to obtain triple structure data.

[0034] By extracting a knowledge database of a specific field on the computer Internet, where the specific field can be the medical field, the geological field, the legal field or the financial field, in this embodiment, the knowledge database of the medical field is extracted. The knowledge database is composed of text materials extracted from journals, papers or professional terminology collections in different professional fields. Among them, in this embodiment, the professional fields include basic medicine, clinical medicine, oral medicine, traditional Chinese medicine, traditional Chinese medicine and nursing.

[0035] By installing and configuring the DeepDive open source knowledge extraction tool, we perform entity extraction, relationship extraction, and attribute extraction on all text materials in the knowledge database, and obtain entity information, relationships between entities, and attribute information of entities in the text materials in the knowledge database. Among them, entity extraction, relationship extraction, and attribute extraction are all well-known technologies and will not be described in detail.

[0036] Furthermore, the entity information extracted from the information and the relationships between entities are subjected to triple extraction using the DeepDive open source knowledge extraction tool to obtain all triple structure data in the text materials in the knowledge database. The triple structure data is used to characterize the entities in the text materials and the relationships between entities.

[0037] Step 2: Vectorize each triple structure data to extract each bag-of-words vector, and then perform cluster analysis. Based on the average level of bag-of-words vector similarity between each triple structure data and other triple structure data in the text cluster, as well as the difference in bag-of-words vector similarity, the deviation interference outlier value of each triple structure data is obtained.

[0038] Due to the diversity of information sources in professional fields, there is a large amount of redundant information fragments in the entity information extracted. It is necessary to fully consider the error interference caused by the redundant information characteristics in the entity information, and to test and correct the entity information extracted to avoid inconsistent expressions of the same entity, improve the quality of subsequent knowledge graph construction, and thereby improve the credibility of knowledge exchange based on knowledge graphs.

[0039] Furthermore, in order to accurately verify and correct the entity information extracted from the information, each triple structure data in the text data of the knowledge database is vectorized through the bag-of-words model to obtain the bag-of-words vector of each triple structure data, and the bag-of-words vectors of all triple structure data are used as the input of the K-means clustering algorithm, wherein the elbow rule is used to obtain the optimal number of clusters K, and all triple structure data are clustered through the Euclidean distance between the bag-of-words vectors of the triple structure data to obtain the text clustering results of all triple structure data. The text clustering results contain K text clustering clusters. The K-means clustering algorithm is a well-known technology and will not be elaborated on.

[0040] Since there is a certain degree of similarity between different triple structure data within the same text clustering cluster, if the similarity between a certain triple structure data and other triple structure data within the same text clustering cluster shows a high abnormal deviation, it is characterized by the abnormal distribution of the similarity between the triple structure data and other triple structure data, that is, the more likely it is that information redundancy will occur between the triple structure data and other triple structure data, the more it is necessary to verify and correct the entity information represented by the triple structure data, so as to improve the credibility of subsequent knowledge exchange based on the knowledge graph.

[0041] Furthermore, the mean of the bag-of-words vector similarities between each triple structure data and all other triple structure data in each text cluster is calculated and used as the word frequency similarity of each triple structure data in each text cluster. The similarity can be measured by cosine similarity or Jaccard similarity. In this embodiment, cosine similarity is used to measure the similarity between vectors. The greater the similarity, the higher the similarity between the bag-of-words vectors.

[0042] In order to analyze the error interference of redundant information characteristics on the entity information represented by triple structure data, the absolute value of the difference between the word frequency similarity of each triple structure data and the word frequency similarity of other triple structure data in each text cluster is calculated, and the vector composed of all absolute values in ascending order is recorded as the similarity deviation sequence of each triple structure data in each text cluster. The greater the difference in the change of adjacent data in the similarity deviation sequence, and the higher the average level of similarity deviation in the similarity deviation sequence, the more likely it is that information redundancy will occur between the triple structure data and other triple structure data, that is, the entity information represented by the triple structure data is more likely to be inconsistent in expression, and the more necessary it is to check and correct the entity information represented by the triple structure data to avoid adverse effects on the consistency of entity information.

[0043] Therefore, based on the above analysis, the deviation interference outlier value of each triple structure data in each text cluster is calculated:

[0044] ;

[0045] Where, is the deviation interference outlier of the i-th triple structure data in the k-th text cluster, is the data mean in the similarity deviation sequence of the i-th triple structure data in the k-th text cluster, is the number of data in the similarity deviation sequence of the i-th triple structure data in the k-th text cluster, and They are respectively the jth and j-1th data in the similarity deviation sequence of the i-th triple structure data in the k-th text cluster.

[0046] The deviation interference outlier reflects the abnormal deviation characteristics caused by the interference of redundant information features on the similarity between triple structure data. The larger the deviation interference outlier value, the greater the error interference effect of redundant information features on the representation. The entity information represented by the triple structure data should be checked and corrected to avoid inconsistent representation of entity instances and improve the credibility of subsequent knowledge exchange based on knowledge graphs.

[0047] Step 3: Divide the triple structure data based on the deviation interference outlier value to obtain the triple test set, extract the attribute word frequency vectors of the two entities of each triple structure data in the triple test set, analyze the correlation between the two attribute word frequency vectors, and combine the deviation interference outlier value of each triple structure data to obtain the abnormal correlation feature value between the two entities of each triple structure data.

[0048] Furthermore, the deviation interference anomaly values of the triple structure data in all text clustering clusters are input into the maximum inter-class variance algorithm, and the maximum inter-class variance algorithm is used to obtain the abnormal segmentation threshold. The set consisting of the triple structure data corresponding to the deviation interference anomaly degree greater than the abnormal segmentation threshold in all text clustering clusters is used as the triple test set for information extraction. The triple test set represents the triple structure data set that is greatly affected by the error interference of redundant information features. The triple structure data in the triple test set needs to be tested and corrected. The maximum inter-class variance algorithm is a well-known technology and will not be elaborated on.

[0049] Since triple structure data is composed of subject, predicate and object, and the subject and object represent different entities respectively, and the predicate represents the relationship between different entities, if the deviation interference abnormality of the triple structure data is greater and the mutual dependence of the attribute information between different entities in the triple structure data is smaller, it can be more clearly explained that the entity information in the triple structure data is greatly interfered by the error of redundant information characteristics, so that the correlation between the attribute information of two entities with a dependent relationship is abnormally disturbed. At this time, it is more necessary to accurately correct the entity information.

[0050] Therefore, a statistical method is used to extract the two entities of each triple structure data in the triple test set, and the attribute information of the two entities is obtained respectively. In order to extract the abnormal correlation features between the two entity information, the attribute information of the two entities is input into the bag-of-words model respectively, and the bag-of-words model is used to obtain the attribute word frequency vectors of the two entities in each triple structure data.

[0051] Furthermore, the attribute word frequency vectors between the two entities are input into a mutual information algorithm to obtain the mutual information value between the two entities. Mutual information algorithms are well-known technologies and will not be described in detail here. The smaller the mutual information value, the weaker the dependency between the two entities' information. This indicates a greater likelihood of interference from redundant information features during information extraction, leading to abnormal correlations between the entities.

[0052] Based on the above analysis, the abnormal correlation feature value between the two entities of each triple structure data in the triple test set is calculated:

[0053] Where, is the abnormal correlation feature value between the two entities of the s-th triple structure data, is the deviation interference outlier of the s-th triple structure data, is the mutual information value between the two entity attribute word frequency vectors of the sth triple structure data, To avoid a constant with a denominator of 0, the value range is 0.01-0.05, and in this embodiment, the value is 0.05.

[0054] It can be understood that the abnormal correlation feature value reflects the abnormal correlation feature of the entity information in the triple structure data under the interference of the redundant information feature error. The larger the abnormal correlation feature, the more likely the corresponding two entity information are to be interfered with by the error of the redundant information feature, and the more necessary it is to accurately correct the inconsistent entity information. By improving the consistency of the entity instance representation, it is used to improve the credibility of subsequent knowledge exchange based on the knowledge graph.

[0055] Step 4: Determine the target entity of each triple structure data according to the change of the attribute word frequency vectors of the two entities of each triple structure data in the triple test set, and obtain the redundant information interference degree of each triple structure data target entity in the triple test set according to the degree of difference in the change of the attribute word frequency vector between the target entity and the other entity, combined with the abnormal correlation feature value.

[0056] Furthermore, considering the phenomenon of abnormal correlation characteristics that will be generated when an entity in the triple structure data is interfered with by redundant information, if the random uncertainty of the attribute word frequency vector of an entity in the triple structure data is higher and the random uncertainty difference between it and another entity is greater, it represents that the intensity of the redundant information interference of the entity in the triple structure data is greater, and the entity is more likely to have inconsistent expressions. In order to improve the credibility of subsequent knowledge exchange based on the knowledge graph, the entity information needs to be accurately corrected.

[0057] Therefore, we calculate the permutation entropy of the attribute word frequency vectors of the two entities in each triple structure data in the triple test set. The larger the permutation entropy, the higher the random uncertainty representing the relative arrangement within the attribute word frequency vector. Furthermore, the entity with the largest permutation entropy is used as the target entity for each triple structure data.

[0058] Furthermore, in this embodiment, the difference in permutation entropy between the target entity and another entity is calculated, and the normalized result of the product of the difference and the abnormal associated eigenvalue is used as the redundant information interference degree of each triple structure data target entity in the triple test set.

[0059] According to the above process of this embodiment, it can be understood that the greater the interference degree of redundant information, the more likely it is that the target entity will be interfered with by the error of redundant information when the representation information is extracted, making it more likely that the target entity will be inconsistently expressed, and the more necessary it is to accurately correct the entity information.

[0060] Step 5: Divide the target entities of all triple structure data in the triple test set based on the redundant information interference degree to obtain the entities to be corrected after information extraction. Correct the entities to be corrected according to the similarity between each entity to be corrected and other entities. Combine the corrected entity information to construct a knowledge graph and build a knowledge exchange system to realize knowledge exchange based on the knowledge graph.

[0061] In order to correct abnormal entity information more accurately, the redundant information interference degree of the target entity of all triple structure data in the triple test set is input into the maximum inter-class variance algorithm. The maximum inter-class variance algorithm is used to obtain the redundant segmentation threshold. The target entity corresponding to the redundant information interference degree greater than the redundant segmentation threshold is recorded as the entity to be corrected in information extraction.

[0062] Furthermore, in order to correct the entities to be corrected in the information extraction, the attribute word frequency vectors of all entities to be corrected are used as the input of the K-means clustering algorithm, and the elbow rule is used to obtain the optimal number of clusters. All entities are clustered by calculating the Euclidean distance of the attribute word frequency vectors between entities to obtain each entity set. The cosine similarity method is used to calculate the similarity between the attribute word frequency vectors of the entity to be corrected and other entities in its entity set. The entity to be corrected in the information extraction is replaced with the entity with the greatest similarity, and the entity information extracted is inspected and corrected to avoid the consistency of entity information, which is used to improve the quality of knowledge graph construction.

[0063] Furthermore, knowledge fusion is performed on the entity information, the relationship between entities, and the attribute information of the entities in the text data after verification and correction. In this embodiment, preferably, the construction of the knowledge graph is completed with the help of protege open source software to obtain the knowledge graph in the medical field, and the knowledge graph in the medical field is imported into the Neo4j graph database for storage in the manner of graph data storage. The construction of the knowledge graph is a well-known technology and will not be elaborated on in detail.

[0064] By verifying and correcting the entity information extracted, and using the knowledge graph construction method to process and integrate information and professional knowledge in multiple professional fields, different nodes in the knowledge graph in the medical field represent entity instances in different medical professional fields, and the edges between nodes represent the relationship between entity instances in different medical professional fields. This can realize knowledge exchange in different medical professional fields and improve the reliability of knowledge exchange.

[0065] Furthermore, to enable knowledge exchange across different medical fields, preferably, in this embodiment, a knowledge exchange system is constructed using the SpringBoot framework and a medical knowledge graph stored in a Neo4j graph database. The specific construction process can be implemented using existing technologies, and this embodiment does not impose any particular limitations on this process and will not be described in detail here. The constructed knowledge exchange system can provide knowledge retrieval, knowledge download, knowledge visualization, and knowledge exchange services across different medical fields.

[0066] Based on the same inventive concept as the above method, an embodiment of the present application also provides a knowledge exchange system based on a knowledge graph, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-mentioned knowledge exchange methods based on the knowledge graph.

[0067] It should be understood that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0068] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0069] The above content is only an implementation method of the present application and is not intended to limit the scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the scope of protection of the present application.

Claims

1. The knowledge exchange method based on knowledge graph is characterized by: The following steps are involved: Extract entity information, relationship between entities and attribute information of entities from text data in knowledge database to obtain triple structure data; Each triple structure data is vectorized to extract each bag-of-words vector, and then cluster analysis is performed. Based on the average level of bag-of-words vector similarity between each triple structure data and other triple structure data in the text cluster, as well as the difference in bag-of-words vector similarity, the deviation interference outlier value of each triple structure data is obtained; The triple structure data is divided based on the deviation interference outlier value to obtain a triple test set, the attribute word frequency vectors of the two entities of each triple structure data in the triple test set are extracted, the correlation between the two attribute word frequency vectors is analyzed, and the deviation interference outlier value of each triple structure data is combined to obtain the abnormal correlation feature value between the two entities of each triple structure data; Determine the target entity of each triple structure data according to the change of the attribute word frequency vectors of the two entities in each triple structure data in the triple test set, and obtain the redundant information interference degree of the target entity of each triple structure data in the triple test set according to the difference degree of change of the attribute word frequency vector between the target entity and the other entity, combined with the abnormal correlation feature value; Based on the redundant information interference degree, the target entities of all triple structure data in the triple test set are divided to obtain the entities to be corrected after information extraction. The entities to be corrected are corrected according to the similarity between each entity to be corrected and other entities. The knowledge graph is constructed based on the corrected entity information and a knowledge exchange system is built to realize knowledge exchange based on the knowledge graph.

2. The knowledge exchange method based on knowledge graph according to claim 1, characterized in that: The calculation method of the deviation interference outlier value of each triple structure data is: Where, is the deviation interference outlier of the i-th triple structure data in the k-th text cluster, is the data mean in the similarity deviation sequence of the i-th triple structure data in the k-th text cluster, is the number of data in the similarity deviation sequence of the i-th triple structure data in the k-th text cluster, and They are the jth and j-1th data in the similarity deviation sequence of the i-th triple structure data in the k-th text cluster respectively; Among them, the mean of the bag-of-words vector similarity between each triple structure data in each text cluster and all other triple structure data is calculated as each triple structure data in each text cluster; the absolute value of the difference in word frequency similarity between each triple structure data in each text cluster and other triple structure data is calculated, and all absolute values are arranged from small to large to form a similarity deviation sequence for each triple structure data in each text cluster.

3. The knowledge exchange method based on knowledge graph according to claim 1, characterized in that: The method for obtaining the triple test set is as follows: performing threshold segmentation on the deviation interference anomaly values of the triple structure data in all text clustering clusters, obtaining the anomaly segmentation threshold, and forming a triple test set from the triple structure data corresponding to the deviation interference anomaly degree greater than the anomaly segmentation threshold in all text clustering clusters.

4. The knowledge exchange method based on knowledge graph according to claim 1, characterized in that: The method for extracting the attribute word frequency vectors of the two entities of each triple structure data in the triple test set is: extracting the two entities of each triple structure data in the triple test set, inputting the attribute information of the two entities into the bag-of-words model respectively, and obtaining the attribute word frequency vectors of the two entities of each triple structure data in the triple test set.

5. The knowledge exchange method based on knowledge graph according to claim 1, characterized in that: The calculation method of the abnormal correlation feature value between two entities of each triple structure data is: Where, is the abnormal correlation feature value between the two entities of the s-th triple structure data, is the deviation interference outlier of the s-th triple structure data, is the mutual information value between the two entity attribute word frequency vectors of the sth triple structure data, To avoid constants with denominators equal to 0.

6. The knowledge exchange method based on knowledge graph according to claim 1, characterized in that: The method for determining the target entity of each triple structure data is: respectively calculating the permutation entropy of the attribute word frequency vectors of the two entities in each triple structure data in the triple test set, and taking the entity with the largest permutation entropy as the target entity of each triple structure data in the triple test set.

7. The knowledge exchange method based on knowledge graph according to claim 6, characterized in that: The method for obtaining the redundant information interference degree of each triple structure data target entity in the triple test set is: calculating the difference in the permutation entropy between each triple structure data target entity and another entity, and taking the normalized result of the product of the difference and the abnormal association feature value as the redundant information interference degree of each triple structure data target entity in the triple test set.

8. The knowledge exchange method based on knowledge graph according to claim 1, characterized in that: The method for obtaining the entity to be corrected is: performing threshold segmentation on the redundant information interference degree of the target entity of all triple structure data in the triple test set to obtain a redundant segmentation threshold, and taking the target entity corresponding to the redundant information interference degree greater than the redundant segmentation threshold as the entity to be corrected for information extraction.

9. The knowledge exchange method based on knowledge graph according to claim 1, characterized in that: The method for correcting the entity to be corrected is as follows: clustering the attribute word frequency vectors of all entities to be corrected to obtain entity sets, calculating the similarity between the entity to be corrected and other entities in the entity set in terms of attribute word frequency vectors, and replacing the entity to be corrected with the entity corresponding to the maximum similarity, thereby completing the correction processing of the entity to be corrected.

10. A knowledge exchange system based on a knowledge graph, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the knowledge exchange method based on the knowledge graph as described in any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • An automatic construction method for a big data knowledge graph in the field of public security

    CN109710701A

  • Attribute knowledge graph-based patent authorization prediction and evaluation method and system

    CN117972113A