Knowledge exchange method and system based on knowledge graph
By vectorized processing and clustering analysis of entity information in the knowledge graph, redundant information is identified and corrected, the diversity of information sources in multiple professional fields is solved, and the credibility and consistency of knowledge exchange is improved.
Patent Information
- Application Number
- CN202510764642.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The prior art fails to effectively deal with the redundancy of entity information caused by the diversity of information sources in multiple professional fields when building a knowledge graph, affecting the credibility and consistency of knowledge exchange.
By extracting entity information and attribute information of the knowledge database, performing vectorization processing and clustering analysis, identifying deviation interference outliers, obtaining triple inspection sets, analyzing the abnormal correlation characteristics and redundant information interference between entities, correcting entity information and building a knowledge graph.
It improves the quality and credibility of knowledge graph construction, ensures the consistency of entity information, and improves the reliability of knowledge exchange.
Smart Images

Figure CN120278252A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of knowledge graph analysis and exchange, and specifically relates to a knowledge exchange method and system based on a knowledge graph. Background Art
[0002] With the continuous in-depth application of natural language processing technology in various professional fields, the demand for realizing knowledge exchange between multiple professional fields is increasing, in order to meet the communication and collaboration under multiple professional fields, and finally realize knowledge exchange and information sharing between professional fields. However, since the knowledge exchange process involves complex interactions of information and professional knowledge in multiple professional fields, the complexity of knowledge exchange between different professional fields is relatively high, and it will strongly hinder the knowledge exchange under different professional fields. At present, by adopting the method of constructing a knowledge graph to process and integrate information and professional knowledge in multiple professional fields, and building a knowledge exchange system through the knowledge graph, the efficiency of knowledge exchange between different professional fields can be effectively improved, and at the same time, the rigor of knowledge exchange can be enhanced.
[0003] The prior art adopts the method of constructing a knowledge graph to process and integrate information and professional knowledge in multiple professional fields, so as to effectively improve the efficiency of knowledge exchange between different professional fields. However, in the process of constructing a knowledge graph, due to the diversity of information sources in professional fields, there are easily a large number of redundant information fragments in the entity information content extracted. The prior art does not fully consider the error interference caused by the redundant information characteristics in the entity information to test and correct the entity information extracted, which will affect the consistency of the entity information, thereby reducing the quality of the knowledge graph construction, and further affecting the credibility of knowledge exchange based on the knowledge graph. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of this application is to provide a knowledge exchange method and system based on a knowledge graph, and the specific technical solutions adopted are as follows: The embodiment of this application provides a knowledge exchange method based on a knowledge graph, including the following steps: Extract the entity information, the relationships between entities, and the attribute information of entities in the text materials in the knowledge database to obtain triple structure data; Vectorize each triple structure data to extract each bag-of-words vector, and then perform clustering analysis. According to the average level of the bag-of-words vector similarity between each triple structure data in the text clustering cluster and other triple structure data, and the difference situation of the bag-of-words vector similarity, obtain the deviation interference outliers of each triple structure data; Partition the triple structure data based on the deviation interference outliers to obtain a triple test set, extract the attribute word frequency vectors of the two entities in each triple structure data in the triple test set, analyze the correlation between the two attribute word frequency vectors, and combine the deviation interference outliers of each triple structure data to obtain the abnormal correlation eigenvalue between the two entities of each triple structure data; Determine the target entity of each triple structure data in the triple test set according to the change of the attribute word frequency vectors of the two entities in each triple structure data in the triple test set, and obtain the redundancy information interference degree of the target entity of each triple structure data in the triple test set according to the change difference degree of the target entity and another entity with respect to the attribute word frequency vector, in combination with the abnormal correlation eigenvalue; Partition the target entities of all triple structure data in the triple test set based on the redundancy information interference degree to obtain the entities to be corrected for information extraction, correct the entities to be corrected according to the similarity between each entity to be corrected and other entities, and construct a knowledge graph and build a knowledge exchange system in combination with the corrected entity information to realize knowledge exchange based on the knowledge graph.
[0005] Preferably, the calculation method of the deviation interference outliers of each triple structure data is as follows: ; where is the deviation interference outlier of the i-th triple structure data in the k-th text clustering cluster, is the data mean in the similarity deviation sequence of the i-th triple structure data in the k-th text clustering cluster, is the number of data in the similarity deviation sequence of the i-th triple structure data in the k-th text clustering cluster, and are the j-th and (j - 1)-th data in the similarity deviation sequence of the i-th triple structure data in the k-th text clustering cluster respectively; Among them, calculate the mean of the bag-of-words vector similarities between each triple structure data in each text clustering cluster and all other triple structure data as each triple structure data in each text clustering cluster; calculate the absolute value of the difference in word frequency similarity between each triple structure data in each text clustering cluster and other triple structure data, and arrange all the absolute values in ascending order to form the similarity deviation sequence of each triple structure data in each text clustering cluster.
[0006] Preferably, the method for obtaining the triple test set is: perform threshold segmentation on the deviation interference outliers of the triple structure data in all text clustering clusters to obtain an abnormal segmentation threshold, and form a triple test set with the triple structure data corresponding to the deviation interference abnormality degree greater than the abnormal segmentation threshold in all text clustering clusters.
[0007] Preferably, the method for extracting the attribute word frequency vectors of the two entities in each triple structure data in the triple test set is as follows: extract the two entities in each triple structure data in the triple test set, input the attribute information of the two entities into the bag-of-words model respectively, and obtain the attribute word frequency vectors of the two entities in each triple structure data in the triple test set.
[0008] Preferably, the method for calculating the abnormal association eigenvalue between the two entities of each triple structure data is as follows: ; where is the abnormal association eigenvalue between the two entities of the s-th triple structure data, is the deviation interference outlier of the s-th triple structure data, is the mutual information value between the attribute word frequency vectors of the two entities of the s-th triple structure data, is a constant to avoid the denominator being zero.
[0009] Preferably, the method for determining the target entity of each triple structure data is as follows: calculate the permutation entropy of the attribute word frequency vectors of the two entities in each triple structure data in the triple test set respectively, and use the entity with the largest permutation entropy as the target entity of each triple structure data in the triple test set.
[0010] Preferably, the method for obtaining the redundancy information interference degree of the target entity of each triple structure data in the triple test set is as follows: calculate the difference of the permutation entropy between the target entity of each triple structure data and another entity, and use the normalized result of the product of the difference and the abnormal association eigenvalue as the redundancy information interference degree of the target entity of each triple structure data in the triple test set.
[0011] Preferably, the method for obtaining the entity to be corrected is as follows: perform threshold segmentation on the redundancy information interference degrees of the target entities of all triple structure data in the triple test set to obtain a redundancy segmentation threshold, and use the target entity corresponding to the redundancy information interference degree greater than the redundancy segmentation threshold as the entity to be corrected for information extraction.
[0012] Preferably, the method for correcting the entity to be corrected is as follows: perform clustering on the attribute word frequency vectors of all entities to be corrected to obtain each entity set, calculate the similarity between the entity to be corrected and other entities in its entity set regarding the attribute word frequency vectors, and replace the entity to be corrected with the entity corresponding to the maximum similarity to complete the correction process of the entity to be corrected.
[0013] The embodiments of the present application also provide a knowledge exchange system based on a knowledge graph, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the knowledge exchange method based on the knowledge graph described in any one of the above are implemented.
[0014] As can be seen from the above, the knowledge exchange method and system based on the knowledge graph provided by the present application have at least the following beneficial effects: The present application considers that the existing technology does not fully analyze the error interference caused by the redundant information characteristics in the entity information to check and correct the entity information extracted from the information, which will affect the consistency of the entity information, thus reducing the quality of the knowledge graph construction, and further affecting the credibility of the knowledge exchange based on the knowledge graph. Therefore, the present application uses a text clustering algorithm to analyze the abnormal deviation of the similarity between the triple structure data, and accurately measures the abnormal deviation characteristics of the deviation interference based on the abnormal deviation of the similarity, more accurately reflecting the influence of the error interference of the redundant information characteristics on the triple structure data, and is used for more accurate inspection and correction of the entity information in the follow-up; Furthermore, based on the abnormal deviation characteristics of the deviation interference and combined with the mutual information characteristics between different entity attributes, the present application accurately measures the abnormal association characteristics between different entities, and is used to accurately extract the entities to be corrected in the information extraction in the follow-up, improving the expression consistency of the same entity in the information extraction; At the same time, based on the abnormal association characteristics between different entities and combined with the random uncertainty differences between entities, the present application more accurately extracts the redundant information interference degree of the target entity, and more accurately identifies the entities to be corrected in the information extraction through the redundant information interference degree of the target entity. Furthermore, by correcting the entities to be corrected in the information extraction and constructing a knowledge graph, the quality of the knowledge graph construction is improved, and finally a knowledge exchange system is built through the knowledge graph, enhancing the credibility of the knowledge exchange based on the knowledge graph. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a flowchart of the steps of the knowledge exchange method based on the knowledge graph provided by the present application. Detailed Embodiments
[0017] In order to further elaborate on the technical means and effects adopted by this application to achieve the intended invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, details the specific implementation manner, structure, features, and effects of the knowledge exchange method and system based on a knowledge graph proposed according to this application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0018] Unless otherwise specified and limited, terms such as "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element. Additionally, the term "and / or" used herein includes any and all combinations of one or more of the related listed items. All technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.
[0019] The following specifically describes the specific solution of the knowledge exchange method and system based on a knowledge graph provided by this application in combination with the accompanying drawings.
[0020] Please refer to Figure 1 , which shows the step flow chart of the knowledge exchange method based on a knowledge graph provided by an embodiment of this application, including the following steps: Step 1: Extract the entity information, relationships between entities, and attribute information of the text materials in the knowledge database to obtain triple-structured data.
[0021] Extract a knowledge database in a specific field from the computer Internet. Among them, the specific field can be the medical field, the geological field, the legal field, or the financial field. In this embodiment, a knowledge database in the medical field is extracted. The knowledge database is composed of text materials extracted from periodicals, papers, or specialized term collections in different professional fields. In this embodiment, the professional fields include basic medicine, clinical medicine, stomatology, traditional Chinese medicine, traditional Chinese pharmacy, and nursing.
[0022] Install and configure the DeepDive open-source knowledge extraction tool to perform entity extraction, relationship extraction, and attribute extraction on all text materials in the knowledge database, obtaining the entity information, relationships between entities, and attribute information of the text materials in the knowledge database. Among them, entity extraction, relationship extraction, and attribute extraction are all well-known technologies and will not be elaborated further.
[0023] Further, the DeepDive open-source knowledge extraction tool is used to extract triples from the entity information and the relationships between entities in the information extraction, obtaining all the triple structure data in the text materials of the knowledge database. The triple structure data is used to represent the entities in the text materials and the relationships between the entities.
[0024] Step 2: Vectorize each triple structure data to extract each bag-of-words vector, and then perform clustering analysis. According to the average level of the similarity of the bag-of-words vectors between each triple structure data in the text clustering cluster and other triple structure data, as well as the difference in the similarity of the bag-of-words vectors, the deviation interference outliers of each triple structure data are obtained.
[0025] Due to the diversity of information sources in the professional field, there are likely to be a large number of redundant information fragments in the entity information extracted from the information. It is necessary to fully consider the error interference caused by the redundant information characteristics in the entity information, check and correct the entity information extracted from the information, avoid the phenomenon of inconsistent expressions of the same entity, improve the quality of the subsequent knowledge graph construction, and thus improve the credibility of knowledge exchange based on the knowledge graph.
[0026] Further, in order to accurately check and correct the entity information extracted from the information, each triple structure data in the text materials of the knowledge database is vectorized through the bag-of-words model to obtain the bag-of-words vector of each triple structure data, and the bag-of-words vectors of all triple structure data are used as the input of the K-means clustering algorithm. Among them, the elbow method is used to obtain the optimal number of clustering clusters K, and all triple structure data are clustered through the Euclidean distance between the bag-of-words vectors of the triple structure data, obtaining the text clustering result of all triple structure data. The text clustering result contains K text clustering clusters. The K-means clustering algorithm is a well-known technology and will not be elaborated further.
[0027] Since there is a certain degree of similarity between different triple structure data in the same text clustering cluster, but if the similarity between a certain triple structure data and other triple structure data in the same text clustering cluster shows a relatively high abnormal deviation, it indicates that the similarity between this triple structure data and other triple structure data shows an abnormal distribution characteristic, that is, the more likely there is an information redundancy phenomenon between this triple structure data and other triple structure data, the more necessary it is to check and correct the entity information represented by the triple structure data to improve the credibility of subsequent knowledge exchange based on the knowledge graph.
[0028] Further, calculate the mean of the bag-of-words vector similarities between each triple structure data within each text clustering cluster and all other triple structure data, and use it as the word frequency similarity of each triple structure data within each text clustering cluster. The similarity measurement method can be cosine similarity or Jaccard similarity. In this embodiment, cosine similarity is used to measure the similarity between vectors. The greater the similarity, the higher the similarity between the bag-of-words vectors is represented.
[0029] To analyze the error interference of the redundant information features on the entity information represented by the triple structure data, calculate the absolute value of the difference between the word frequency similarity of each triple structure data within each text clustering cluster and the word frequency similarity of other triple structure data. Denote the vector composed of all absolute values in ascending order as the similarity deviation sequence of each triple structure data within each text clustering cluster. If the difference in the change of adjacent data within the similarity deviation sequence is greater, and the average level of similarity deviation within the similarity deviation sequence is higher, it is more likely that there is information redundancy between this triple structure data and other triple structure data, that is, the entity information represented by the triple structure data is more likely to show inconsistent expressions, and it is more necessary to test and correct the entity information represented by the triple structure data to avoid having an adverse impact on the consistency of entity information.
[0030] Therefore, based on the above analysis, calculate the deviation interference outlier of each triple structure data within each text clustering cluster: ; In the formula, is the deviation interference outlier of the i-th triple structure data in the k-th text clustering cluster, is the data mean within the similarity deviation sequence of the i-th triple structure data in the k-th text clustering cluster, is the number of data within the similarity deviation sequence of the i-th triple structure data in the k-th text clustering cluster, and are the j-th and (j - 1)-th data within the similarity deviation sequence of the i-th triple structure data in the k-th text clustering cluster respectively.
[0031] The deviation interference outlier reflects the abnormal deviation characteristics generated by the interference of redundant information features on the similarity between triple structure data. The greater the deviation interference outlier, the greater the influence of the error interference of redundant information features, and the more the entity information represented by this triple structure data should be tested and corrected to avoid the phenomenon of inconsistent expressions of entity instances and improve the credibility of subsequent knowledge exchange based on the knowledge graph.
[0032] Step 3: Divide the triple structure data based on the deviation interference outliers to obtain a triple test set, extract the attribute word frequency vectors of the two entities of each triple structure data in the triple test set, analyze the correlation between the two attribute word frequency vectors, and combine the deviation interference outliers of each triple structure data to obtain the abnormal correlation eigenvalue between the two entities of each triple structure data.
[0033] Further, input the deviation interference outliers of the triple structure data within all text clustering clusters into the maximum inter-class variance algorithm, use the maximum inter-class variance algorithm to obtain the abnormal segmentation threshold, and use the set composed of the triple structure data corresponding to the deviation interference abnormality degree greater than the abnormal segmentation threshold within all text clustering clusters as the triple test set for information extraction. The triple test set represents a set of triple structure data that is greatly affected by the error interference of redundant information features and needs to be tested and corrected for the triple structure data within the triple test set. The maximum inter-class variance algorithm is a well-known technology and will not be elaborated further.
[0034] Since the triple structure data is composed of a subject, a predicate, and an object, where the subject and the object represent different entities respectively, and the predicate represents the relationship between different entities. If the deviation interference abnormality degree of the triple structure data is greater, and the degree of mutual dependence of the attribute information between different entities within the triple structure data is smaller, it can more clearly indicate that the entity information within the triple structure data is greatly affected by the error interference of redundant information features, resulting in an abnormal interference in the correlation between the attribute information of the two entities with a dependency relationship. At this time, it is more necessary to accurately correct the entity information.
[0035] Therefore, a statistical method is used to extract the two entities of each triple structure data in the triple test set, and the attribute information of each of the two entities is obtained respectively. In order to extract the abnormal correlation eigenvalue between the two entity information, the attribute information of each of the two entities is input into the bag-of-words model respectively, and the bag-of-words model is used to obtain the attribute word frequency vectors of the two entities of each triple structure data.
[0036] Further, input the attribute word frequency vectors between the two entities into the mutual information algorithm, and use the mutual information algorithm to obtain the mutual information value between the two entities. The mutual information algorithm is a well-known technology and will not be elaborated further. The smaller the mutual information value, the weaker the dependence between the two entity information, and the more likely it is to be affected by the error interference of redundant information features during information extraction, resulting in an abnormal situation in the correlation between entities.
[0037] Based on the above analysis, calculate the abnormal correlation eigenvalue between the two entities of each triple structure data in the triple test set: ; where, is the abnormal association eigenvalue between two entities of the s-th triple structure data, is the deviation interference outlier of the s-th triple structure data, is the mutual information value between the word frequency vectors of the attributes of two entities of the s-th triple structure data, is a constant to avoid a zero denominator, with a value range of 0.01 - 0.05, and the value in this embodiment is 0.05.
[0038] It can be understood that the abnormal association eigenvalue reflects the abnormal association characteristics of the entity information in the triple structure data under the interference of redundant information feature errors. The greater the abnormal association feature, the more likely the information of the corresponding two entities is affected by the errors of redundant information features, and the more necessary it is to accurately correct the inconsistent entity information. By improving the consistency of entity instance expressions, it is used to improve the credibility of subsequent knowledge exchange based on the knowledge graph.
[0039] Step 4: Determine the target entity of each triple structure data in the triple test set according to the change situation of the word frequency vectors of the attributes of the two entities in each triple structure data. According to the degree of change difference of the word frequency vectors of the target entity and another entity, combined with the abnormal association eigenvalue, obtain the redundant information interference degree of the target entity of each triple structure data in the triple test set.
[0040] Furthermore, considering the phenomenon of abnormal association characteristics that occur when an entity in the triple structure data is interfered by redundant information, if the random uncertainty of the word frequency vector of an entity's attributes in the triple structure data is higher, and the random uncertainty difference with another entity is greater, it indicates that the intensity of the entity being interfered by redundant information in the triple structure data is greater, and the entity is more likely to have inconsistent expressions. In order to improve the credibility of subsequent knowledge exchange based on the knowledge graph, it is necessary to accurately correct the entity information.
[0041] Therefore, calculate the permutation entropy of the word frequency vectors of the attributes of the two entities in each triple structure data in the triple test set respectively. The greater the permutation entropy, the higher the random uncertainty of the relative arrangement within the word frequency vector. Further, take the entity with the largest permutation entropy as the target entity of each triple structure data.
[0042] Furthermore, in this embodiment, calculate the difference in permutation entropy between the target entity and another entity, and take the normalized result of the product of the difference and the abnormal association eigenvalue as the redundant information interference degree of the target entity of each triple structure data in the triple test set.
[0043] According to the above process of this embodiment, it can be understood that the greater the interference degree of redundant information, the more likely it is that the target entity for information extraction is interfered by the error of redundant information when representing the information extraction target entity, making it more likely that the target entity has inconsistent expressions, and the more necessary it is to accurately correct the entity information.
[0044] Step 5: Divide the target entities in all triple structure data in the triple test set based on the interference degree of redundant information to obtain the entities to be corrected for information extraction. Correct the entities to be corrected according to the similarity between each entity to be corrected and other entities, and construct a knowledge graph and build a knowledge exchange system in combination with the corrected entity information to achieve knowledge exchange based on the knowledge graph.
[0045] In order to be able to correct the abnormal entity information more accurately, input the interference degree of redundant information of the target entities in all triple structure data in the triple test set into the maximum inter-class variance algorithm, use the maximum inter-class variance algorithm to obtain the redundant segmentation threshold, and record the target entities corresponding to the interference degree of redundant information greater than the redundant segmentation threshold as the entities to be corrected for information extraction.
[0046] Furthermore, in order to correct the entities to be corrected for information extraction, use the attribute word frequency vectors of all entities to be corrected as the input of the K-means clustering algorithm, and use the elbow method to obtain the optimal number of clustering clusters. Cluster all entities by calculating the Euclidean distance of the attribute word frequency vectors between entities to obtain each entity set. Calculate the similarity between the entities to be corrected for information extraction and other entities in their respective entity sets with respect to the attribute word frequency vectors using the method of cosine similarity, and replace the entities to be corrected for information extraction with the entity with the greatest similarity to complete the inspection and correction of the entity information for information extraction, avoiding the consistency of entity information and improving the quality of knowledge graph construction.
[0047] Furthermore, perform knowledge fusion on the entity information, the relationships between entities, and the attribute information of entities in the text materials after inspection and correction. In this embodiment, preferably, use the protege open-source software to complete the construction of the knowledge graph to obtain the knowledge graph in the medical field, and import the knowledge graph in the medical field into the Neo4j graph database for storage in the form of graph data storage. The construction of the knowledge graph is a well-known technology and will not be elaborated further.
[0048] By inspecting and correcting the entity information for information extraction and using the method of knowledge graph construction to process and integrate information and professional knowledge in multiple professional fields, different nodes in the knowledge graph in the medical field represent entity instances in different medical professional fields, and the edges between nodes represent the relationships between entity instances in different medical professional fields, enabling knowledge exchange in different medical professional fields and improving the reliability of knowledge exchange.
[0049] Furthermore, in order to achieve knowledge exchange in different medical specialty fields, preferably, in this embodiment, a knowledge exchange system is built through the Spring Boot framework and the knowledge graph of the medical field stored in the Neo4j graph database. The specific building process can be realized by the prior art, and this embodiment does not make special restrictions on it and will not be elaborated here. The built knowledge exchange system can face different medical specialty fields and provide services such as knowledge retrieval, knowledge download, knowledge visualization, and knowledge exchange.
[0050] Based on the same inventive concept as the above method, an embodiment of the present application also provides a knowledge exchange system based on a knowledge graph, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-mentioned knowledge exchange methods based on a knowledge graph.
[0051] It can be understood that the above sequence of embodiments of the present application is only for description and does not represent the advantages or disadvantages of the embodiments. And the above description of specific embodiments of this specification has been made. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0052] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.
[0053] The above content is only the implementation manner of the present application and is not used to limit the scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, shall be included in the protection scope of the present application in the same way.
Claims
1. A knowledge exchange method based on a knowledge graph, characterized in that It includes the following steps: Extract the entity information, relationships between entities, and attribute information of entities in the knowledge database to obtain triple-structured data; Vectorize each triple-structured data to extract each bag-of-words vector, and then perform clustering analysis. Based on the average level of the bag-of-words vector similarity between each triple-structured data in the text clustering cluster and other triple-structured data, as well as the difference in the bag-of-words vector similarity, obtain the deviation interference outliers of each triple-structured data; Based on the deviation interference outliers, divide the triple-structured data to obtain a triple test set. Extract the attribute word frequency vectors of the two entities of each triple-structured data in the triple test set, analyze the correlation between the two attribute word frequency vectors, and combine the deviation interference outliers of each triple structure data to obtain the abnormal association eigenvalue between the two entities of each triple-structured data; Determine the target entity of each triple-structured data in the triple test set according to the change of the attribute word frequency vectors of the two entities of each triple-structured data in the triple test set. According to the degree of change difference of the attribute word frequency vectors between the target entity and another entity, combined with the abnormal association eigenvalue, obtain the redundancy information interference degree of the target entity of each triple-structured data in the triple test set; Based on the redundancy information interference degree, divide the target entities of all triple-structured data in the triple test set to obtain the entities to be corrected for information extraction. Correct the entities to be corrected according to the similarity between each entity to be corrected and other entities, and construct a knowledge graph and build a knowledge exchange system based on the corrected entity information to realize knowledge exchange based on the knowledge graph.
2. The knowledge exchange method based on a knowledge graph according to claim 1, wherein The calculation method of the deviation interference outliers of each triple-structured data is as follows: ; wherein, is the deviation interference outlier of the i-th triple structure data in the k-th text clustering cluster, is the data mean in the similarity deviation sequence of the i-th triple structure data in the k-th text clustering cluster, is the number of data in the similarity deviation sequence of the i-th triple structure data in the k-th text clustering cluster, and are respectively the j-th and (j-1)-th data in the similarity deviation sequence of the i-th triple structure data in the k-th text clustering cluster; Among them, calculate the mean value of the bag-of-words vector similarity between each triple-structured data in each text clustering cluster and all other triple-structured data as each triple-structured data in each text clustering cluster; calculate the absolute value of the difference in word frequency similarity between each triple-structured data in each text clustering cluster and other triple-structured data, and arrange all the absolute values from small to large to form the similarity deviation sequence of each triple-structured data in each text clustering cluster.
3. The knowledge exchange method based on a knowledge graph according to claim 1, wherein The method for obtaining the triple test set is: perform threshold segmentation on the deviation interference outliers of the triple-structured data in all text clustering clusters to obtain an abnormal segmentation threshold, and form a triple test set with the triple-structured data corresponding to the deviation interference degree greater than the abnormal segmentation threshold in all text clustering clusters.
4. The knowledge exchange method based on a knowledge graph according to claim 1, wherein The method for extracting the attribute word frequency vectors of the two entities of each triple-structured data in the triple test set is: extract the two entities of each triple-structured data in the triple test set, and input the attribute information of the two entities into the bag-of-words model respectively to obtain the attribute word frequency vectors of the two entities of each triple-structured data in the triple test set.
5. The knowledge exchange method based on a knowledge graph according to claim 1, wherein The calculation method of the abnormal association eigenvalue between the two entities of each triple-structured data is as follows: ; wherein, is the abnormal association feature value between two entities of the s-th triple structure data, is the deviation interference outlier of the s-th triple structure data, is the mutual information value between the two entity attribute word frequency vectors of the s-th triple structure data, is a constant to avoid a zero denominator.
6. The knowledge exchange method based on a knowledge graph according to claim 1, wherein The method for determining the target entity of each triple structure data is as follows: calculate the permutation entropy of the attribute word frequency vectors of the two entities in each triple structure data in the triple test set respectively, and take the entity with the largest permutation entropy as the target entity of each triple structure data in the triple test set.
7. The knowledge exchange method based on a knowledge graph according to claim 6, wherein The method for obtaining the redundancy information interference degree of the target entity of each triple structure data in the triple test set is as follows: calculate the difference of the permutation entropy between the target entity of each triple structure data and another entity, and take the normalized result of the product of the difference and the abnormal association eigenvalue as the redundancy information interference degree of the target entity of each triple structure data in the triple test set.
8. The knowledge exchange method based on a knowledge graph according to claim 1, wherein The method for obtaining the entity to be corrected is as follows: perform threshold segmentation on the redundancy information interference degree of the target entity of all triple structure data in the triple test set to obtain a redundancy segmentation threshold, and take the target entity corresponding to the redundancy information interference degree greater than the redundancy segmentation threshold as the entity to be corrected for information extraction.
9. The knowledge exchange method based on a knowledge graph according to claim 1, characterized in that, The method for correcting the entity to be corrected is as follows: after clustering the attribute word frequency vectors of all entities to be corrected, obtain each entity set, calculate the similarity between the entity to be corrected and other entities in its entity set with respect to the attribute word frequency vector, and replace the entity to be corrected with the entity corresponding to the maximum similarity to complete the correction process of the entity to be corrected.
10. A knowledge exchange system based on a knowledge graph, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the knowledge exchange method based on the knowledge graph according to any one of claims 1-9.
Citation Information
Patent Citations
An automatic construction method for a big data knowledge graph in the field of public security
CN109710701A
Attribute knowledge graph-based patent authorization prediction and evaluation method and system
CN117972113A
Text error correction method, apparatus, and device, and storage medium
WO2023005293A1
Cited By
Equipment operation and maintenance knowledge graph construction method and system based on AI
CN120806100A
An AI-based device operation and maintenance knowledge graph construction method and system
CN120806100B
Power grid professional knowledge graph construction method for green supply chain platform
CN121766416A