Multi-source heterogeneous knowledge graph data fusion method and system

Through the methods of data cleaning, semantic alignment and structure mapping, the structure, semantics, quality and update evolution problems in multi-source heterogeneous knowledge graph data fusion are solved, and high-quality knowledge graph data fusion and real-time update are achieved.

CN120067984APending Publication Date: 2025-05-30BEIJING SCI & TECH PATENT OFFICE

Patent Information

Application Number
CN202510138237.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively integrate multi-source heterogeneous knowledge graph data, and faces problems of structure, semantics, quality and update evolution.

Method used

Through data cleaning, semantic alignment and structure mapping, data from different data sources are converted into a unified format, and semantic models and structure mapping rules are established to achieve data fusion. At the same time, through data monitoring mechanisms and machine learning algorithms, the knowledge graphs are updated and evolved in real time.

Benefits of technology

Effectively reduce noise, repetition and semantic inconsistency problems, improve the accuracy and consistency of knowledge graphs, adapt to the development of domain knowledge, and maintain validity and practicality in the long run.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067984A_ABST
    Figure CN120067984A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source heterogeneous knowledge graph data fusion method and system, and the method specifically comprises the following steps: S1, collecting data from a plurality of heterogeneous data sources, cleaning the collected data, removing noise data and repeated data, and converting the data in different formats into a unified intermediate format; s2, a semantic model is established, and semantic heterogeneous concepts and relations in different data sources are recognized and aligned through a natural language processing technology and domain ontology knowledge. The invention relates to the technical field of knowledge graph data processing, and the multi-source heterogeneous knowledge graph data fusion method and system effectively reduce the problems of noise, repetition and semantic inconsistency in data through data cleaning, semantic alignment and conflict resolution, and ensure the accuracy and consistency of fused data, thereby improving the quality of a knowledge graph and improving the reliability of the knowledge graph. Reliable knowledge support is provided for application in the fields of medical treatment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph data processing, and specifically provides a method and system for multi-source heterogeneous knowledge graph data fusion. Background Art

[0002] With the rapid development of information technology, knowledge graphs have been widely applied in various fields. For example, in the medical field, they are used for disease diagnosis assistance; in the financial field, for risk assessment; and in the e-commerce field, for personalized recommendation, etc. However, the data sources of knowledge graphs are often diverse, including but not limited to databases with different structures, text files, web data collection, etc. The data from these different sources has the characteristics of multi-source heterogeneity.

[0003] Currently, in fields such as medicine, knowledge graphs are widely used but face challenges in multi-source heterogeneous data fusion. First, the data sources are diverse, including relational databases, medical literature, and online medical forums, etc., and their structures are different. For example, relational databases are in tabular form, while text data has an irregular structure and is difficult to integrate. Second, there are semantic heterogeneities in different data sources, with different expressions for the same concept or relationship, which easily causes confusion in the knowledge graph information. Third, the data quality is uneven, some has noise and errors, and some is not updated in a timely manner. Fourth, the update frequencies and methods of each data source are different, and the knowledge graph requires an effective update and evolution mechanism to adapt to changes. For this reason, the present invention provides a method and system for multi-source heterogeneous knowledge graph data fusion. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a method and system for multi-source heterogeneous knowledge graph data fusion, which overcomes the problems of structure, semantics, quality, and update evolution in the existing data fusion process.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for multi-source heterogeneous knowledge graph data fusion, characterized in that it specifically includes the following steps:

[0006] S1. Collect data from multiple heterogeneous data sources, clean the collected data to remove noise data and duplicate data, and at the same time convert data in different formats into a unified intermediate format;

[0007] S2. Establish a semantic model, and identify and align the semantically heterogeneous concepts and relationships in different data sources through natural language processing technology and domain ontology knowledge;

[0008] S3. Analyze the data structure characteristics of different data sources, establish structure mapping rules, and map data with different structures into a unified knowledge graph structure model;

[0009] S4. Integrate the preprocessed data into the knowledge graph according to the results of semantic alignment and structure mapping, and handle data conflicts by formulating conflict resolution strategies;

[0010] S5. Establish a data monitoring mechanism to monitor the changes of data sources in real time or regularly. When data updates are detected, update the knowledge graph accordingly based on the updated content. At the same time, mine new knowledge through machine learning algorithms to evolve the knowledge graph.

[0011] Preferably, in S1, different acquisition methods are adopted for different types of data sources, including acquiring relational database data through database connection drivers, acquiring text file data through file reading operations, and acquiring network data sources data through web crawler technology.

[0012] Preferably, in S2, when calculating the word vector similarity between concepts, determine semantically equivalent concepts and relationships by combining medical domain dictionaries and knowledge ontologies, where the word vector model uses a pre-trained general word vector model or a word vector model trained on specific domain data.

[0013] Preferably, in S3, for the mapping from a relational database to the knowledge graph structure, map the records in the table to nodes in the knowledge graph, and map the data in different tables to edges in the knowledge graph through foreign key relationships; for the triples in text data, directly map the entities in the triples to nodes in the knowledge graph and the relationships to edges.

[0014] Preferably, the conflict resolution strategy includes determining the ultimately retained data according to the credibility of the data source and the data update time, where the credibility evaluation of the data source is based on factors such as the authority of the data source and the reliability of the data acquisition method.

[0015] Preferably, the fusion system of the multi-source heterogeneous knowledge graph data fusion method includes:

[0016] A data acquisition module for connecting to multiple heterogeneous data sources to obtain data;

[0017] A preprocessing module for cleaning and format conversion operations on the acquired data;

[0018] A semantic alignment module for implementing the construction of a semantic model and the semantic alignment function;

[0019] A structure mapping module for creating mapping rules according to the data structure characteristics and performing structure mapping;

[0020] A data fusion and conflict resolution module for completing data fusion and conflict handling;

[0021] Knowledge Graph Update and Evolution Module, which monitors changes in data sources and updates the Knowledge Graph, and at the same time realizes the evolution of the Knowledge Graph.

[0022] Preferably, the data acquisition module includes a database connection sub-module, a file reading sub-module, and a web crawling sub-module, which are respectively used to collect data from relational databases, text file data, and web data sources.

[0023] Preferably, the Knowledge Graph Update and Evolution Module includes a data monitoring sub-module and a knowledge mining sub-module. The data monitoring sub-module is used to monitor changes in data sources in real time or regularly, and the knowledge mining sub-module is used to mine new knowledge through machine learning algorithms to realize the evolution of the Knowledge Graph.

[0024] Beneficial Effects

[0025] The present invention provides a method and system for multi-source heterogeneous Knowledge Graph data fusion. Compared with the prior art, it has the following beneficial effects:

[0026] (1) This method and system for multi-source heterogeneous Knowledge Graph data fusion effectively reduce noise, duplication, and semantic inconsistency problems in data through data cleaning, semantic alignment, and conflict resolution, ensuring the accuracy and consistency of the fused data, thereby improving the quality of the Knowledge Graph and providing reliable knowledge support for applications in fields such as healthcare.

[0027] (2) This method and system for multi-source heterogeneous Knowledge Graph data fusion can update the Knowledge Graph in a timely manner to reflect changes in data sources and mine new knowledge through data monitoring and knowledge mining mechanisms, enabling the Knowledge Graph to adapt to the development of domain knowledge and maintain effectiveness and practicality in the long term. Description of the Drawings

[0028] Figure 1 It is the method flow chart of the present invention.

[0029] Figure 2 It is the system block diagram of the present invention. Detailed Embodiments

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] Please refer to Figure 1 - Figure 2 The present invention provides two technical solutions:

[0032] Embodiment 1, as Figure 1As shown in the figure, the present invention discloses a method for fusing multi-source heterogeneous knowledge graph data, specifically including the following steps:

[0033] S1. Data collection and preprocessing;

[0034] S11. Data collection:

[0035] Suppose there are three heterogeneous data sources: the first is a relational database in the medical field (including patient information tables, medical record tables, doctor information tables, etc.), the second is a medical literature text file, and the third is data from an online medical forum.

[0036] For the relational database, establish a connection through a database connection driver (such as JDBC), and use SQL query statements to obtain data from relevant tables. For example, obtain information such as the patient's name, age, gender, and contact information from the patient information table, obtain the patient's medical history, diagnosis results, treatment process, etc. from the medical record table, and obtain data such as the doctor's name, title, and department from the doctor information table.

[0037] For the medical literature text file, read the text content line by line through file reading operations. You can use the file reading function in a programming language to read the text content into memory for subsequent processing.

[0038] For the data from the online medical forum, use web crawler technology to extract information such as the post title, content, author, and publication time according to the forum's page structure and relevant rules. For example, by analyzing the HTML tag structure where the post is located, locate the elements containing key information, and then extract the text content.

[0039] S12. Data processing:

[0040] Among the data obtained from the relational database, check whether the ID number field in the patient information table meets the format requirements, and delete the records that do not meet the requirements. At the same time, check whether the content in the medical record table is complete, and mark or delete the records lacking key diagnosis information. In addition, remove duplicate medical record records with exactly the same content in the medical record table.

[0041] In the medical literature text file, remove noise information such as punctuation marks and stop words in the text. You can use the stop word list and punctuation mark processing functions in a natural language processing library to achieve this. At the same time, remove the garbled characters generated due to format conversion and other reasons.

[0042] In the data from the online medical forum, delete the content of posts with an advertising nature and obvious false information that does not conform to medical logic. Determine the validity of the post content by setting keyword filtering rules and simple semantic analysis.

[0043] S13. Format conversion:

[0044] Convert the data in the relational database into intermediate data in JSON format. For example, convert a record {"name": "Zhang San", "age": 30, "gender": "male"} in the patient information table into the form of {"entity_type": "patient", "properties": {"name": "Zhang San", "age": 30, "gender": "male"}}. For the data in the medical record table, associate each medical record with the patient information to form a more complex JSON structure, such as {"patient": {"name": "Zhang San", "age": 30}, "medical_record": {"diagnosis": "cold", "treatment": "taking medicine"}}.

[0045] For medical literature texts, split them into sentences through natural language processing tools, and perform part-of-speech tagging and named entity recognition on each sentence. For example, for the sentence "The patient has symptoms such as fever and cough", identify "patient" as the entity, "has" as the relational word, and "fever" and "cough" as symptom entities. Convert the identified entity and relational information into a JSON-like format, such as {"entity": {"patient": "not mentioned by name", "symptoms": ["fever", "cough"]}, "relation": {"patient-symptom": "not mentioned by name-fever; not mentioned by name-cough"}}.

[0046] For online medical forum data, extract the entities and relationships in the post content and convert them into JSON format as well. For example, for the post "My friend has diabetes and doesn't know what to do", extract it as {"entity": {"patient": "friend (not mentioned by name)", "disease": "diabetes"}, "relation": {"patient-disease": "friend (not mentioned by name)-diabetes"}}.

[0047] S2. Semantic alignment;

[0048] S21. Build a semantic model: Based on the authoritative dictionary in the medical field (such as the Medical Subject Headings) and the existing medical knowledge ontology, build a semantic model. For example, in the medical knowledge ontology, concepts such as diseases, symptoms, treatment methods, examination items, patients, doctors, etc. are defined, as well as the relationships between them (such as diseases have symptoms, diseases have treatment methods, doctors treat patients, examination items are used to diagnose diseases, etc.). Ontology modeling tools can be used to create and manage the semantic model, representing each concept and relationship in a structured way for subsequent semantic alignment operations.

[0049] S22. Semantic alignment operation: Calculate the word vector similarity between concepts in JSON format data obtained from different data sources. For example, for "coronary atherosclerotic heart disease" mentioned in medical literature and "coronary heart disease" mentioned in an online medical forum, calculate their similarity through a word vector model. A pre-trained word vector model (such as Word2Vec, GloVe, etc.) or a word vector model trained on a medical domain-specific corpus can be used. Combining a medical domain dictionary and knowledge ontology, determine that they are semantically equivalent concepts. For the alignment of relationships, such as the relationship between a doctor and a department in a relational database being "affiliated department" and being expressed as "working in the department" in medical literature, through analyzing sentence structure and semantic role annotation, and combining medical domain knowledge, determine that these two relationships are the same and establish a mapping relationship in the semantic model. Tools for syntactic analysis and semantic role annotation in natural language processing technology can be used to assist in the analysis.

[0050] S3. Structural mapping;

[0051] S31. Analyze the characteristics of the data structure: The data structure of a relational database is in tabular form, where patient information tables, medical record tables, doctor information tables, etc. are associated through foreign keys. For example, the patient ID field in the medical record table is associated with the primary key ID in the patient information table to establish the correspondence between patient information and medical records. The data formed after processing medical literature text is in the form of entity-relationship-entity triples, and the data extracted from online medical forum data is also in a similar triple form, but the types of entities and relationships may be different from those in medical literature. For example, the triples in medical literature may focus more on academic disease-treatment relationships, while the triples in online medical forums may be more related to patients' individual experiences and problems;

[0052] S32. Establish structure mapping rules: For the mapping from a relational database to the knowledge graph structure, each record in the patient information table is mapped to a patient node in the knowledge graph. The attributes of the node include information such as name, age, and gender extracted from the patient information table. The records in the medical record table are associated with the patient nodes according to the patient ID. The diagnostic information, treatment process, etc. in the medical records are used as attributes of the patient nodes or are connected to the patient nodes through relationship edges. For example, the diagnostic information can be used as an attribute value of the patient node, and the treatment process can be connected to the patient node through a relationship edge such as "received treatment", and the other end of the edge points to the node representing the treatment method. The records in the doctor information table are mapped to doctor nodes, and the node attributes include information such as the doctor's name and title. The doctor nodes are connected to the department nodes through the "belongs to department" relationship, and the doctor nodes are connected to the patient nodes through the "treatment" relationship (if the doctor has treated the patient). For the mapping of triples in medical literature and online medical forum data to the knowledge graph structure; directly map the entities in the triples to nodes in the knowledge graph and the relationships to edges. For example, for the triple ("Patient A", "has", "diabetes") in medical literature, create a patient A node and a diabetes node in the knowledge graph and connect them through a "has" relationship edge. For the triple ("Patient B", "asks", "diabetes treatment cost") in an online medical forum, create a patient B node, a "diabetes treatment cost" node, and connect them through an "asks" relationship edge.

[0053] S4. Data fusion and conflict resolution;

[0054] S41. Data fusion: According to the results of semantic alignment and structure mapping, fuse the data from different preprocessed data sources into the knowledge graph. For example, fuse the medical record information of patient Zhang San obtained from the relational database, the latest research results on the diseases suffered by Zhang San obtained from medical literature, and the feedback information of patient Zhang San on treatment obtained from the online medical forum into the patient node corresponding to Zhang San in the knowledge graph and its related relationships and attributes. During the fusion process, pay attention to the data integration method to ensure the consistency of the structure and semantics of the knowledge graph. For example, for multiple diagnostic information of the same patient, it should be reasonably merged into the diagnostic attribute of the patient node, or represented by a time series relationship edge to indicate the diagnostic situations at different times.

[0055] S42. Conflict Resolution: Suppose during the fusion process, the age record of patient Zhang San in the relational database is 30 years old, while someone mentions in the online medical forum that patient Zhang San is 32 years old. At this time, based on the credibility assessment of the data sources (the relational database is official hospital data with relatively high credibility) and the data update time (assuming the relational database was updated recently), it is determined to retain the age information in the relational database. For other possible conflict situations, such as inconsistent diagnostic criteria for the same disease in different data sources, factors such as the authority of the data sources, the integrity of the data, and the consistency with other relevant data need to be comprehensively considered to determine the ultimately retained data. A conflict resolution rule library can be established, and corresponding resolution strategies can be set according to different types of conflicts.

[0056] S5. Knowledge Graph Update and Evolution;

[0057] S51 Data Monitoring and Update: Set up a scheduled task to check whether the data in the relational database, medical literature, and online medical forum are updated every certain period (such as one day). If there is new patient medical record information in the relational database, analyze the entities and relationships in the new medical records and add them to the knowledge graph. For example, if the new medical record records the new diagnosis results and treatment plans of patient Li Si, fuse this information into the relevant attributes and relationships of the Li Si patient node in the knowledge graph. If there are new disease research results in the medical literature, update the attributes and relationships of the relevant diseases in the knowledge graph. For example, if a new study discovers a new pathogenic gene for a certain disease, add this information to the relevant attributes of the disease node in the knowledge graph and update the relationships related to the diagnosis and treatment of this disease. For the newly published content of users in the online medical forum, extract the entities and relationships from it and update the knowledge graph accordingly. For example, if a user posts a new experience about the side effects of a certain drug, integrate this information into the relevant attributes and relationships of the drug node in the knowledge graph.

[0058] S52. Knowledge Graph Evolution: Analyze the patient data in the knowledge graph through machine learning algorithms (such as clustering algorithms) to discover patient groups with similar symptoms. For example, use the K-Means clustering algorithm to cluster patients with similar symptoms (such as fever, cough, fatigue, etc.) together. Based on this, new knowledge is mined, such as these patient groups may have a better response to a certain new treatment method. Add this new knowledge to the knowledge graph to achieve the evolution of the knowledge graph. For example, add new attributes or relationships to the patient group nodes after clustering in the knowledge graph to indicate their potential response to a certain new treatment method, so as to better support medical decision-making.

[0059] Example 2, as Figure 2 shown, the difference compared with Example 1 is:

[0060] The present invention also discloses a multi-source heterogeneous knowledge graph data fusion system;

[0061] In an embodiment of the present invention, the data acquisition module includes:

[0062] A1. Database connection sub-module: Responsible for establishing a connection with the relational database in the medical field. Using specific database connection technologies (such as JDBC), a stable connection is established according to the configuration information of the database (including database address, port, username, password, etc.). A query interface is provided to allow other modules to obtain the required data from the relational database through SQL query statements. Parameterized processing can be performed on the query statements to improve the flexibility and security of the query.

[0063] A2. File reading sub-module: For medical literature text files, implement the file reading function. Support the reading operations of multiple text file formats (such as.txt,.pdf, etc.). Appropriate reading methods can be selected according to the file type. For.txt files, the text content can be directly read line by line, and for.pdf files, a dedicated PDF text extraction library can be used to obtain the text content. The read text content is passed to the subsequent preprocessing module for processing.

[0064] A3. Web crawler sub-module: Used to collect data from online medical forums. Using web crawler technology, customized development is carried out according to the website address and page structure characteristics of the forum. The parameters of the crawler can be configured, such as crawling depth, crawling frequency, user agent, etc., to ensure legal and efficient acquisition of forum data. The obtained data is initially sorted and filtered, and some obvious invalid data (such as navigation bars, advertisements, and other irrelevant information on the web page) is removed, and then the valid data is passed to the preprocessing module.

[0065] In an embodiment of the present invention, the preprocessing module includes:

[0066] B1. Data cleaning unit: Receive data from different data sources of the data acquisition module and perform cleaning operations on it. For the data obtained from the relational database, check the integrity and legality of the data. For example, check whether the required fields are empty and whether the data format conforms to the preset rules (such as date format, number format, etc.). Repair or delete the data that does not meet the requirements. For the data in medical literature text files, remove the noise information, such as punctuation marks, stop words, garbled characters, etc. It can be implemented using text preprocessing algorithms in natural language processing technology. For online medical forum data, delete invalid content such as advertisements and false information according to the preset rules. For example, judge whether the post content is an advertisement or false information through keyword matching and simple semantic analysis.

[0067] B2. Format Conversion Unit: Convert the data from different data sources after cleaning into a unified intermediate format (such as JSON). For the data in a relational database, according to the database table structure and the pre-designed mapping rules, convert the tabular data into JSON-format objects. For example, convert a record in the patient information table into a JSON object, where each field serves as an attribute of the object. For the medical literature text data, further analyze and convert the text content after cleaning and preprocessing. Through the named entity recognition and relation extraction algorithms in natural language processing technology, convert the entity and relation information in the text into JSON format. For the online medical forum data, also convert the extracted entity and relation information into JSON format for subsequent semantic alignment and structure mapping operations.

[0068] In the embodiment of the present invention, the semantic alignment module includes:

[0069] C1. Semantic Model Construction Unit: Construct a semantic model based on the authoritative materials in the medical field (such as the Medical Subject Headings and the existing medical knowledge ontology). Use an ontology modeling tool or a dedicated semantic model construction software to create a hierarchical structure of concepts and relationships. For example, classify and define concepts such as diseases, symptoms, and treatment methods, and clarify the hierarchical relationships and semantic connections between them. Continuously update and improve the semantic model to adapt to the development of medical field knowledge and the addition of new data sources.

[0070] C2. Semantic Alignment Execution Unit: Receive the JSON-format data from the preprocessing module and perform semantic alignment operations according to the semantic model. For the concepts in the data, determine the semantically equivalent concepts by calculating the word vector similarity and combining the medical field dictionary and knowledge ontology. Efficient word vector calculation algorithms and semantic matching algorithms can be used to improve the accuracy and speed of alignment. For the alignment of relationships, use the syntactic analysis and semantic role annotation methods in natural language processing technology to analyze the relationship expressions in the sentences and match them with the relationship definitions in the semantic model to determine the equivalent relationships. Pass the alignment results to the structure mapping module.

[0071] In the embodiment of the present invention, the structure mapping module includes:

[0072] D1. Data Structure Analysis Unit: Analyze the data structure characteristics of different data sources. For a relational database, obtain the table structure information, including table names, column names, primary keys, foreign keys, etc., by querying the database system tables (such as the relevant tables in the information_schema library in MySQL), so as to understand the association method between the data. For the medical literature text and online medical forum data, analyze the structure characteristics of the JSON-format data after preprocessing and semantic alignment, including the types and distributions of entities and relationships, etc.

[0073] D2. Structural mapping rule execution unit: formulate and execute structural mapping rules according to the data structure analysis results. For the mapping of relational database to knowledge graph structure, write a conversion program to convert the table data in the database into nodes and edges in the knowledge graph according to the mapping rules. For example, convert the records in the patient information table into patient nodes, associate the data in the medical record table with the patient node through the foreign key relationship, and create the corresponding relationship edges. For the mapping of triples of medical literature and online medical forum data to the knowledge graph structure, map the entities in the JSON format data into nodes and the relationships into edges. Ensure the accuracy and consistency of the mapping process, and pass the structural mapping results to the data fusion and conflict resolution module.

[0074] In the embodiment of the present invention, the data fusion and conflict resolution module includes:

[0075] E1. Data fusion unit: Receive data from the structure mapping module and fuse it into the knowledge graph. According to the results of semantic alignment and structure mapping, determine the location and fusion method of the data in the knowledge graph. For example, suppose we are fusing information about the patient "Zhang San" from a relational database, medical literature, and an online medical forum. From the relational database, we get Zhang San's age as 30 years old and the diagnosis information as "cold". The credibility of this data is assessed as 80% (based on its official system from the hospital, it has high authority); from the medical literature, we get an analysis of Zhang San's condition, mentioning "complications that may be caused by colds and corresponding preventive measures", and the credibility is assessed as 70% (because it is based on medical research, but there may be some deviations due to the limitations of the research); from the online medical forum, there are users describing Zhang San's symptoms and personal treatment experience, and the credibility is assessed as 60% (due to the uncertainty of its information source). When the information from these different data sources is integrated into the nodes and relationships corresponding to Zhang San in the knowledge graph, for entities of the same family (such as the patient entity "Zhang San"), attributes (such as "age", "diagnosis information", etc.) or relationships (such as the relationship "Zhang San has a cold"), as long as their credibility reaches 50% (the set credibility threshold), they will be included in the fusion scope to achieve inclusiveness. Using the operation interface of the graph database (such as Neo4j), the data is inserted into the knowledge graph through Cypher query statements, and associated operations are established to ensure the integrity and consistency of the fused knowledge graph data. Through this fusion method, data of different credibility can be reasonably organized in the knowledge graph, providing richer data information for subsequent knowledge graph applications, and avoiding the loss of potentially valuable data due to a single data source or a single credibility judgment.

[0076] E2. Conflict Resolution Unit: Handles possible conflict situations. First, a clear credibility threshold of 75% is set, and this threshold can be dynamically adjusted according to specific application scenarios and data characteristics. When inconsistent data for the same entity or relationship is detected in different data sources, a decision is made based on this credibility threshold. For example, for the age information of patient "Li Si", the relational database records it as 45 years old with a credibility of 85% (because it is the latest record from the hospital), while someone mentioned on an online medical forum that Li Si is 47 years old with a credibility of 60%. Since the data in the relational database has a credibility exceeding the 75% threshold, the age information in the relational database is finally retained. If the credibility of multiple data exceeds the threshold, for example, for the diagnosis information of patient "Wang Wu", the relational database shows "hypertension" with a credibility of 82%, and the medical literature indicates it may be "primary hypertension" with a credibility of 78%, then these data are all retained in the knowledge graph, and their respective credibilities are stored as key attributes. For the credibility assessment of data sources, in addition to considering the authority of the data source itself (such as relational databases are generally considered authoritative sources with higher credibility; medical literature is evaluated based on factors such as the impact factor of the publishing journal and the authority of the authors; online medical forums are evaluated according to user authentication and information rationality), factors such as data update time (e.g., more recently updated data may be more credible) and data integrity (data with complete information has higher credibility) are also comprehensively considered. To better manage conflict resolution strategies, a complete conflict resolution rule library is established, which contains detailed rules for different types of data conflicts. For example, for age conflicts, data with high credibility and recent update time are given priority; for diagnosis information conflicts, the authority, integrity of the data source and consistency with other relevant data are comprehensively considered. At the same time, a corresponding evaluation mechanism is constructed, which can automatically call the corresponding rules for processing according to different types of conflict scenarios, such as conflicts of different severity levels for age conflicts, diagnosis information conflicts, treatment plan conflicts, etc., and the importance of the entities or relationships involved in the conflict data (e.g., key diagnosis information is more important than minor symptom information), so as to ensure the quality of the knowledge graph, provide reliable data support for subsequent knowledge graph applications, and further provide better data basis for optimizing services.

[0077] In the embodiments of the present invention, the knowledge graph update and evolution module includes:

[0078] F1. Data Monitoring Sub-module: Establish a data monitoring mechanism to check the changes in each data source in real-time or periodically. For relational databases, the built-in functions of the database (such as triggers, logs, etc.) or custom query mechanisms can be used to detect data updates. For example, by setting database triggers, when new patient records are inserted or existing records are modified, corresponding events are triggered to notify the data monitoring sub-module. For medical literature and online medical forum data, the update detection function of web crawlers or the method of re-crawling periodically is used to check for data updates. For example, by comparing the hash values or update timestamps of web page content to determine whether new content has been published on medical literature web pages. When data updates are detected in the data source, other sub-modules of the knowledge graph update and evolution module are notified for corresponding processing.

[0079] F2. Knowledge Mining Sub-module: Based on machine learning and data mining techniques, analyze and mine the data in the knowledge graph to achieve the evolution of the knowledge graph. Use clustering algorithms (such as DBSCAN, hierarchical clustering, etc.) to perform clustering analysis on the patient data in the knowledge graph, and discover potential patient group patterns according to the characteristics of patients such as symptoms, diagnosis results, treatment responses, etc. Analyze the association patterns between different entities and relationships in the knowledge graph through association rule mining algorithms (such as Apriori algorithm, FP-Growth algorithm, etc.). For example, mine the frequent association patterns between diseases, symptoms, and treatment methods, discover a group of symptoms that often accompany a certain disease, or the effectiveness of a certain treatment method for specific disease and symptom combinations. Based on the newly mined knowledge, update the structure and semantic information of the knowledge graph. For example, if a new disease-symptom association pattern is discovered, add the corresponding relationship edge in the knowledge graph; if there is new information on the effectiveness of a treatment method, update the relevant attributes of the treatment method node and its relationships with other nodes.

[0080] In addition, use classification algorithms (such as decision tree algorithms, support vector machines, etc.) to classify and predict the data newly added to the knowledge graph, further improving the knowledge system of the knowledge graph. For example, predict the possible disease categories that a patient may have based on the patient's symptoms, examination results, etc., and use the prediction results as a potential attribute of the patient node in the knowledge graph to provide more reference information for medical diagnosis. At the same time, combined with the graph neural network (GNN) technology in deep learning, perform representation learning on the knowledge graph to better capture the complex semantic information between entities and relationships, further enhancing the ability of the knowledge graph in knowledge mining and evolution to adapt to the integration needs of constantly changing medical domain knowledge and multi-source heterogeneous data. Through such a complete data fusion system and method, multi-source heterogeneous medical knowledge graph data can be effectively fused, and the high-quality update and evolution of the knowledge graph can be ensured, providing a solid data support and knowledge foundation for intelligent applications in the medical field.

[0081] Furthermore, in the evolution process of the knowledge graph, in order to better manage and maintain the knowledge graph, a dynamic update and traceability mechanism is introduced:

[0082] Regarding dynamic update: The dynamic update mechanism of the knowledge graph mainly involves data source monitoring to trigger updates and the coordination of knowledge graph updates. For data source monitoring, in a relational database, database triggers can be utilized, which will be automatically activated during INSERT, UPDATE, and DELETE operations, and notify the data monitoring sub-module through stored procedure or function calls. At the same time, a custom query mechanism can also be used to periodically query the timestamp or version number of the latest records and compare them with the previous results. Once there is a difference, an update is triggered. For medical literature and online medical forum data, web crawlers will judge updates by calculating the hash values of web page elements (such as MD5, SHA-256) or comparing update timestamps, and can also adopt the method of periodically re-crawling, comparing new content with old content to discover updates and trigger knowledge graph updates. In terms of the coordination of knowledge graph updates, when new data enters, entity and relationship matching and conflict handling will be carried out. Whether an entity already exists is judged based on the unique identifier or attribute similarity, and strategies such as giving priority to the latest data or weighted averaging based on the credibility of the data source are adopted to resolve relationship conflicts. In addition, the knowledge mining sub-module will be updated accordingly. The clustering algorithm updates the clustering results based on the feature vector distance, the association rule mining algorithm recalculates the frequent item sets to update the association, the classification algorithm updates the model and classifies new data to add the results, and the graph neural network re-learns the node and edge representations to integrate new elements into the updated semantics.

[0083] Regarding traceability: The traceability mechanism provides convenience for the maintenance of the knowledge graph. It is mainly implemented through version control, transaction mechanism, and logging system. In version control, a unique identifier (such as a version number or timestamp) is assigned to each update state of the knowledge graph, and the update state is stored in an independent location so that the corresponding state can be restored according to the identifier, which is similar to the snapshot technology in software version control. The transaction mechanism encapsulates update operations in database transactions to ensure the atomicity, consistency, and isolation of operations. Once an error occurs during the update, a rollback operation can be performed to restore the previous state. The logging system will record the information of update operations in detail, including the operation type (such as adding nodes, updating relationships, deleting nodes), operation time, data content of the operation, and the source of the operation. These information are stored in chronological order, which helps to find problems and trace back the historical state of the knowledge graph.

[0084] In summary, through data cleaning, semantic alignment, and conflict resolution, the problems of noise, duplication, and semantic inconsistency in the data are effectively reduced, ensuring the accuracy and consistency of the fused data, thereby improving the quality of the knowledge graph and providing reliable knowledge support for applications in fields such as healthcare. Moreover, through the data monitoring and knowledge mining mechanisms, the knowledge graph can be updated in a timely manner to reflect changes in the data sources and new knowledge can be mined, enabling the knowledge graph to adapt to the development of domain knowledge and maintain its effectiveness and practicality in the long term.

[0085] Meanwhile, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

[0086] It should be noted that, in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device.

[0087] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multi-source heterogeneous knowledge graph data fusion method, characterized by: The specific steps include: S1. Collect data from multiple heterogeneous data sources, clean the collected data, remove noise data and duplicate data, and convert data in different formats into a unified intermediate format; S2. Establish a semantic model to identify and align semantically heterogeneous concepts and relationships in different data sources through natural language processing technology and domain ontology knowledge; S3. Analyze the data structure characteristics of different data sources, establish structure mapping rules, and map data with different structures into a unified knowledge graph structure model; S4. According to the results of semantic alignment and structural mapping, the preprocessed data is integrated into the knowledge graph, and data conflicts are handled by formulating conflict resolution strategies; S5. Establish a data monitoring mechanism to monitor changes in data sources in real time or regularly. When data updates are detected, the knowledge graph is updated accordingly based on the updated content. At the same time, new knowledge is mined through machine learning algorithms to evolve the knowledge graph.

2. According to claim 1, a multi-source heterogeneous knowledge graph data fusion method is characterized by: In S1, different collection methods are used for different types of data sources, including collecting relational database data through database connection driver, collecting text file data through file reading operation, and collecting network data source data through network crawler technology.

3. A multi-source heterogeneous knowledge graph data fusion method according to claim 2, characterized in that: In S2, when calculating the word vector similarity between concepts, the domain dictionary and knowledge ontology are combined to determine semantically equivalent concepts and relationships.

4. A multi-source heterogeneous knowledge graph data fusion method according to claim 3, characterized in that: In the S3, for the mapping of relational database to knowledge graph structure, the records in the table are mapped to nodes in the knowledge graph, and the data in different tables are mapped to edges in the knowledge graph through foreign key relationships; for triples in text data, the entities in the triples are directly mapped to nodes in the knowledge graph, and the relationships are mapped to edges.

5. The multi-source heterogeneous knowledge graph data fusion method according to claim 1 is characterized by: The conflict resolution strategy includes determining the data to be finally retained according to the credibility of the data source and the data update time.

6. A fusion system for a multi-source heterogeneous knowledge graph data fusion method according to any one of claims 1 to 5, characterized in that: include: Data acquisition module, used to connect with multiple heterogeneous data sources to obtain data; The preprocessing module cleans and converts the collected data; Semantic alignment module, which realizes the construction of semantic model and semantic alignment function; The structure mapping module creates mapping rules and performs structure mapping according to the characteristics of data structure; Data fusion and conflict resolution module, completes data fusion and conflict processing; The knowledge graph update and evolution module monitors changes in data sources and updates the knowledge graph, while also realizing the evolution of the knowledge graph.

7. A multi-source heterogeneous knowledge graph data fusion system according to claim 6, characterized in that: The data acquisition module includes a database connection submodule, a file reading submodule and a network crawler submodule, which are respectively used to collect relational database data, text file data and network data source data.

8. A multi-source heterogeneous knowledge graph data fusion system according to claim 6, characterized in that: The knowledge graph update and evolution module includes a data monitoring submodule and a knowledge mining submodule. The data monitoring submodule is used to monitor the changes of the data source in real time or periodically, and the knowledge mining submodule is used to mine new knowledge through a machine learning algorithm to realize the evolution of the knowledge graph.

Citation Information

Patent Citations

  • Knowledge graph construction method and system of power system

    CN110727741A

  • Multi-source data investigation and real-time analysis method

    CN116894152A

  • Knowledge graph construction method and device, computer equipment and storage medium

    CN117271802A

  • Non-residual knowledge graph construction method based on multi-source heterogeneous data fusion

    CN117851609A

  • Clinical examination result analysis method based on medical knowledge graph

    CN118538429A

Cited By

  • Method and device for tracing beef producing area based on artificial intelligence, and electronic equipment

    CN120258845A

  • Multi-source heterogeneous astronomical data management method and device and medium

    CN120295970A

  • Territorial space planning multi-source heterogeneous data intelligent integration system based on semantic analysis

    CN120336940A

  • Intelligent integration system of multi-source heterogeneous data for national land spatial planning based on semantic analysis

    CN120336940B

  • Multi-source heterogeneous data alignment method and system

    CN120561612A