Heterogeneous Knowledge Graphing Method and System Based on Fusion Mapping
By acquiring and analyzing heterogeneous knowledge data sources and generating and integrating knowledge graphs, the integration difficulties caused by the diversity of knowledge data are solved, and efficient knowledge acquisition and utilization are achieved.
Patent Information
- Application Number
- CN202411018875.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2044-07-29
AI Technical Summary
In the prior art, due to the wide and diverse sources of knowledge data, the format, structure and semantic differences are huge, making it difficult to directly integrate, lack of association, inefficient search Q&A, difficulty in updating and expanding, and limited use of cross-system knowledge.
By acquiring multiple heterogeneous knowledge data sources, analyzing the entities, relationships and attributes of the acquired knowledge data, outputting the initial knowledge graph, and collecting heterogeneous fusion models for fusion, generating a fusion knowledge graph, connecting Q&A and search modules, realizing efficient integration and utilization of data sources.
It enables users to obtain and utilize the required knowledge more accurately, timely and conveniently, and promotes the efficient use and sharing of knowledge.
Smart Images

Figure CN118885625B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a heterogeneous knowledge graphing method and system based on fusion mapping. Background Art
[0002] In the field of knowledge management and utilization, the problems caused by the diversification and heterogeneity of knowledge sources are extremely prominent, and the demand contradiction of knowledge management and utilization is becoming more and more prominent. Achieving the effective integration and efficient utilization of heterogeneous knowledge has become a crucial link in promoting the development of knowledge management and utilization. Traditional knowledge management methods are often relatively limited and scattered, only focusing on specific data formats or single data sources, lacking the overall integration of multiple heterogeneous knowledge data sources, and lacking a comprehensive analysis and utilization of knowledge entities, relationships, and attributes. It is difficult to clearly and accurately construct the association structure between knowledge, and there are inaccurate situations in knowledge integration, resulting in an imperfect construction of the knowledge graph, and it is difficult to update and expand knowledge, and it cannot well cope with the changing knowledge demand situation.
[0003] In the current related technologies, there are technical problems such as huge differences in format, structure, and semantics caused by the wide range and diversity of knowledge data sources, making it difficult to directly integrate, missing associations, low efficiency of search and question answering, difficult to update and expand, and limited cross-system knowledge utilization. Summary of the Invention
[0004] This application provides a heterogeneous knowledge graphing method and system based on fusion mapping. By obtaining multiple heterogeneous knowledge data sources including multiple heterogeneous knowledge data tables, analyzing and obtaining the entities, relationships, and attributes of knowledge data, mapping and outputting an initial knowledge graph with these, collecting a heterogeneous fusion model to fuse it into a fused knowledge graph, and connecting this graph with a question answering and search module, which perform question answering and search on the data source based on the graph, it achieves the technical effect of enabling users to obtain and utilize the required knowledge more accurately, timely, and conveniently, and promoting the efficient utilization and sharing of knowledge.
[0005] This application provides a heterogeneous knowledge graphing method based on fusion mapping, including:
[0006] Obtain multiple heterogeneous knowledge data sources, where the multiple heterogeneous knowledge data sources include multiple heterogeneous knowledge data tables; analyze the multiple heterogeneous knowledge data sources to obtain knowledge data entities, knowledge data relationships, and knowledge data attributes; map the multiple heterogeneous knowledge data tables through the knowledge data entities, the knowledge data relationships, and the knowledge data attributes to output an initial knowledge graph; collect a heterogeneous fusion model to perform heterogeneous fusion on the initial knowledge graph to obtain a fused knowledge graph, where the fused knowledge graph is connected to a question-and-answer module and a search module; the question-and-answer module and the search module perform question-and-answer and search on the multiple heterogeneous knowledge data sources based on the fused knowledge graph.
[0007] This application also provides a heterogeneous knowledge graph system based on fusion mapping, including:
[0008] A heterogeneous knowledge data source acquisition module, which is used to obtain multiple heterogeneous knowledge data sources, where the multiple heterogeneous knowledge data sources include multiple heterogeneous knowledge data tables; a data source analysis module, which is used to analyze the multiple heterogeneous knowledge data sources to obtain knowledge data entities, knowledge data relationships, and knowledge data attributes; an initial knowledge graph output module, which is used to map the multiple heterogeneous knowledge data tables through the knowledge data entities, the knowledge data relationships, and the knowledge data attributes to output an initial knowledge graph; a heterogeneous fusion module, which is used to collect a heterogeneous fusion model to perform heterogeneous fusion on the initial knowledge graph to obtain a fused knowledge graph, where the fused knowledge graph is connected to a question-and-answer module and a search module; a question-and-answer search module, which is used for the question-and-answer module and the search module to perform question-and-answer and search on the multiple heterogeneous knowledge data sources based on the fused knowledge graph.
[0009] It is intended to use the heterogeneous knowledge graph method and system based on fusion mapping proposed in this application. First, obtain multiple heterogeneous knowledge data sources including multiple heterogeneous knowledge data tables, analyze and obtain the entities, relationships, and attributes of knowledge data, use these to map the data tables to output an initial knowledge graph, collect a heterogeneous fusion model to fuse it to obtain a fused knowledge graph, which is connected to the question-and-answer and search modules. These two modules perform question-and-answer and search on the data sources based on the graph, achieving the technical effect of enabling users to obtain and utilize the required knowledge more accurately, timely, and conveniently, and promoting the efficient utilization and sharing of knowledge. Brief Description of the Drawings
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the operations described above or below do not necessarily need to be performed precisely in sequence. On the contrary, as needed, various steps can be performed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or more steps can be removed from these processes.
[0011] Figure 1 It is a schematic flowchart of the heterogeneous knowledge graphing method based on fusion mapping provided by the embodiments of the present application;
[0012] Figure 2 It is a schematic structural diagram of the heterogeneous knowledge graphing system based on fusion mapping provided by the embodiments of the present application.
[0013] Explanation of reference numerals: Heterogeneous knowledge data source acquisition module 10, data source analysis module 20, initial knowledge graph output module 30, heterogeneous fusion module 40, question and answer search module 50. Detailed implementation manners
[0014] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically gives the detailed implementation manners of this application.
[0015] In order to make the purpose, technical solutions and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. The described embodiments should not be regarded as limitations of this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0016] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The terms "first / second" involved are only used to distinguish similar objects and do not represent a specific order for the objects. The terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application.
[0017] Embodiments of this application provide a heterogeneous knowledge graphing method based on fusion mapping, as Figure 1 shown, the method includes:
[0018] Step S100, obtain multiple heterogeneous knowledge data sources, where the multiple heterogeneous knowledge data sources include multiple heterogeneous knowledge data tables. Specifically, obtaining multiple heterogeneous knowledge data sources means resources containing various knowledge data collected from different channels, where these multiple heterogeneous knowledge data sources cover numerous knowledge data tables in various forms. It is necessary to clarify the scope and field of the required knowledge data and carry out the collection work accordingly. Establish connections with various database systems, data warehouses, file storage systems, etc. For database systems, build corresponding connections, use specific query languages to extract the required data tables, and be familiar with their structures, relationships between tables, and data storage methods. When dealing with data warehouses, use specialized tools or technologies to select relevant data tables from large-scale data sets according to set rules and conditions. Facing file storage systems, handle various file formats such as CSV, XML, JSON, etc., use corresponding parsers and reading tools to read and identify the data therein, and pick out the data tables containing heterogeneous knowledge. In addition, it is also necessary to communicate and cooperate with external data source providers to obtain the data resources they own, and ensure the legality, security, and availability of the data. After obtaining multiple heterogeneous knowledge data tables, perform preliminary regularization and classification on these data to facilitate the more efficient and orderly development of subsequent analysis and processing work.
[0019] In a possible implementation, multiple heterogeneous knowledge data sources are obtained. Among them, the multiple heterogeneous knowledge data sources include multiple heterogeneous knowledge data tables. Step S100 further includes step S110 of inputting the multiple heterogeneous knowledge data sources into a status verification module to obtain multiple status verification results corresponding to the multiple heterogeneous knowledge data sources, where each heterogeneous knowledge data source corresponds to one status verification result. Specifically, multiple heterogeneous knowledge data sources are obtained and input into the status verification module one by one. During the input process, the status verification module will evaluate each heterogeneous knowledge data source, considering multiple dimensions, such as checking whether the data format conforms to the standard specification, whether the data content is complete without omission, and whether the data logical relationship is reasonable and accurate. Since the characteristics and content of each heterogeneous knowledge data source are different, after being processed by the status verification module, a dedicated status verification result will be generated for each data source, clarifying its current status.
[0020] Step S120 is to obtain the heterogeneous knowledge data sources with the status verification result of verification passed according to the multiple status verification results. Specifically, the status verification results are sorted and analyzed. Only when the status verification result shows "verified", it indicates that the data source has passed strict inspection and is recognized as a heterogeneous knowledge data source with verification passed. Data sources with statuses of "not verified", "verifying", or "verification exception" will be excluded because they may have problems in aspects such as data quality, integrity, or accuracy and are not suitable for subsequent operations for the time. For example, assume there are multiple heterogeneous knowledge data sources, including employee information tables, sales records, and product specifications from different systems. After being input into the status verification module, for the employee information table, it may be checked whether all required fields, such as name, employee number, and department, are included and whether the formats of this information are correct; for the sales records, it will be verified whether the amount, date, and customer information of each transaction are accurate; for the product specifications, it will be confirmed whether the technical parameters and descriptions are clear and complete. Only data sources with the status shown as "verified", such as an employee information table with correct format and complete information and accurate sales records, will be selected for subsequent knowledge processing and analysis work.
[0021] Step S200 analyzes the multiple heterogeneous knowledge data sources to obtain knowledge data entities, knowledge data relationships, and knowledge data attributes. Specifically, the analysis of the multiple heterogeneous knowledge data sources aims to obtain knowledge data entities, knowledge data relationships, and knowledge data attributes. First, the architecture and format of the heterogeneous knowledge data sources are interpreted. For structured data, such as database tables, attention is paid to elements such as column names, data types, and constraints. For semi-structured data, such as XML or JSON formatted files, tags and nested structures are analyzed. For unstructured data, such as text files, pre-processing operations such as word segmentation and part-of-speech tagging are performed. Natural language processing techniques are then used to semantically interpret the text data, using lexical analysis, syntactic analysis, and semantic analysis to extract potential key information. This process focuses on identifying knowledge data entities, which can be specific people, objects, events, concepts, etc. For example, in a data source related to sales, "customer," "product," and "order" might be identified as knowledge data entities. Simultaneously, the relationships between entities in the data, namely, knowledge data relationships, are analyzed. For example, there's a purchase association between "customer" and "order," and a containment relationship between "product" and "order." Furthermore, the data extraction process extracts the characteristics and descriptive information of each knowledge data entity, known as knowledge data attributes. For example, the attributes of a "customer" might include "name," "age," and "address," while the attributes of a "product" might include "name," "price," and "specifications." To more accurately capture this information, existing knowledge graphs, ontologies, or domain dictionaries may be used for reference and comparison. Furthermore, data cleaning and preprocessing techniques are employed to remove noise and invalid data, improving the accuracy and reliability of analysis results. For example, when analyzing heterogeneous knowledge data sources in the medical field, knowledge data entities such as "patient," "disease," and "treatment plan" might be identified from medical records. Knowledge data relationships, such as the illness relationship between "patient" and "disease" and the treatment relationship between "disease" and "treatment plan," can be determined. Furthermore, knowledge data attributes such as the patient's age, gender, and symptoms, as well as the disease's symptom manifestations and severity, can be extracted.
[0022] In a possible implementation, the multiple heterogeneous knowledge data sources are analyzed to obtain knowledge data entities, knowledge data relationships, and knowledge data attributes. Step S200 further includes step S210 of defining entity type samples, relationships between entity type samples, and attribute samples for each entity type sample. Specifically, in the definition stage, according to specific business requirements and domain knowledge, various entity type samples are clarified. The entity type samples can be specific objects, such as "employee", "product", "order", etc. It is also necessary to determine the relationship samples between entity type samples. For example, there may be a "processing" relationship between an "employee" and an "order", and a "containing" relationship between a "product" and an "order". For each entity type sample, its attribute samples are defined. For example, the attribute samples of an "employee" can be "name", "age", "position", and the attribute samples of a "product" can be "name", "price", "specification", etc.
[0023] Step S220, according to the entity type samples, relationships between entity type samples, and attribute samples for each entity type sample, obtain a graph schema model. Specifically, based on the already defined entity type samples, relationship samples, and attribute samples, a graph schema model is constructed. The model will organize the definitions in a structured manner to form a framework, and various modeling tools and technologies will be used to ensure that the model can accurately reflect the logic and structure among entities, relationships, and attributes.
[0024] Step S230, based on the graph schema model, analyze the multiple heterogeneous knowledge data sources to obtain knowledge data entities, knowledge data relationships, and knowledge data attributes. Specifically, based on the constructed graph schema model, in-depth analysis of the multiple heterogeneous knowledge data sources is carried out. During the analysis process, the data in the data sources is matched and compared with the definitions in the graph schema model. Through operations such as data cleaning, transformation, and parsing, knowledge data entities that conform to the model definitions are identified from the data sources. For example, specific data items representing "employee", "product", "order" are found, the relationships between entities are determined, that is, which data items have relationships similar to the "processing", "containing", etc. defined in the model, and the attribute values corresponding to each entity are extracted. For example, the specific "name", "age", "position", etc. attribute information of a certain "employee" is obtained. For example, in an application in the e-commerce field, entity type samples such as "product", "customer", "order" are defined, the relationship sample of "a customer purchases a product to form an order", and the attribute samples of "product" such as "name", "price", "inventory", etc. After constructing a graph schema model according to this definition, multiple heterogeneous knowledge data sources including product information, customer purchase records, etc. are analyzed to obtain specific product knowledge data entities, the purchase relationship between customers and products, and the attribute data of each product.
[0025] Step S300, map the multiple heterogeneous knowledge data tables through the knowledge data entities, the knowledge data relationships, and the knowledge data attributes, and output an initial knowledge graph. Specifically, map the multiple heterogeneous knowledge data tables through the knowledge data entities, the knowledge data relationships, and the knowledge data attributes to output an initial knowledge graph. First, conduct a detailed exploration and analysis of the multiple heterogeneous knowledge data tables to check their field names, data types, and data contents. Then, adapt the knowledge data entities to the relevant data items in the data tables. For example, if the knowledge data entity is "employee", search for the corresponding column in the data table that contains employee information. For the knowledge data relationships, identify which columns or data in the data table have such relationships. For example, for the ownership relationship between "employee" and "department", find the data that can demonstrate this association. For the knowledge data attributes, correspond them to the specific field values in the data table. For example, for the "age" attribute of the "employee" entity, find the corresponding field in the data table that records the ages of employees. After achieving the matching and correspondence, use specific mapping rules and algorithms to convert these corresponding relationships into nodes and edges in the knowledge graph. The knowledge data entities become nodes, the knowledge data relationships become the edges connecting the nodes, and the knowledge data attributes are marked as the attributes of the nodes. Through such a mapping process, gradually integrate the data in the multiple heterogeneous knowledge data tables into an initial knowledge graph framework. During this process, it may be necessary to continuously adjust and optimize the mapping rules to ensure the accuracy and integrity of the knowledge graph. For example, in a human resources management system of an enterprise, there are multiple heterogeneous knowledge data tables such as an employee basic information table, a department information table, and an employee performance table. Through the above mapping process, adapt the knowledge data entity of "employee" to the relevant columns in the employee basic information table, reflect the relationship between "employee" and "department" through the associated data in the department information table, and obtain the "performance" attribute of "employee" from the employee performance table, and finally output an initial knowledge graph that can display the knowledge related to the enterprise's human resources.
[0026] In a possible implementation, the multiple heterogeneous knowledge data tables are mapped through the knowledge data entity, the knowledge data relationship, and the knowledge data attribute to output an initial knowledge graph. Step S300 further includes step S310 of analyzing the entity field information, relationship field information, and attribute field information of each heterogeneous data table in the multiple heterogeneous knowledge data tables. Specifically, when analyzing the entity field information, the data fields representing specific entity objects are studied. For example, in an employee data table, fields such as "employee number", "name", and "position" may be regarded as entity fields. For the relationship field information, attention is paid to the data fields indicating the association between different entities. For instance, in an order data table, the associated field between "customer number" and "order number". And for the analysis of the attribute field information, emphasis is placed on the data fields describing the specific characteristics or properties of the entity, such as "age" and "working years" in the employee table.
[0027] Step S320, define a distribution mapping code. The graph architecture model maps the entity field information, relationship field information, and attribute field information corresponding to each heterogeneous data table according to the distribution mapping code to output an initial knowledge graph. Specifically, defining the distribution mapping code provides clear rules and methods for subsequent mapping operations. The graph architecture model starts working based on the defined distribution mapping code, reads the entity field information, relationship field information, and attribute field information in each heterogeneous data table one by one. For the entity fields, the model will, according to the indication of the code, map them to the corresponding nodes in the knowledge graph. The relationship fields are used to establish connections between these nodes to form a clear relationship network. The information of the attribute fields is attached to the corresponding entity nodes to enrich the description and characteristics of the nodes. Through the mapping process, an initial knowledge graph is finally output. The initial knowledge graph integrates the key information in multiple heterogeneous data tables, presenting an organic structure among entities, relationships, and attributes. For example, in an enterprise management system, there are heterogeneous knowledge data tables such as employee data tables, department data tables, and project data tables. In the analysis stage, it is clear that "employee number" in the employee data table is an entity field, "department number" is a relationship field, and "skills and specialties" is an attribute field. After defining the distribution mapping code, the graph architecture model maps this information according to the code. For example, "employee number" is mapped to the employee node in the knowledge graph, "department number" establishes the connection between the employee node and the department node, and "skills and specialties" is used as the attribute of the employee node, thus outputting an initial knowledge graph reflecting the enterprise personnel structure and related information.
[0028] Step S400, the heterogeneous fusion model is collected to perform heterogeneous fusion on the initial knowledge graph to obtain a fused knowledge graph, where the fused knowledge graph is connected to the question-answering module and the search module. Specifically, the heterogeneous fusion model is collected to perform heterogeneous fusion operations on the initial knowledge graph to obtain a fused knowledge graph, where the fused knowledge graph is connected to the question-answering module and the search module. First, prepare the initial knowledge graph, which is the result obtained by processing and mapping multiple heterogeneous knowledge data tables before. Then, start the operation of the heterogeneous fusion model. This model is designed to process and integrate knowledge data with different structures and sources. When the model starts working, it will deeply analyze various elements in the initial knowledge graph. It will identify different types of nodes, edges, and the entities, relationships, and attributes they represent. Based on this analysis, the model will use specific algorithms and strategies to perform heterogeneous fusion operations. This may include eliminating duplicate nodes and relationships, integrating similar but not identical entities, and coordinating differences between data from different sources. For example, if there are two nodes in the initial knowledge graph that represent the same concept but have slightly different names, the model will identify and merge them into a unified node. For conflicting or inconsistent relationships, the model will adjust and integrate them according to preset rules and priorities. Through such a fusion process, a more accurate, complete, and consistent fused knowledge graph is finally obtained. Moreover, this fused knowledge graph is connected to the question-answering module and the search module. This means that when the user asks a question through the question-answering module or uses the search module to search, these two modules can directly access and utilize the information in the fused knowledge graph to provide accurate and useful answers and search results. For example, when the user asks a question about a specific topic in the question-answering module, the question-answering module will extract relevant knowledge and information from the fused knowledge graph and answer the user's question in a clear and understandable way. Similarly, when the user enters keywords or query conditions in the search module, the search module can quickly locate and return relevant content in the fused knowledge graph. In short, through the processing of the initial knowledge graph by the heterogeneous fusion model and the connection with the question-answering module and the search module, the effective integration and utilization of knowledge are achieved, providing better knowledge services for users.
[0029] In a possible implementation, the heterogeneous fusion model collects and performs heterogeneous fusion on the initial knowledge graph to obtain a fused knowledge graph. The fused knowledge graph is connected to the question-answering module and the search module. Step S400 further includes step S410. The heterogeneous fusion model includes a support merging fusion processing channel and a priority fusion processing channel. Specifically, two different but complementary processing channels are built inside the heterogeneous fusion model. The support merging fusion processing channel aims to process conflicting heterogeneous data sources with similar characteristics in terms of quality, reliability, etc. When faced with multiple data sources and after evaluation, if they are found to perform similarly in key metrics, this channel will come into play. For example, in a fusion scenario of product sales data, if there are two data sources from different sales channels but their data quality, accuracy, and integrity are comparable. In this case, the support merging fusion processing channel will be activated to conduct a detailed comparison and analysis of these two data sources, checking each field in the data, such as product name, sales quantity, sales price, etc. For exactly the same data fields, they are directly retained; for data fields with differences that can be merged through certain rules, such as sales quantity, a summation calculation may be performed; for some minor differences, such as slightly different expressions of the product name but with the same essence, they are unified and standardized. The priority fusion processing channel is to solve the situation when there are differences in quality metrics between conflicting heterogeneous data sources. It decides which data source's data to use as the main based on pre-set priority rules. The basis for setting priorities can be diverse. For example, the authority of the data source. If one data source comes from an official authoritative institution and the other from an unofficial channel, then the official authoritative data source may have a higher priority. Or, according to the data update frequency, if one data source can update data more frequently to ensure data timeliness, then it may be given a higher priority. When the priority is determined, the data of the higher-priority data source is used as the backbone, and important information that exists in the lower-priority data sources but is missing in the higher-priority data source is supplemented and fused to form a more complete and accurate fusion result. The support merging fusion processing channel and the priority fusion processing channel together constitute the heterogeneous fusion model, enabling flexible and accurate fusion processing when facing complex and diverse heterogeneous data sources, thus providing high-quality and consistent data support for the knowledge graph.
[0030] Step S420: Identify the data source quality metrics of the conflicting heterogeneous data sources in the initial knowledge graph. If the data source quality metrics of the conflicting heterogeneous data sources are the same, activate the support merge and fusion processing channel. Specifically, analyze the conflicting heterogeneous data sources in the initial knowledge graph to accurately identify their data source quality metrics. The identification process may comprehensively consider multiple key factors, such as data accuracy, that is, whether the data is accurate and error-free; data integrity, whether all necessary information is covered; data consistency, whether different parts of the data are coordinated and consistent with each other; and data timeliness, whether the data is up-to-date and valid. When it is found that the data source quality metrics of the conflicting heterogeneous data sources are the same, the support merge and fusion processing channel will be activated. During this process, special attention should be paid to the situation where there are partially different mapped field values between the two data sources to be mapped and fused. For this problem of inconsistent partial field values, a series of strategies and algorithms will be adopted during the support merge and fusion processing. First, compare and evaluate these inconsistent field values to determine which values are more accurate or more in line with the actual situation. If a direct judgment cannot be made, some comprehensive methods may be used, such as taking the average value, referring to authoritative data sources, or determining the final value to be adopted according to specific business rules.
[0031] Step S430: Perform data source merging according to the support merge and fusion processing channel, and store the merged data source in the form of a string. Specifically, after completing the merge operation, in order to facilitate subsequent storage, processing, and use, the merged data source will be converted into a string form for storage. The string form of storage has the characteristics of generality and easy readability, and can facilitate other modules or systems to quickly and accurately obtain and parse the data when needed. For example, in a knowledge graph about product information, there are two conflicting data sources from different suppliers. When identifying the data source quality metrics, it is found that their qualities are the same, but there are differences in the two mapped fields of product price and inventory quantity. Activate the support merge and fusion processing channel. By comparing factors such as the reputation of the two suppliers and the accuracy of historical data, determine the final values of the price and inventory quantity. Then, save the merged complete product information, including name, specifications, price, inventory, etc., in the form of a string with a specific format, so that these data can be efficiently utilized in subsequent scenarios such as querying, analysis, or transaction processing. When the quality metrics of the conflicting heterogeneous data sources are different, the priority fusion processing channel will be adopted. According to the pre-set priority rules, for example, based on the authority of the data source, update frequency, or credibility of the data, determine which data source's information to adopt first. The information of the selected data source will be used as the main basis and fused with the supplementary or corrected information in other data sources, and finally stored in the form of a string.
[0032] In a possible implementation, identify the data source quality metrics of the conflicting heterogeneous data sources in the initial knowledge graph. If the data source quality metrics of the conflicting heterogeneous data sources are the same, activate the support for the merging and fusion processing channel. Step S420 further includes step S421 of constructing a data source evaluation model and connecting the data source evaluation model to the heterogeneous fusion model. Specifically, when constructing a data source evaluation model, relevant technologies and methods such as statistics, data mining, and machine learning will be applied. In the model, a series of rules, algorithms, and parameters for evaluating data quality are defined. After completing the construction of the data source evaluation model, connect it to the heterogeneous fusion model to achieve the circulation and interaction of data and information.
[0033] Step S422: Conduct a data source quality assessment on each of the multiple heterogeneous knowledge data sources according to the data source evaluation model, including data reliability, data integrity, and data accuracy. Specifically, use the constructed data source evaluation model to conduct a quality assessment on each of the multiple heterogeneous knowledge data sources. When evaluating data reliability, check whether the data source is trustworthy, whether the data collection method is scientific and reasonable, whether there are outliers or error values, etc. For the assessment of data integrity, confirm whether the data covers all necessary fields and information, and whether there are missing key parts. When examining data accuracy, compare the data with the actual situation and check the logical consistency and numerical rationality of the data.
[0034] Step S423: Obtain the data source quality metrics of each heterogeneous knowledge data source based on the data reliability, data integrity, and data accuracy. Specifically, based on the evaluation results of data reliability, integrity, and accuracy, comprehensively calculate and determine the data source quality metrics of each heterogeneous knowledge data source. The quality metric is a quantitative value or level that can comprehensively reflect the quality level of the data source.
[0035] Step S424: Input the data source quality metrics of each heterogeneous knowledge data source into the heterogeneous fusion model for storage. Specifically, input the data source quality metrics of each heterogeneous knowledge data source into the heterogeneous fusion model for storage. When the heterogeneous fusion model performs fusion processing, it can select appropriate processing channels (such as the support merge fusion processing channel or the priority fusion processing channel) based on these stored quality metrics and make more informed and accurate fusion decisions. For example, in an e-commerce data analysis scenario, there are multiple data sources of product information from different suppliers. Through the data source evaluation model, the reliability (such as the reputation of the supplier), integrity (whether all specification parameters are included), and accuracy (whether the price is reasonable and the inventory quantity is accurate) of data such as product prices, inventory, and descriptions in each data source are evaluated. After obtaining the quality metrics, they are input into the heterogeneous fusion model. When fusing product information, the model can determine how to fuse the information from different data sources based on these metrics to provide users with accurate and complete product knowledge.
[0036] In a possible implementation, merge the data sources according to the support merge fusion processing channel and store the merged data source in the form of a string. Step S430 further includes step S431: If the data source quality metrics of the conflicting heterogeneous data sources are different, activate the priority fusion processing channel. Specifically, when it is found that there are differences in the data source quality metrics of the conflicting heterogeneous data sources, start the priority fusion processing channel to handle.
[0037] Step S432, the priority fusion processing channel identifies priorities according to the magnitudes of the data source quality metrics, and selects the data source with a higher priority for storage in the form of a string. Specifically, a detailed comparison and analysis are performed on different data source quality metrics. The magnitude of the quality metric becomes the key basis for judging priorities. The comparison process comprehensively considers multiple factors, such as data accuracy, integrity, reliability, update frequency, etc. Based on the priority identification of each data source, the data source with a higher quality metric will be given a higher priority. After determining the priorities, the data source with a higher priority is selected as the main data source. During the fusion process, the data in this data source will be preferentially adopted. For the selected data source with a higher priority, it is stored in the form of a string. For example, suppose there are two conflicting heterogeneous data sources. One data source has very high data accuracy but slightly poor integrity, and the other data source has better integrity but relatively lower accuracy. After comprehensive evaluation, it is determined that the data source with high accuracy has a higher quality metric and a higher priority. Then, the data in this data source with a higher priority is converted into a string format for storage. When processing weather data in different regions, if a data source has a higher update frequency and can better reflect the latest weather conditions, even if it may not be as complete as another data source in some details, it will be given a higher priority due to its higher timeliness, and its data will be stored in the form of a string.
[0038] In a possible implementation, the heterogeneous fusion model collects and performs heterogeneous fusion on the initial knowledge graph to obtain a fused knowledge graph. The fused knowledge graph is connected to the question-answering module and the search module. Step S400 further includes step S440. The fused knowledge graph includes a data generation graph structure that supports hierarchical layout, network structure layout, automatic layout, horizontal and vertical line layout, and vertical line layout. Specifically, the fused knowledge graph covers a data generation graph structure that supports hierarchical layout, network structure layout, automatic layout, horizontal and vertical line layout, and vertical line layout. First is the hierarchical layout. In this layout, the nodes and relationships in the knowledge graph are arranged in a clear hierarchical structure. For example, perhaps the core and important nodes are placed at a higher level, and the related secondary nodes are distributed at a lower level, forming a tree-like structure that clearly shows the primary-secondary and inclusion relationships between the data. The network structure layout focuses on presenting the complex connection relationships between nodes. Through this layout, the mutual associations and interactions between each node can be intuitively observed, which is suitable for displaying highly interconnected data. The automatic layout is an intelligent choice. The system will automatically calculate and generate a relatively reasonable and beautiful layout method according to the characteristics of the data and the complexity of the relationships, saving the time and effort of manual adjustment. The horizontal and vertical line layout organizes the nodes and relationships with horizontal and vertical lines, forming a regular arrangement, which is suitable for the situation where the data relationships are relatively simple and linear, resulting in a simple and clear graph structure. The vertical line layout emphasizes the vertical arrangement and may be suitable for certain specific types of data or can better display the logic and hierarchy of the data in a specific display environment. In practical applications, users can flexibly select the appropriate layout method according to the specific content of the knowledge graph and the usage purpose. For example, when analyzing the organizational structure of an enterprise, the hierarchical layout may be selected to highlight the relationship between the senior and junior levels; while when studying the interpersonal relationships in a social network, the network structure layout may better reflect the complex social interactions. In short, by supporting multiple layout methods, the fused knowledge graph provides users with more flexible, intuitive, and effective means of data display and analysis.
[0039] Step S500, the Q&A module and the search module perform Q&A and search on the multiple heterogeneous knowledge data sources based on the fused knowledge graph. Specifically, the Q&A module and the search module perform Q&A and search on the multiple heterogeneous knowledge data sources based on the fused knowledge graph. First, in the Q&A scenario, when the user inputs a question, the Q&A module immediately analyzes and understands the question. It extracts the key information and keywords in the question, and then searches and matches them in the fused knowledge graph. The fused knowledge graph contains the knowledge and relationships after integrating multiple heterogeneous knowledge data sources. The Q&A module will use the nodes, edges, and related attribute information in the graph to find the answer most relevant to the question. For example, if the user asks "What is the latest price of a certain product", the Q&A module will search for the node related to the product in the fused knowledge graph, find the associated price attribute, and obtain the latest price data as the answer. In the search scenario, when the user inputs search keywords or conditions, the search module also analyzes and processes the input. It compares and matches these keywords with the nodes and relationships in the fused knowledge graph. Then, it sorts and filters the search results according to relevance and importance. For example, when the user searches for "products related to healthy eating", the search module will search for the product nodes related to healthy eating in the fused knowledge graph, as well as other related information such as product features and user reviews, and organize these relevant contents into search results and return them to the user. In the whole process, the Q&A module and the search module continuously optimize the search and matching algorithms to improve the accuracy of the answers and the quality of the search results. At the same time, they will also continuously improve and perfect their own functions according to the user's feedback and usage habits to provide better services. In short, relying on the powerful integration ability of the fused knowledge graph, the Q&A module and the search module can quickly and accurately provide users with the required information and answers from multiple heterogeneous knowledge data sources.
[0040] In the embodiment of the present application, multiple heterogeneous knowledge data sources including multiple heterogeneous knowledge data tables are acquired, the entities, relationships, and attributes of the knowledge data are analyzed and acquired, and these are used to map and output an initial knowledge graph. The heterogeneous fusion model is used to fuse it to obtain a fused knowledge graph. This graph is connected to the Q&A and search modules, and these two modules perform Q&A and search on the data sources based on the graph, achieving the technical effect of enabling users to obtain and utilize the required knowledge more accurately, timely, and conveniently, and promoting the efficient utilization and sharing of knowledge.
[0041] In the above text, reference is made to Figure 1 which describes in detail the heterogeneous knowledge graphing method based on fusion mapping according to the embodiment of the present invention. Next, reference will be made to Figure 2 to describe the heterogeneous knowledge graphing system based on fusion mapping according to the embodiment of the present invention.
[0042] The heterogeneous knowledge graph system based on fusion mapping according to an embodiment of the present invention is used to solve the technical problems in the prior art, such as the huge differences in format, structure and semantics caused by the wide and diverse sources of knowledge data, which make it difficult to directly integrate, have missing associations, low efficiency in search and question answering, difficult to update and expand, and limited cross-system knowledge utilization. It achieves the technical effect of enabling users to obtain and utilize the required knowledge more accurately, timely and conveniently, and promoting the efficient utilization and sharing of knowledge. The heterogeneous knowledge graph system based on fusion mapping includes: a heterogeneous knowledge data source acquisition module 10, a data source analysis module 20, an initial knowledge graph output module 30, a heterogeneous fusion module 40, and a question answering and search module 50.
[0043] The heterogeneous knowledge data source acquisition module 10 is used to acquire multiple heterogeneous knowledge data sources, where the multiple heterogeneous knowledge data sources include multiple heterogeneous knowledge data tables;
[0044] The data source analysis module 20 is used to analyze the multiple heterogeneous knowledge data sources to obtain knowledge data entities, knowledge data relationships and knowledge data attributes;
[0045] The initial knowledge graph output module 30 is used to map the multiple heterogeneous knowledge data tables through the knowledge data entities, the knowledge data relationships and the knowledge data attributes, and output an initial knowledge graph;
[0046] The heterogeneous fusion module 40 is used to collect a heterogeneous fusion model to perform heterogeneous fusion on the initial knowledge graph to obtain a fused knowledge graph, where the fused knowledge graph is connected to the question answering module and the search module;
[0047] The question answering and search module 50 is used for the question answering module and the search module to perform question answering and search on the multiple heterogeneous knowledge data sources based on the fused knowledge graph.
[0048] Next, the specific configuration of the heterogeneous knowledge data source acquisition module 10 will be described in detail. As described above, multiple heterogeneous knowledge data sources are acquired, where the multiple heterogeneous knowledge data sources include multiple heterogeneous knowledge data tables. The heterogeneous knowledge data source acquisition module 10 may further include: a status verification result acquisition unit for inputting the multiple heterogeneous knowledge data sources into a status verification module to obtain multiple status verification results corresponding to the multiple heterogeneous knowledge data sources, where each heterogeneous knowledge data source corresponds to a status verification result; a heterogeneous knowledge data source passing unit for obtaining the heterogeneous knowledge data sources with the status verification results being verified as passed according to the multiple status verification results.
[0049] Next, the specific configuration of the data source analysis module 20 will be described in detail. As described above, the multiple heterogeneous knowledge data sources are analyzed to obtain knowledge data entities, knowledge data relationships, and knowledge data attributes. The data source analysis module 20 may further include: a sample definition unit for defining entity type samples, relationship samples between entity type samples, and attribute samples for each entity type sample; a graph schema model acquisition unit for obtaining a graph schema model according to the entity type samples, relationship samples between entity type samples, and attribute samples for each entity type sample; and a knowledge data acquisition unit for analyzing the multiple heterogeneous knowledge data sources based on the graph schema model to obtain knowledge data entities, knowledge data relationships, and knowledge data attributes.
[0050] Next, the specific configuration of the initial knowledge graph output module 30 will be described in detail. As described above, the multiple heterogeneous knowledge data tables are mapped through the knowledge data entities, the knowledge data relationships, and the knowledge data attributes to output an initial knowledge graph. The initial knowledge graph output module 30 may further include: a field information analysis unit for analyzing the entity field information, relationship field information, and attribute field information of each heterogeneous data table in the multiple heterogeneous knowledge data tables; and a distribution mapping code definition unit for defining a distribution mapping code, and the graph schema model maps the entity field information, relationship field information, and attribute field information corresponding to each heterogeneous data table according to the distribution mapping code to output an initial knowledge graph.
[0051] Next, the specific configuration of the heterogeneous fusion module 40 will be described in detail. As described above, a heterogeneous fusion model is collected to perform heterogeneous fusion on the initial knowledge graph to obtain a fused knowledge graph. The fused knowledge graph is connected to the question answering module and the search module. The heterogeneous fusion module 40 may further include: a heterogeneous fusion model composition unit, where the heterogeneous fusion model includes a support merging fusion processing channel and a priority fusion processing channel; a data source quality index identification unit for identifying the data source quality index of the conflicting heterogeneous data sources in the initial knowledge graph. If the data source quality indexes of the conflicting heterogeneous data sources are the same, the support merging fusion processing channel is activated; and a data source merging unit for performing data source merging according to the support merging fusion processing channel and storing the merged data source in the form of a string.
[0052] Among them, the data source quality indicators of the conflicting heterogeneous data sources in the initial knowledge graph are identified. If the data source quality indicators of the conflicting heterogeneous data sources are the same, the support for the merge and fusion processing channel is activated. The data source quality indicator identification unit may further include: a data source evaluation model construction subunit for constructing a data source evaluation model and connecting the data source evaluation model to the heterogeneous fusion model; a data source quality evaluation subunit for evaluating the data source quality of each heterogeneous knowledge data source in the multiple heterogeneous knowledge data sources according to the data source evaluation model, including data reliability, data integrity, and data accuracy; a data source quality indicator acquisition subunit for obtaining the data source quality indicators of each heterogeneous knowledge data source according to the data reliability, data integrity, and data accuracy; and a storage subunit for inputting the data source quality indicators of each heterogeneous knowledge data source into the heterogeneous fusion model for storage.
[0053] Among them, the heterogeneous fusion module 40 may further include: a fusion knowledge graph composition unit, where the fusion knowledge graph includes a data generation graph structure and supports hierarchical layout, network structure layout, automatic layout, horizontal and vertical line layout, and vertical line layout.
[0054] The heterogeneous knowledge graph system based on fusion mapping provided by the embodiments of the present invention can execute the heterogeneous knowledge graph method based on fusion mapping provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0055] Although this application makes various references to certain modules in the system according to the embodiments of this application, however, any number of different modules can be used and run on the user terminal and / or server. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0056] The above specific implementation manners do not constitute a limitation to the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application. In some cases, the actions or steps recorded in this application can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A heterogeneous knowledge graphing method based on fusion mapping, characterized by: The method comprises: Acquire multiple heterogeneous knowledge data sources, wherein the multiple heterogeneous knowledge data sources include multiple heterogeneous knowledge data tables; Analyzing the plurality of heterogeneous knowledge data sources to obtain knowledge data entities, knowledge data relationships, and knowledge data attributes; Mapping the plurality of heterogeneous knowledge data tables through the knowledge data entities, the knowledge data relationships, and the knowledge data attributes to output an initial knowledge graph; Acquire a heterogeneous fusion model to perform heterogeneous fusion on the initial knowledge graph to obtain a fused knowledge graph, wherein the fused knowledge graph is connected to a question-answering module and a search module; The question-answering module and the search module perform question-answering and search on the multiple heterogeneous knowledge data sources based on the fused knowledge graph; The method of acquiring a heterogeneous fusion model and performing heterogeneous fusion on the initial knowledge graph to obtain a fused knowledge graph includes: The heterogeneous fusion model includes support for merging fusion processing channels and priority fusion processing channels; Identifying data source quality indicators of conflicting heterogeneous data sources in the initial knowledge graph, and activating the merging and fusion processing channel if the data source quality indicators of the conflicting heterogeneous data sources are the same; Merging data sources according to the supporting merging and fusion processing channel, and storing the merged data source in the form of a character string; Among them, the fused knowledge graph includes a data generation graph structure, which supports hierarchical layout, network structure layout, automatic layout, horizontal and vertical line layout and vertical line layout.
2. The heterogeneous knowledge graphing method based on fusion mapping according to claim 1 is characterized in that: Analyzing the multiple heterogeneous knowledge data sources, the method includes: Define entity type samples, relationship samples between entity type samples, and attribute samples of each entity type sample; Acquire a graph architecture model according to the entity type samples, the relationship samples between the entity type samples, and the attribute samples of each entity type sample; The plurality of heterogeneous knowledge data sources are analyzed based on the graph architecture model to obtain knowledge data entities, knowledge data relationships and knowledge data attributes.
3. The heterogeneous knowledge graphing method based on fusion mapping according to claim 2 is characterized in that: The method for mapping the plurality of heterogeneous knowledge data tables by using the knowledge data entities, the knowledge data relationships and the knowledge data attributes includes: Analyzing entity field information, relationship field information, and attribute field information of each heterogeneous data table in the plurality of heterogeneous knowledge data tables; A distribution mapping code is defined, and the graph architecture model maps the entity field information, relationship field information, and attribute field information corresponding to each heterogeneous data table according to the distribution mapping code, and outputs an initial knowledge graph.
4. The heterogeneous knowledge graphing method based on fusion mapping according to claim 1 is characterized in that: Acquiring multiple heterogeneous knowledge data sources, the method further includes: Inputting the plurality of heterogeneous knowledge data sources into a state verification module, and obtaining a plurality of state verification results corresponding to the plurality of heterogeneous knowledge data sources, wherein each heterogeneous knowledge data source corresponds to one state verification result; According to the multiple status verification results, a heterogeneous knowledge data source whose status verification result passes the verification is obtained.
5. The heterogeneous knowledge graphing method based on fusion mapping according to claim 1 is characterized in that: If the data source quality indicators of the conflicting heterogeneous data sources are different, the priority fusion processing channel is activated: The priority fusion processing channel performs priority identification according to the size of the data source quality index, and selects the data source with the highest priority and stores it in the form of a character string.
6. The heterogeneous knowledge graphing method based on fusion mapping according to claim 1 is characterized in that: Identifying data source quality indicators of conflicting heterogeneous data sources in the initial knowledge graph includes: Constructing a data source evaluation model, and connecting the data source evaluation model with the heterogeneous fusion model; Performing data source quality assessment, data reliability, data integrity, and data accuracy on each of the plurality of heterogeneous knowledge data sources according to a data source assessment model; Obtaining data source quality indicators of each heterogeneous knowledge data source based on the data reliability, data integrity, and data accuracy; The data source quality indicators of each heterogeneous knowledge data source are input into the heterogeneous fusion model for storage.
7. A heterogeneous knowledge graphing system based on fusion mapping, characterized by: The system is used to implement the heterogeneous knowledge graphing method based on fusion mapping according to any one of claims 1 to 6, and the system includes: A heterogeneous knowledge data source acquisition module, wherein the heterogeneous knowledge data source acquisition module is used to acquire multiple heterogeneous knowledge data sources, wherein the multiple heterogeneous knowledge data sources include multiple heterogeneous knowledge data tables; A data source analysis module, configured to analyze the plurality of heterogeneous knowledge data sources to obtain knowledge data entities, knowledge data relationships, and knowledge data attributes; An initial knowledge graph output module, configured to map the plurality of heterogeneous knowledge data tables through the knowledge data entities, the knowledge data relationships, and the knowledge data attributes, and output an initial knowledge graph; A heterogeneous fusion module, which is used to collect heterogeneous fusion models and perform heterogeneous fusion on the initial knowledge graph to obtain a fused knowledge graph, wherein the fused knowledge graph is connected to the question-answer module and the search module; A question-answering search module is used by the question-answering module and the search module to perform question-answering and search on the multiple heterogeneous knowledge data sources based on the fused knowledge graph.
Citation Information
Patent Citations
Construction method for multi-source heterogeneous data knowledge graph of power distribution network
CN117273133A
Question and answer interaction method, device and equipment based on multi-source heterogeneous mapping knowledge domain
CN117610650A