A multi-source heterogeneous data fusion and knowledge graph automatic construction system
By using a multi-source heterogeneous data fusion and knowledge graph automatic construction system, and employing semantic extraction, data cleaning, disambiguation calibration, and graph fusion modules, the system solves the problems of standardized semantic alignment and knowledge updating of multi-source heterogeneous data, and achieves efficient and reliable knowledge fusion and graph updating.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI SHENHE INFORMATION TECH CO LTD
- Filing Date
- 2026-04-03
- Publication Date
- 2026-06-23
AI Technical Summary
Traditional multi-source heterogeneous data fusion and knowledge graph automatic construction systems lack unified semantic governance and lightweight ontology modeling capabilities, making it difficult to perform standardized semantic alignment on multi-source heterogeneous data. They suffer from data noise, redundancy, and semantic ambiguity, and their cross-source fusion accuracy and stability are insufficient. Entity disambiguation often relies on literal matching or simple rules, knowledge updates cannot achieve incremental expansion, and manual correction is costly and inefficient.
The semantic extraction module extracts and binds the smallest semantic fragments, and completes cross-source alignment through symbol similarity and local logical implication to generate the smallest connected ontology; the data purification module performs symbolic predicate transformation and density filtering; the ternary generation module extracts entities, relations and attributes; the disambiguation calibration module uses topological fingerprint disambiguation and alignment; the graph fusion module performs integrated fusion and verification; and the incremental update module realizes local merging and updating.
It achieves unified semantic governance and lightweight modeling of multi-source heterogeneous data, improves the accuracy of knowledge fusion and the reliability of graph quality, reduces the cost of data governance and semantic alignment, supports lightweight incremental updates and global consistency self-maintenance of knowledge graphs, and adapts to complex scenarios of continuous access to multi-source data.
Smart Images

Figure CN121981232B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of knowledge graph construction technology, specifically a multi-source heterogeneous data fusion and knowledge graph automatic construction system. Background Technology
[0002] Traditional multi-source heterogeneous data fusion and knowledge graph automatic construction systems generally lack unified semantic governance and lightweight ontology modeling capabilities, making it difficult to perform standardized semantic alignment on multi-source heterogeneous data. Data noise, redundancy, and semantic ambiguity are prominent issues, resulting in insufficient accuracy and stability in cross-source fusion. Entity disambiguation often relies on literal matching or simple rules, failing to consider entity topological association features, easily leading to misjudgments of homonyms and heteronyms, resulting in low knowledge reliability. The fusion process lacks end-to-end consistency verification and closed-loop correction, leading to frequent conflicts among entities, relationships, and attributes, and high costs and low efficiency for manual correction. Knowledge updates often employ global reconstruction methods, failing to achieve incremental expansion. Adding new data is time-consuming and resource-intensive, making it difficult to support dynamic business expansion and long-term stable evolution. Summary of the Invention
[0003] To overcome the shortcomings of existing technologies, this invention proposes a multi-source heterogeneous data fusion and knowledge graph automatic construction system. This invention primarily addresses the challenges of ambiguity, conflict, and inefficient construction in multi-source heterogeneous data fusion.
[0004] This invention provides a multi-source heterogeneous data fusion and knowledge graph automatic construction system, comprising:
[0005] The semantic extraction module is used to automatically extract the smallest semantic fragments from various data sources and bind constraints and examples. It completes cross-source alignment through symbol similarity and local logical implication, generates the smallest connected ontology containing only the necessary semantics, and obtains a symbolic semantic framework.
[0006] The data cleansing module is used to perform symbolic predicate transformation on multi-source heterogeneous data according to the symbolic semantic framework, calculate the data occurrence density under isomorphic semantics, marginalize low evidence data according to the threshold, and output a cleaned symbolic predicate set.
[0007] The triple generation module is used to input the purified symbolic predicate set into the automatically induced minimal rule grammar set, and directly extract entities, relations and attributes through the grammar parse tree to generate an initial candidate set of triples.
[0008] The disambiguation calibration module is used to generate a topological fingerprint of entity two-hop relationships using an initial set of candidate triples, perform entity disambiguation and alignment based on fingerprint similarity, verify and correct conflicting triples, and output unambiguous structured knowledge.
[0009] The graph fusion module is used to integrate unambiguous structured knowledge with the minimum connected ontology, simultaneously unifying data semantics and constructing the knowledge graph within the symbol space, verifying fusion consistency, backtracking and correcting conflicting information, and forming an initial knowledge graph.
[0010] The incremental update module is used to generate corresponding new ontology fragments and topological fingerprints for newly added data based on the initial knowledge graph, incrementally merge them with the local subgraphs of the knowledge graph, automatically maintain consistency constraint chains, and form a complete knowledge graph.
[0011] According to the present invention, a multi-source heterogeneous data fusion and knowledge graph automatic construction system includes a semantic extraction module comprising:
[0012] The semantic extraction unit is used to clean and delineate boundaries of multi-source data, extract the smallest indivisible semantic fragments, and form a basic semantic set.
[0013] The constraint binding unit is used to bind corresponding logical constraints and practical application examples to each semantic fragment based on the basic semantic set, clarify the semantic boundaries and usage scenarios, and generate a standardized semantic set.
[0014] The cross-source alignment unit is used to standardize the semantic set by using symbol similarity calculation and local logical implication judgment to match and associate semantic units across data sources, establish semantic correspondence between different sources, and obtain semantic association results.
[0015] The minimal ontology is used to eliminate redundant semantics and unnecessary relationships based on semantic association results, retaining the minimum number of semantic nodes that support the core logic, and constructing a minimal connected ontology.
[0016] The symbolic modeling unit is used to standardize and logically structure the smallest connected ontology, unify the symbol system and the expression of connected relationships, and generate a symbolic semantic framework.
[0017] According to the multi-source heterogeneous data fusion and knowledge graph automatic construction system provided by the present invention, the specific steps for obtaining semantic association results in the cross-source alignment unit are as follows:
[0018] We perform feature vectorization representation on cross-data source semantics in a standardized semantic set, extract the symbolic attributes, logical structure and constraint information of semantic fragments, and generate semantic feature vectors.
[0019] Based on the semantic feature vectors, a symbolic similarity algorithm is used to calculate the degree of similarity between pairs of semantics, resulting in a semantic similarity matrix.
[0020] Based on the semantic similarity matrix and the local logical implication judgment rule, the antecedent and consequent inference relationship between semantics is verified, and candidate matching pairs are selected.
[0021] Based on candidate matching pairs, establish mapping associations between semantic units across data sources, form stable semantic correspondences, and output semantic association results.
[0022] According to the present invention, a multi-source heterogeneous data fusion and knowledge graph automatic construction system includes a data cleansing module comprising:
[0023] The predicate mapping unit is used to map multi-source heterogeneous data into standardized predicate expressions according to the predicate rules defined by the symbolic semantic framework, adapt the data type to the predicate transformation logic, determine the semantic predicate and parameters of each data unit, and output the initial symbolic predicate set.
[0024] The frequency statistics unit is used to count the frequency of each symbolic predicate based on the initial symbolic predicate set, grouped by isomorphic semantics, and calculate the frequency of predicates by combining the total amount of data involved in the transformation, and output a frequency statistics table.
[0025] The threshold marginalization unit is used to set a density threshold according to the business scenario, filter low-evidence predicates with a density lower than the threshold according to the frequency statistics table, remove the corresponding entries from the initial symbolic predicate set, and output the intermediate predicate set.
[0026] The semantic cleanup unit is used to perform semantic consistency checks based on the intermediate predicate set, eliminate predicate logic conflicts, merge predicate entries with duplicate content, and output a cleaned symbolic predicate set.
[0027] According to the multi-source heterogeneous data fusion and knowledge graph automatic construction system provided by the present invention, the specific steps for outputting the intermediate predicate set in the threshold marginalization unit are as follows:
[0028] Based on the knowledge reliability and data distribution characteristics of the business scenario, a density threshold is set to adapt to the current semantic environment.
[0029] Based on the density threshold, the density values of each symbolic predicate in the frequency statistics table are traversed, and predicates with densities less than the threshold are selected and summarized to generate threshold-marginalized low evidence data.
[0030] Based on the generated threshold marginalization of low evidence data, match and remove predicate entries from the initial symbolic predicate set, and output the intermediate predicate set.
[0031] According to the present invention, a multi-source heterogeneous data fusion and knowledge graph automatic construction system includes a ternary generation module comprising:
[0032] The grammar parsing unit is used to adapt the cleaned symbolic predicate set to the automatically induced minimal rule grammar set, perform structured parsing of each predicate expression according to the grammar rules, and output the predicate parsing structure.
[0033] The tree-labeling unit is used to construct a grammar parse tree based on the predicate parsing structure, determine the semantic role of each node in the parse tree, locate the positions of entity nodes, relation nodes and attribute nodes, and output the labeled parse tree structure.
[0034] The ternary generation unit is used to extract the three core elements of entities, relations and attributes from the tree according to the labeled parse tree structure, and combine them in the format of entity-relation-attribute triples to output the initial candidate set of triples.
[0035] According to the present invention, a multi-source heterogeneous data fusion and knowledge graph automatic construction system includes a disambiguation calibration module comprising:
[0036] The topology coding unit is used to sort out the relationships between entities based on the initial triplet candidate set. Taking each entity as the core, it expands the entities and relationships associated with its one-hop and two-hop relationships, extracts and encodes the topology features, generates a unique two-hop relationship topology fingerprint for each entity, and outputs a set of entity topology fingerprints.
[0037] The fingerprint alignment unit is used to calculate the topological fingerprint similarity of different entities based on the entity topological fingerprint set, filter entities whose similarity reaches a preset threshold, determine them as the same entity, and perform disambiguation and alignment, and output the entity mapping result and associated triplet.
[0038] The conflict calibration unit is used to perform logical verification on all triples based on the entity mapping results and associated triples, identify entries with conflicting entities, relations or attributes, eliminate ambiguity by correcting conflicting content and retaining high-confidence information, and output unambiguous structured knowledge.
[0039] According to the multi-source heterogeneous data fusion and knowledge graph automatic construction system provided by the present invention, the specific steps for outputting the entity topological fingerprint set in the topological coding unit are as follows:
[0040] Iterate through all entities and associated edges in the initial triplet candidate set, build a direct association list for each entity, extract one-hop adjacent entities and their corresponding relationships, and output the one-hop association set of the entities.
[0041] Based on the one-hop association set of entities, extend outward to obtain indirectly related entities and relationships, forming a complete relationship subgraph covering two hops, and outputting a set of entity two-hop topological subgraphs.
[0042] Structural features are extracted and serialized from the set of two-hop topological subgraphs of entities to generate hash codes that uniquely identify the structure, thereby obtaining and summarizing the topological fingerprints of the two-hop relationships.
[0043] According to the present invention, a multi-source heterogeneous data fusion and knowledge graph automatic construction system is provided, wherein the graph fusion module includes:
[0044] The semantic fusion unit is used to establish the semantic mapping relationship between unambiguous structured knowledge and minimal connected ontology in the symbol space. It integrates the entities, relations, and attributes of unambiguous structured knowledge with the semantic specifications of minimal connected ontology, promotes the unification of data semantics and the initial construction of knowledge graph, and outputs intermediate graph.
[0045] The consistency verification unit is used to perform fusion consistency verification based on the intermediate graph, comprehensively detect conflict information in entity semantics, relational logic and attribute constraints, record the conflict location and specific content, and output graph fragments.
[0046] The conflict correction unit is used to backtrack and analyze conflict information based on graph fragments, correct conflict content in conjunction with the semantic specifications of the minimum connected ontology, form a closed loop of fusion and construction, and output the initial knowledge graph.
[0047] According to the present invention, a multi-source heterogeneous data fusion and knowledge graph automatic construction system includes an incremental update module comprising:
[0048] The ontology generation unit is used to perform semantic parsing and entity extraction on new data based on the ontology specifications and topology of the initial knowledge graph, generate lightweight new ontology fragments compatible with the graph, and output the new ontology fragments.
[0049] The fingerprint addition unit is used to construct a local association subgraph based on entities in the newly added ontology fragment according to the two-hop relationship rule and extract structural features to generate a new entity topological fingerprint consistent with the graph format, and output a set of new topological fingerprints.
[0050] The incremental merging unit is used to match the newly added ontology fragments and the newly added topological fingerprints to the corresponding local subgraphs of the initial knowledge graph, perform incremental node and edge merging, and output a preliminary merged graph.
[0051] The consistency maintenance unit is used to check the consistency constraint chain of the initially merged graph, detect conflicts in entities, relations and attributes, correct conflict items according to ontology rules, automatically update the constraint chain, and output a complete knowledge graph.
[0052] This invention provides a system for multi-source heterogeneous data fusion and automatic knowledge graph construction, with the following beneficial effects:
[0053] 1. This invention performs unified cleaning, minimum semantic fragment extraction, constraint and instance binding on multi-source heterogeneous data such as text, structured, and semi-structured data, and completes cross-source semantic alignment using a two-layer mechanism of symbol similarity and local logical implication. Then, through redundancy removal, weak association pruning and structural optimization, a minimum connectivity ontology is generated that retains only the core semantics and necessary connectivity relationships. This achieves unified semantic governance and lightweight symbolic modeling of multi-source heterogeneous data, effectively solving the pain points of traditional solutions such as inconsistent cross-source data standards, semantic ambiguity, redundancy expansion and low reasoning efficiency. It significantly reduces the cost of data governance and semantic alignment, improves the standardization, reusability and reasoning efficiency of symbolic frameworks, and provides a stable, unified and lightweight semantic foundation for subsequent knowledge fusion.
[0054] 2. This invention generates a unique topological fingerprint based on the two-hop association topology of entities, achieves entity disambiguation and alignment based on fingerprint similarity, and performs entity resolution, relation normalization, attribute standardization and ontology semantic mapping during the semantic fusion stage. Then, a consistency verification unit comprehensively detects conflicts in entity semantics, relation logic and attribute constraints, and a conflict correction unit completes source correction and closed-loop verification, achieving highly robust entity disambiguation and end-to-end consistency assurance. It effectively distinguishes homonymous and heteronymous entities, significantly reduces logical conflicts and alignment errors, improves the accuracy of knowledge fusion and the reliability of the graph quality, reduces human intervention and subjective bias, and outputs semantically consistent, logically self-consistent and directly usable structured knowledge for reasoning.
[0055] 3. This invention generates lightweight ontology fragments and topological fingerprints from newly added data, and uses incremental local subgraph matching to achieve accurate merging without triggering global graph reconstruction. It also automatically maintains consistency constraint chains to achieve real-time conflict detection and correction, realizing lightweight incremental updates and global consistency self-maintenance of the knowledge graph. This significantly improves update efficiency, reduces computing resources and long-term operation and maintenance costs, and ensures that the graph maintains semantic unity, structural stability and evolvability throughout continuous expansion and business iteration. It balances initial construction efficiency with long-term dynamic expansion capabilities and adapts to complex scenarios with continuous access to multi-source data. Attached Figure Description
[0056] The invention will now be further described with reference to the accompanying drawings.
[0057] Figure 1 This is a module diagram of a multi-source heterogeneous data fusion and knowledge graph automatic construction system provided in an embodiment of the present invention;
[0058] Figure 2 This is a flowchart of a multi-source heterogeneous data fusion and knowledge graph automatic construction system provided by an embodiment of the present invention;
[0059] Figure 3 This is a flowchart illustrating the steps for obtaining semantic association results in a cross-source alignment unit, as provided in an embodiment of the present invention. Detailed Implementation
[0060] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below according to specific embodiments.
[0061] like Figures 1 to 3 As shown in the figure, an embodiment of the present invention provides a multi-source heterogeneous data fusion and knowledge graph automatic construction system, the system comprising:
[0062] The semantic extraction module is used to automatically extract the smallest semantic fragments from various data sources and bind constraints and examples. Cross-source alignment is completed by symbolic similarity and local logical implication, generating a minimal connected ontology containing only the necessary semantics, thus obtaining a symbolic semantic framework.
[0063] The semantic extraction unit is used to perform unified preprocessing on multi-source heterogeneous data from different systems, structures, and formats. First, it cleans and filters noisy text, incorrect format, duplicate content, and invalid fields in the original data. Then, it divides the data into boundaries and segments according to the principles of semantic integrity and minimum independent expression. It automatically identifies and extracts the smallest semantic segments that cannot be further divided, ensuring that each segment has independent and complete basic semantic information and can be understood without relying on external context. Finally, it summarizes and deduplicates all the extracted smallest semantic segments to form a basic semantic set covering all data source information, providing clean, unified, and granular original semantic data for subsequent constraint binding and standardization processing.
[0064] The constraint binding unit is used to traverse each semantic fragment in the basic semantic set. Based on the semantic content, data type, application scenario, and business rules, it automatically matches and binds the corresponding logical constraints, attribute restrictions, relational conditions, value ranges, and typical application examples. The constraint information clarifies the usage boundaries, effective conditions, legal values, and abnormal situations of each semantic fragment. The application examples enhance the interpretability and identifiability of the semantics, transforming the semantic fragments that originally only had basic meanings into standardized information that is normalized, decidable, and verifiable. Finally, all semantic fragments that have completed constraint and example binding are integrated to generate a standardized semantic set that is complete in information, uniform in format, and has clear boundaries, providing standardized input for cross-source semantic alignment.
[0065] The cross-source alignment unit is used to simultaneously perform symbol similarity calculation and local logical implication determination on semantic fragments from different data sources based on a standardized semantic set. It performs feature vectorization representation of cross-data source semantics in the standardized semantic set, extracts symbolic attributes, logical structure and constraint information of semantic fragments, and generates semantic feature vectors. Based on semantic feature vectors, a symbolic similarity algorithm is used to calculate the similarity between pairs of semantics, resulting in a semantic similarity matrix. Symbolic similarity calculation is used to determine the similarity of semantic segments from surface features such as literal symbols, structural composition, and keyword matching. Local logical implication judgment is used to determine whether there are equivalent, inclusive, or deductive relationships between segments from deeper features such as logical relations, constraints, and inference results. This two-layer judgment mechanism achieves accurate matching and association of semantic segments across data sources, establishing clear correspondences between semantic segments expressing the same or similar meanings from different sources, filtering invalid matches and weakly associated matches, verifying the antecedent and consequent deductive relationships between semantics based on the semantic similarity matrix and local logical implication judgment rules, screening candidate matching pairs, and establishing mapping associations between semantic units across data sources based on the candidate matching pairs, forming stable semantic correspondences, and ultimately generating stable, reliable, and reusable semantic association results.
[0066] The minimal ontology is used to identify and optimize redundancy in aligned semantic fragments and relational structures based on semantic association results. It eliminates repetitive expressions, redundant semantics, unnecessary nodes, and weak relationships that do not support the core logic. It retains the minimum number of semantic nodes and key connectivity relationships necessary for functional implementation, reasoning support, and structural composition. Without sacrificing the core semantics and logical integrity, it compresses the ontology size to the maximum extent. By reconstructing the connectivity paths and dependencies between semantic nodes, it forms a compact, non-redundant, logically complete, and computationally efficient minimal connectivity ontology, providing a lightweight and high-quality ontology foundation for symbolic modeling.
[0067] The symbolic modeling unit is used to perform unified symbolic transformation and standardized definition of semantic nodes, logical relationships, connectivity paths, and constraints in the minimum connected ontology. It adopts a unified symbol system, logical symbols, relational symbols, and structural rules to formally express the semantic content in natural language form, transforming the fuzzy and flexible natural language semantics into a rigorous, standardized, machine-readable, computable, and reasonable symbolic representation. At the same time, it solidifies the hierarchical relationships, connectivity relationships, and constraint rules between semantics, ensuring that the entire structure has consistency, scalability, and reusability. Finally, it generates a symbolic semantic framework with a standardized structure, rigorous logic, concise expression, and suitability for semantic reasoning and knowledge retrieval, providing a unified symbolic foundation for upper-level semantic understanding, knowledge fusion, and intelligent decision-making.
[0068] The data cleansing module is used to perform symbolic predicate structure transformation on multi-source heterogeneous data according to the symbolic semantic framework, calculate the occurrence density of each data in the isomorphic semantics, automatically marginalize low evidence data according to the density threshold, and output a cleaned symbolic predicate set.
[0069] The predicate mapping unit, based on a symbolic semantic framework, performs unified transformation of heterogeneous data from multiple sources, including text, structured, and semi-structured data, according to preset predicate rules and an adaptive matching mechanism for data types. First, it performs type identification and preprocessing on the input data. For structured data, it parses field names, data types, and values; for semi-structured data, it extracts tags, attributes, and hierarchical relationships; and for unstructured data, it performs preprocessing operations such as word segmentation, entity recognition, and relation extraction. Then, it matches the most suitable predicate template based on the preprocessed data structure and semantic connotation. Next, it fills the template parameter slots with entities, attribute values, or relational objects extracted from the original data, while standardizing date formats and units, replacing synonyms, and marking missing parameters. Finally, it summarizes all the transformed and standardized predicate expressions to form an initial symbolic predicate set covering all input data.
[0070] The frequency statistics unit is used to group predicates based on their semantic meaning and type, using isomorphic semantics as the grouping dimension. Predicates expressing the same semantic logic are grouped into the same category. Then, within each group, the initial symbolic predicate set is traversed to count the frequency of each unique predicate expression. Combined with the total amount of data involved in the transformation, the frequency of each predicate is calculated using a frequency formula. Finally, the statistical results are compiled into a frequency statistics table containing fields such as predicate identifier, semantic group, frequency, and relative frequency. This provides a quantitative basis for the threshold marginalization unit, enabling a calculable expression of the strength of predicate evidence. Its output directly serves as the core basis for the threshold marginalization unit to screen predicates with low evidence.
[0071] The threshold marginalization unit is used to set a density threshold based on the business scenario. It filters low-evidence predicates with densities below the threshold using a frequency statistics table, removes corresponding entries from the initial symbolic predicate set, and outputs an intermediate predicate set. It sets a density threshold adapted to the current semantic environment, considering the knowledge reliability and data distribution characteristics of the business scenario. Based on the density threshold, it iterates through the density values of each symbolic predicate in the frequency statistics table, filters out predicates with densities below the threshold, and aggregates them to generate threshold marginalized low-evidence data. Based on the generated threshold marginalized low-evidence data, it matches and removes predicate entries from the initial symbolic predicate set, outputting an intermediate predicate set.
[0072] The semantic cleansing unit performs semantic consistency checks on the intermediate predicate set, eliminates predicate logical conflicts, merges predicate entries with duplicate content, and outputs a cleaned symbolic predicate set. A comprehensive semantic consistency check is performed to examine whether there are conflicts in the logical relationships between predicates, such as contradictory attributes of the same entity or parameter values violating constraints. Then, the causes of the discovered logical conflicts are analyzed, and corrections are made by deleting conflicting entries, adjusting parameter values, or introducing new predicates. Subsequently, the intermediate predicate set is traversed, identifying and merging predicate entries with duplicate content or semantic equivalence. The structure of the retained predicates is optimized, and complex predicates are split or simple predicates are merged when necessary, generating a cleaned symbolic predicate set that is structurally regular, semantically consistent, and free of conflicts and redundancy.
[0073] The triple generation module is used to input the purified symbolic predicate set into the automatically induced minimal rule grammar set, and directly extract entities, relations and attributes through the grammar parse tree to generate an initial candidate set of triples.
[0074] The grammar parsing unit uses an automatically induced minimal rule grammar set as the parsing benchmark. It loads a cleaned set of symbolic predicates and a minimal rule grammar set, establishing a bidirectional adaptation mechanism. Based on the semantic type, number of parameters, and structural features of each predicate expression, it accurately matches the corresponding grammar rules, eliminating mismatched and redundant grammar entries to ensure the relevance of the parsing rules. Following the successfully matched grammar rules, it performs hierarchical structured parsing on each predicate expression, breaking it down into key parts such as the predicate core, parameter components, and logical connectors. It clarifies the semantic attributes, value ranges, and mutual constraints of each component, and standardizes and marks the decomposed components to avoid semantic ambiguity, forming a standardized and machine-readable predicate parsing structure.
[0075] The tree-labeling unit, based on the hierarchical logic of the automatically induced minimal rule grammar set, uses a hierarchical mapping method to map each normalized component in the parsing structure to a corresponding node in the grammar parse tree. It clarifies the hierarchical relationships and structure between root nodes, child nodes, and leaf nodes, ensuring a complete match between the parse tree structure and the predicate logic structure. Next, it loads preset semantic role labeling rules, which are highly compatible with the minimal rule grammar set. These rules accurately determine the semantic role of each node based on its type, parameter characteristics, and semantic connotation, clearly distinguishing between entity nodes, relation nodes, and attribute nodes. The identified three types of nodes are explicitly labeled, recording their semantic type, specific values, and associated node information. This accurately locates each node's position in the parse tree and outputs a labeled parse tree structure carrying complete semantic roles, node positions, and associated relationships.
[0076] The ternary generation unit is used to traverse all nodes of the labeled parse tree structure. Based on the node annotation information, it accurately extracts specific information about the three core elements—entities, relations, and attributes—including the specific values of entities, the semantic connotations of relations, the feature descriptions of attributes, and the relationships between various elements. Redundant nodes without actual semantic value are filtered out to ensure the effectiveness of the extracted elements. Subsequently, the semantic relevance of the extracted three types of elements is verified to confirm the logical correspondence between entities and relations, and between entities and attributes. Invalid element combinations with logical mismatches or weak connections are excluded to ensure the rationality and semantic integrity of the element combinations. Next, strictly following the standard triplet format of entity-relationship-entity and entity-attribute-attribute value, the verified elements are combined in an orderly manner, clarifying the semantic role and position of each element in each triplet to ensure that the triplet format is standardized and semantically clear. All standardized triplet combinations are summarized, and incomplete or non-standard invalid triplets are filtered out to form an initial candidate set of triplets containing multiple sets of valid semantic triplets.
[0077] The disambiguation calibration module is used to generate a topological fingerprint by using the relation structure type sequence within two hops around each entity in the initial triple candidate set. Based on the fingerprint similarity, it performs entity disambiguation and alignment, performs structural isomorphism verification and correction on candidate triples with conflicts, and outputs unambiguous structured knowledge.
[0078] The topology coding unit is used to sort out the relationships between entities based on the initial candidate set of triples. It performs structured parsing on all triples in the initial candidate set, extracting the unique identifier of each entity, the relationship type between entities, and the one-hop entity directly associated with the core entity. Based on this, it constructs an entity association network, clarifying the direct association boundaries and relationship distribution of each entity. Using each entity as a core node, it further expands outward along the identified one-hop relationships to obtain the two-hop associated entities indirectly connected to the one-hop entities and their corresponding relationships. This fully covers the entity set and relationship set within one to two hops around the core entity, forming a local topology structure centered on the core entity and containing two layers of association depth. After completing the topology expansion, feature extraction is performed on the local topology. Specifically, this includes the correlation density statistics between the core entity and one-hop and two-hop entities, the types and frequency distribution of relation types, the combination patterns of entity types, and the length of the relation path from the core entity to each hop entity and the relation sequence features on the path. These topological features are encoded using numerical methods such as hash mapping and vector embedding to generate a topological fingerprint that can uniquely represent the two-hop relation topological pattern of the entity. This ensures that the fingerprints of different entities can effectively distinguish their correlation structure differences. Finally, an entity topological fingerprint set containing the topological fingerprints of all entities and their corresponding two-hop relation topological fingerprints is output.
[0079] Iterate through all entities and associated edges in the initial triplet candidate set, build a direct association list for each entity, extract one-hop adjacent entities and their corresponding relationships, and output the one-hop association set of the entities.
[0080] Based on the one-hop association set of entities, extend outward to obtain indirectly related entities and relationships, forming a complete relationship subgraph covering two hops, and outputting a set of entity two-hop topological subgraphs.
[0081] Structural features are extracted and serialized from the set of two-hop topological subgraphs of entities to generate hash codes that uniquely identify the structure, thereby obtaining and summarizing the topological fingerprints of the two-hop relationships.
[0082] The fingerprint alignment unit calculates the topological fingerprint similarity of different entities based on the entity topological fingerprint set. It selects similarity calculation algorithms suitable for high-dimensional vectors, such as cosine similarity, Euclidean distance, and Manhattan distance, to perform dimension-by-dimensional comparison and quantification of the topological fingerprints of any two entities in the entity topological fingerprint set, obtaining numerical values of their topological structural similarity. Simultaneously, a reasonable similarity threshold is preset based on the actual application scenario and data distribution characteristics. This threshold can be dynamically optimized and adjusted through historical data verification and cross-testing to balance the accuracy and recall of entity alignment. Entity pairs with similarity values reaching or exceeding the preset threshold are selected and determined to be different denotations referring to the same real-world entity. Entity disambiguation and alignment operations are then performed to unify multiple denotations of the same entity, establish a mapping relationship between denotation forms and unique entity identifiers, and eliminate redundant and ambiguous entity denotations. During this process, all associated triples related to the aligned entity are simultaneously organized to ensure that the alignment operation does not destroy the integrity of the original association relationship. The final output is a clear and unambiguous entity mapping result, as well as all associated triples corresponding to the aligned entity, providing a unified entity basis and association data support for the subsequent conflict calibration process.
[0083] The conflict calibration unit performs logical verification on all triples based on entity mapping results and associated triples. First, it aggregates and categorizes all associated triples pointing to the same entity, identifying the relationship types, attribute values, and associated object information of the same entity in different triples, forming a set of association information at the entity level. Then, based on predefined semantic rules, the constraints of the minimum connected ontology, and the logical requirements of the business scenario, it conducts a full-dimensional logical verification on the aggregated associated triples, focusing on identifying inconsistencies between entity identifiers and entity mapping results, mismatches between relationship types and entity types, contradictory values for the same attribute of the same entity, violations of logical axioms such as transitivity and mutual exclusion between relationships, and conflicts arising from attribute values violating constraints such as data type, value range, and unit specifications. For each identified conflict, it further performs source analysis to clarify the root cause of the conflict, distinguishing between entity alignment deviations, errors in the initial candidate triple set, and problems caused by feature extraction deviations during topology encoding. Based on the source tracing results, a targeted correction strategy is adopted to precisely adjust conflicting content. This includes correcting erroneous attribute values, inconsistent relationship types, and supplementing missing key association information. Simultaneously, by considering indicators such as the credibility of the data source, the frequency of triple occurrences, and the weight of entity associations, high-credibility information is retained while low-credibility conflicting data is discarded, thereby completely eliminating ambiguity in the triples. After correction, the adjusted triples are logically verified again to ensure that all conflicts have been resolved and no new logical contradictions have been introduced. The final output is unambiguous, semantically consistent, and logically self-consistent structured knowledge.
[0084] The graph fusion module is used to integrate unambiguous structured knowledge with the minimum connected ontology, and simultaneously complete the unification of data semantics and the construction of knowledge graph in symbol space. It verifies the consistency of fusion, backtracks and corrects conflicting information, and forms an initial knowledge graph with a closed-loop output of fusion and construction.
[0085] The semantic fusion unit is used to establish the semantic mapping relationship between unambiguous structured knowledge and the minimum connected ontology in the symbol space. First, it performs entity resolution, relation normalization, and attribute standardization preprocessing on the unambiguous structured knowledge, converging heterogeneous entity references to unique identifiers, mapping business-side custom relations to ontology preset predicates, and aligning attribute names, data types, units, enumeration values with ontology constraints. Then, based on the concept hierarchy, relation axioms, and attribute cardinality constraints of the minimum connected ontology, it performs entity-to-ontology concept attribution matching, relation-to-ontology predicate semantic binding, and attribute-to-ontology specification format alignment, forming a traceable mapping index in the symbol space. It reorganizes multi-source heterogeneous structured knowledge into a unified semantic representation according to ontology specifications, completing data semantic unification and preliminary knowledge graph construction, and outputting an intermediate graph.
[0086] The consistency verification unit is used to perform a full consistency check based on the intermediate graph and the semantic rules of the minimum connected ontology. In the entity semantic dimension, it verifies the uniqueness of entity classification, the legality of concept attribution, and the absence of duplicate and ambiguous identifiers. In the relational logic dimension, it verifies the matching of relation domain and value domain, compliance with axioms such as relation transitivity / symmetry / mutual exclusion, and the absence of suspended relations and logical loops. In the attribute constraint dimension, it verifies the compliance of data type, value range, mandatory content, uniqueness, and unit precision. For each type of conflict, it marks the node ID, edge ID, conflict type, violated clause, conflict content, and evidence fragments to form a graph fragment with conflict location and details.
[0087] The conflict correction unit is used to trace the source and analyze the influence domain of conflicts based on graph fragments, distinguishing between types such as entity classification errors, misuse of relation predicates, illegal attribute values, and structural logical contradictions. It performs targeted corrections according to the minimum connectivity ontology specification, rematching the optimal concept for entity classification conflicts, adjusting predicates or reconstructing associations for relational logical conflicts, and cleaning or supplementing information for attribute constraint conflicts. After correction, a closed-loop verification is performed to ensure that the conflicts are completely eliminated without introducing new problems. Finally, it outputs an initial knowledge graph that is semantically consistent, logically self-consistent, and conforms to the ontology specification.
[0088] The incremental update module is used to generate corresponding new ontology fragments and topological fingerprints for newly added data based on the initial knowledge graph, incrementally merge them with the local subgraphs of the knowledge graph, automatically maintain consistency constraint chains, and form an initial complete knowledge graph.
[0089] The ontology generation unit, based on the ontology specifications and topology of the initial knowledge graph, first preprocesses and semantically parses the newly added data from multiple sources, removing redundant and invalid data and accurately extracting entity, relation, and attribute information. Then, strictly adhering to the specifications of the initial ontology, it normalizes the extracted semantic elements, binds them with constraint rules consistent with the initial ontology, extracts and integrates the smallest semantic units, generating lightweight, non-redundant, and fully compatible new ontology fragments. Finally, it outputs these new ontology fragments, providing a compliant and standardized semantic foundation for subsequent fingerprint generation and graph merging.
[0090] The fingerprint addition unit is used to extract all core entities from the newly added ontology fragment. Utilizing the mature two-hop relationship rules from the disambiguation calibration module, it constructs a local association subgraph for each new entity, expanding the entity's one-hop and two-hop associated entities and their corresponding relationships. Subsequently, key topological features such as node type, relationship path, hierarchical structure, and association density are extracted from the subgraph. Using the same encoding algorithm as the initial knowledge graph, these topological features are transformed into computable and comparable unique codes, generating a set of new entity topological fingerprints with a format consistent with the original graph and comparable features. This set is then output for subsequent accurate matching with the local subgraphs of the initial graph.
[0091] The incremental merging unit reads the newly added ontology fragments output by the ontology generation unit and the newly added topological fingerprint set output by the fingerprint addition unit. Through topological fingerprint comparison and semantic similarity matching, it accurately locates the newly added ontology fragments and new topological fingerprints to the corresponding local subgraphs in the initial knowledge graph. Subsequently, an incremental merging strategy is adopted to complete the insertion and merging of new nodes, edges, and attributes in a minimally invasive manner, strictly preserving the original structure and semantic relationships of the initial knowledge graph. Only the matched local subgraphs are updated, avoiding global graph reconstruction. Finally, a preliminary merged graph is output, achieving efficient and lightweight expansion of the knowledge graph.
[0092] The consistency maintenance unit traverses the consistency constraint chains in the initially merged knowledge graph, comprehensively detecting potential issues such as entity alignment conflicts, relational logic contradictions, and abnormal attribute values, accurately locating the conflict's position, type, and specific content. Subsequently, based on the ontology specifications of the initial knowledge graph and pre-defined conflict handling rules, it automatically corrects various conflict items and synchronously updates the consistency constraint chains, ensuring the corrected graph's global semantic consistency and structural integrity. Finally, after completing all conflict corrections and constraint updates, it outputs a structurally stable, semantically consistent, and unambiguous complete knowledge graph, achieving a closed loop for incremental updates.
[0093] In summary, this embodiment provides a multi-source heterogeneous data fusion and knowledge graph automatic construction system. It generates unique topological fingerprints based on the two-hop association topology of entities, achieves entity disambiguation and alignment based on fingerprint similarity, and performs entity resolution, relation normalization, attribute standardization and ontology semantic mapping in the semantic fusion stage. Then, the consistency verification unit comprehensively detects conflicts in entity semantics, relation logic and attribute constraints, and the conflict correction unit completes source correction and closed-loop verification. This achieves highly robust entity disambiguation and end-to-end consistency assurance, effectively distinguishes homonymous and heteronymous entities, significantly reduces logical conflicts and alignment errors, improves knowledge fusion accuracy and graph quality reliability, reduces human intervention and subjective bias, and outputs semantically consistent, logically self-consistent, and directly usable structured knowledge for reasoning.
[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to depart from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A system for multi-source heterogeneous data fusion and automatic knowledge graph construction, characterized in that, include: The semantic extraction module is used to automatically extract the smallest semantic fragments from various data sources and bind constraints and examples. It uses symbolic similarity calculation to standardize the semantic set, matches and associates semantic units across data sources, establishes semantic correspondences between different sources, completes cross-source alignment through symbolic similarity and local logical implication, generates the smallest connected ontology, and obtains a symbolic semantic framework. The data purification module is used to perform symbolic predicate transformation on multi-source heterogeneous data according to the symbolic semantic framework, calculate the data occurrence density under isomorphic semantics, marginalize low evidence data according to the threshold, and output a purified symbolic predicate set. The triple generation module is used to input the purified symbolic predicate set into the rule grammar set, extract entities, relations and attributes, and generate an initial triple candidate set. The disambiguation calibration module is used to organize the association relationships of each entity based on the initial triplet candidate set, expand the entities and relationships of one-hop and two-hop associations with each entity as the core, extract and encode the topological structure features, generate a unique two-hop relationship topological fingerprint for each entity, complete entity disambiguation and alignment with fingerprint similarity, and output unambiguous structured knowledge. The graph fusion module is used to fuse unambiguous structured knowledge with the minimum connected ontology, perform data semantic unification and knowledge graph construction in the symbol space, and form an initial knowledge graph. The incremental update module is used to incrementally merge new data based on the initial knowledge graph to form a complete knowledge graph.
2. The multi-source heterogeneous data fusion and knowledge graph automatic construction system according to claim 1, characterized in that: The semantic extraction module includes: The semantic extraction unit is used to clean and delineate boundaries of multi-source data, extract the smallest indivisible semantic fragments, and form a basic semantic set. The constraint binding unit is used to bind corresponding logical constraints and practical application examples to each semantic fragment according to the basic semantic set, clarify the semantic boundaries and usage scenarios, and generate a standardized semantic set. The cross-source alignment unit is used to simultaneously perform symbol similarity calculation and local logical implication determination on semantic fragments from different data sources based on a standardized semantic set, match and associate semantic units across data sources, establish semantic correspondence between different sources, and obtain semantic association results. The minimal ontology is used to eliminate redundant semantics and unnecessary associations based on the semantic association results, retain the fewest semantic nodes that support the core logic, and construct the minimal connected ontology. The symbolic modeling unit is used to perform symbolic standardization and logical structuring on the minimum connected ontology, unify the symbol system and the expression of connected relationships, and generate a symbolic semantic framework.
3. The multi-source heterogeneous data fusion and knowledge graph automatic construction system according to claim 2, characterized in that: The specific steps to obtain semantic association results in the cross-source alignment unit are as follows: The semantics across data sources in the standardized semantic set are represented by feature vectorization, and the symbolic attributes, logical structure and constraint information of semantic fragments are extracted to generate semantic feature vectors; The semantic similarity matrix is obtained by calculating the similarity between pairs of semantic features using a symbolic similarity algorithm based on the semantic feature vectors. Based on the semantic similarity matrix and the local logical implication judgment rule, the antecedent and consequent inference relationship between semantics is verified, and candidate matching pairs are screened. Based on the candidate matching pairs, establish a mapping association between semantic units across data sources, form a stable semantic correspondence, and output the semantic association result.
4. The multi-source heterogeneous data fusion and knowledge graph automatic construction system according to claim 1, characterized in that: The data cleansing module includes: The predicate mapping unit is used to map multi-source heterogeneous data into standardized predicate expressions according to the predicate rules defined by the symbolic semantic framework, adapt the data types of the predicate transformation logic, determine the semantic predicate and parameters of each data unit, and output the initial symbolic predicate set. The frequency statistics unit is used to count the occurrence frequency of each symbolic predicate based on the initial symbolic predicate set, using isomorphic semantics as the grouping dimension, and calculate the occurrence frequency of the predicates in combination with the total amount of data involved in the conversion, and output a frequency statistics table. The threshold marginalization unit is used to set a density threshold according to the business scenario, filter low evidence predicates with a density lower than the threshold according to the frequency statistics table, remove the corresponding entries from the initial symbolic predicate set, and output the intermediate predicate set. The semantic purification unit is used to perform semantic consistency verification based on the intermediate predicate set, eliminate predicate logic conflicts, merge predicate entries with duplicate content, and output a purified symbolic predicate set.
5. The multi-source heterogeneous data fusion and knowledge graph automatic construction system according to claim 4, characterized in that: The specific steps for outputting the intermediate predicate set in the threshold marginalization unit are as follows: Based on the knowledge reliability and data distribution characteristics of the business scenario, set a density threshold that adapts to the current semantic environment; Based on the density threshold, the density values of each symbolic predicate in the frequency statistics table are traversed, predicates with densities less than the threshold are selected and summarized to generate threshold-marginalized low evidence data. Based on the generated threshold-margined low-evidence data, predicate entries are matched and removed from the initial symbolic predicate set, and an intermediate predicate set is output.
6. The multi-source heterogeneous data fusion and knowledge graph automatic construction system according to claim 1, characterized in that: The ternary generation module includes: The grammar parsing unit is used to adapt the cleaned symbolic predicate set to the automatically induced minimal rule grammar set, perform structured parsing of each predicate expression according to the grammar rules, and output the predicate parsing structure; The tree labeling and positioning unit is used to construct a grammar parse tree based on the predicate parsing structure, determine the semantic role of each node in the parse tree, locate the positions of entity nodes, relation nodes and attribute nodes, and output the labeled parse tree structure. The ternary generation unit is used to extract three core elements—entities, relations, and attributes—from the tree according to the labeled parsing tree structure, combine them in the format of entity-relation-attribute triples, and output an initial candidate set of triples.
7. The multi-source heterogeneous data fusion and knowledge graph automatic construction system according to claim 1, characterized in that: The disambiguation calibration module includes: The topology coding unit is used to sort out the association relationships of each entity based on the initial triplet candidate set, expand the entities and relationships of its one-hop and two-hop associations with each entity as the core, extract the topology structure features and encode them, generate a unique two-hop relationship topology fingerprint for each entity, and output the entity topology fingerprint set. The fingerprint alignment unit is used to calculate the topological fingerprint similarity of different entities based on the entity topological fingerprint set, filter entities whose similarity reaches a preset threshold, determine them as the same entity, perform disambiguation and alignment, and output entity mapping results and associated triplets. The conflict calibration unit is used to perform logical verification on all triples based on the entity mapping results and associated triples, identify entries with conflicting entities, relations or attributes, eliminate ambiguity by correcting conflicting content and retaining high-confidence information, and output unambiguous structured knowledge.
8. The multi-source heterogeneous data fusion and knowledge graph automatic construction system according to claim 7, characterized in that: The specific steps for outputting the entity topological fingerprint set in the topological coding unit are as follows: Iterate through all entities and associated edges in the initial triplet candidate set, build a direct association list for each entity, extract one-hop adjacent entities and their corresponding relationships, and output the one-hop association set of the entity. Based on the entity one-hop association set, indirectly associated entities and relationships are obtained by extending outward, forming a complete relationship subgraph covering two hops, and outputting a set of entity two-hop topology subgraphs; Structural features are extracted and serialized into the set of two-hop topological subgraphs of the entities to generate hash codes that uniquely identify the structure, thereby obtaining and summarizing the two-hop relationship topological fingerprints.
9. The system for multi-source heterogeneous data fusion and automatic knowledge graph construction according to claim 1, characterized in that: The map fusion module includes: The semantic fusion unit is used to establish the semantic mapping relationship between the unambiguous structured knowledge and the minimum connected ontology in the symbol space, integrate the entities, relations, and attributes of the unambiguous structured knowledge with the semantic specifications of the minimum connected ontology, promote the unification of data semantics and the initial construction of knowledge graph, and output intermediate graph. The consistency verification unit is used to perform fusion consistency verification based on the intermediate graph, comprehensively detect conflict information in entity semantics, relational logic and attribute constraints, record the conflict location and specific content, and output graph fragments. The conflict correction unit is used to perform backtracking analysis on the conflict information based on the graph fragments, correct the conflict content in combination with the semantic specifications of the minimum connected ontology, form a closed loop of fusion and construction, and output the initial knowledge graph.
10. The multi-source heterogeneous data fusion and knowledge graph automatic construction system according to claim 1, characterized in that: The incremental update module includes: The ontology generation unit is used to perform semantic parsing and entity extraction on the new data according to the ontology specification and topology of the initial knowledge graph, generate a lightweight new ontology fragment compatible with the graph, and output the new ontology fragment. The fingerprint addition unit is used to construct a local association subgraph and extract structural features based on the entities in the newly added ontology fragment according to the two-hop relationship rule, generate a new entity topological fingerprint consistent with the graph format, and output a set of new topological fingerprints. The incremental merging unit is used to match the newly added ontology fragment and the newly added topological fingerprint to the corresponding local subgraph of the initial knowledge graph, perform incremental node and edge merging, and output a preliminary merged graph. The consistency maintenance unit is used to check the consistency constraint chain of the preliminary merged graph, detect conflicts in entities, relations and attributes, correct conflict items according to ontology rules, automatically update the constraint chain, and output a complete knowledge graph.
Citation Information
Patent Citations
Mine ventilation knowledge graph construction method based on large model
CN120706527A
SQL (Structured Query Language) statement structure verification system based on large model and knowledge graph fusion enhancement
CN121326958A