Knowledge fusion method and device, machine readable storage medium and electronic equipment
By generating candidate interpretation hypotheses through mapping rule base and topological attribute analysis, this study solves the translation dilemma in the research of integrated traditional Chinese and Western medicine, realizes the modern interpretation of TCM concepts and the generation of scientific hypotheses, and promotes in-depth research and application of integrated traditional Chinese and Western medicine.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING SHIJITAN HOSPITAL CAPITAL MEDICAL UNIVERSITY
- Filing Date
- 2025-11-25
- Publication Date
- 2026-04-24
AI Technical Summary
In the study of integrating traditional Chinese and Western medicine, the unique concepts of traditional Chinese medicine are difficult to directly correspond with modern medical terminology, leading to translation difficulties and hindering the integration of research from different knowledge fields.
By using a mapping rule base, the set of mismatched nodes in the terminology standardization transition network is determined. Candidate interpretation hypotheses are generated using topological attributes and reasoning principles. The mapping rule base is then adjusted by expert arbitration until preset conditions are met, and the final mapping rule base is output.
It alleviated the translation dilemma of mismatched nodes, promoted the integration of traditional Chinese and Western medicine research, generated verifiable scientific hypotheses, and guided clinical research and drug development.
Smart Images

Figure CN121920495A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a knowledge fusion method and apparatus, a machine-readable storage medium, and an electronic device. Background Technology
[0002] Traditional Chinese medicine is a complex knowledge system based on experience. Its theoretical concepts (such as "qi" and "yin fire") are abstract and difficult to directly correspond to modern medical terminology.
[0003] Current research on the integration of traditional Chinese and Western medicine faces fundamental challenges, such as the translation dilemma. One-to-one concept mapping often fails due to differences in terminology systems, and a large number of unique TCM concepts become "isolated points".
[0004] Therefore, how to alleviate the translation difficulties between different knowledge domains and promote collaborative research between them has become a technical problem that needs to be solved in this field. Summary of the Invention
[0005] In view of this, this application proposes a knowledge fusion method and apparatus, a machine-readable storage medium and an electronic device to alleviate the translation difficulties between different knowledge domains and promote the integration of research between different knowledge domains.
[0006] In a first aspect, this application provides a knowledge fusion method, which includes: determining a set of mismatched nodes in a terminology standardization transition network based on a mapping rule base, wherein the terminology standardization transition network is determined based on a first knowledge domain network, and the set of mismatched nodes includes first nodes in the terminology standardization transition network that do not have a matching second node in the second knowledge domain network; for each first node in the determined set of mismatched nodes, determining at least one candidate interpretation hypothesis, wherein the candidate interpretation hypothesis is related to the potential meaning of the first node in the second knowledge domain; verifying the at least one candidate interpretation hypothesis and adjusting the mapping rule base according to the verification result; repeatedly executing the steps of determining the set of mismatched nodes, determining at least one candidate interpretation hypothesis, verifying at least one candidate interpretation hypothesis, and adjusting the mapping rule base until a preset condition is met; and outputting the final mapping rule base.
[0007] Optionally, determining the set of mismatched nodes in the terminology standardization transition network based on the mapping rule base includes: determining the largest common subgraph between the terminology standardization transition network and the second knowledge domain network based on the mapping rule base; and determining the first node in the terminology standardization transition network that is not included in the determined largest common subgraph, thereby determining the set of mismatched nodes.
[0008] Optionally, for each first node in the determined set of mismatched nodes, at least one candidate interpretation hypothesis is determined, including: determining the topological properties of the first node in the terminology standardization transition network; and determining the at least one candidate interpretation hypothesis based on the determined topological properties.
[0009] Optionally, the topological properties include at least one of the following: degree centrality, betweenness centrality, structural hole index, and egocentric network structure.
[0010] Optionally, determining the at least one candidate interpretation hypothesis based on the determined topological attributes includes: determining the structural role of the first node in the terminology standardization transition network based on the determined topological attributes; and determining the at least one candidate interpretation hypothesis based on the determined structural role.
[0011] Optionally, based on the determined structural role, the at least one candidate interpretive hypothesis is determined, including: determining the at least one candidate interpretive hypothesis based on the determined structural role and at least one of the following principles: functional equivalence principle, mediator variable principle, and system emergent property principle.
[0012] Secondly, this application also provides a knowledge fusion apparatus, comprising: a processing module configured to: determine a set of mismatched nodes in a terminology standardization transition network based on a mapping rule base, wherein the terminology standardization transition network is determined based on a first knowledge domain network, and the set of mismatched nodes includes first nodes in the terminology standardization transition network that do not have a matching second node in the second knowledge domain network; for each first node in the determined set of mismatched nodes, determine at least one candidate interpretation hypothesis, wherein the candidate interpretation hypothesis is related to the potential meaning of the first node in the second knowledge domain; verify the at least one candidate interpretation hypothesis and adjust the mapping rule base according to the verification result; repeatedly execute the steps of determining the set of mismatched nodes, determining at least one candidate interpretation hypothesis, verifying at least one candidate interpretation hypothesis, and adjusting the mapping rule base until a preset condition is met; and output the final mapping rule base.
[0013] Optionally, determining the set of mismatched nodes in the terminology standardization transition network based on the mapping rule base includes: determining the largest common subgraph between the terminology standardization transition network and the second knowledge domain network based on the mapping rule base; and determining the first node in the terminology standardization transition network that is not included in the determined largest common subgraph, thereby determining the set of mismatched nodes.
[0014] Optionally, for each first node in the determined set of mismatched nodes, at least one candidate interpretation hypothesis is determined, including: determining the topological properties of the first node in the terminology standardization transition network; and determining the at least one candidate interpretation hypothesis based on the determined topological properties.
[0015] Optionally, the topological properties include at least one of the following: degree centrality, betweenness centrality, structural hole index, and egocentric network structure.
[0016] Optionally, determining the at least one candidate interpretation hypothesis based on the determined topological attributes includes: determining the structural role of the first node in the terminology standardization transition network based on the determined topological attributes; and determining the at least one candidate interpretation hypothesis based on the determined structural role.
[0017] Optionally, based on the determined structural role, the at least one candidate interpretive hypothesis is determined, including: determining the at least one candidate interpretive hypothesis based on the determined structural role and at least one of the following principles: functional equivalence principle, mediator variable principle, and system emergent property principle.
[0018] Thirdly, this application also provides a machine-readable storage medium storing instructions that cause a machine to execute the knowledge fusion method described above.
[0019] Fourthly, this application also provides an electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the executable instructions to implement the knowledge fusion method described above.
[0020] According to the technical solution of this application, a set of mismatched nodes in the terminology standardization transition network is determined based on a mapping rule base; for each first node in the determined set of mismatched nodes, at least one candidate interpretation hypothesis is determined; the at least one candidate interpretation hypothesis is verified, and the mapping rule base is adjusted according to the verification results; the process of determining the set of mismatched nodes, determining at least one candidate interpretation hypothesis, verifying at least one candidate interpretation hypothesis, and adjusting the mapping rule base is repeated until a preset condition is met; the final mapping rule base is output; thus, the final mapping rule base adds the mapping relationship between the mismatched nodes in the terminology standardization transition network and their corresponding interpretations in the second knowledge domain, alleviating the translation dilemma of mismatched nodes to the second knowledge domain, and promoting the combined research of the first and second knowledge domains.
[0021] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description
[0022] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application, and the illustrative embodiments and descriptions thereof are used to explain this application. In the drawings: Figure 1 A flowchart of a knowledge fusion method according to a preferred embodiment of this application; Figure 2 This is a schematic diagram of the maximum common subgraph and mismatched nodes after network alignment according to a preferred embodiment of this application; Figure 3 This is a schematic diagram of the egocentric network analysis of mismatched nodes according to a preferred embodiment of this application; Figure 4 A schematic diagram of the human-computer interaction interface for expert arbitration is generated based on the candidate assumptions of the preferred embodiments of this application. Detailed Implementation
[0023] The technical solution of this application will now be described in detail with reference to the accompanying drawings and embodiments.
[0024] Firstly, this application provides a knowledge fusion method.
[0025] Figure 1 This is a flowchart of a knowledge fusion method according to a preferred embodiment of this application. For example... Figure 1 As shown, this knowledge fusion method includes the following:
[0026] In step S10, based on the mapping rule base, the set of mismatched nodes in the terminology standardization transition network is determined. The terminology standardization transition network is determined based on the first knowledge domain network, and the set of mismatched nodes includes first nodes in the terminology standardization transition network that do not have a matching second node in the second knowledge domain network; these are the mismatched nodes.
[0027] The first and second knowledge domain networks correspond to different knowledge domains, and are labeled as the first knowledge domain and the second knowledge domain, respectively. For example, the first knowledge domain network could be a Traditional Chinese Medicine (TCM) pathogenesis network, while the second knowledge domain network could be a modern pathology network. The network consists of nodes and edges; nodes represent concepts, and edges represent relationships between concepts.
[0028] The terminology standardization transition network, also known as the vernacular network (G_Vernacular), is a key bridge for achieving efficient alignment in this application. It inherits the entire topology of the first knowledge domain network and describes the nodes of the first knowledge domain network using vernacular terminology.
[0029] For example, the terminology standardization transition network corresponding to the Traditional Chinese Medicine Pathogenesis Network (G_TCM) can be determined based on the following:
[0030] (1) Node generation. Each node in G_TCM is reviewed by multiple domain experts (e.g., 3), and one or more standardized vernacular terms are established for it, and the corresponding node in the terminology standardization transition network is determined. For example, “Qingyang” is interpreted as “the function of light and clear rising Yang Qi”, and “spleen and stomach deficiency” is mapped as “reduced digestive and absorptive functions”.
[0031] (2) Structural inheritance. G_Vernacular inherits all the topological structures of G_TCM, meaning that the two are completely isomorphic in terms of edge connections. Essentially, it is a semantically translated version of G_TCM, aiming to preserve the original logical relationships while reducing the semantic complexity of subsequent cross-domain alignment.
[0032] (3) Unmapped processing. For extremely abstract concepts that experts cannot reach a consensus on or that cannot be interpreted in modern language (such as “chonghe zhi wei qi”), their original node names are preserved in G_Vernacular and they are marked as high-value mismatch candidates.
[0033] The mapping rule base shows the correspondence between nodes in networks of two different knowledge domains, including definite mapping pairs that can be constructed by domain experts.
[0034] Based on the mapping rule base, the first node in the terminology standardization transition network is matched with the second node in the second knowledge domain network. The set of first nodes that do not have a matching second node is determined as the set of mismatched nodes.
[0035] The second knowledge domain network can be a modern medical network (G_Modern). Specifically, it is independently constructed from modern medical textbooks, research literature, and biomedical databases. The nodes are explicit modern pathological and physiological concepts (such as "decreased blood volume" and "increased core body temperature"), and the edges represent proven physiological and pathological mechanisms.
[0036] The mapping rule base M is constructed by domain experts based on the plaintext nodes of G_Vernacular and the nodes of G_Modern, and contains exact mapping pairs. For example: "Reduced digestive and absorptive function" (G_Vernacular) → "Intestinal malabsorption" (G_Modern) [Strong mapping] "Insufficient blood volume" (G_Vernacular) → "Hypovolemia" (G_Modern) [Strong mapping] In step S11, for each first node in the determined set of mismatched nodes, at least one candidate interpretation hypothesis is determined. The candidate interpretation hypothesis is related to the potential meaning of the first node in the second knowledge domain.
[0037] In step S12, at least one candidate interpretation hypothesis is verified, and the mapping rule base is adjusted based on the verification results.
[0038] In step S13, it is determined whether the preset conditions are met. If yes, step S14 is executed; otherwise, step S10 is executed.
[0039] The preset conditions can be determined according to specific circumstances. For example, a preset condition could be that there are no mismatched nodes. Alternatively, a preset condition could be that the running time reaches a preset time value. Or, a preset number of iterations could be a preset number of iterations.
[0040] In step S14, the final mapping rule base is output. The mapping rule base is continuously adjusted during the iteration process. When a preset condition is met, the adjustment of the mapping rule base stops, and the final mapping rule base is determined.
[0041] Optionally, in an embodiment of this application, determining the set of mismatched nodes in the terminology standardization transition network based on the mapping rule base may include the following:
[0042] Based on a mapping rule base, the maximum common subgraph (MCS) of the terminology standardization transition network and the second knowledge domain network is determined. The two networks are then structurally aligned using the mapping rule base, and their MCS is calculated. The core alignment and discovery process of this application occurs between the terminology standardization transition network and the second knowledge domain network.
[0043] To identify the set of mismatched nodes, the first node in the terminology standardization transition network that is not included in the identified largest common subgraph is determined. Figure 2 As shown, the TCM pathogenesis network and the modern pathology network are examples.
[0044] In this application embodiment, nodes in the terminology-standardized transition network that fail to be included in the maximum common subgraph are identified and output, and are defined as "mismatched node set".
[0045] Identification and Origin Tracing of Mismatched Nodes. 1) After the algorithm is completed, all nodes in G_Vernacular that were not included in the MCS are defined as the set of mismatched nodes in this alignment. 2) This set of mismatched nodes has a clear origin path: some may originate from highly abstract TCM concepts in the first knowledge domain network (e.g., G_TCM) that retained their original names when constructing the terminology standardization transition network G_Vernacular; others may be concepts that, although they have vernacular interpretations, cannot find structural counterparts in the second knowledge domain network (e.g., the modern medicine network G_Modern). These nodes are the core targets for subsequent topological attribute analysis and generation of candidate scientific hypotheses.
[0046] Optionally, in this embodiment of the application, determining the maximum common subgraph (MCS) between the terminology standardization transition network and the second knowledge domain network based on the mapping rule base can be achieved by using a graph isomorphism algorithm based on backtracking search and structural constraints (such as the McGregor algorithm or its optimized variants) to compute the maximum common subgraph (MCS). Specifically, the MCS is determined according to the following:
[0047] 1) Seed matching initialization. Based on the mapping pair with the highest confidence in the mapping rule base, the initial matching node pair is determined as the "seed" for the growth of the common subgraph.
[0048] 2) Neighborhood Collaborative Expansion. Starting from the two matched seed nodes, synchronously traverse their respective direct neighbor nodes in the terminology standardization transition network and the second knowledge domain network. For each unmatched neighbor node u in the terminology standardization transition network, according to the mapping rule base M, find all candidate matching nodes v of the unmatched neighbor node u in the second knowledge domain network.
[0049] 3) Dual Constraint Verification. A rigorous dual feasibility verification is performed on each candidate matching pair (u, v), namely semantic constraints and structural consistency verification.
[0050] For semantic constraints, candidate pairings (u, v) must conform to or closely approximate the definition of the mapping rule base M.
[0051] For structural consistency constraints, the connection relationship between an unmatched neighbor node u and its matched neighbor nodes in the terminology standardization transition network must be completely consistent with the connection relationship between a candidate matching node v and its corresponding matched neighbor nodes in the second knowledge domain network. For example, if there is an edge (u, u1) in the terminology standardization transition network and u1 is matched with v1, then there must be an edge (v, v1) in the second knowledge domain network.
[0052] 4) Recursive backtracking and optimization. Verified candidate pairs are formally added to the common subgraph, and this becomes the new starting point for recursive expansion. When multiple candidate paths exist, the algorithm employs a depth-first or best-first strategy for exploration, and backtracks when encountering dead ends to ensure that the final common subgraph (MCS) is the maximally connected subgraph that satisfies the constraints.
[0053] 5) Approximate calculation option. When the network size is extremely large, approximation methods such as greedy algorithms can be used to find a sufficiently large common subgraph within a reasonable time to balance computational efficiency and accuracy.
[0054] Optionally, in an embodiment of this application, for each first node in the determined set of mismatched nodes, at least one candidate interpretation hypothesis is determined, which may include the following.
[0055] Determine the topological properties of the first node in the terminology standardization transition network. Based on the determined topological properties, identify at least one candidate interpretation hypothesis.
[0056] Optionally, in embodiments of this application, the topological properties include at least one of the following: degree centrality, betweenness centrality, structural hole index, and egocentric network structure.
[0057] For each mismatched node in the set of mismatched nodes, a series of topological properties are computed in the terminology standardization transition network to quantify its function and status within the network. Key properties include: a) Degree centrality. Calculate the degree of a mismatched node, which is the number of its direct connecting edges. Height centrality indicates that the node is a hub in the network, directly related to many other concepts, and may correspond to a core physiological function or pathological hub.
[0058] b) Betweenness centrality. Measures the frequency with which a mismatched node appears on the shortest path between any two other nodes in a terminology normalization transition network. The formula is: .in, It is a node in the terminology standardization transition network. To the node The number of shortest paths, It is one of the mismatches The number of paths. High betweenness centrality indicates that the node is a key bridge or information relay station between different modules in the network, and its dysfunction may cause widespread systemic effects.
[0059] c) Structural void index. Assessing mismatched nodes. The ability to connect different communities and act as an "intermediary." We use a constraint coefficient to measure this; the lower the value, the more important the structural hole the node occupies. This means that the node controls the flow of information or resources between different groups.
[0060] d) Egocentric network structure. Extract terminology from the transition network, focusing on mismatched nodes. A local network centered on a node, within one or more steps (i.e., its direct neighbors and the connections between these neighbors). This reveals the node's direct functional context and interaction patterns.
[0061] Optionally, in embodiments of this application, determining at least one candidate interpretation hypothesis based on the determined topological properties may include the following.
[0062] Based on the determined topological properties, the structural role of the first node in the terminology standardization transition network is determined.
[0063] Based on the predefined correspondence between topological attributes and structural roles, and combined with the determined topological attributes, the structural role corresponding to the first node (i.e., the mismatched node) in the set of mismatched nodes is determined. The correspondence between topological attributes and structural roles includes: "height value" corresponding to "core hub," "high betweenness coefficient" corresponding to "critical bridge," "low constraint coefficient" corresponding to "community mediator," and "specific self-network structure" corresponding to "local regulator."
[0064] Based on the determined structural roles, at least one candidate interpretive hypothesis is identified. For example... Figure 3 As shown, taking the TCM pathogenesis network and the modern pathology network as examples, combined with... Figure 2 The content shown.
[0065] In this application, candidate interpretation hypothesis generation is the core of achieving automated knowledge discovery. This process automatically infers the potential correspondence mechanism in the second knowledge domain for each mismatched node based on its structural role in the terminology standardization transition network.
[0066] Optionally, in embodiments of this application, determining the at least one candidate interpretation hypothesis based on the determined structural role includes: Based on the determined structural role and at least one of the following principles, at least one candidate interpretive hypothesis is determined: the functional equivalence principle, the mediator variable principle, and the system emergent property principle.
[0067] In this application embodiment, structural roles are mapped to one or more scientific reasoning principles, and candidate interpretation hypotheses for natural language description are generated.
[0068] (1) Functional equivalence principle. If the structural role corresponding to the mismatched node u is a "core hub", the system will search for a second node in the second knowledge domain network that has similarity centrality and has an existing association with the neighboring nodes of the mismatched node in the terminology standardization transition network, i.e., the second node v corresponding to the mismatched node. Based on this, the hypothesis is generated: "the concept corresponding to the mismatched node u" may play a similar systemic core functional role in the second knowledge domain corresponding to the second knowledge domain network as "the second node v corresponding to the mismatched node".
[0069] (2) Mediation Variable Principle. If the structural role corresponding to the mismatched node u is a "key bridge" or "community mediator", the system will analyze the neighbor node pairs (first node A, second node B) that it successfully maps to in the second knowledge domain network in the terminology standardization transition network. Based on this, the hypothesis is generated: "The concept corresponding to the mismatched node u" may be an unknown intermediate mechanism or pathological process connecting known concept A (the concept corresponding to first node A) and concept B (the concept corresponding to second node B). For example, for "failure of clear yang to rise", the system finds that it is highly between the two connections of "spleen and stomach deficiency" (mapped to "absorption disorder") and "insufficient body fluids" (mapped to "decreased blood volume"). Therefore, the hypothesis is generated: "'Failure of clear yang to rise' may be 'microcirculatory dynamics or skin blood flow regulation dysfunction' connecting 'absorption disorder' and 'decreased blood volume'."
[0070] (3) Emergent Attribute Principle. If the egocentric network structure of mismatch node u exhibits a specific, dense or modular connection pattern, but is not mapped itself, the system will generate the hypothesis that "the concept corresponding to mismatch node u" may not be an independent entity, but a comprehensive functional state or clinical syndrome that emerges from the interaction of known mechanisms represented by its neighboring nodes.
[0071] In this embodiment, the output format of candidate interpretation hypotheses can be set. Ultimately, the system outputs one or more structured candidate interpretation hypotheses for each mismatched node, in the format: "[The concept corresponding to the mismatched node] exhibits [core topological properties, such as high betweenness] in the terminology standardization transition network. Based on [reasoning principles, such as the mediator variable principle], we hypothesize that it may correspond to [a specific functional description or mechanism conjecture] in the second knowledge domain." These candidate interpretation hypotheses serve as the original material driving the next step of expert arbitration.
[0072] Optionally, in embodiments of this application, verifying at least one candidate interpretation hypothesis and adjusting the mapping rule base based on the verification results may include the following:
[0073] Submit at least one candidate interpretation hypothesis to the expert arbitration system for verification, and receive the verification results returned by the expert arbitration system.
[0074] Optionally, the expert arbitration system adopts a multi-person back-to-back review, conference arbitration and voting mechanism, and outputs the acceptance, rejection or modification opinions of candidate hypotheses, and can add new enhanced mapping or weak mapping rules.
[0075] Specifically, the operational process of an expert arbitration system may include the following:
[0076] (1) Input the generated candidate interpretation hypothesis set, each hypothesis is accompanied by the basis for its generation (such as topological attribute data, reasoning principles).
[0077] (2) Arbitration Mode. A structured human-computer interaction interface is adopted, such as... Figure 4 As shown, one or more of the following modes are supported: (a) Back-to-back independent review, where multiple domain experts (such as experts in modern medicine, traditional Chinese medicine, and computer science) independently review each candidate interpretation hypothesis. (b) Conference arbitration, where expert teams hold online or offline meetings to discuss controversial hypotheses. (c) Voting mechanism, where each candidate interpretation hypothesis is voted on, and a passing threshold is set (such as a two-thirds majority agreement).
[0078] Arbitration decision output. For each candidate interpretation hypothesis, the expert system outputs one of the following four decision categories: (1) Accept: The hypothesis is considered scientifically reasonable and adopted. (2) Reject: The hypothesis is considered invalid and rejected. (3) Accept after modification: The hypothesis is accepted after adjusting its description or correspondence. (4) Shelve: Due to insufficient information, no decision is made for the time being, and the decision will be left for subsequent iterations.
[0079] Optionally, in this embodiment of the application, adjusting the mapping rule base based on the verification results may include the following: Based on the verification results, the mapping rule base may be dynamically adjusted as follows.
[0080] a) Add new mapping rules. The application scenario is for candidate interpretation hypotheses whose verification results are "accepted" or "accepted after modification". Specifically, a new mapping relationship is established between the mismatched node u (from the terminology standardization transition network) and the concept v proposed in the interpretation hypothesis (from the second knowledge domain network), and this relationship is added to the mapping rule base.
[0081] Rule grading. Newly added mappings are labeled as "weak mappings" or "exploratory mappings" and accompanied by a confidence score (the initial score can be based on the proportion of votes in favor). For example: New rule: "Qingyang not rising" -> "Skin blood flow regulation dysfunction" [Mapping type: weak mapping | Confidence: 0.8 | Source: Algorithm generation - expert arbitration].
[0082] b) Strengthen existing mapping rules. The application scenario is that in subsequent iterations, if a "weak mapping" rule is repeatedly and independently verified and accepted by arbitration, the specific operation involves automatically or with expert confirmation upgrading it to a "strong mapping" and increasing its confidence level. This signifies that a hypothesis is gradually transforming into relatively stable knowledge.
[0083] c) Modifying and discarding mapping rules. Application scenarios include when the arbitration result is "rejected," or when existing rules are found to cause structural conflicts in multiple alignments. Specifically, this involves "downgrading" the mapping rule (e.g., from a strong mapping to a weak mapping); or marking it as "invalid" or "questionable" in the mapping rule base, reducing its weight in subsequent calculations, or discontinuing its use.
[0084] d) Enrich rule attributes. Specifically, adjusting the mapping rule base involves not only adding or removing rules, but also enriching its metadata. For example, adding the following to a rule: 1) Contextual constraints, specifying the pathological context in which the mapping holds; 2) Evidence sources, linking to the original topological attributes and arbitration records that generated the hypothesis; 3) Proposed iteration rounds, recording in which iteration the rule was added or modified.
[0085] In this application embodiment, the process of adjusting the mapping rule base is iterative. Each adjusted mapping rule base will serve as input for determining the set of mismatched nodes. This means the following: (1) In the first round, alignment is performed using a small number of initial, expert-manufactured "strong mapping" rules to discover the first batch of mismatched nodes. (2) In the second round, alignment is performed again using a new rule base containing "weak mappings" generated in the first round of arbitration. These new rules may help previously misaligned nodes find their positions or expose new, deeper mismatches. (3) Multiple rounds of iteration. This cycle continues, and the system's knowledge base (mapping rule base) grows and is optimized. The aligned common subgraph becomes larger and larger, and the discovered mismatched nodes become more abstract and core, driving knowledge discovery to develop in depth.
[0086] Optionally, in embodiments of this application, new scientific hypotheses can be determined based on the final mapping rule base. For example, new scientific hypotheses can be determined using newly added mappings in the final mapping rule base (e.g., mappings between mismatched nodes and their corresponding validated candidate interpretative hypotheses).
[0087] Preferably, new scientific hypotheses can be determined by combining the largest common subgraph corresponding to the final mapping rule base. For example, the mapping rule base can be adjusted (by adding a weak mapping rule for "failure of clear yang to rise" → "dysregulation of skin blood flow"), and a new scientific hypothesis can be generated based on this: "The prescriptions for treating 'yin fire' in the *Treatise on the Spleen and Stomach* (such as Buzhong Yiqi Tang and Qingshu Yiqi Tang) and their treatment theories can be used to guide the development of drugs that improve skin microcirculation and treat related modern diseases (such as chronic skin ulcers and microcirculation-related diseases)." The new scientific hypothesis provides a clear direction for subsequent pharmacological experiments and clinical research.
[0088] Optionally, in this embodiment, the "newly generated scientific hypotheses" are not generated out of thin air, but rather are the integration and refinement of all the computational analysis and expert arbitration results from the aforementioned steps. These hypotheses are mainly generated automatically through cross-domain, in-depth correlation analysis of the adjusted mapping rule base and the continuously expanding public subgraph. The specific generation mechanism is as follows: (1) Data foundation for hypothesis generation. 1) Input 1: The refined mapping rule base. This is a rich knowledge system containing confidence and source information for mappings from "strong mappings" to "weak mappings". 2) Input 2: A stable common subgraph formed after multiple iterations. This subgraph represents the core transformation process with structural similarity between the first and second knowledge domains. For example, for traditional Chinese medicine and modern medicine, the common subgraph represents the core pathophysiological processes with structural similarity between traditional Chinese medicine and modern medicine. 3) Input 3: The mismatched nodes confirmed by arbitration and their final interpretation, i.e., the mismatched nodes and the verified candidate interpretation hypotheses. These are the most valuable "crystallizations" in the knowledge fusion process.
[0089] (2) The generation path and types of scientific hypotheses.
[0090] Based on the results of iterative fusion, the system generates verifiable scientific hypotheses through the following core pathways. Examples from traditional Chinese medicine and modern medicine are provided below.
[0091] (a) Path 1: Hypothesis on the definition of the research population based on the transformation of diagnostic criteria Generation logic: When the system confirms through network alignment that an abstract concept (such as "yin fire") is highly similar to a syndrome with mature diagnostic criteria (such as "spleen and stomach damp-heat syndrome") in terms of pathogenesis network structure, it automatically proposes a proxy scheme to use the diagnostic criteria of the syndrome as the research object for the abstract concept.
[0092] Hypothesis output: "Based on the network alignment showing that concept A and syndrome B are highly similar in structure, it is recommended to adopt the diagnostic criteria of the Expert Consensus on Syndrome B as a standard proxy for recruiting study subjects of A in clinical research." (b) Path 2: Hypothesis of non-invasive diagnostic biomarkers based on symptom pattern quantification Generation logic: By aligning with ancient texts to accurately define the differences in key symptoms (such as the distribution range of "spontaneous sweating" and "micro-sweating" on the body surface), and then aligning with modern physiological knowledge (such as the differences in sweat gland function in different skin areas), quantifiable physiological indicators are proposed as identification markers.
[0093] Hypothesis output: "Based on the precise definition of symptom patterns P1 and P2 and their alignment with modern physiological mechanisms M, a verifiable hypothesis is proposed: the quantitative indicator Q can serve as a non-invasive biomarker to distinguish the target syndrome from healthy / other syndromes." (c) Path 3: Drug repositioning hypothesis based on therapy-mechanism mapping Generation logic: The system identifies the pathogenesis module intervened by a certain TCM prescription, and after mapping it to the modern pathology module through mapping rules, it automatically retrieves the modern diseases involved in the pathology module and proposes a new indication hypothesis for the therapy.
[0094] Hypothesis output: "Traditional Chinese medicine formula F, because it acts on the pathogenesis module M_TCM (corresponding to the modern pathology module M_Modern), may have therapeutic potential for modern diseases D_Modern." (d) Path 4: Hypothesis of early diagnostic biomarkers based on reverse translation Generation logic: The system associates pathological states in modern medicine that lack early diagnostic indicators with specific signs or symptoms in traditional Chinese medicine, and proposes to use traditional Chinese medicine diagnostic indicators as early non-invasive biomarkers.
[0095] Hypothesis output: "The TCM diagnostic indicator D_TCM may serve as an early, non-invasive biomarker for modern diseases P_Modern." (e) Path 5: New Mechanism Hypothesis Based on Network Topology Generative logic: The system discovers that the modern molecule corresponding to a successfully interpreted traditional Chinese medicine concept occupies a key topological position (such as a hub or bridge) in the overall network, thereby inferring that it may have important regulatory functions that have not been fully recognized.
[0096] Hypothesis output: "Molecular X (a modern interpretation of the TCM concept Y) may play a key regulatory role in the molecular network of disease Z and is a potential new therapeutic target." (f) Pathway Six: Hypothesis of Integrative Pathological Model Based on Cross-Systemic Associations Generative logic: By analyzing the overall structure of the common subgraph, it was found that the specific transmission rules of traditional Chinese medicine and the inter-organ dialogue mechanism of modern medicine are highly consistent in network topology, thus proposing a theoretical framework hypothesis to support the research of this modern mechanism.
[0097] Hypothesis Output: "The traditional Chinese medicine theoretical model T provides a network - based and computable theoretical framework and intervention ideas for studying the modern medical mechanism M (such as organ crosstalk)." Optionally, in the embodiments of the present application, when outputting a newly generated scientific hypothesis, it can be structured. The finally output scientific hypothesis will be presented in a standardized structure to ensure its clarity, operability, and verifiability. Specifically, it includes a hypothesis title, hypothesis content, basis, verification path suggestions, and confidence rating. Hypothesis title: A concise statement. Hypothesis content: Describe its scientific connotation in detail. Basis: List the mapping rules, common sub - picture segments, and topological analysis results on which this hypothesis depends. Verification path suggestions: Propose a preliminary experimental verification plan (such as "the efficacy of prescription F on disease D can be verified through an animal model" or "the correlation between diagnostic index D and pathology P can be verified through a clinical cross - sectional study"). Confidence rating: Based on the confidence of the mapping rules on which it depends, conduct a preliminary rating on the reliability of this hypothesis.
[0098] In the embodiments of the present application, the finally confirmed mapping rule library, the maximum common sub - graph corresponding to the finally confirmed mapping rule library, and the newly generated scientific hypothesis can be output.
[0099] The following combines Figures 2 to 4 to give an exemplary illustration of the technical solutions provided by the embodiments of the present application. Among them, take the alignment of the pathogenesis of "the failure of clear yang to ascend → yin fire" in traditional Chinese medicine and the pathological mechanism of "evaporative heat dissipation disorder" in modern medicine as an example.
[0100] S1: Construct a traditional Chinese medicine pathogenesis network (nodes: spleen - stomach deficiency, insufficient body fluid, yin fire, etc.; edges: causal relationships) by text - mining "Treatise on the Spleen and Stomach". Convert the traditional Chinese medicine pathogenesis network into a standardized transition network. Construct a modern pathological network (nodes: water absorption disorder, insufficient blood volume, etc.; edges: causal relationships) from the literature.
[0101] S2: Initialize mapping rules (such as: "gastrointestinal water absorption disorder" → "water absorption disorder" (strong mapping)).
[0102] S3 - S4: Use the McGregor algorithm to align the networks and extract the MCS. It is found that nodes such as "the failure of clear yang to ascend" cannot be matched.
[0103] S5: Systematically analyze the node of "the failure of clear yang to ascend": It is found that it has a high betweenness centrality and is the key mediator between "gastrointestinal water absorption disorder" and "insufficient water entering the pulse to transform into blood". Based on this, generate candidate hypotheses: 1) "The 'failure of clear yang to ascend' is related to the disorder of skin blood flow control function; 2) "The 'failure of clear yang to ascend' is related to insufficient microcirculation power".
[0104] S6: The expert arbitrators believed that "dysfunction of skin blood flow control" was more appropriate in terms of function, and voted to add it as a weak mapping rule.
[0105] S7: Outputs a refined mapping rule base (with the addition of a weak mapping rule for "Qingyang not rising" → "skin blood flow control dysfunction"), and based on this, generates a new scientific hypothesis: "The prescriptions for treating 'yin fire' in the *Treatise on the Spleen and Stomach* (such as Buzhong Yiqi Tang and Qingshu Yiqi Tang) and their treatment theories can be used to guide the development of drugs that improve skin microcirculation and treat related modern diseases (such as chronic skin ulcers and microcirculation-related diseases)." This hypothesis provides a clear direction for subsequent pharmacological experiments and clinical research.
[0106] This application relates to the interdisciplinary fields of artificial intelligence, computational medicine, and information technology in traditional Chinese medicine (TCM). Specifically, it relates to a method and system that utilizes complex network analysis, graph algorithms, and human-computer collaborative reasoning to identify and analyze mismatches in network alignment, thereby achieving a modern interpretation of TCM theory and the discovery of new knowledge. This transforms mismatches in network alignment from technical obstacles into a driving force for scientific discovery. This application not only verifies the structural similarities between TCM and modern medicine but also automatically generates modern interpretation hypotheses about TCM-specific concepts. Through iterative refinement via a human-computer collaborative closed loop, it ultimately produces experimentally verifiable scientific hypotheses.
[0107] The beneficial effects achievable by this application include the following aspects: 1) Paradigm innovation. It creates a research paradigm of "discovery based on mismatch," deeply integrating computational linguistics and complex network analysis, solving the problem of cross-domain interpretation of abstract concepts. 2) Automation and intelligence. By automatically generating candidate hypotheses through computational topological attributes, it greatly reduces the arbitrariness and workload of manual discovery, making large-scale, systematic knowledge discovery possible. 3) Scientific rigor. Through a human-machine collaborative closed loop of "computational generation - expert arbitration," it ensures both the efficiency of the discovery process and the scientific authority of the final results. 4) Outstanding application value. The final output is not a simple mapping table, but a structured public subnet and verifiable scientific hypotheses, which can directly guide clinical research design, drug repurposing, and diagnostic standard optimization.
[0108] Secondly, this application also provides a knowledge fusion apparatus, comprising: a processing module configured to: determine a set of mismatched nodes in a terminology standardization transition network based on a mapping rule base, wherein the terminology standardization transition network is determined based on a first knowledge domain network, and the set of mismatched nodes includes first nodes in the terminology standardization transition network that do not have a matching second node in the second knowledge domain network; for each first node in the determined set of mismatched nodes, determine at least one candidate interpretation hypothesis, wherein the candidate interpretation hypothesis is related to the potential meaning of the first node in the second knowledge domain; verify the at least one candidate interpretation hypothesis and adjust the mapping rule base according to the verification result; repeatedly execute the steps of determining the set of mismatched nodes, determining at least one candidate interpretation hypothesis, verifying at least one candidate interpretation hypothesis, and adjusting the mapping rule base until a preset condition is met; and output the final mapping rule base.
[0109] Optionally, based on the mapping rule base, the set of mismatched nodes in the terminology standardization transition network is determined, including: based on the mapping rule base, determining the largest common subgraph between the terminology standardization transition network and the second knowledge domain network; and determining the first node in the terminology standardization transition network that is not included in the determined largest common subgraph, so as to determine the set of mismatched nodes.
[0110] Optionally, for each first node in the determined set of mismatched nodes, at least one candidate interpretation hypothesis is determined, including: determining the topological properties of the first node in the terminology standardization transition network; and determining at least one candidate interpretation hypothesis based on the determined topological properties.
[0111] Optionally, the topological properties include at least one of the following: degree centrality, betweenness centrality, structural hole index, and egocentric network structure.
[0112] Optionally, based on the determined topological properties, at least one candidate interpretation hypothesis is determined, including: based on the determined topological properties, determining the structural role of the first node in the terminology standardization transition network; and based on the determined structural role, determining at least one candidate interpretation hypothesis.
[0113] Optionally, based on the determined structural role, at least one candidate interpretive hypothesis is determined, including: determining at least one candidate interpretive hypothesis based on the determined structural role and at least one of the following principles: functional equivalence principle, mediator variable principle, system emergent property principle.
[0114] The specific working principle and benefits of the knowledge fusion device provided in this application are similar to those of the knowledge fusion method provided in this application, and will not be repeated here.
[0115] Thirdly, this application also provides a machine-readable storage medium storing instructions that cause a machine to execute the knowledge fusion method described above.
[0116] Fourthly, this application also provides an electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the executable instructions to implement the knowledge fusion method described above.
[0117] The preferred embodiments of this application have been described in detail above. However, this application is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this application, various simple modifications can be made to the technical solution of this application, and these simple modifications all fall within the protection scope of this application.
[0118] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, this application will not describe the various possible combinations separately.
[0119] Furthermore, various different implementations of this application can be combined in any way, as long as they do not violate the spirit of this application, they should also be regarded as the content disclosed in this application.
Claims
1. A knowledge fusion method, characterized in that, This knowledge fusion method includes: Based on the mapping rule base, a set of mismatched nodes in the terminology standardization transition network is determined, wherein the terminology standardization transition network is determined based on a first knowledge domain network, and the set of mismatched nodes includes a first node in the terminology standardization transition network that does not have a second node to match in the second knowledge domain network. For each first node in the determined set of mismatched nodes, at least one candidate interpretation hypothesis is determined, wherein the candidate interpretation hypothesis is related to the potential meaning of the first node in the second knowledge domain; The at least one candidate interpretation hypothesis is verified, and the mapping rule base is adjusted based on the verification results; Repeat the process of determining the set of mismatched nodes, determining at least one candidate interpretation hypothesis, verifying at least one candidate interpretation hypothesis, and adjusting the mapping rule base until the preset conditions are met. Output the final mapping rule library.
2. The knowledge fusion method according to claim 1, characterized in that, Based on the mapping rule base, the set of mismatched nodes in the terminology standardization transition network is determined, including: Based on the mapping rule base, determine the largest common subgraph between the terminology standardization transition network and the second knowledge domain network; and The first node in the terminology standardization transition network that is not included in the determined largest common subgraph is identified to determine the set of mismatched nodes.
3. The knowledge fusion method according to claim 1, characterized in that, For each first node in the determined set of mismatched nodes, at least one candidate interpretation hypothesis is determined, including: Determine the topological properties of the first node in the terminology standardization transition network; and Based on the determined topological properties, at least one candidate interpretation hypothesis is determined.
4. The knowledge fusion method according to claim 3, characterized in that, The topological properties include at least one of the following: degree centrality, betweenness centrality, structural hole index, and egocentric network structure.
5. The knowledge fusion method according to claim 3, characterized in that, Based on the determined topological properties, the at least one candidate interpretive hypothesis is determined, including: Based on the determined topological properties, the structural role of the first node in the terminology standardization transition network is determined; Based on the determined structural roles, at least one candidate interpretation hypothesis is determined.
6. The knowledge fusion method according to claim 5, characterized in that, Based on the determined structural roles, the at least one candidate interpretive hypothesis is determined, including: The at least one candidate interpretive hypothesis is determined based on the identified structural role and at least one of the following principles: functional equivalence principle, mediator variable principle, and system emergent property principle.
7. A knowledge fusion device, characterized in that, The knowledge fusion device includes: Processing module, used for: Based on the mapping rule base, a set of mismatched nodes in the terminology standardization transition network is determined, wherein the terminology standardization transition network is determined based on a first knowledge domain network, and the set of mismatched nodes includes a first node in the terminology standardization transition network that does not have a second node to match in the second knowledge domain network. For each first node in the determined set of mismatched nodes, at least one candidate interpretation hypothesis is determined, wherein the candidate interpretation hypothesis is related to the potential meaning of the first node in the second knowledge domain; The at least one candidate interpretation hypothesis is verified, and the mapping rule base is adjusted based on the verification results; Repeat the process of determining the set of mismatched nodes, determining at least one candidate interpretation hypothesis, verifying at least one candidate interpretation hypothesis, and adjusting the mapping rule base until the preset conditions are met. Output the final mapping rule library.
8. The knowledge fusion device according to claim 7, characterized in that, Based on the mapping rule base, the set of mismatched nodes in the terminology standardization transition network is determined, including: Based on the mapping rule base, determine the largest common subgraph between the terminology standardization transition network and the second knowledge domain network; and The first node in the terminology standardization transition network that is not included in the determined largest common subgraph is identified to determine the set of mismatched nodes.
9. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores instructions for causing the machine to perform the knowledge fusion method according to any one of claims 1-6.
10. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the knowledge fusion method according to any one of claims 1-6.