Pharmacological analysis system for traditional Chinese medicine prescriptions

By constructing a component synergy network and a disease association pathway discovery system, and combining machine learning, the problems of unclear components and insufficient mechanism explanation in the research of traditional Chinese medicine prescriptions have been solved. Multi-level network association analysis has been achieved, which has improved the depth and efficiency of research and provided comprehensive data fusion and intelligent analysis capabilities.

CN121983249APending Publication Date: 2026-05-05ANTON HEALTH TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANTON HEALTH TECH CO LTD
Filing Date
2025-12-11
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies in the pharmacological research of traditional Chinese medicine prescriptions suffer from problems such as complex chemical composition, unclear effective components, insufficient mechanism explanation, lack of standardized analytical methods, difficulty in achieving multi-level network correlation analysis, and insufficient performance of existing digital platforms in terms of multi-source data fusion and intelligent analysis.

Method used

The system employs a component synergy network construction subsystem, a disease association path discovery subsystem, and a prescription optimization decision generation subsystem. Through multi-omics feature fusion, knowledge graph construction, machine learning, and other technologies, it integrates multi-source data, constructs a component-target dynamic interaction network, discovers potential disease association paths, and generates TCM prescription optimization decisions.

Benefits of technology

It enables systematic analysis of multiple components and targets in traditional Chinese medicine prescriptions, improving the depth and efficiency of research, providing comprehensive multi-source data fusion capabilities and powerful knowledge discovery and visualization capabilities, and supporting the intuitive presentation of complex pharmacological relationships and intelligent question answering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983249A_ABST
    Figure CN121983249A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information retrieval, and provides a pharmacological analysis system for traditional Chinese medicine prescriptions, the system comprises a multi-omics feature fusion subsystem, a multi-omics data set is subjected to feature extraction and dimension reduction operation, and a fused multi-omics feature vector is obtained; the multi-omics feature vector captures the comprehensive influence of the traditional Chinese medicine prescription on genome, proteome and metabolome levels; performing clustering and community detection on a component-target dynamic interaction network of the component collaborative network construction subsystem to obtain a component collaborative sub-network; a prescription-disease association map of the disease association path discovery subsystem is subjected to path mining to obtain a potential disease association path set; a novel prescription-disease association path is obtained; and the traditional Chinese medicine prescription optimization candidate schemes of the prescription optimization decision generation subsystem are subjected to multi-objective optimization to obtain a traditional Chinese medicine prescription optimization decision set. According to the invention, full-chain correlation analysis of chemical components, action targets, biological pathways and disease treatment of traditional Chinese medicine prescriptions is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information retrieval technology, and in particular to a system for pharmacological analysis of traditional Chinese medicine prescriptions. Background Technology

[0002] Traditional Chinese medicine (TCM) formulas, as the core carriers of TCM theory and practice, are characterized by multiple components, multiple targets, and holistic regulation. However, they face numerous technical challenges in modern pharmacological research. Firstly, the chemical composition of TCM formulas is complex, the effective components are unclear, and the material basis of their efficacy is difficult to define. Secondly, the interaction between TCM formulas and the body involves multiple biological levels, from molecules, cells, tissues to organs, and the mechanisms of action are not fully elucidated. Thirdly, traditional TCM formula research relies heavily on the personal experience of physicians, lacking standardized and quantitative analytical methods. While existing digital TCM platforms have made some progress in data integration and simple analysis with the continuous development and advancement of science and technology, they still have significant shortcomings in multi-source data fusion, intelligent analysis model construction, and systematic analysis of mechanisms of action. For example, existing platforms often struggle to achieve a comprehensive correlation analysis from the chemical components of the formula to disease treatment, lacking the ability to deeply mine and visualize the multi-level network of formula-component-target-pathway-disease. Existing technology 1, Publication No. CN120561281A, relates to an intelligent matching system for traditional Chinese medicine prescriptions, including a data acquisition module, a data processing module, a model training module, and a prescription matching module. The data acquisition module collects medicinal material and prescription information from traditional Chinese medicine literature, medical databases, experimental research results, and clinical cases through data interfaces and web crawling technology, storing it in a unified database. The data processing module uses NLP technology to structure the collected unstructured text data, extracting medicinal material names, indications, dosages, pharmacological effects, and contraindications. The model training module, based on machine learning algorithms, classifies and performs association analysis on prescription combinations, generating model parameters for prescription matching. The prescription matching module allows users to search based on various filtering conditions such as purpose, dosage form, mechanism of action, and region. Although matching algorithms and training models are used to select the most suitable prescription combinations from the database, recommending existing, suitable prescription combinations to users through conditional selection and matching algorithms is a system based on historical experience and data retrieval. Its essence is to find known information, without delving into why the prescriptions are effective or what their underlying biological mechanisms are. Existing technology two, publication number CN113486231A, relates to a method, system, device, and storage medium for analyzing traditional Chinese medicine prescriptions and diseases. It combines compound information, drug target information sets, disease target information sets, and a preset target interaction matrix to form an analysis mechanism for targets in the prescription-disease system, obtaining analysis results and ranking them to obtain the target information with the highest relevance. While this improves the accuracy and reliability of network pharmacology research mechanisms, it focuses on relatively simple linear or static network relationships between components, targets, and diseases. Its core analysis is to find common targets and construct network graphs based on them; it often treats prescriptions as a collection of components, lacking in-depth analysis of how components synergistically influence the dynamic processes of the disease network.Existing technology three, publication number CN111241164A, relates to a systemic pharmacology analysis platform and method for traditional Chinese medicine (TCM). It includes a data mining module for mining relevant TCM information from related TCM databases; a data processing module for processing the collected data; a target identification module for identifying common targets between TCM and diseases, and common targets between TCM active ingredients and diseases; a Venn diagram generation module for drawing Venn diagrams of TCM and disease targets, TCM and active ingredients, and TCM active ingredients and disease targets; a network construction module for using Cytoscape to draw a regulatory network diagram between TCM, TCM active ingredients, targets, and diseases; and a database for storing relevant data. While it can analyze TCM active ingredients and target data, obtain the intersection targets between TCM and diseases, construct regulatory networks, and determine the mechanisms of action of TCM active ingredients and disease targets, its main output is the analysis results, such as top-ranked targets and regulatory network diagrams. Its purpose is to explain the possible mechanisms of action of existing prescriptions. Current technologies 1, 2, and 3 suffer from superficial analytical methods, limited functional endpoints, and weak data support. Therefore, this invention provides a pharmacological analysis system for traditional Chinese medicine formulas. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention provides a pharmacological analysis system for traditional Chinese medicine prescriptions, comprising: The component synergy network construction subsystem is used to process multi-omics feature vectors containing structured and unstructured data through network analysis, calculate the correlation strength between chemical components, and construct a component-target dynamic interaction network; the component-target dynamic interaction network is then subjected to clustering and community detection to obtain the component synergy subnetwork; The disease association path discovery subsystem is used to integrate the component synergy subnetwork with the disease knowledge base after processing by the knowledge graph construction module to form a prescription-disease association graph. The prescription-disease association graph is then used for path mining to obtain a set of potential disease association paths. The set of potential disease association paths is then filtered and scored to obtain novel prescription-disease association paths. The formula optimization decision generation subsystem is used to obtain candidate schemes for TCM formula optimization by processing novel formula-disease association paths through machine learning and combining formula optimization functions, through simulating TCM formula compatibility adjustment and effect prediction; the candidate schemes for TCM formula optimization are then optimized through multi-objective optimization to obtain the TCM formula optimization decision set.

[0004] Compared with traditional formula analysis methods, this invention enables systematic analysis of multiple components and targets in traditional Chinese medicine (TCM) formulas. Traditional methods are often limited to single-component or single-target studies, failing to comprehensively reflect the overall regulatory characteristics of the formula. This invention, by integrating network pharmacology, multi-omics analysis, and artificial intelligence algorithms, can systematically analyze the complex interaction networks between multiple active ingredients and multiple biological targets in the formula, better aligning with the holistic view of TCM. It also enhances the depth and efficiency of TCM formula mechanism research: traditional experimental research methods are time-consuming and costly, making high-throughput analysis difficult. This invention, through a combination of computational model prediction and experimental verification, can rapidly screen potential active ingredients, predict targets, and analyze signaling pathways, significantly improving research efficiency and providing important directions for guiding experimental research. Compared with existing digital TCM platforms, this invention offers more comprehensive multi-source data fusion capabilities: existing digital TCM platforms are mostly based on limited data sources, resulting in incomplete knowledge coverage. By integrating formula databases, chemical component databases, target protein databases, disease databases, and multi-omics data, and through unified data standards and quality management, a more comprehensive and accurate TCM knowledge graph is constructed. For example, the system can integrate information on 500 kinds of traditional Chinese medicine and 30,069 compounds from the TCMSP database, as well as 144,410 entities and 3.6 million relationships from the TCMKD platform. A more advanced intelligent analysis model: Employing a multi-level, multi-modal intelligent analysis engine, it integrates various algorithms such as network analysis, machine learning, and deep learning. For example, the system can identify high-frequency combinations of medicinal materials based on the Apriori algorithm, perform association analysis based on the Bron-Kerbosch algorithm, and combine clustering analysis and decision tree analysis tools to quantitatively assess the strength of the association between medicinal materials and diseases. More powerful knowledge discovery and visualization capabilities: This invention not only provides basic data query functions but also supports complex network analysis, multi-dimensional visualization, and intelligent question answering; for example, the system can construct a multi-level association network of medicinal materials, components, targets, and diseases, intuitively presenting complex pharmacological relationships; it also integrates a natural language question answering system, facilitating efficient information retrieval and knowledge acquisition for users.

[0005] Table 1. Comparison of key technical indicators between the present invention and existing digital TCM platforms Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. The technical solutions of the invention will now be described in further detail with reference to the accompanying drawings and embodiments. Attached Figure Description

[0006] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a block diagram of the pharmacological analysis system for traditional Chinese medicine prescriptions in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the pharmacological analysis system for traditional Chinese medicine prescriptions in Embodiment 1 of the present invention; Figure 3 This is a block diagram of the multi-omics feature fusion subsystem in Embodiment 2 of the present invention; Figure 4 This is a block diagram of the component synergy network construction subsystem in Embodiment 6 of the present invention; Figure 5 This is a block diagram of the disease association path discovery subsystem in Embodiment 9 of the present invention; Figure 6 This is a block diagram of the agent optimization decision generation subsystem in Embodiment 15 of the present invention. Detailed Implementation

[0007] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the present invention. The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more associated listed items. When the following description relates to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application. In the description of this application, it should be understood that the terms “first,” “second,” “third,” etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0008] Example 1: As Figure 1 As shown, this embodiment of the invention provides a pharmacological analysis system for traditional Chinese medicine prescriptions, comprising: The multi-omics feature fusion subsystem provides multi-source data, including both structured and unstructured data, which are then integrated and analyzed using multi-omics methods to obtain a multi-omics dataset. This dataset undergoes feature extraction and dimensionality reduction to generate a fused multi-omics feature vector. This feature vector captures the comprehensive impact of traditional Chinese medicine formulas at the genomic, proteomic, and metabolomic levels. The structured data includes the chemical components and target proteins of the formulas, while the unstructured data includes the full text of relevant literature. The component synergy network construction subsystem is used to process multi-omics feature vectors through network analysis, calculate the correlation strength between chemical components, and construct a component-target dynamic interaction network. The component-target dynamic interaction network is then subjected to clustering and community detection to obtain a component synergy subnetwork. The component synergy subnetwork reveals the synergistic relationships between multiple chemical components in traditional Chinese medicine prescriptions and identifies key synergistic groups. The disease association path discovery subsystem is used to integrate the component synergy subnetwork with the disease knowledge base after processing by the knowledge graph construction module to form a formula-disease association graph. The formula-disease association graph is then mined to obtain a set of potential disease association paths. The set of potential disease association paths is filtered and scored to obtain novel formula-disease association paths, which represent the potential mechanism of action of traditional Chinese medicine formulas on specific diseases. The formula optimization decision generation subsystem is used to generate candidate schemes for TCM formula optimization by processing novel formula-disease association paths through machine learning and combining them with formula optimization functions, through simulating TCM formula compatibility adjustments and effect predictions. The candidate schemes for TCM formula optimization are then optimized through multi-objectives to obtain a TCM formula optimization decision set. The TCM formula optimization decision set includes suggestions for formula component adjustments and dosage recommendations, directly supporting visualization and report generation.

[0009] The working principle and beneficial effects of the above technical solution are as follows: The multi-omics feature fusion subsystem of this embodiment is used to provide multi-source data containing structured and unstructured data. After multi-omics integration and analysis, a multi-omics dataset is obtained. The multi-omics dataset is then subjected to feature extraction and dimensionality reduction to obtain a fused multi-omics feature vector. The multi-omics feature vector captures the comprehensive influence of traditional Chinese medicine formulas at the genomic, proteomic, and metabolomic levels. Among them, the structured data includes the chemical components and target proteins of the formula; the unstructured data includes the full text of the literature. The component synergy network construction subsystem is used to process the multi-omics feature vectors through network analysis, calculate the correlation strength between chemical components, and construct a component-target dynamic interaction network. The component-target dynamic interaction network is subjected to clustering and community detection to obtain a component synergy subnetwork. The component synergy subnetwork reveals the synergistic relationship between multiple chemical components in traditional Chinese medicine formulas and labels them. The system identifies key synergistic groups; the disease association path discovery subsystem integrates the component synergistic effect subnetwork with the disease knowledge base after processing by the knowledge graph construction module, forming a formula-disease association graph; the formula-disease association graph undergoes path mining to obtain a set of potential disease association paths; the set of potential disease association paths is filtered and scored to obtain novel formula-disease association paths, representing the potential mechanism of action of traditional Chinese medicine formulas on specific diseases; the formula optimization decision generation subsystem uses novel formula-disease association paths, processed by machine learning, combined with formula optimization functions, to obtain candidate schemes for traditional Chinese medicine formula optimization by simulating the adjustment of traditional Chinese medicine formula compatibility and effect prediction; the candidate schemes for traditional Chinese medicine formula optimization undergo multi-objective optimization to obtain a set of traditional Chinese medicine formula optimization decisions; the set of traditional Chinese medicine formula optimization decisions includes suggestions for formula component adjustment and dosage recommendations, directly supporting visualization and report generation (see appendix for details). Figure 2The above scheme integrates structured chemical components and target information with unstructured full-text literature through a multi-omics feature fusion subsystem, extracts features, and reduces dimensionality to form a feature vector that reflects the comprehensive influence of the prescription at the genomic, proteomic, and metabolomic levels. Subsequently, a component synergy network construction subsystem performs network analysis on this feature vector, calculates the correlation strength between components, and constructs a dynamic interaction network. Through clustering and community detection, key synergistic subnetworks are extracted to reveal the synergistic action patterns within the prescription. Next, a disease association path discovery subsystem integrates the synergistic subnetworks with a disease knowledge base to construct a prescription-disease association map. Through path mining, filtering, and scoring, potential prescription-disease association paths are obtained, clarifying the possible mechanisms of action of the prescription on specific diseases. Finally, a prescription optimization decision generation subsystem uses machine learning to simulate compatibility and predict effects on these association paths. Combined with multi-objective optimization, an optimization decision set for component adjustment and dosage recommendation is generated to support the generation of a visual report. This embodiment achieves a closed-loop platform for the pharmacological analysis and decision-making of traditional Chinese medicine (TCM) formulas, encompassing unified characterization of multi-source omics data, precise delineation of collaborative networks, systematic discovery of disease association mechanisms, and intelligent optimization of formula compatibility. This embodiment provides a systematic and intelligent TCM formula pharmacological analysis system; by integrating multi-source heterogeneous data, constructing knowledge graphs, and introducing multimodal intelligent analysis models, it achieves full-chain correlation analysis of the chemical components, targets, biological pathways, and disease treatment of TCM formulas, providing core technical support for revealing the mechanisms of action of TCM formulas, optimizing compatibility rules, discovering new indications, and developing new TCM drugs.

[0010] In this embodiment, the various subsystems adopt a layered architecture design, including a data layer, analysis layer, application layer, and presentation layer. The layers communicate with each other through standard interfaces to ensure the system's flexibility, scalability, and maintainability. The data layer is responsible for the collection, storage, and management of multi-source data. It uses a distributed database system to store structured data, such as basic information about prescriptions, chemical components, and target proteins, as well as unstructured data, such as full-text literature and experimental data. The data layer cleans, standardizes, and integrates the raw data through extraction, transformation, and loading processes, providing a high-quality data foundation for upper-level analysis. The analysis layer is the core of the system, including a knowledge graph construction module, a network analysis module, a machine learning module, and a multi-omics integrated analysis module. These modules work collaboratively to transform basic data into deep knowledge. The application layer, based on the results of the analysis layer, provides specific application functions such as prescription mechanism analysis, compatibility rule mining, new indication discovery, and prescription optimization. The presentation layer provides a web-based interactive interface, including data query, result visualization, and analysis report generation functions, supporting multi-terminal access. The system's workflow is as follows: Users submit analysis tasks through the presentation layer, such as analyzing the mechanism of action of traditional Chinese medicine prescriptions; the system obtains relevant data from the data layer and performs processing such as knowledge graph query, network analysis, and machine learning model calculation in the analysis layer; the analysis results are transformed into specific application function outputs in the application layer; and the final results are displayed to users in an interactive and visual form through the presentation layer.

[0011] Example 2: Figure 3 As shown, based on Embodiment 1, the multi-omics feature fusion subsystem provided in this embodiment of the invention includes: The heterogeneous data semantic association component is used to process structured and unstructured data from multiple sources through deep semantic parsing based on concept entity recognition, resulting in a set of concept entities containing chemical components, biological pathways, and pharmacological effects. The concept entity set is then used to extract contextual semantic relationships to construct a heterogeneous semantic network across data types. The heterogeneous semantic network establishes deep semantic associations between concepts in data from different sources and of different types. The synergy effect quantification mapping component is used to obtain a quantified correlation matrix between concept entities by calculating the correlation strength based on the network topology of the heterogeneous semantic network obtained from the heterogeneous data semantic association component. The quantified correlation matrix is ​​weighted and calibrated by introducing multi-omics prior knowledge to form a synergy effect matrix with enhanced omics context. The synergy effect matrix quantifies the potential synergy strength between chemical components and maps it onto a specific genomic and proteomic background. The multidimensional feature fusion and dimensionality reduction component is used to align and concatenate the synergistic effect matrix with the features of the original multi-omics data to generate a high-dimensional fusion feature set. The high-dimensional fusion feature set is then processed by feature dimensionality reduction based on nonlinear manifold learning to obtain a low-dimensional multi-omics feature vector that can characterize the overall influence pattern of the prescription at multiple biological levels. The multi-omics feature vector captures the comprehensive biological features enhanced by semantic association and synergistic effect.

[0012] The working principle and beneficial effects of the above technical solution are as follows: The heterogeneous data semantic association component in this embodiment is used to process structured and unstructured data from multiple sources through deep semantic parsing based on concept entity recognition, resulting in a set of concept entities containing chemical components, biological pathways, and pharmacological effects; the concept entity set is then used to extract contextual semantic relationships to construct a heterogeneous semantic network across data types; the heterogeneous semantic network establishes deep semantic associations between concepts in data from different sources and of different types; the synergistic effect quantification mapping component is used to calculate the quantified association moments between concept entities through association strength calculation based on network topology of the heterogeneous semantic network obtained by the heterogeneous data semantic association component. The matrix quantifies the correlation degree matrix by introducing multi-omics prior knowledge for weight calibration, forming an omics context-enhanced synergistic effect matrix. The synergistic effect matrix quantifies the potential synergistic strength between chemical components and maps it onto a specific genomic and proteomic background. The multi-dimensional feature fusion and dimensionality reduction component is used to align and concatenate the synergistic effect matrix with the features of the original multi-omics data to generate a high-dimensional fusion feature set. The high-dimensional fusion feature set is processed by feature dimensionality reduction based on nonlinear manifold learning to obtain a low-dimensional multi-omics feature vector that can characterize the overall influence pattern of the prescription at multiple biological levels. The multi-omics feature vector captures the comprehensive biological features enhanced by semantic association and synergistic effect. The above scheme extracts conceptual entities such as chemical components, pathways, and pharmacological effects from structured and unstructured data through deep semantic parsing, and constructs a heterogeneous semantic network across data types to achieve deep associations between information from different sources. Subsequently, the association strength between conceptual entities is calculated based on the network topology, and weight calibration is performed using multi-omics prior knowledge to obtain a synergistic effect matrix enhanced in the context of genomics and proteomics, quantifying the potential synergistic strength of chemical components. Finally, this synergistic effect matrix is ​​aligned and concatenated with the original multi-omics features to form a high-dimensional fusion feature set. Dimensionality reduction is then achieved through nonlinear manifold learning to generate a low-dimensional feature vector that comprehensively reflects the multi-level biological effects of the formula at the gene, protein, and metabolic levels. This embodiment achieves semantic unification of heterogeneous data, quantitative enhancement of synergistic effects, and effective dimensionality reduction of high-dimensional features, forming a precise and comprehensive multi-omics feature representation.

[0013] In this embodiment, the data includes: Traditional Chinese Medicine (TCM) prescriptions: Prescriptions from classic ancient texts such as *Shang Han Lun* and *Jin Gui Yao Lue*, as well as prescriptions recorded in modern Chinese Pharmacopoes and textbooks. This includes information such as prescription name, source, constituent medicinal materials, dosage, efficacy, and indications. For example, the system can integrate 48,644 prescriptions from the TCMKD platform. TCM and chemical component data: Basic information on TCMs (nature, flavor, meridian tropism, efficacy, etc.) and their chemical components are collected. This includes the name, structure, physicochemical properties, and pharmacokinetic parameters (such as oral bioavailability, blood-brain barrier permeability, etc.). For example, information on 500 TCMs and 30,069 compounds from the TCMSP database can be integrated. Target and pathway data: Protein information (gene name, amino acid sequence, function, etc.), drug-target interaction data, and biological pathway data (such as the KEGG pathway) are integrated. Disease data: Information on disease names, classifications, symptoms, and related genes is collected. For example, data on 23,750 diseases from the TCMKD platform can be integrated. Multi-omics data: Integrating transcriptomics, proteomics, metabolomics and other data to reflect state changes at different biological levels.

[0014] Example 3: Based on Example 2, the synergistic effect quantification mapping component provided in this embodiment of the invention includes: The topological potential field calculation sub-component is used to calculate the topological potential field distribution of each concept entity in the heterogeneous semantic network based on its topological location after global influence calculation. The topological potential field distribution is then used to calculate the potential gradient between entities to form a preliminary global association strength matrix between entities. The global association strength matrix reflects the intrinsic association between concept entities in the entire semantic network topology. The prior knowledge injection sub-component is used to identify association patterns in the global association strength matrix that match known biological pathways by performing pattern matching with an external multi-omics prior knowledge base. The identified association patterns are then subjected to a constraint-optimized weight allocation process to calibrate the corresponding elements in the global association strength matrix, generating a calibrated association network. The collaborative context mapping sub-component is used to extract and aggregate the association subgraphs between chemical component entities in the association network to obtain a chemical component-specific association matrix. The chemical component-specific association matrix is ​​then mapped to the genomic variation background and proteome interaction background to form an omics context-enhanced synergistic effect matrix. The synergistic effect matrix quantifies the potential synergistic effect strength of chemical components in a specific molecular biological context.

[0015] The working principle and beneficial effects of the above technical solution are as follows: The topological potential field calculation subcomponent of this embodiment is used to calculate the topological potential field distribution of each concept entity in the heterogeneous semantic network based on its topological position through global influence calculation; the topological potential field distribution is used to calculate the potential gradient between entities to form a preliminary global association strength matrix between entities; the global association strength matrix reflects the intrinsic association of concept entities in the entire semantic network topology; the prior knowledge injection subcomponent is used to perform pattern matching between the global association strength matrix between entities and an external multi-omics prior knowledge base to identify association patterns in the global association strength matrix that are consistent with known biological pathways; the identified association patterns are used to perform a weight allocation process based on constraint optimization to calibrate the corresponding elements in the global association strength matrix and generate a calibrated association network; the cooperative context mapping subcomponent is used to extract and aggregate the association network for chemical component entity association subgraphs to obtain a chemical component-specific association matrix; the chemical component-specific association matrix is ​​used to map its elements to the genomic variation background and proteome interaction background to form an omics context-enhanced synergistic effect matrix; the synergistic effect matrix quantifies the potential synergistic effect strength of chemical components in a specific molecular biological background. The above scheme obtains the positional potential energy distribution of conceptual entities in the semantic network through global topological potential field calculation, and generates a global association strength matrix between entities based on the potential energy gradient to capture the intrinsic associations in the overall network structure. Then, the matrix is ​​matched with a multi-omics prior knowledge base to identify association patterns that conform to known biological pathways, and the corresponding elements are weighted and calibrated through constraint optimization to obtain a calibrated association network. Finally, the association subgraphs of chemical components are extracted and aggregated, and mapped onto the background of genomic variation and protein interaction to form a synergistic effect matrix in a specific molecular biological environment, thereby achieving precise quantification of the potential synergistic effect strength of chemical components.

[0016] Example 4: Based on Example 3, the topological potential field calculation sub-component provided in this embodiment of the invention includes: The potential energy propagation range definition module is used to obtain the influence attenuation distribution of each concept entity on other entities after the heterogeneous semantic network is processed by the attenuation coefficient based on the shortest path distance between network nodes. The influence attenuation distribution is then fused with the entity's own attribute weights to form an independent potential energy field for each concept entity. The independent potential energy field defines the original influence range and intensity of a single entity in the network. The potential field gradient field construction module is used to construct the independent potential fields of all conceptual entities obtained by the potential energy propagation range definition module. After vector superposition processing based on field theory principles, the synthetic potential field distribution of the entire heterogeneous semantic network is obtained. The synthetic potential field distribution is processed by the potential energy difference between any two adjacent entities to construct the network potential field gradient field. The network potential field gradient field describes the rate and direction of potential energy change of conceptual entities in the network topology space. The association strength tensor generation module is used to generate the network potential gradient field obtained from the potential gradient field construction module. After integrating the potential gradient paths between non-adjacent entities, the equivalent association gradient between any pair of entities is obtained. The equivalent association gradient is normalized and symmetricized to form a preliminary global association strength matrix between entities. The global association strength matrix encapsulates the intrinsic association between entities derived from the potential gradient in tensor form.

[0017] The working principle and beneficial effects of the above technical solution are as follows: The potential energy propagation range definition module in this embodiment is used to process the heterogeneous semantic network using an attenuation coefficient based on the shortest path distance between network nodes, obtaining the attenuation distribution of the influence of each concept entity on the remaining entities; the influence attenuation distribution is then fused with the entity's own attribute weights to form an independent potential energy field for each concept entity. This independent potential energy field defines the original influence range and intensity of a single entity in the network; the potential field gradient field construction module is used to process the independent potential energy fields of all concept entities obtained by the potential energy propagation range definition module, using vector superposition based on field theory principles, to obtain the composite potential of the entire heterogeneous semantic network. Field distribution; the synthetic potential field distribution, after processing the potential energy difference between any two adjacent entities, constructs the network potential field gradient field; the network potential field gradient field describes the rate and direction of potential energy change of conceptual entities in the network topology space; the association strength tensor generation module is used to generate the network potential field gradient field obtained from the potential field gradient field construction module, and after integrating the potential field gradient paths between non-adjacent entities, obtains the equivalent association gradient between any pair of entities; the equivalent association gradient is normalized and symmetricized to form a preliminary global association strength matrix between entities; the global association strength matrix encapsulates the intrinsic association between entities derived from the potential field gradient in tensor form. The above scheme limits the influence range of each conceptual entity in the network by using the shortest path attenuation coefficient and forms an independent potential energy field by combining entity attribute weights. Then, all independent potential energy fields are vector-superimposed to obtain the synthetic potential energy field of the whole network, and the potential energy difference between adjacent entities is calculated to construct the potential energy gradient field, which describes the rate and direction of change of potential energy in the topological space. Finally, the potential energy gradient paths between non-adjacent entities are integrated, normalized and symmetricized to generate a global correlation strength matrix encapsulated in tensor form, realizing a quantitative expression of the intrinsic correlation between any pair of entities in the network.

[0018] Example 5: Based on Example 3, the collaborative context mapping sub-component provided in this embodiment of the invention includes: The genomic variation background sensitivity module is used to process the chemical component-specific association matrix through pattern matching with external genomic variation background data to obtain the genomic variation influence weight vector; the genomic variation influence weight vector is then subjected to context adjustment of the elements of the chemical component-specific association matrix to obtain the genomic variation background sensitivity association matrix; the association matrix reflects the adjustment strength of chemical component associations under specific genomic variation backgrounds. The protein interaction background integration module is used to align the correlation matrix sensitive to genomic variation background with external proteomics interaction background data to obtain protein interaction network constraints. The protein interaction network constraints are then topologically optimized on the correlation matrix sensitive to genomic variation background to obtain a protein interaction background integrated correlation matrix. The protein interaction background integrated correlation matrix contains the structural modulation of the protein interaction network on the association of chemical components. The multi-background synergistic effect fusion module is used to process the correlation matrix of protein interaction background integration by multi-background weight fusion to obtain the multi-background fusion weight matrix. The multi-background fusion weight matrix is ​​then used to finally quantify and integrate the correlation matrix of protein interaction background integration to form the omics context-enhanced synergistic effect matrix. The synergistic effect matrix quantifies the potential synergistic effect strength of chemical components in the dual background of genomic variation and proteomics interaction.

[0019] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, the genomic variation background sensitivity module is used to obtain a genomic variation influence weight vector by pattern matching the chemical component-specific association matrix with external genomic variation background data; the genomic variation influence weight vector is then used to obtain a genomic variation background sensitivity association matrix by contextual adjustment of the elements of the chemical component-specific association matrix; the association matrix reflects the adjustment strength of chemical component associations under a specific genomic variation background; the protein interaction background integration module is used to obtain protein interaction network constraints by structurally aligning the genomic variation background sensitivity association matrix with external proteomics interaction background data; protein interactions… The network constraints are optimized by topology of the association matrix sensitive to genomic variation background to obtain an association matrix integrated with the protein interaction background. The association matrix integrated with the protein interaction background contains the structural modulation of chemical component associations by the protein interaction network. The multi-background synergistic effect fusion module is used to process the synergistic strength of the association matrix integrated with the protein interaction background based on multi-background weight fusion to obtain a multi-background fusion weight matrix. The multi-background fusion weight matrix is ​​then used to finally quantify and integrate the association matrix integrated with the protein interaction background to form a synergistic effect matrix enhanced with omics context. The synergistic effect matrix quantifies the potential synergistic effect strength of chemical components in the dual background of genomic variation and proteomics interaction. The above scheme generates a weight vector reflecting the influence of gene variation by matching the chemical component-specific association matrix with genomic variation data, and then performs context-adjustment on the association matrix accordingly to obtain a sensitive association matrix under a specific genomic background. Subsequently, this matrix is ​​aligned with a protein interaction network, and the matrix is ​​topologically optimized using network constraints to form an association matrix that integrates protein interaction modulation. Finally, the association matrices under multiple backgrounds are weighted and fused to generate a multi-background fused weight matrix, which is then finally quantified and integrated to obtain a synergistic effect matrix under the dual backgrounds of genomic variation and protein interaction, thus achieving accurate and context-aware quantification of the potential synergistic effect strength of chemical components.

[0020] Example 6: As Figure 4 As shown, based on Example 1, the component synergy network construction subsystem provided in this embodiment of the invention includes: The dynamic interaction relationship inference component is used to obtain the conditional association probability distribution between chemical components and target proteins by performing conditional dependency analysis based on multi-omics context on multi-omics feature vectors; the conditional association probability distribution is then filtered for significance based on statistical potential to construct the component-target initial association framework; the component-target initial association framework defines the potential interaction relationships between chemical components and targets in multiple biological contexts. A network topology generation component is used to generate the initial component-target association skeleton obtained from the dynamic interaction relationship inference component. After introducing the temporal or dose response patterns contained in the multi-omics feature vectors, the connection strength is assigned to form a weighted component-target dynamic interaction network. The edge weights in the weighted component-target dynamic interaction network quantify the strength and dynamic characteristics of the interaction relationship. The functional community discovery component is used to generate a weighted component-target dynamic interaction network from the network topology generation component. After functional module boundary detection based on network flow betweenness centrality, potential functional communities are obtained in the network. The potential functional communities are then optimized and screened based on module cohesion and inter-module segregation to form component synergy sub-networks. The component synergy sub-networks reveal chemical component synergistic groups with specific biological functions as the core and tight internal connections.

[0021] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, the dynamic interaction relationship inference component is used to obtain the conditional association probability distribution between chemical components and target proteins by performing conditional dependency analysis based on multi-omics context on multi-omics feature vectors; the conditional association probability distribution is then filtered for significance based on statistical potential to construct an initial component-target association framework; the initial component-target association framework defines the potential interaction relationships between chemical components and targets in multiple biological contexts; the network topology generation component is used to further refine the initial component-target association framework obtained by the dynamic interaction relationship inference component by introducing temporal or... The dose-response model assigns connection strength values ​​to form a weighted component-target dynamic interaction network. The edge weights in the weighted component-target dynamic interaction network quantify the strength and dynamic characteristics of the interaction relationships. The functional community discovery component is used to generate the weighted component-target dynamic interaction network from the network topology component. After functional module boundary detection based on network flow betweenness centrality, the potential functional communities in the network are divided. The potential functional communities are optimized and screened based on module cohesion and inter-module segregation to form component synergy sub-networks. The component synergy sub-networks reveal chemical component synergistic groups with specific biological functions as the core and tight internal connections. The above scheme obtains the conditional association probability distribution between chemical components and target proteins through multi-omics contextual conditional dependency analysis, and uses statistical potential for significance filtering to construct an initial component-target association framework reflecting multiple biological backgrounds. Subsequently, time-series or dose-response patterns are combined to assign connection strengths to this framework, forming a weighted component-target dynamic interaction network, with edge weights quantifying the strength and dynamic characteristics of the interaction. Finally, based on the network flow betweenness centrality, the boundaries of functional modules are detected, and the cohesion and segregation between modules are optimized to delineate functional communities and screen out component synergistic sub-networks, revealing chemical component synergistic groups with specific biological functions as the core and tight internal connections. This embodiment realizes the layer-by-layer abstraction and quantification from conditional association probability to weighted dynamic networks and then to functional communities, forming a network model that can accurately characterize the internal synergistic mechanism of the prescription.

[0022] This embodiment showcases the system's core analytical capabilities, integrating multiple computational models and algorithms for mining deep-level knowledge from knowledge graphs. Network pharmacology analysis includes: Formula-component-target network construction: Based on the knowledge graph, a multi-layered formula-component-target network is automatically generated for specific formulas. Network topology analysis, such as degree centrality and betweenness centrality, identifies key nodes, core components, and key targets within the network. Pathway enrichment analysis: The targets of formula action are mapped to biological pathway databases such as KEGG and GO. Statistical methods, such as hypergeometric tests, are used to identify significantly enriched pathways, revealing the biological pathways through which the formula may act. For example, the TCMKD platform provides a functional enrichment analysis module, supporting enrichment analysis based on GO annotation and KEGG pathways. Modular analysis: Community detection algorithms, such as the Louvain algorithm, are used to decompose complex networks into functional modules, identifying the functional division of different drug groups within the formula and elucidating the scientific implications of formula compatibility. Association Rule Mining: Utilizing association rule mining algorithms such as Apriori and FP-Growth, this approach identifies frequently occurring combinations of medicinal materials in prescriptions, revealing commonly used drug pairs and groups. For example, the TCMKD platform uses the Apriori algorithm to identify high-frequency combinations of medicinal materials; it also employs the Bron-Kerbosch algorithm for maximal clique mining to discover closely related sets of medicinal materials within prescriptions, analyzing the correlation between core drug combinations and efficacy. Machine Learning and Deep Learning Models; Efficacy Prediction Models: Based on known prescription components, targets, and efficacy data, machine learning models, such as support vector machines, random forests, and graph neural networks, are trained to predict the potential efficacy of new prescriptions or components; Target Prediction Models: Deep learning models, such as multilayer perceptrons and graph convolutional networks, are used to predict the interaction between traditional Chinese medicine components and target proteins. For example, the SYSTCM platform constructs grayscale maps to represent the relationship between components and pharmacological effects and trains an identification model; the IPE model identifies the pharmacological effects of traditional Chinese medicine and its components; Prescription Optimization Models: Combining reinforcement learning or evolutionary algorithms, this approach optimizes the composition or dosage of prescriptions with the goal of maximizing efficacy and minimizing side effects. Multi-omics integrated analysis: Integrating transcriptomic, proteomic, metabolomic and other multi-omics data to analyze changes at different molecular levels in the body after herbal intervention from a systemic perspective, and to more comprehensively reveal the mechanism of action of the herbal intervention; using pathway analysis, network analysis and other methods to identify key biomarkers and affected core biological processes under the intervention of herbal intervention.

[0023] Example 7: Based on Example 6, the functional community discovery component provided in this embodiment of the invention includes: The network flow bottleneck identification sub-component is used to simulate the weighted component-target dynamic interaction network based on the principle of maximum flow and minimum cut to obtain the flow medium values ​​of all edges and nodes in the weighted component-target dynamic interaction network. After detecting and filtering local maxima, the flow medium values ​​form a set of key boundary nodes. The set of key boundary nodes identifies the key bottlenecks and potential community boundaries of information flow in the network. The community skeleton generation sub-component is used to perform an iterative boundary removal operation from the original network to obtain a series of mutually separated connected components from the set of key boundary nodes; the connected components are then evaluated based on the connectivity density within the components to form a set of candidate communities; the set of candidate communities constitutes the preliminary skeleton of potential functional communities. The community efficiency optimization sub-component is used to calculate the cohesion of the candidate community set based on the aggregation strength of all edge weights within the module, and the separation based on the sum of cross-community connection weights, to obtain the synergistic efficiency index of each candidate community. The synergistic efficiency index is then optimized through threshold screening and local merging to form a component synergistic effect sub-network. The component synergistic effect sub-network is a group of chemical component functional groups with high cohesion and good separation.

[0024] The working principle and beneficial effects of the above technical solution are as follows: The network flow bottleneck identification sub-component of this embodiment is used to simulate the weighted component-target dynamic interaction network based on the principle of maximum flow and minimum cut to obtain the flow medium values ​​of all edges and nodes in the weighted component-target dynamic interaction network; the flow medium values ​​are used to detect and filter local maxima to form a set of key boundary nodes; the set of key boundary nodes identifies the key bottlenecks and potential community boundaries of information flow in the network; the community skeleton generation sub-component is used to perform iterative boundary removal operations from the original network to obtain a series of mutually separated connected components; the connected components are used to form a set of candidate communities based on the preliminary evaluation of the connectivity density within the components; the set of candidate communities constitutes the preliminary skeleton of potential functional communities; the community efficiency optimization sub-component is used to calculate the cohesion of the candidate community set based on the aggregation strength of all edge weights within the module and the separation based on the sum of cross-community connection weights to obtain the collaborative efficiency index of each candidate community; the collaborative efficiency index is used to form a component synergy sub-network through threshold screening and local merging optimization, which is a chemical component functional group with high cohesion and good separation. The above scheme obtains the flow medium value of each edge and each node in the network through maximum flow-minimum cut simulation, and selects local maxima to form a set of key boundary nodes, which identify bottlenecks and potential community boundaries of information flow. Then, these key boundary nodes are iteratively removed from the original network to obtain several mutually separated connected components. Based on the internal connectivity density of the components, a set of candidate communities is formed to construct the preliminary skeleton of the functional community. Finally, the internal edge weight aggregation strength of the candidate community is calculated to evaluate cohesion, and the sum of cross-community connection weights is calculated to evaluate separation. Based on the synergistic efficiency index, threshold screening and local merging optimization are performed to obtain a component synergistic subnetwork with high cohesion and good separation, thereby achieving accurate division and efficiency improvement of chemical component functional groups.

[0025] Example 8: Based on Example 7, the community performance optimization sub-component provided in this embodiment of the invention includes: The cohesion strength coefficient generation module is used to evaluate the cohesion of candidate communities based on the product of the kurtosis of the edge weight distribution and the sum of the edge weights within the module, and to obtain the original cohesion strength of each community. The original cohesion strength is then standardized based on the average path length between nodes within the community to form the cohesion strength coefficient. The cohesion strength coefficient characterizes the concentration and tightness of the connections within the community, and its calculation is derived from the statistical concept of kurtosis and the average path length in network topology. The separation resistance coefficient generation module is used to evaluate the separation of candidate communities based on the ratio of the sum of cross-community connection weights to the variance of inter-community connection weights, and obtain the original separation resistance of each community. The original separation resistance is normalized based on the global network inter-community connection weights to form the separation resistance coefficient. The separation resistance coefficient characterizes the sparsity and interference of inter-community connections, and its calculation comes from the connection weight variance and global weight normalization in network analysis. The collaborative efficiency index calculation module is used to calculate the collaborative efficiency index of each candidate community by fusing the cohesion strength coefficient and the separation resistance coefficient based on the hyperbolic tangent function and the weight balance parameter. The collaborative efficiency index is used to quantify the efficiency of the community as a collaborative functional unit. Its calculation comes from the hyperbolic tangent function in mathematics to smooth the fusion of the two coefficients, and the weight balance parameter comes from the trade-off setting in optimization theory.

[0026] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, the cohesion strength coefficient generation module is used to evaluate the cohesion of the candidate community set based on the product of the kurtosis of the edge weight distribution and the sum of the edge weights within the module, obtaining the original cohesion strength of each community; the original cohesion strength is then standardized based on the average path length between nodes within the community to form a cohesion strength coefficient; the cohesion strength coefficient characterizes the concentration and tightness of connections within the community, and its calculation originates from the statistical concept of kurtosis and the average path length in network topology; the separation resistance coefficient generation module is used to evaluate the separation of the candidate community set based on the ratio of the sum of cross-community connection weights to the variance of inter-community connection weights, obtaining the original cohesion strength of each community. The original separation resistance is normalized based on the connection weights between global network communities to form a separation resistance coefficient. This coefficient characterizes the sparsity and interference of connections between communities, and its calculation originates from the connection weight variance and global weight normalization in network analysis. A collaborative efficiency index calculation module is used to calculate the collaborative efficiency index for each candidate community by fusing the cohesion strength coefficient and the separation resistance coefficient based on the hyperbolic tangent function and weight balance parameters. The collaborative efficiency index quantifies the effectiveness of a community as a collaborative functional unit; its calculation originates from the hyperbolic tangent function used in mathematics to smooth the fusion of the two coefficients, and the weight balance parameters are derived from the trade-off settings in optimization theory. The above scheme evaluates the product of the kurtosis and sum of the edge weight distribution within the candidate community to obtain the original cohesion strength. Then, it is standardized using the average path length of nodes within the community to generate a cohesion strength coefficient, quantifying the concentration and tightness of connections within the community. Subsequently, the original separation resistance is evaluated based on the ratio of the sum of cross-community connection weights to the weight variance, and the separation resistance coefficient is obtained by normalizing the global community connection weights, characterizing the sparsity and interference of connections between communities. Finally, the cohesion strength coefficient and separation resistance coefficient are fused using the hyperbolic tangent function and combined with the weight balance parameter to calculate the synergistic effectiveness index, thereby achieving a smooth and comprehensive quantitative evaluation of the effectiveness of each candidate community as a synergistic functional unit.

[0027] Used to measure cohesive strength coefficient With separation drag coefficient Integration into a synergistic effectiveness index The mathematical expression is as follows:

[0028] in: Indicates the first The collaborative effectiveness index of each candidate community has a range of (-1, 1). The closer the value is to 1, the higher the overall effectiveness of the community as a collaborative functional unit. The hyperbolic tangent function is a smooth, monotonically increasing sigmoid function that maps a linear weighted sum to a bounded, standardized interval. This mapping ensures that extreme values ​​do not cause exponential explosion or disappearance, thereby enhancing the stability and comparability of the values. Indicates the first The cohesion strength coefficient of each candidate community is evaluated by the product of the kurtosis and sum of the edge weight distributions within the community, and then standardized by the average path length. This coefficient is used to quantify the tightness and structure of the connections within the community. Indicates the first The separation resistance coefficient of each candidate community is evaluated by the ratio of the sum of the connection weights between communities to the variance, and then normalized by the global weights. This coefficient is used to quantify the sparsity and anti-interference ability of the community's connection to the external network. and These represent the weighting balance parameters assigned to the cohesive strength coefficient and the separation resistance coefficient, respectively. + =1); the two parameters originate from the trade-off concept in optimization theory, used to adjust the relative emphasis of the system on the two dimensions of "internal tightness" and "external independence". For example, increasing This will make the analysis focus more on the collaboration within the module, and increase... This emphasizes the functional independence between modules. The core principle is based on linear weighted fusion and nonlinear standardization. First, two coefficients that characterize community features from different dimensions ( and A preliminary comprehensive evaluation value is formed by linear fusion through weighted summation. * + * This approach embodies the fundamental idea of ​​comprehensively considering both internal cohesion and external separation, with weight parameters allowing for flexible adjustment based on specific analytical objectives. Subsequently, a nonlinear transformation is applied to this linear combination using the hyperbolic tangent function tanh. This transformation plays three key roles: standardization and boundedness: it smoothly compresses the input values ​​into a fixed interval of (-1, 1), ensuring that the final exponent is constrained within this range regardless of the original weighted sum, facilitating unified comparison and threshold selection across different communities. Enhanced discriminability: near 0, the tanh function is approximately linear, effectively reflecting changes in the weighted sum; at both ends, its growth gradually saturates, moderately suppressing the dominance of individual communities with extremely high weighted sums, resulting in a more balanced distribution of outcomes.

[0029] Example 9: As Figure 5 As shown, based on Example 1, the disease association path discovery subsystem provided in this embodiment of the invention includes: The heterogeneous knowledge network collaborative enhancement component is used to process the component synergy subnetwork through the knowledge graph construction module, aligning it with entities and mapping relationships in the disease knowledge base to form a preliminary integrated heterogeneous graph. The heterogeneous graph, by introducing relationship completion and collaborative embedding learning based on graph neural networks, models and completes the potential high-order relationships between component-target synergy groups and target-disease associations, generating a collaboratively enhanced knowledge network. The collaboratively enhanced knowledge network not only contains explicit associations but also encodes implicit, structured association features between component synergy groups and diseases. The causal mechanism path reasoning component is used to process the synergistically enhanced knowledge network through a hybrid symbolic and numerical reasoning engine. Combining rule exploration through inductive logic programming with pattern recognition capabilities of graph attention networks, it traverses and generates candidate action path sequences that connect specific formula synergistic groups with disease nodes in the knowledge network. The candidate action path sequences are filtered by biological rationality constraints based on prior knowledge of biological pathways to obtain a series of formula-disease candidate mechanism paths with potential biological significance. A multi-dimensional evidence aggregation scoring and screening component is used to process candidate herbal formula-disease mechanisms. The components are quantitatively evaluated and weighted across four dimensions: pathway tightness, significant enrichment of genes involved in the pathway in relevant disease pathways, perturbation correlation of omics data at key nodes of the pathway, and literature co-occurrence support. The comprehensive confidence score of each pathway is calculated. All pathways are ranked according to their comprehensive confidence scores and screened based on dynamic thresholds. High-confidence novel herbal formula-disease association pathways are output, representing potential novel mechanisms of action supported by multi-dimensional evidence.

[0030] The working principle and beneficial effects of the above technical solution are as follows: This embodiment aligns the component synergistic subnetwork with the disease knowledge base and constructs a heterogeneous graph. Then, it utilizes graph neural networks for relation completion and synergistic embedding learning to generate a synergistically enhanced knowledge network that includes both explicit associations and implicit higher-order associations. On this network, symbolic reasoning and numerical reasoning, inductive logic programming rules, and graph attention pattern recognition are combined to traverse and generate candidate action paths from formula synergistic groups to disease nodes. Based on prior knowledge of biological pathways, rationality filtering is performed, retaining mechanism paths with potential biological significance. Finally, these paths are quantitatively scored and weighted across four dimensions: path density, gene pathway enrichment significance, correlation of perturbations at key omics nodes, and literature co-occurrence. A comprehensive confidence score is calculated, and novel formula-disease association paths with high confidence are selected and output based on dynamic thresholds. This embodiment achieves a comprehensive evaluation from the synergistic enhancement of heterogeneous knowledge networks and multimodal reasoning of causal mechanisms to multidimensional evidence, forming a reliable and structured potential mechanism of action between formulas and diseases.

[0031] In this embodiment, the knowledge graph is constructed as follows: Ontology construction: Defining the semantic model of the formula knowledge graph, including core concepts such as formula, traditional Chinese medicine, ingredients, targets, diseases, pathways, and their relationships, such as composition, containment, action on, treatment, and participation. Entity recognition and relation extraction: A combination of rule-based and deep learning methods is used to extract entities and relations from structured data and unstructured text. For structured data, it is directly converted into RDF triples through mapping rules. For unstructured text, such as literature abstracts, a BERT-BiLSTM-CRF model finely tuned on traditional Chinese medicine corpus is used for named entity recognition, and then a relation extraction model, such as the BERT classification model, is used to identify the relations between entities. Knowledge fusion and storage: Solving the conflict and inconsistency problems between multi-source data to form a unified and consistent knowledge graph. A graph database, such as Neo4j, is used to store the knowledge graph for efficient querying and relation reasoning.

[0032] Example 10: Based on Example 9, the causal mechanism path reasoning component provided in this embodiment of the invention includes: The causal hypothesis generation sub-component is used to generate a set of initial meta-path instances that connect the synergistic groups of prescriptions and diseases through matching and instantiation based on high-order relation templates in the collaborative enhanced knowledge network. The initial meta-path instance set is then filled and validated based on network embedding similarity to form a set of initial causal path hypotheses with fine granularity down to specific entities. The path optimization sub-component is used to process the initial set of causal path hypotheses through the neural symbolic joint inference engine. Using the symbolic inferencer, the logical consistency and completeness of the path are verified and deduced based on formal biological logic rules. At the same time, the neural network inferencer evaluates the local confidence and semantic coherence of each step in the path based on the graph attention mechanism. The deduction results of symbolic inference and the confidence information evaluated by the neural network are fused and iteratively adjusted through a collaborative optimization algorithm to output an optimized set of causal paths with logical reinforcement and confidence improvement. The biological pruning component is used to optimize the set of causal paths through pruning based on dynamic biological constraints. These constraints include, but are not limited to: organelle colocalization constraints, requiring molecular interactions in the path to be cellularly spatially possible; and biological process temporal constraints, requiring the path to conform to the temporal order of physiological or pathological processes. Each optimized causal path is assigned a constraint conflict score based on the degree to which it violates the dynamic constraints. All paths are sorted and thresholded according to their constraint conflict scores, ultimately outputting a formula-disease candidate mechanism path.

[0033] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, the causal hypothesis generation sub-component, through matching and instantiation based on high-order relation templates, generates a set of initial meta-path instances connecting formula synergistic groups and diseases in a collaboratively enhanced knowledge network. This initial meta-path instance set undergoes entity filling and verification based on network embedding similarity to form a fine-grained initial causal path hypothesis set down to specific entities. The path optimization sub-component processes this initial causal path hypothesis set through a neural symbolic joint reasoning engine. Using a symbolic inference engine, the logical consistency and completeness of the path are verified and deduced based on formalized biological logic rules. Simultaneously, the neural network inference engine, based on a graph attention mechanism, evaluates the local correlation of each step in the path. The confidence and semantic coherence of the system are considered. The deductive results of symbolic reasoning and the confidence information of neural network evaluation are fused and iteratively adjusted through a collaborative optimization algorithm to output an optimized causal path set that has undergone logical reinforcement and confidence enhancement. The biological pruning sub-component is used to optimize the causal path set through pruning based on dynamic biological constraints. The constraints include, but are not limited to: organelle colocalization constraints, which require that molecular interactions in the path have cellular spatial possibilities; biological process temporal constraints, which require that the path conforms to the temporal order of physiological or pathological processes; each optimized causal path is calculated with a constraint conflict score based on the degree to which it violates the dynamic constraints; all paths are sorted and thresholded according to their constraint conflict scores, and finally, the formula-disease candidate mechanism path is output. The above scheme generates meta-paths connecting formula synergistic groups and diseases in a collaboratively enhanced knowledge network through high-order relation template matching and instantiation, and fills them with specific entities to form fine-grained initial causal path hypotheses. Subsequently, a neural symbolic joint reasoning engine is used to perform logical verification and deductive extension of formal biological rules on the hypothetical paths, and the local confidence of each step is evaluated through a graph attention mechanism. Logical consistency and confidence information are fused for iterative optimization to obtain a set of causal paths with strengthened logic and improved confidence. Finally, the paths are pruned according to dynamic biological constraints such as organelle colocalization and biological process temporality, conflict scores are calculated and filtered according to thresholds, and high-confidence formula-disease candidate mechanism paths that meet spatial and temporal constraints are output, realizing a systematic derivation from coarse-grained hypotheses to fine-grained, constraint-consistent causal mechanisms.

[0034] Example 11: Based on Example 10, the biological pruning component provided in this embodiment of the invention includes: The spatiotemporal dependency graph construction module is used for organelle colocalization data and biological process temporal data in dynamic biological constraints. After ontology-based relation standardization processing, standardized sets of spatial colocalization relations and temporal sequence relations are formed. The two relation sets are then fused based on graph structure to construct a spatiotemporal dependency graph. Nodes in the spatiotemporal dependency graph represent biological entities or processes, and edges are labeled with their spatial colocalization probabilities or temporal sequence relations, providing a unified quantitative judgment benchmark for path verification. The path spatiotemporal consistency verification module is used to perform parallel verification of an optimized causal path and its spatiotemporal dependency graph. In the spatial verification channel, the interaction relationship between adjacent entity pairs in the optimized causal path is compared with the spatial co-location probability between corresponding entities in the spatiotemporal dependency graph to generate a series of spatial consistency verification results. In the temporal verification channel, the biological processes or state changes described in the path are compared with the standard temporal sequence relationships recorded in the spatiotemporal dependency graph to generate a series of temporal sequence verification results. The verification results are classified and recorded to form a set of spatiotemporal constraint conflict instances of the optimized causal path. The conflict severity quantification and integration module is used to evaluate and process the set of spatiotemporal constraint conflict instances. Each spatial conflict instance is assigned a severity value based on the co-location probability value of the entity pairs involved, with lower probabilities indicating more severe conflicts. Each temporal conflict instance is assigned a severity value based on the logical level of its violation of temporal order, with conflicts violating the core causal order being more severe. All the assigned conflict instances are then calculated based on a nonlinear aggregation function to simulate the effect of multiple minor conflicts potentially causing overall path failure, and integrated to generate a constraint conflict score that characterizes the overall path's violation of biological rationality.

[0035] The working principle and beneficial effects of the above technical solution are as follows: The spatiotemporal dependency graph construction module of this embodiment is used for organelle colocalization data and biological process temporal data in dynamic biological constraints. After ontology-based relation standardization processing, a standardized set of spatial colocalization relations and a set of temporal sequence relations are formed. The two relation sets are fused based on graph structure to construct a spatiotemporal dependency graph. Nodes in the spatiotemporal dependency graph represent biological entities or processes, and edges are labeled with their spatial colocalization probabilities or temporal sequence relations, providing a unified quantitative judgment benchmark for path verification. The path spatiotemporal consistency verification module is used for parallel verification processing of an optimized causal path and the spatiotemporal dependency graph. In the spatial verification channel, the interaction relationship between adjacent entity pairs in the optimized causal path is compared with the spatial colocalization probability between corresponding entities in the spatiotemporal dependency graph to generate a series of spatial consistency verification results. In the time-verification channel, the biological processes or state changes described in the path are compared with the standard temporal relationships recorded in the spatiotemporal dependency map to generate a series of time-series verification results. The verification results are classified and recorded to form a set of spatiotemporal constraint conflict instances for optimizing the causal path. The conflict severity quantification and integration module is used to evaluate the set of spatiotemporal constraint conflict instances. Each spatial conflict instance is assigned a severity value based on the co-location probability value of the entity pairs involved, with lower probabilities indicating more severe conflicts. Each temporal conflict instance is assigned a severity value based on the logical level of its violation of the temporal relationship, with conflicts violating the core causal order being more severe. All the assigned conflict instances are calculated based on a nonlinear aggregation function to simulate the effect of multiple minor conflicts potentially causing the overall path failure, and are integrated to generate a constraint conflict score characterizing the overall violation of biological rationality of the path. The above scheme unifies organelle colocalization and biological process temporal data into a set of spatial colocalization relationships and a set of temporal sequence relationships through ontological standardization. It then fuses these into a spatiotemporal dependency graph, providing a unified quantitative benchmark for subsequent path verification. Subsequently, it compares and optimizes the interactions between adjacent entities in causal paths with the colocalization probabilities in the graph, generating spatial consistency verification results. In the temporal channel, it compares the biological processes or state changes described in the paths with the standard temporal relationships in the graph, generating temporal sequence verification results. These two types of verification results are recorded as a set of spatiotemporal constraint conflict instances. Finally, it assigns severity to spatial conflicts based on colocalization probabilities and to temporal conflicts based on temporal hierarchy. Through comprehensive calculation using a nonlinear aggregation function based on cascading effects, it obtains the overall constraint conflict score for each path, achieving a quantitative assessment and screening of the spatiotemporal rationality of the paths.

[0036] Example 12: Based on Example 11, the conflict severity quantification and integration module provided in this embodiment of the invention includes: The conflict dependency network construction submodule is used for the set of spatiotemporal constraint conflict instances that have been assigned values. Through analysis based on the causal and conditional dependencies between conflicts, it identifies the dependencies between instances, such as triggering, aggravating, or concurrent. The identified dependencies are modeled using a graph structure, transforming each conflict instance into a node and the dependencies into directed edges, thus constructing a conflict dependency network. The cascading failure effect simulation submodule is used for conflict-dependent network simulation. Starting from the initial conflict node in the network, it iteratively calculates the propagation process and energy accumulation of the conflict impact along the network topology based on the dependency strength and type represented by the edges. After multiple rounds of simulation, each conflict node not only retains its initial severity assignment but also obtains a cascading impact value superimposed from the upstream conflict propagation. All nodes carry their initial values ​​and cascading impact values, forming a conflict state update set. The comprehensive conflict score generation submodule is used to weight and fuse the initial severity assignment of each conflict node with its cascaded impact value to obtain the comprehensive conflict intensity of the node. It reacts mildly to the sum of low-intensity conflicts, reacts sensitively to the sum of high-intensity conflicts, and the rate of increase decreases as it approaches the theoretical limit, generating the final constraint conflict score.

[0037] The working principle and beneficial effects of the above technical solution are as follows: The conflict dependency network construction submodule of this embodiment is used to analyze the set of spatiotemporal constraint conflict instances that have been assigned values. Based on the analysis of causal and conditional dependencies between conflicts, it identifies the initiation, aggravation, or concurrency dependencies between instances. The identified dependencies are then modeled using a graph structure, transforming each conflict instance into a node and the dependencies into directed edges, thus constructing a conflict dependency network. The cascading failure effect simulation submodule is used for conflict dependency network simulation. Starting from the initial conflict node in the network, it iteratively calculates the conflict impact based on the dependency strength and type represented by the edges. The propagation process and energy accumulation along the network topology are analyzed. After multiple rounds of simulation, each conflict node retains its initial severity assignment and obtains a cascaded impact value resulting from the propagation of upstream conflicts. All nodes carry their initial values ​​and cascaded impact values, forming a conflict state update set. The comprehensive conflict score generation submodule is used to weight and fuse the initial severity assignment and cascaded impact value of each conflict node to obtain the comprehensive conflict intensity of the node. The response to the sum of low-intensity conflicts is mild, while the response to the sum of high-intensity conflicts is sensitive, and the rate of increase decreases as it approaches the theoretical limit, generating the final constrained conflict score. The above scheme uses a conflict dependency network construction submodule to perform causal and conditional dependency analysis on assigned spatiotemporal constraint conflict instances, identify and model the initiation, aggravation, or concurrent relationships between conflicts, and form a conflict dependency network with conflict instances as nodes and dependencies as directed edges. Subsequently, in the cascading failure effect simulation submodule, based on the network topology, starting from the initial conflict node, the conflict influence is iteratively propagated and energy is accumulated according to the dependency strength and type of the edges to obtain the cascading influence value of each node. Finally, in the comprehensive conflict score generation submodule, the initial severity assignment of the node and its cascading influence value are integrated in a weighted manner to generate a constraint conflict score that responds smoothly to low-intensity conflicts, is sensitive to high-intensity conflicts, and whose rate of increase decreases when approaching the theoretical limit, thereby realizing a quantitative assessment of the degree of violation of the overall path spatiotemporal constraints.

[0038] Example 13: Based on Example 12, the comprehensive conflict score generation submodule provided in this embodiment of the invention includes: The conflict intensity level division unit is used to determine the comprehensive conflict intensity of each node in the conflict state update set. After level determination processing based on dynamic threshold, it is divided into multiple predefined conflict intensity levels. Each node is assigned a level identifier according to the threshold range to which its intensity value belongs. All nodes are classified according to their level identifiers to form a conflict node set with a level structure. The cross-level strength transfer unit is used for a set of conflict nodes with a hierarchical structure. After processing by the cross-level strength transfer function, the strength value of the higher-level conflict node will be linearly superimposed on the sum of the strengths of all levels lower than its own according to a specific transfer coefficient. The strength value of the lower-level conflict node will not produce an upward transfer effect. The sum of the original strengths of the nodes in each level, together with the transferred strength received from the higher level, is synthesized by vector addition to generate a set of transferred and corrected level strength vectors. A saturated nonlinear aggregation unit is used to calculate the magnitude of the grade intensity vector as the aggregation input value. The aggregation input value is substituted into a monotonically increasing function with an upper asymptote for mapping. When the input value is much smaller than the saturation threshold, the first derivative of the output of the monotonically increasing function remains constant, exhibiting linear characteristics. When the input value approaches the saturation threshold, the first derivative of the function output approaches zero, exhibiting saturation characteristics. The mapped output value is standardized and scaled to generate the final constraint conflict score.

[0039] The working principle and beneficial effects of the above technical solution are as follows: The conflict intensity level division unit in this embodiment is used to determine the comprehensive conflict intensity of each node in the conflict state update set. After level determination processing based on dynamic thresholds, it is divided into multiple predefined conflict intensity levels. Each node is assigned a level identifier according to the threshold range to which its intensity value belongs. All nodes are classified according to their level identifiers, forming a set of conflict nodes with a level structure. The cross-level intensity transfer unit is used for the set of conflict nodes with a level structure. After processing by the cross-level intensity transfer function, the intensity value of high-level conflict nodes will be linearly superimposed on the sum of the intensities of all levels lower than its own according to a specific transfer coefficient. The strength value of the synaptic node does not produce an upward propagation effect; the sum of the original strengths of nodes within each level, and the received propagation strengths from higher levels, are synthesized by vector addition to generate a set of propagation-corrected level strength vectors; saturated nonlinear aggregation units are used to calculate the magnitude of the level strength vectors as the aggregation input value; the aggregation input value is substituted into a monotonically increasing function with an upper asymptote for mapping; when the input value is much smaller than the saturation threshold, the first derivative of the output of the monotonically increasing function remains constant, exhibiting linear characteristics; when the input value approaches the saturation threshold, the first derivative of the function's output approaches zero, exhibiting saturation characteristics; the mapped output value is standardized and scaled to generate the final constraint conflict score. The above scheme classifies the comprehensive conflict intensity of each node in the conflict state update set according to a dynamic threshold, assigns a level label, and forms a hierarchical set of conflict nodes. Then, in the cross-level intensity transfer unit, the intensity of high-level nodes is linearly superimposed onto all levels below it using a preset coefficient, achieving downward propagation of high-risk conflicts to the overall intensity, while low-level conflicts do not affect upwards. Finally, in the saturated nonlinear aggregation unit, the modulus of the modified intensity vectors of each level is calculated as the aggregation input and substituted into a monotonically increasing function with an upper asymptote for mapping. The function maintains linear growth in the low input region and tends to level off near the saturation threshold. The output is then standardized and scaled to obtain the final constraint conflict score. Overall, this scheme achieves hierarchical evaluation of conflict intensity, intensity transfer between levels, and nonlinear aggregation based on a saturation function, thereby providing a refined and adjustable quantitative evaluation of the overall spatiotemporal constraint violation degree of the path.

[0040] Example 14: Based on Example 13, the saturated nonlinear aggregation unit provided in this embodiment of the invention includes: The scaling benchmark construction subunit is used to construct the theoretical upper reference limit and statistical reference distribution required for the standardization process through analysis based on the theoretical maximum output value and empirical distribution percentiles. The theoretical upper reference limit is derived from the mathematical relationship between the domain and range of the monotonically increasing function, while the statistical reference distribution comes from the statistical analysis results of the historical conflict score dataset. The two reference benchmarks together define the calibration framework for the standardization operation. The dynamic scale alignment sub-unit is used to combine the output value of the monotonically increasing function with the theoretical reference upper limit and statistical reference distribution generated by the scale benchmark construction unit, and then undergoes a scale transformation based on nonlinear interpolation. This process first converts the original output value into a cumulative probability value according to the statistical reference distribution, and then, in conjunction with the theoretical reference upper limit, maps the cumulative probability value to an intermediate scale value with a clear theoretical boundary through nonlinear interpolation. The fractional interval calibration subunit is used to process the intermediate scale values ​​generated by the dynamic scale alignment unit through a linear affine transformation based on a preset target interval. Through translation and scaling operations, the value range of the intermediate scale values ​​is precisely adjusted to the preset target value interval. The final adjusted value is then used to generate a standardized constraint conflict score with comparative significance.

[0041] The working principle and beneficial effects of the above technical solution are as follows: In this embodiment, the scale benchmark construction subunit is used to construct the theoretical reference upper limit and statistical reference distribution required for the standardization process through analysis based on the theoretical maximum output value and empirical distribution percentiles. The theoretical reference upper limit is derived from the mathematical relationship between the domain and range of the monotonically increasing function, while the statistical reference distribution originates from the statistical analysis results of the historical conflict score dataset. The two reference benchmarks jointly define the calibration framework for the standardization operation. The dynamic scale alignment subunit is used to process the output value of the monotonically increasing function and the theoretical reference upper limit and statistical reference distribution generated by the scale benchmark construction unit through a scale transformation based on nonlinear interpolation. This process first converts the original output value into a cumulative probability value according to the statistical reference distribution, and then, combined with the theoretical reference upper limit, maps the cumulative probability value to an intermediate scale value with a clear theoretical boundary through a nonlinear interpolation algorithm. The score interval calibration subunit is used to process the intermediate scale value generated by the dynamic scale alignment unit through a linear affine transformation based on a preset target interval. Through translation and scaling operations, the range of the intermediate scale value is precisely adjusted to the preset target value interval. The final adjusted value is then used to generate a standardized constraint conflict score with comparative significance. The above scheme constructs a unified theoretical upper limit and statistical distribution benchmark by performing percentile analysis of the theoretical maximum output value and historical conflict scores, providing a reference framework for subsequent standardization. Then, the original output value of the monotonically increasing function is mapped to the cumulative probability, and nonlinear interpolation is used to perform scaling transformation between the theoretical upper limit and the statistical distribution to obtain an intermediate scale value with a clear theoretical boundary. Finally, a linear affine transformation is performed on the intermediate scale value to shift and scale its value range to a preset target interval, generating a comparable standardized constraint conflict score, thus achieving unified calibration and quantification of conflict scores on both theoretical and empirical scales.

[0042] Example 15: As Figure 6As shown, based on Example 1, the prescription optimization decision generation subsystem provided in this embodiment of the invention includes: The candidate perturbation generation component, used for novel prescription-disease association pathways, identifies target groups and biological processes with core regulatory effects on the target disease mechanism through parsing based on key nodes of the pathway. The identification results are compared with the component synergistic subnetwork of the original prescription. By calculating the differences in network structure and the degree of overlap of node functions, a set of initial operation instructions for increasing, decreasing, and adjusting prescription components is generated. The set of operation instructions is initially filtered based on the knowledge base of traditional Chinese medicine compatibility contraindications to form a set of prescription perturbation schemes that conform to basic compatibility principles, i.e., a set of primary optimization schemes. The efficacy simulation and evaluation component is used to process a set of primary optimized regimens through three core evaluation channels of a multimodal deep efficacy simulator: In the pharmacological effect channel, a graph neural network model trained based on biomedical knowledge graphs and molecular docking prediction data simulates changes in downstream biological effects after acting on disease-related pathways, outputting predicted efficacy intensity; in the chemical stability channel, a model based on computational chemistry rules evaluates potential new compound interactions and stability risks; in the compatibility and synergy channel, the overall compatibility rationality is evaluated by analyzing the spectral similarity between the component combination and classic prescription compatibility patterns; the outputs of the three channels are fused to generate a quantitative evaluation report containing multidimensional predictive indicators for each primary optimized regimen. The Pareto front navigation component carries all schemes from the quantitative assessment report into a multi-objective optimization space. Within this space, the defined core optimization objectives include: maximizing predicted efficacy, minimizing potential stability risks, and minimizing deviations from the core compatibility characteristics of the original formula. The optimization process excludes non-compliant schemes based on hard constraints such as clinical safety thresholds. Among the remaining schemes, those that cannot be further improved on any one objective without harming other objectives are identified, forming the Pareto optimal frontier. Through a frontier search based on decision-maker preferences, a final subset of schemes is selected from the optimal frontier, namely the formula optimization decision set.

[0043] The working principle and beneficial effects of the above technical solution are as follows: This embodiment analyzes the key targets and biological processes of the novel prescription-disease association pathway, compares the component synergy network of the original prescription, generates addition-subtraction-adjustment operation instructions that conform to the contraindications of traditional Chinese medicine, and forms a primary optimization scheme; then, in a multimodal efficacy simulator, it predicts the downstream pharmacological effects, chemical stability risks, and compatibility synergy rationality, and integrates the three types of evaluation results into a multidimensional quantitative report for each scheme; finally, in the multi-objective optimization space, with the core objectives of maximizing efficacy, minimizing stability risks, and minimizing compatibility deviations, it eliminates non-compliant schemes based on safety constraints, identifies the set of schemes that cannot be further improved on any objective without sacrificing other objectives, constitutes the Pareto optimal frontier, and selects a subset of the final schemes from the frontier according to the decision-maker's preferences to generate the prescription optimization decision set. Overall, it realizes a closed-loop decision-making process from key target identification, compatibility constraint filtering, cross-dimensional efficacy and risk simulation, to multi-objective Pareto optimization, ensuring that the optimized prescription achieves an optimal state of balance in efficacy, stability, and compatibility rationality.

[0044] This embodiment also includes a knowledge discovery and visualization module, which presents the analysis results to users in an intuitive and interactive way to assist in knowledge discovery and decision support. Interactive network visualization: Employing force-directed graphs, adjacency matrices, and other visualization methods, it displays a multi-layered network of formula-component-target-disease. Users can interact with the network by clicking, dragging, and zooming nodes to view detailed information. It provides network simplification functions, such as filtering nodes and edges based on importance scores to highlight the core network structure. Multi-dimensional data display: It provides various visualization charts such as scatter plots, heatmaps, bubble charts, and pathway diagrams to display the analysis results from different dimensions. For example, a heatmap can show the changes in the content of formula components in different samples, and a pathway diagram can mark the key targets of the formula's action. It supports multi-view linkage; when a user selects an element of interest in one view, other views automatically display related information. Intelligent question answering and report generation: It integrates a natural language question answering system, allowing users to ask questions in natural language, such as "What are the main targets of Ephedra Decoction?" The system automatically analyzes questions and retrieves answers from the knowledge graph; it provides a one-click function to generate analysis reports, automatically organizing analysis results such as network diagrams, enrichment analysis results, and key target lists into structured reports, and supports exporting in PDF, Word and other formats.

[0045] The system application and output module transforms the system's analytical capabilities into practical application scenarios, providing users with specific solutions.

[0046] **Formula Mechanism of Action Analysis:** Inputting a formula name automatically generates a component-target-pathway network for that formula, identifying potential active ingredients, key targets, and significantly enriched biological pathways, presented in a visual format to provide clues for elucidating the formula's mechanism of action. Output includes: a list of core components, a list of key targets, an enriched pathway chart, a network visualization, and explanatory text on the mechanism. **Compatibility Pattern Mining:** Based on association rule mining and network modular analysis, it discovers high-frequency combinations of medicinal materials and functional modules within the formula, revealing synergistic and adjuvant relationships among the medicinal materials. Output includes: a list of high-frequency combinations of medicinal materials, association rules (support, confidence), a network module decomposition diagram, and a compatibility pattern analysis report. **New Indication Discovery:** Based on network association analysis between the formula and diseases, it calculates the association strength between the formula and diseases, predicting potential new indications for the formula. Output includes: a list of candidate new indications, association score, an evidence network diagram (showing which targets and pathways the formula associates with the disease), and an explanation of the prediction basis. Formula optimization and new formula design: Based on the efficacy prediction model and optimization algorithm, we propose optimization suggestions for the composition or dosage of existing formulas, or design new formulas based on specific efficacy goals; the output includes: the composition and dosage ratio of the optimized formula, the predicted efficacy evaluation, and the comparative analysis with the original formula.

[0047] This embodiment can be deployed using a cloud computing architecture. The front end uses Web technologies, such as HTML5 and JavaScript, to implement the interactive interface, while the back end uses a microservice architecture, encapsulating different functional modules into independent services and communicating through RESTful APIs. The knowledge graph is stored in the Neo4j graph database, and other structured data can be stored in relational databases, such as MySQL, or non-relational databases.

[0048] The system development in this embodiment can follow an agile development process, first building the core data model and basic analysis functions; then gradually expanding the data sources and analysis modules; system maintenance includes regularly updating data, optimizing algorithm models, fixing vulnerabilities, and improving functions based on user feedback.

[0049] To demonstrate the practicality of the system, the following example uses the analysis of Ephedra Decoction to illustrate the workflow of the system of the present invention: First step: The user inputs Mahuang Decoction in the search interface, and the system retrieves the constituent herbs of this formula from the knowledge graph, namely Ephedra sinica, Cinnamomum cassia, Armeniaca vulgaris, and Glycyrrhiza uralensis. Second step: The system automatically constructs a Mahuang Decoction - ingredient - target network, identifies potential active ingredients such as ephedrine and cinnamaldehyde, as well as key targets such as ADRA2A and PTGS2. Third step: Conduct pathway enrichment analysis and find that the action targets of Mahuang Decoction are significantly enriched in pathways related to inflammatory response, fever, etc. Fourth step: Through network topology analysis and modular analysis, identify the core status of Ephedra sinica and Cinnamomum cassia in the formula, as well as the compatibility relationship between herbs. Fifth step: The results are presented to the user in the form of interactive network diagrams, enrichment analysis charts, data tables, etc., and an analysis report can be generated with one click.

[0050] Through the above implementation manners, the present invention can effectively support the modern research of traditional Chinese medicine formulas and contribute to the inheritance and innovation of traditional Chinese medicine.

[0051] This embodiment integrates various types of data, including TCM formulas, chemical components, targets, and diseases, through multi-source data fusion technology; constructs a TCM knowledge graph based on a multi-dimensional semantic network of formulas, herbs, components, targets, pathways, and diseases; develops an intelligent analysis engine using network analysis, machine learning, and deep learning algorithms; and achieves multi-dimensional data display and knowledge discovery through an interactive visualization interface; ultimately forming a comprehensive solution supporting the analysis of formula mechanisms of action, the mining of compatibility rules, the discovery of new indications, and the development of new drugs. This embodiment achieves deep fusion of multi-source data, integration of multi-level analysis models, and intuitive visualization of complex relationships, solving key problems in traditional formula pharmacological analysis such as data silos, single methods, and unclear mechanisms, providing a powerful tool for the modernization of TCM research. This invention belongs to the technical solution combining computational systems pharmacology and TCM informatics. Its core lies in treating TCM formulas as a complex chemical-biological network system for processing. In modern systems science, many theoretical models used to analyze complex engineering systems such as power networks and information networks, such as network flow theory, cascade failure models, conflict detection, and multi-objective optimization, have been widely adopted and successfully applied to biological network analysis, such as protein-protein interaction networks, metabolic pathways, and gene regulatory networks, due to their effectiveness in handling general problems such as node association, dynamic propagation, constraint satisfaction, and resource allocation. This invention also modifies the computational model within this interdisciplinary paradigm to quantify the complex mechanisms of synergistic effects of multiple components, targets, and pathways in traditional Chinese medicine formulas. In network flow theory, betweenness centrality measures the frequency with which a node acts as a bridge for the shortest path; high betweenness centrality nodes are usually key hubs or flow bottlenecks in the network. In component-target dynamic interaction networks, information flow or regulatory effects are considered as network flow; a high betweenness centrality of a target protein indicates that it plays a crucial signal transfer or regulatory hub role in connecting different groups of chemical components or different functional modules; such targets are often key to maintaining the overall function of the network or communicating different biological processes. The rationale for its application in functional module discovery: Community detection based on network flow betweenness numbers aims to identify relatively sparsely connected boundaries with low flow rates within a network. In biological networks, this corresponds to functional modules, such as coupling points between specific signaling pathways and metabolic modules. By identifying these boundaries, we can more accurately delineate tightly connected, functionally synergistic subnetworks, aligning closely with the concept of functional modularization in biology. Compared to simple topological clustering, this method better integrates network dynamics and identifies modules with more explicit biological significance. In network security or reliability engineering, conflict-dependent networks are used to model causal chains between failures or security events. Cascade failures describe the process where an initial failure propagates and amplifies along the dependency chain, leading to widespread system collapse.In this invention, conflict does not refer to security attacks, but rather to inconsistencies between the inferred formula-disease candidate mechanism pathway and known, inviolable biological spatiotemporal constraints, such as protein colocalization and the sequence of biological processes. These inconsistencies or conflicts may be logically related. For example, a hypothesis that protein A and B interact on the cell membrane, if a step in the pathway conflicts with the constraint that protein A is only expressed in the nucleus, will further cause all subsequent pathway steps dependent on this interaction, such as B activating C, to lose their basis, forming a chain of conflict propagation. The rationality of the simulation: Cascade failure effect simulation is used to quantitatively assess the severity of a candidate pathway's overall violation of fundamental biological laws. A small conflict or initial failure with basic cell biology facts may severely impair the rationality of the entire pathway or cause cascade failure through dependencies. By constructing conflict dependency networks and simulating cascade effects, a more comprehensive and robust constraint conflict score can be calculated, thereby prioritizing mechanism pathways with higher biological rationality and avoiding the support of hypotheses built on a chain of biological fallacies.

[0052] In multi-objective optimization, Pareto optimality refers to the set of states where no further improvement is possible without compromising any of the objectives, forming the Pareto front. Formula optimization is essentially a multi-objective decision-making problem, often involving trade-offs between objectives. For example, enhancing efficacy (objective one) might require increasing the dosage of certain components or introducing new ones, but this could increase potential chemical instability risks (objective two) or deviate from the original formula's core principles of harmony (objective three). No single solution can achieve absolute optimality across all three objectives simultaneously. Introducing Pareto front navigation is a scientific approach to handling these multi-objective trade-offs. Instead of subjectively providing an absolute optimal solution, it reveals the set or frontier of all optimal trade-off solutions or Pareto optimal solutions through computation. Decision-makers, such as pharmacists and clinical experts, can make a final choice from this set of optimal solutions objectively provided by the algorithm and constrained by biological and chemical principles, based on actual clinical needs, such as prioritizing efficacy or safety. This approach is more in line with the holistic view and precise regulation requirements of traditional Chinese medicine than single-objective optimization or subjective weighting.

[0053] An exemplary electronic device suitable for implementing embodiments of the present invention. The electronic device may include a central processing unit / microprocessor / main control chip; a storage medium coupled to the central processing unit / microprocessor / main control chip and storing computer-executable instructions therein for performing steps of various methods of embodiments of the present invention when executed by a processor. The central processing unit / microprocessor / main control chip may include, but is not limited to, one or more processors or microprocessors. The storage medium may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (e.g., hard disk, floppy disk, solid-state drive, removable disk, CD-ROM, DVD-ROM, Blu-ray disc, etc.). In addition, the electronic device may also include (but is not limited to) a data bus, an input / output bus / external bus / device bus, a display, and input / output devices (e.g., keyboard, mouse, speaker, etc.). The central processing unit / microprocessor / main control chip may communicate with external devices via the input / output bus / external bus / device bus through a wired or wireless network. The storage medium may also store at least one computer-executable instruction for performing the steps of various functions and / or methods in the embodiments described herein when run by a central processing unit / microprocessor / main control chip. In one embodiment, the at least one computer-executable instruction may also be compiled into or comprise a software product, wherein one or more computer-executable instructions are run by a processor to perform the steps of various functions and / or methods in the embodiments described herein. A computer-readable storage medium according to an embodiment of the present invention. Instructions, such as computer-readable instructions, are stored on a non-transitory computer-readable storage medium. When the computer-readable instructions are run by a processor, the various methods described above can be performed. The non-transitory computer-readable storage medium includes, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-transitory non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, the non-transitory computer-readable storage medium can be connected to a computing device such as a computer, and then, when the computing device runs the computer-readable instructions stored on the non-transitory computer-readable storage medium, the various methods described above can be performed. Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of equivalents of this invention, this invention is also intended to include these modifications and variations.

Claims

1. A pharmacological analysis system for traditional Chinese medicine prescriptions, characterized in that, Include: The component synergy network construction subsystem is used to process multi-omics feature vectors containing structured and unstructured data through network analysis, calculate the correlation strength between chemical components, and construct a component-target dynamic interaction network; the component-target dynamic interaction network is then subjected to clustering and community detection to obtain the component synergy subnetwork; The disease association path discovery subsystem is used to integrate the component synergy subnetwork with the disease knowledge base after processing by the knowledge graph construction module to form a prescription-disease association graph. The prescription-disease association graph is then used for path mining to obtain a set of potential disease association paths. The set of potential disease association paths is then filtered and scored to obtain novel prescription-disease association paths. The formula optimization decision generation subsystem is used to obtain candidate schemes for TCM formula optimization by processing novel formula-disease association paths through machine learning and combining formula optimization functions, through simulating TCM formula compatibility adjustment and effect prediction; the candidate schemes for TCM formula optimization are then optimized through multi-objective optimization to obtain the TCM formula optimization decision set.

2. The pharmacological analysis system for traditional Chinese medicine prescriptions as described in claim 1, characterized in that, The component synergistic network construction subsystem includes: The dynamic interaction relationship inference component is used to obtain the conditional association probability distribution between chemical components and target proteins by performing conditional dependency analysis based on multi-omics context on multi-omics feature vectors; the conditional association probability distribution is then filtered for significance based on statistical potential to construct the initial component-target association backbone. The network topology generation component is used to infer the initial component-target association skeleton obtained from the dynamic interaction relationship inference component. After introducing the temporal or dose response patterns contained in the multi-omics feature vectors, the connection strength is assigned to form a weighted component-target dynamic interaction network. The functional community discovery component is used to obtain the potential functional community partitioning in the network by performing functional module boundary detection based on network flow betweenness centrality on the weighted component-target dynamic interaction network obtained from the network topology generation component. Potential functional communities are optimized and screened based on module cohesion and inter-module separation to form component synergistic subnetworks.

3. The pharmacological analysis system for traditional Chinese medicine prescriptions as described in claim 1, characterized in that, The disease association path discovery subsystem includes: The heterogeneous knowledge network collaborative enhancement component is used to process the component synergy subnetwork through the knowledge graph construction module, align it with the entity and map the relationship with the disease knowledge base, and form a preliminary integrated heterogeneous graph. The heterogeneous graph introduces relation completion and collaborative embedding learning based on graph neural network to model and complete the potential high-order relationship between component-target synergy groups and target-disease association, and generate a collaboratively enhanced knowledge network. The causal mechanism path reasoning component is used to process the synergistic enhanced knowledge network through a symbolic and numerical hybrid reasoning engine, combined with the rule exploration of inductive logic programming and the pattern recognition capabilities of graph attention networks, to traverse and generate candidate action path sequences that connect specific prescription synergistic groups and disease nodes in the knowledge network. The multi-dimensional evidence aggregation scoring and screening component is used to process the prescription-disease candidate mechanism pathway. It performs quantitative evaluation and weighted fusion from four dimensions: pathway tightness, enrichment significance of genes involved in the pathway in related disease pathways, perturbation correlation of omics data at key nodes of the pathway, and literature co-occurrence support, and calculates the comprehensive confidence score of each pathway. All paths are sorted according to their overall confidence score and filtered based on dynamic thresholds to output novel prescription-disease association paths with high confidence.

4. The pharmacological analysis system for traditional Chinese medicine prescriptions as described in claim 3, characterized in that, The causal mechanism path reasoning component includes: The causal hypothesis generation sub-component is used to generate a set of initial meta-path instances that connect the synergistic groups of prescriptions and diseases through matching and instantiation based on high-order relation templates in the collaborative enhanced knowledge network. The initial meta-path instance set is then filled and validated based on network embedding similarity to form a set of initial causal path hypotheses with fine granularity down to specific entities. The path optimization sub-component is used to process the initial set of causal path hypotheses through the neural symbolic joint inference engine. Using the symbolic inferencer, the logical consistency and completeness of the path are verified and deduced based on formal biological logic rules. At the same time, the neural network inferencer evaluates the local confidence and semantic coherence of each step in the path based on the graph attention mechanism. The deduction results of symbolic inference and the confidence information evaluated by the neural network are fused and iteratively adjusted through a collaborative optimization algorithm to output an optimized set of causal paths with logical reinforcement and confidence improvement. Biological pruning sub-component, used to optimize the set of causal paths after pruning based on dynamic biological constraints; Each optimized causal path is assigned a constraint conflict score based on the degree to which it violates dynamic constraints. All paths are sorted and thresholded based on their constraint conflict scores, and the final output is the prescription-disease candidate mechanism path.

5. The pharmacological analysis system for traditional Chinese medicine prescriptions as described in claim 4, characterized in that, Biological pruning component, including: The spatiotemporal dependency graph construction module is used for organelle colocalization data and biological process temporal data in dynamic biological constraints. After ontology-based relation standardization processing, it forms a standardized set of spatial colocalization relations and a set of temporal sequence relations. Two sets of relations are fused using a graph structure to construct a spatiotemporal dependency graph; The path spatiotemporal consistency verification module is used to perform parallel verification of an optimized causal path and spatiotemporal dependency graph. In the spatial verification channel, the interaction relationship between adjacent entity pairs in the optimized causal path is compared with the spatial colocation probability between corresponding entities in the spatiotemporal dependency graph to generate a series of spatial consistency verification results. The verification results are classified and recorded to form a set of spatiotemporal constraint conflict instances for optimizing causal paths; The conflict severity quantification and integration module is used to evaluate and process the set of spatiotemporal constraint conflict instances. Each spatial conflict instance is assigned a severity value based on the co-location probability value of the entity pairs involved. All the conflict instances with assigned values ​​are integrated to generate a constraint conflict score that represents the overall path's violation of biological rationality through calculation based on a nonlinear aggregation function.

6. The pharmacological analysis system for traditional Chinese medicine prescriptions as described in claim 5, characterized in that, The conflict severity quantification and integration module includes: The conflict dependency network construction submodule is used to identify the triggering, aggravating, or concurrent dependencies between instances of a set of spatiotemporal constraint conflict instances that have been assigned values, based on the analysis of causal and conditional dependencies between conflicts. The identified dependencies are modeled using a graph structure, transforming each conflict instance into a node and the dependency into a directed edge, thus constructing a conflict-dependency network. The cascading failure effect simulation submodule is used for conflict-dependent network simulation. Starting from the initial conflict node in the network, it iteratively calculates the propagation process and energy accumulation of the conflict effect along the network topology based on the dependency strength and type represented by the edges. After multiple rounds of simulation, a cascading impact value resulting from the superposition of upstream network conflict propagation was obtained; All nodes carry their initial values ​​and cascading impact values, forming a conflict state update set; The comprehensive conflict score generation submodule is used to weight and fuse the initial severity assignment of each conflict node with its cascaded impact value to obtain the comprehensive conflict intensity of the node. Generate a constraint conflict score.

7. The pharmacological analysis system for traditional Chinese medicine prescriptions as described in claim 6, characterized in that, The comprehensive conflict score generation submodule includes: The conflict intensity level division unit is used to determine the comprehensive conflict intensity of each node in the conflict state update set. After the level determination process based on dynamic threshold, it is divided into multiple predefined conflict intensity levels. Each node is assigned a level identifier according to the threshold range to which its intensity value belongs. All nodes are categorized according to their level identifiers, forming a set of conflicting nodes with a hierarchical structure. The cross-level strength transfer unit is used for a set of conflict nodes with a hierarchical structure. After processing by the cross-level strength transfer function, the sum of the original strengths of the nodes in each level and the received transfer strengths from higher levels are combined by vector addition to generate a set of transfer-corrected level strength vectors. A saturated nonlinear aggregation unit is used to calculate the magnitude of the grade intensity vector as the aggregation input value. The aggregation input value is then substituted into a monotonically increasing function with an upper asymptote for mapping. The mapped output value is then standardized and scaled to generate the final constraint conflict score.

8. The pharmacological analysis system for traditional Chinese medicine prescriptions as described in claim 7, characterized in that, Saturated nonlinear aggregation unit, comprising: The scaling benchmark construction sub-unit is used to construct the theoretical reference upper limit and statistical reference distribution required for the standardization process through analysis based on the theoretical maximum output value and empirical distribution percentiles; The dynamic scale alignment sub-unit is used to compare the output value of the monotonically increasing function with the theoretical reference upper limit and statistical reference distribution generated by the scale benchmark construction unit, after scale transformation based on nonlinear interpolation. The original output value is converted into a cumulative probability value based on the statistical reference distribution. Then, combined with the theoretical reference upper limit, the cumulative probability value is mapped to an intermediate scale value with a clear theoretical boundary through nonlinear interpolation. The fractional interval calibration sub-unit is used to process the intermediate scale values ​​generated by the dynamic scale alignment unit through a linear affine transformation based on a preset target interval. By using translation and scaling operations, the range of intermediate scale values ​​is precisely adjusted to the preset target value range; The adjusted final value is then used to generate a standardized constraint conflict score that is meaningful for comparison.

9. The pharmacological analysis system for traditional Chinese medicine prescriptions as described in claim 1, characterized in that, The prescription optimization decision generation subsystem includes: The candidate perturbation generation component is used to identify target groups and biological processes that play a core regulatory role in the mechanism of the target disease by analyzing the path based on key nodes of the path through the analysis of novel prescription-disease association pathways. The identification results are compared with the component synergy subnetwork of the original prescription. By calculating the differences in network structure and the degree of overlap of node functions, a set of initial operation instructions for adding, subtracting and adjusting the prescription components is generated. The set of operation instructions is then filtered based on the knowledge base of traditional Chinese medicine compatibility contraindications to form a set of primary optimization schemes that conform to the basic compatibility principles. The efficacy simulation and evaluation component is used to process a set of primary optimized regimens through three core evaluation channels of a multimodal deep efficacy simulator: In the pharmacological effect channel, a graph neural network model trained based on biomedical knowledge graphs and molecular docking prediction data simulates changes in downstream biological effects after acting on disease-related pathways, outputting predicted efficacy intensity; in the chemical stability channel, a model based on computational chemistry rules evaluates potential new compound interactions and stability risks; in the compatibility and synergy channel, the overall compatibility rationality is evaluated by analyzing the spectral similarity between the component combination and classic prescription compatibility patterns; the outputs of the three channels are fused to generate a quantitative evaluation report containing multidimensional predictive indicators for each primary optimized regimen. The Pareto front navigation component is used to carry all schemes with quantitative evaluation reports into the multi-objective optimization space; Within the space, the defined core optimization objectives include: maximizing the predicted efficacy, minimizing potential stability risks, and minimizing deviations from the core compatibility characteristics of the original formula; the optimization process excludes non-compliant schemes based on the hard constraints of clinical safety thresholds; among the remaining schemes, those schemes that cannot be further improved on any one objective without harming other objectives are found, forming the Pareto optimal frontier; through a frontier search based on decision-maker preferences, the formula optimization decision set is selected from the optimal frontier.

10. The pharmacological analysis system for traditional Chinese medicine prescriptions as described in claim 1, characterized in that, It also includes a multi-omics feature fusion subsystem, which provides multi-source data containing both structured and unstructured data. After multi-omics integration and analysis, a multi-omics dataset is obtained. The multi-omics dataset is then subjected to feature extraction and dimensionality reduction operations to obtain a fused multi-omics feature vector.

Citation Information

Patent Citations

  • Traditional Chinese medicine system pharmacology analysis platform and analysis method

    CN111241164A

  • Traditional Chinese medicine prescription and disease analysis method and system, equipment and storage medium

    CN113486231A

  • Intelligent matching system for traditional Chinese medicine formula

    CN120561281A