Traditional Chinese medicine screening system and method based on AI assistance

By constructing a hierarchical semantic relation knowledge graph and performing representation learning in the dual quaternion space, the problem of insufficient characterization and prediction in existing technologies for traditional Chinese medicine drug screening is solved, and high precision and efficiency in traditional Chinese medicine drug screening are achieved.

CN120977433APending Publication Date: 2025-11-18SHANGHAI UNIV OF T C M
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511108814.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies for screening traditional Chinese medicine rely on knowledge graphs with relatively flat relationships, which are difficult to effectively represent the complex hierarchy and correlations in traditional Chinese medicine theory. Conventional knowledge representation learning methods have limited expressive power when dealing with such complex relationships, resulting in insufficient accuracy in screening and prediction.

Method used

An AI-based traditional Chinese medicine (TCM) drug screening system is adopted, comprising a data processing module, a knowledge graph construction module, a knowledge representation learning module, and a drug screening prediction module. The data processing module cleans and standardizes multi-source heterogeneous data; the knowledge graph construction module constructs a hierarchical semantic relationship knowledge graph through a multi-agent collaborative framework; the knowledge representation learning module performs embedding representation learning in a dual quaternion space; and the drug screening module calculates the prediction score between TCM and target using dual quaternion interaction operators.

Benefits of technology

It significantly improves the semantic accuracy and information depth of knowledge graphs, enhances the accuracy of drug-target correlation prediction, shortens the R&D cycle, improves the interpretability of prediction results and the scalability of the system, and provides a high-quality data foundation and decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977433A_ABST
    Figure CN120977433A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional Chinese medicine screening system based on AI assistance and a method thereof, belongs to the field of computer technology and biological medicine cross technology, and aims to solve the problem that complex action mechanisms between traditional Chinese medicines and diseases are difficult to accurately characterize and calculate in the prior art. The method comprises the following steps: firstly, automatically constructing a traditional Chinese medicine knowledge graph from multi-source heterogeneous data by utilizing a multi-agent collaboration framework; the key point is that the relationship semantic analysis agent in the framework can define the relationship between the entities as a structure with hierarchical semantics such as'upstream activation ', 'synergistic interaction' and the like, so that the biomedical logic is deeply reflected. Then, the constructed knowledge graph is mapped to a dual quaternion space, and a dual quaternion graph attention network is adopted for knowledge representation learning; according to the method, hierarchical semantics and dual quaternion representation are introduced, so that the drug screening accuracy and the action mechanism revealing ability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer technology and biological medicine, in particular to an AI-assisted traditional Chinese medicine drug screening system and method. BACKGROUND

[0002] Traditional Chinese medicine is the treasure of the Chinese nation. It has accumulated rich prescriptions and medication experience in thousands of years of clinical practice, especially in the treatment of complex chronic diseases such as heart failure, showing the advantages of multi-component, multi-target, and multi-path whole regulation. However, the complexity of traditional Chinese medicine system, especially the compatibility theory of "jun, chen, za and shi" and the ambiguity of its mechanism of action, also makes it difficult for modern pharmacology and drug development system to fully understand and accept it, greatly limiting its modernization and internationalization process.

[0003] In recent years, with the rapid development of artificial intelligence technology, the application of knowledge graph and other technologies to the modernization of traditional Chinese medicine, especially for screening effective components and prescriptions for treating heart failure, has become a promising research direction. The existing technology usually extracts entities and relationships from massive medical literature and databases to construct a knowledge graph Yup containing "traditional Chinese medicine-component-target-disease" associations. Researchers hope to discover new drug-target associations by analyzing the topological structure of this graph, thereby accelerating drug development.

[0004] Although this idea brings new light to the modernization of traditional Chinese medicine, there are still deep-seated limitations in current technical practice. Existing methods often only establish a flat, lack of semantic depth relationship network when constructing knowledge graph. The association between entities is simply abstracted as a nondiscriminatory link, for example, the relationship between a traditional Chinese medicine component and a protein target is labeled as "acting on". This processing method cannot effectively distinguish whether a component directly binds to a target or indirectly regulates the target by activating upstream signaling pathways; it also cannot reflect whether different components are synergistic or antagonistic. This rough simplification of the intrinsic logic of biological systems greatly reduces the real pharmacological information that the graph can carry, and the accuracy and interpretability of the analysis results are therefore fundamentally constrained.

[0005] Furthermore, when transforming this simplified knowledge graph into a machine-computable mathematical model, existing technologies generally employ traditional Euclidean space for vectorization. However, this classic vector space is geometrically ill-suited to the complex structures prevalent in biological networks, such as hierarchy, dependency, and transitivity. It falls short in expressing asymmetric relationships like "upstream-downstream regulation," which involve directional and hierarchical dependencies. This deficiency in mathematical representation causes the model to lose a significant amount of valuable structured information during learning and reasoning, ultimately affecting prediction accuracy and hindering the true understanding of the profound scientific implications of traditional Chinese medicine (TCM) compound prescriptions for treating complex diseases. Therefore, a novel technological solution is urgently needed that can more profoundly characterize the logic of TCM pharmacology and express it using more appropriate mathematical language. Summary of the Invention

[0006] The technical problem that this invention aims to solve is that the existing technologies for screening traditional Chinese medicine rely on relatively flat knowledge graph relationships, which are difficult to effectively represent the complex hierarchy and correlation in traditional Chinese medicine theory. At the same time, conventional knowledge representation learning methods have limited expressive power when dealing with such complex relationships, resulting in insufficient accuracy in screening and prediction.

[0007] To address the aforementioned technical problems, this invention provides an AI-assisted traditional Chinese medicine drug screening system and method.

[0008] The first aspect of this invention provides an AI-assisted traditional Chinese medicine drug screening system, which includes: a data processing module, a knowledge graph construction module, a knowledge representation learning module, and a drug screening prediction module.

[0009] The data processing module is used to collect and preprocess pre-defined multi-source heterogeneous data. This multi-source heterogeneous data may include a traditional Chinese medicine knowledge base, a biomedical database, and a scientific literature database. By cleaning, aligning, and standardizing this data, a high-quality data foundation is provided for subsequent processing.

[0010] The knowledge graph construction module receives the preprocessed multi-source heterogeneous data and constructs a traditional Chinese medicine knowledge graph. To overcome the limitation of traditional knowledge graphs having only single relationships, the relationships in the knowledge graph constructed by this system are hierarchical semantic relationships. Specifically, this module employs a multi-agent collaborative framework, which includes a data processing agent, a knowledge extraction agent, a relational semantic analysis agent, and a knowledge verification agent. Among these, the relational semantic analysis agent is crucial for forming hierarchical semantic relationships. It automatically analyzes and defines the basic relationships (such as "influence") extracted from the text by the knowledge extraction agent, based on a pre-defined medical knowledge rule base, into hierarchical semantic relationships with clear levels or logical directions. Furthermore, this module can also perform a knowledge verification step, logically verifying and correcting the extracted triples through a predefined ontology library and entity type constraints to ensure the accuracy of the graph.

[0011] The knowledge representation learning module is used to learn the embedding representations of entities and hierarchical semantic relationships in the traditional Chinese medicine knowledge graph within a dual quaternion space. The dual quaternion space, due to its ability to uniformly express transformations such as rotation and translation, is naturally suitable for modeling complex relationships such as hierarchy and dependency in the graph. This module preferably employs a dual quaternion graph attention network, aggregating neighbor node information through a graph attention mechanism. Specifically, for an entity... It is a vector that aggregates neighborhood information. It can be calculated in the following ways:

[0012] in, For entities The set of first-order neighbors, For entities With its neighboring three groups Normalized attention weights between them For neighboring entities The module can further nonlinearly fuse the entity's own embedding representation with the vector that aggregates neighborhood information after aggregating neighbor information to generate an updated and more informative entity embedding representation.

[0013] The drug screening prediction module is used to calculate the prediction score between traditional Chinese medicine and heart failure-related targets based on the embedded representation of the entities and relations. To fully utilize the expressive power of the dual quaternion space, this module can employ dual quaternion interaction operators to calculate heart failure-related targets. The final embedding representation With Chinese medicine The final embedding representation The correlation between them is used to obtain the predicted score. The calculation method is as follows:

[0014] Wherein, * represents the dual quaternion interaction operator. The score intuitively reflects the potential association strength between traditional Chinese medicine and target points.

[0015] The second aspect of the application provides an AI-assisted traditional Chinese medicine drug screening method, which comprises the following steps: A data processing step, collecting and preprocessing the preset multi-source heterogeneous data; A knowledge graph construction step, receiving the preprocessed multi-source heterogeneous data and constructing a traditional Chinese medicine knowledge graph, wherein the relationship in the traditional Chinese medicine knowledge graph is a hierarchical semantic relationship; A knowledge representation learning step, embedding representation learning of entities and hierarchical semantic relationships in the traditional Chinese medicine knowledge graph in a dual quaternion space to obtain embedding representations of entities and relationships; A drug screening prediction step, calculating the predicted score between traditional Chinese medicine and heart failure related target points based on the embedding representations of entities and relationships.

[0016] Compared with the prior art, the application can more deeply and accurately capture the complex associations in traditional Chinese medicine theory and modern pharmacological knowledge by constructing a knowledge graph containing hierarchical semantic relationships and using a dual quaternion graph attention network for representation learning and reasoning, thereby providing more reliable decision support for traditional Chinese medicine drug screening.

[0017] The application provides an AI-assisted traditional Chinese medicine drug screening system and method. The application has the following beneficial effects: 1. The application greatly improves the semantic accuracy and information depth of the knowledge graph. By introducing a relationship semantic analysis agent, the technical solution is no longer satisfied with constructing a flat network containing only "acting on" such vague associations. It can automatically upgrade these basic relationships to hierarchical semantics with clear biological logic, such as "upstream activation" or "synergistic inhibition". This makes the constructed knowledge graph not a simple entity link set, but a deep knowledge model that can more realistically simulate the complex causal chain and regulatory hierarchy within the biological system, providing an unprecedented high-quality data foundation for all subsequent analysis.

[0018] 2. This invention innovatively solves the mathematical challenge of effectively representing complex biological relationships. The solution abandons the traditional Euclidean space, creatively embedding entities and relationships into a dual quaternion algebraic space. This space can naturally model transformations such as rotation, translation, and spirals within a unified framework, which forms a profound geometric isomorphism with the structural features of biological networks, such as hierarchical dependencies, transitivity, and asymmetric regulation. This leap in mathematical expressive power allows the model to capture key structural information easily lost in traditional vector spaces, thus generating a more sophisticated and powerful entity embedding representation.

[0019] 3. This invention significantly improves the accuracy of drug-target association prediction. This effect is a direct result of the first two innovations: based on a knowledge graph with precise semantics and more expressive dual quaternion embeddings, the dual quaternion graph attention network used in this invention can more acutely learn deep associations between entities. When aggregating neighborhood information, this network can intelligently assign different weights to relationships of different levels and types, thereby optimizing the information concentration in the final entity representation. This directly translates into higher prediction fidelity in drug screening tasks, effectively reducing false positives and uncovering more non-obvious potential relationships.

[0020] 4. This invention achieves automation and scalability throughout the entire process of constructing a knowledge graph for traditional Chinese medicine by designing a multi-agent collaborative framework. Compared to traditional methods that rely on manual annotation and verification by a large number of domain experts, this solution decouples and assigns tasks such as knowledge extraction, semantic deepening, and logical verification to agents with different roles, forming a highly efficient and self-iterable knowledge production pipeline. This not only significantly shortens the research and development cycle but also ensures that the system can reliably and efficiently expand when faced with massive amounts of new literature and data, maintaining the real-time nature and completeness of the knowledge.

[0021] 5. This invention significantly enhances the interpretability of prediction results, providing clear guidance for drug mechanism research. When the system predicts a potential drug-target association, it is not based on an uninterpretable black box model. Because the underlying relationships of the knowledge graph itself contain clear biological meaning, researchers can trace back along the path with the highest prediction score to clearly see which pathway(s) composed of specific relationships such as "upstream activation" and "downstream inhibition" contribute the most to the final prediction result. This not only answers "what is effective" but also provides a verifiable scientific hypothesis about "how it works," greatly facilitating subsequent experimental design. Attached Figure Description

[0022] Figure 1 This is the overall system architecture of the present invention; Figure 2 This is a schematic diagram of the multi-agent cooperation framework of the present invention; Figure 3 This is a schematic diagram illustrating the generation of hierarchical semantic relationships in this invention; Figure 4 This is a schematic diagram of the dual quaternion graph attention network layer structure of the present invention.

[0023] Figure 5 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0024] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] See attached document Figure 1 , Figure 1 This is a functional module diagram of an AI-assisted traditional Chinese medicine (TCM) drug screening system according to an embodiment of the present invention. The present invention provides an AI-assisted TCM drug screening system that aims to accurately reveal the complex mechanisms of action between TCM and diseases by constructing a knowledge graph with hierarchical semantic relationships and utilizing dual quaternion space for knowledge representation learning.

[0026] The system may include: a) a data processing module, b) a knowledge graph construction module, c) a knowledge representation learning module, and d) a drug screening prediction module. These modules operate with the support of hardware processors and memory, and cooperate with each other to form a complete technical chain from raw data to drug screening prediction.

[0027] In the initial stage of system operation, data processing module A is responsible for the unified collection and preprocessing of heterogeneous data from multiple external sources. These data sources are diverse, including professional databases of traditional Chinese medicine containing information on Chinese herbs, ingredients, and prescriptions; biomedical databases containing information on diseases, genes, targets, and signaling pathways; and massive amounts of unstructured scientific research literature. This module performs data cleaning to remove redundant noise, entity alignment to unify entity naming from different sources, and format standardization, providing a high-quality, standardized data source for subsequent knowledge extraction.

[0028] Subsequently, the preprocessed data is transmitted to the knowledge graph construction module b. This module is one of the core components of the technological breakthrough achieved in this invention; its function is to automatically construct a traditional Chinese medicine knowledge graph. Unlike traditional graphs that only contain flat relationships, the "edges" in this knowledge graph are defined as hierarchical semantic relationships, which can more profoundly reflect the complex logic of the principal, assistant, adjuvant, and guide principles in traditional Chinese medicine theory or the upstream and downstream regulation in modern pharmacology. This module efficiently transforms unstructured text into structured knowledge triples containing hierarchical semantics through an internally integrated multi-agent collaborative framework.

[0029] After the knowledge graph is constructed, the knowledge representation learning module c receives the graph as input. The innovation of this module lies in its mapping of entities (such as traditional Chinese medicine, ingredients, and targets) and complex hierarchical semantic relationships in the graph to a high-dimensional dual quaternion embedding space. By employing a specially designed dual quaternion graph attention network, this module can capture and quantify the transitive, dependent, and spatial coupling relationships between entities, ultimately generating a unique, highly condensed embedding representation for each entity and relationship. This process is a key step in the deep encoding and mathematical abstraction of the graph information.

[0030] Finally, the entity and relation embeddings output by the knowledge representation learning module c are fed into the drug screening prediction module d. This module uses these embeddings to perform specific drug screening tasks. For example, when predicting the association between a specific traditional Chinese medicine and targets related to heart failure, this module calculates a prediction score between the two embeddings using a dual quaternion interaction operator. The calculation method can be expressed as:

[0031] in, Represents targets related to heart failure The final embedding representation, Representative Chinese medicine The final embedding representation is given by , where * represents the predefined dual quaternion interaction operator. This score intuitively quantifies the potential interaction strength between traditional Chinese medicine and the target, providing data-driven decision-making support for drug development.

[0032] In summary, this system forms a complete and logically rigorous closed loop, from data processing and hierarchical knowledge graph construction to representation learning in the dual quaternion space and finally to quantitative prediction. This system deeply integrates the complex theoretical knowledge of traditional Chinese medicine with advanced artificial intelligence technology, effectively addressing the limitations of existing technologies in processing complex traditional Chinese medicine knowledge and significantly improving the accuracy and efficiency of drug screening.

[0033] In one specific embodiment of the present invention, the appendix will be... Figure 1The functions and implementation of data processing module a shown in the diagram are explained in detail. This module is the data entry point and the foundation for quality assurance of the entire system; the quality of its processing results directly affects the accuracy of all subsequent analyses.

[0034] See attached document Figures 2 to 4 The data processing module A first performs the data acquisition task. To ensure the comprehensiveness and authority of the knowledge graph, its acquisition scope covers multiple dimensions and data sources with different structures. Specifically, these data sources include: First, traditional Chinese medicine professional knowledge bases, such as the Traditional Chinese Medicine Database and the Encyclopedia of Traditional Chinese Medicine, from which structured information on Chinese medicine, active ingredients, prescription compatibility, and properties and meridian tropism are obtained; Second, modern biomedical databases, such as the Kyoto Encyclopedia of Genes and Genomes, Human Mendelian Inheritance Online, and GeneCards, from which information on the disease mechanism, related genes, protein targets, and signaling pathways of heart failure is obtained; Third, scientific research literature text databases, such as PubMed and CNKI, from which massive amounts of unstructured latest research papers, clinical reports, and reviews are obtained. These documents are key information carriers for discovering new knowledge and new connections.

[0035] After data acquisition is completed, data processing module A immediately initiates the data preprocessing process, which aims to transform the raw, unrefined data into a clean, standardized, high-quality dataset that can be directly used by subsequent modules. This process includes several key steps.

[0036] The first step is data cleaning. This step aims to improve the intrinsic quality of the data. Specific operations include identifying and removing completely duplicate data entries, correcting common spelling errors in the literature using pre-set dictionaries or simple rules, and marking or filling in records with missing data using predetermined strategies.

[0037] The next crucial step is entity alignment. Because data originates from different databases and literature, the same biomedical entity often suffers from synonymy or homonymy. For example, the same protein target may have different identifiers or names in different databases. The entity alignment step establishes a globally unique entity identifier mapping table, linking names from different data sources that refer to the same objective entity (such as "AKT1" and "protein kinase B") to the same standardized entity ID. This process ensures that each node has a unique and clear referent when constructing the knowledge graph subsequently, avoiding knowledge redundancy and logical confusion.

[0038] The final step in the preprocessing process is data standardization. Data processing module a converts all the data, after the aforementioned cleaning and alignment processes—whether from structured databases or extracted from unstructured text—into a unified, predefined data format, such as JSON or XML. This standardized output ensures that the data can be seamlessly and efficiently received and understood by the knowledge graph construction module b, forming the foundation for efficient information flow between modules within the entire system.

[0039] In one specific embodiment of the present invention, the appendix will be... Figure 1 The knowledge graph construction module b shown is described in detail. This module receives normalized data from the data processing module a, and its core task is to construct a well-structured and semantically rich knowledge graph of traditional Chinese medicine. This graph is the cornerstone of all subsequent knowledge discovery and drug screening tasks.

[0040] See attached document Figures 2 to 4 In this invention, the significant feature of the knowledge graph construction module b is that it does not construct a traditional graph containing only simple, flat relationships, but rather generates an advanced knowledge network whose relationships themselves carry hierarchical semantics. To achieve this goal, this module integrates and runs a multi-agent collaborative framework. This framework achieves a high degree of automation and intelligence throughout the entire process by decomposing the complex graph construction task and assigning it to multiple agents with different specializations based on a large language model.

[0041] In a typical processing flow, the knowledge extraction agent first processes the input text data. This agent is responsible for performing basic named entity recognition and relation extraction tasks, identifying key medical entities from sentences, such as traditional Chinese medicine, chemical components, protein targets, signaling pathways, and diseases, and capturing the preliminary, non-directional associations between them to form basic triples, such as (Astragalus membranaceus, acting on, AKT1).

[0042] These basic triples are then passed to the relational semantic analysis agent, which constitutes a key link in the technological innovation of this module. This agent is specifically responsible for deep semantic mining and upgrading of basic relations. After receiving a basic triple, it combines the context of the triple in the original text with a pre-configured rule base that integrates traditional Chinese medicine theory and modern pharmacology knowledge for reasoning. For example, if the original text describes "studies show that Astragalus can inhibit cardiomyocyte apoptosis by activating the AKT1 signaling pathway", the relational semantic analysis agent will accurately transform the aforementioned broad relation (Astragalus, acts on, AKT1) into a hierarchical semantic relation (Astragalus, upstream activation, AKT1) with clear directionality and mechanism. In this way, rich and computable hierarchical relations such as "direct targeting", "indirect regulation", "synergistic effect", and "antagonistic inhibition" are defined and injected into the knowledge graph.

[0043] To ensure the accuracy and logical consistency of the final generated graph, the framework also includes a knowledge verification agent. This agent performs rigorous quality control. On one hand, it checks type constraints based on a predefined ontology, eliminating triples that do not conform to entity type specifications; for example, there cannot be a "targeting" relationship between two "traditional Chinese medicine" entities. On the other hand, it performs logical consistency checks, comparing the entities in the triples with an existing list of standard entities, and deleting invalid triples containing unrecognized or isolated entities, thereby effectively mitigating the "illusion" problem that may arise from large language models.

[0044] Through the collaborative work and layer-by-layer refinement of the aforementioned multi-agent system, the knowledge graph construction module b ultimately outputs a high-quality knowledge graph of traditional Chinese medicine. This map From entity set hierarchical semantic relation set Composition, represented as a series of triples ,in , It is precisely because of the set of relations The rich hierarchical semantics contained within the knowledge graph enable it to provide an unprecedented knowledge foundation that deeply aligns with biomedical logic for the subsequent knowledge representation learning module c.

[0045] See attached document Figures 2 to 4 In one specific embodiment of the present invention, the appendix will be... Figure 1 The knowledge representation learning module c shown is elaborated in detail. This module receives a knowledge graph containing hierarchical semantic relationships generated by the knowledge graph construction module b. Its core task is to transform the symbolic information of the knowledge graph into a mathematical representation in a continuous, computable vector space. The unique feature of the knowledge representation learning module c in this invention is that it does not use a traditional vector space, but instead innovatively embeds all entities and relations in the knowledge graph into a higher-dimensional dual quaternion algebraic space. The reason for choosing this space is that dual quaternions can naturally model complex spatial transformations such as translation, rotation, and dependencies within a unified mathematical framework, which deeply aligns with the hierarchical semantic relationships such as "upstream and downstream regulation" and "direct or indirect effects" inherent in the knowledge graph. A dual quaternion vector From the real part and duality Composition, its form is ,in Dual units and satisfy .

[0046] To efficiently learn embedded representations of entities and relations in this space, this module employs a specially designed dual quaternion graph attention network. This network, through an information aggregation mechanism, enables the representation of each entity node in the graph to incorporate information carried by its neighboring nodes. In an information aggregation layer, for any entity in the graph... The update process of its embedded representation is as follows: First, network computing entities With each of its first-order neighbor triples Attention scores between This score is designed to measure the relationship. Under the influence of the entity with neighboring entities The strength of the association between them. Then, the Softmax function is used to analyze the entities. The attention scores of all neighbors are normalized to obtain the attention weights. The calculation formula is as follows:

[0047] in, Referent entity The set of first-order neighbor triples.

[0048] Then, using the calculated attention weights, the neighboring entities are... Dual quaternion embedding representation A weighted summation is performed to generate an aggregated entity. Vector of all neighborhood information This polymerization process can be represented by the following formula:

[0049] Finally, the entity Its own primitive embedding representation with vectors that aggregate neighbor information Perform a non-linear fusion to generate the updated embedding representation of the entity at the current network level. A specific fusion method is as follows: LeakyReLU LeakyReLU

[0050] in, and It is the weight matrix that can be used for training in the network. It represents the element-wise product between vectors, while LeakyReLU is a non-linear activation function.

[0051] By stacking the aforementioned information aggregation layers in multiple layers, the knowledge representation learning module c enables information to propagate over longer distances within the knowledge graph, thereby capturing higher-order and more complex structural and semantic relationships. After completing the computation at all levels, this module generates a highly condensed final embedding representation for each entity and relation in the knowledge graph, and passes these representations to the drug screening prediction module d, providing a solid mathematical foundation for subsequent accurate predictions.

[0052] In one specific embodiment of the present invention, the appendix will be... Figure 1 The drug screening prediction module d shown is described in detail. This module is the final execution unit of the entire technical chain of this invention. It receives highly condensed entity embedding representations generated by the knowledge representation learning module c, and performs quantitative prediction of the potential correlation between traditional Chinese medicine and targets related to heart failure based on these representations.

[0053] See attached document Figures 2 to 4 The core function of the drug screening prediction module d is to calculate a prediction score. This is useful when evaluating a traditional Chinese medicine entity. For a target entity related to heart failure When considering its potential applications, this module first obtains the final embedding representations of both elements after learning through a multi-layer dual quaternion graph attention network. and It then employs a pre-defined dual quaternion interaction operator* to compute the association score between the two embedding representations. The calculation process is defined by the following formula:

[0054] The score It is a real value, the magnitude of which directly reflects the properties of traditional Chinese medicine. With target The possibility of interaction between them provides a direct and quantifiable basis for subsequent drug screening and mechanism analysis.

[0055] To ensure that the learned embeddings generate meaningful prediction scores, the entire system requires end-to-end training by optimizing a carefully designed loss function. In this invention, the drug screening prediction module d employs a pairwise learning loss function based on Bayesian personalized ranking. Its main optimization objective is to optimize the target point. The basic principle of this function is that, for a given target point... A known positive sample of Chinese herbal medicine that interacts with it. The predicted score should be significantly higher than that of a randomly sampled negative sample of Chinese herbal medicines with no observed interactions. The score. The specific form of the loss function is as follows:

[0056] in, It represents the set consisting of all known positive sample interaction pairs in the training data. This represents the set of negative sample interaction pairs generated through negative sampling techniques, while It is the standard Sigmoid activation function, used to smoothly map score differences to the (0,1) interval.

[0057] Furthermore, to prevent overfitting due to excessive parameters during model training and thus improve the model's generalization ability, this invention introduces an L2 regularization term. This regularization term penalizes all learnable parameters in the model, and is defined as follows:

[0058] in, It represents the set of all learnable parameters in the model, including the embedded representations of all entities and relations, as well as the parameters in the network weight matrix.

[0059] Ultimately, the system's total loss function From the principal loss function and regularization term Weighted summation yields:

[0060] in, is a preset hyperparameter used to control the regularization strength. During the training process, the system adopts optimization algorithms such as gradient descent to minimize this total loss function as the goal, and iteratively updates all parameters in the model until the model converges.

[0061] A method of an AI-assisted traditional Chinese medicine drug screening system described below can be mutually corresponding and referred to with an AI-assisted traditional Chinese medicine drug screening system described above.

[0062] Refer to the appendix Figure 5 , specific steps: Step 1: Processing of multi-source heterogeneous data The present invention first executes a data processing step, which is the basis for all subsequent analyses. The system automatically collects heterogeneous data related to traditional Chinese medicine and heart failure from multiple preset and publicly available data sources. These data sources mainly include but are not limited to: 1) Biomedical literature databases such as PubMed and CNKI to obtain unstructured text information; 2) Traditional Chinese medicine professional databases such as the Traditional Chinese Medicine Systems Pharmacology Database and Analysis Platform to obtain structured data such as traditional Chinese medicine, chemical components, and their associated targets; 3) General drug and protein databases such as DrugBank and UniProt to supplement and verify entity information.

[0063] After the collection is completed, the data processing module performs a series of preprocessing operations on these raw data. This operation includes data cleaning, such as removing irrelevant information and correcting format errors; entity unification, that is, normalizing different names (such as "Astragalus membranaceus" and "Huangqi") pointing to the same entity into a unique identifier; and preliminary structured conversion, initially organizing unstructured text into a format suitable for information extraction to prepare for the automated construction of a knowledge graph in the next step.

[0064] Step 2: Construction of a knowledge graph based on multi-agent collaboration After obtaining the normalized data, the present invention initiates the knowledge graph construction step, the core of which is to use a multi-agent collaboration framework to automatically construct a traditional Chinese medicine knowledge graph containing hierarchical semantic relationships. This step is specifically composed of the following collaborating sub-processes: First, the knowledge extraction agent analyzes the preprocessed data and performs named entity recognition and basic relationship extraction tasks. This stage aims to identify key entities such as traditional Chinese medicine, chemical components, protein targets, diseases, etc. from the text, and extract the basic and relatively broad associations between them to form basic triples such as (Astragalus membranaceus acts on AKT1).

[0065] Subsequently, these basic triples are passed to a relational semantic analysis agent, which is a key innovation of this invention. This agent does not simply connect entities; instead, it invokes a pre-defined medical knowledge rule base and, combined with the contextual information of the triples, performs deep reasoning and semantic upgrades on the basic relationships. For example, it can analyze that the effect of "Astragalus" on "AKT1" is not a direct binding, but rather achieved by regulating a certain upstream signaling molecule, thus refining the ambiguous "acting on" relationship into a "upstream activation" relationship with a clear biological direction and hierarchy.

[0066] Finally, all semantically upgraded hierarchical triples are sent to a knowledge verification agent. This agent checks the logical consistency and factual accuracy of the triples based on a pre-loaded biomedical ontology library and entity type constraints. For example, it verifies whether the two ends of an "inhibition" relationship conform to the reasonable types of "drug-target" or "protein-protein". Through this verification and correction process, the system can filter out erroneous or illogical knowledge, ensuring that the final output knowledge graph has high accuracy and reliability.

[0067] Step 3: Knowledge Representation Learning Based on Dual Quaternion Graph Attention Network After the knowledge graph is constructed, this invention enters the knowledge representation learning step, aiming to transform the symbolic knowledge in the graph into a machine-processable, information-condensed numerical representation. This invention innovatively chooses to perform this process in the dual quaternion space. First, all entities (such as traditional Chinese medicine and targets) and hierarchical relationships in the knowledge graph are initialized as embedding vectors in the form of dual quaternions.

[0068] The device in this embodiment can be used to execute the above method embodiments, and its principle and technical effects are similar, so they will not be described again here.

[0069] After the knowledge graph is constructed, this invention enters the knowledge representation learning step, aiming to transform the symbolic knowledge in the graph into a machine-processable, information-condensed numerical representation. This invention innovatively chooses to perform this process in the dual quaternion space. First, all entities (such as traditional Chinese medicine and targets) and hierarchical relationships in the knowledge graph are initialized as embedding vectors in the form of dual quaternions. Next, the system employs a multi-layered dual quaternion graph attention network to learn and optimize these initial embeddings. In each layer of the network, for each entity node in the graph, the attention mechanism calculates the association strength between the entity and all its neighboring entities, i.e., the attention weight. This weight is dynamically learned, enabling the model to understand which neighboring nodes are more important under different relational paths. Subsequently, the model performs weighted aggregation of the dual quaternion embeddings of neighboring entities based on the calculated attention weights, forming a context vector that summarizes the neighborhood environment. Finally, through a specific fusion function, the entity's own embedding is combined with its aggregated neighborhood context vector to generate the entity's updated embedding representation in the current network layer. This process iterates layer by layer in the network, enabling entities to capture multi-level structural information from nearest to farthest neighbors, ultimately yielding a final embedded representation containing rich graph structure and semantic information. and .

[0070] Step 4: Drug screening and prediction based on embedding representation This is the final execution step of the invention. The system utilizes the high-quality entity embeddings learned in step three to quantitatively evaluate the potential relationship between traditional Chinese medicine and targets related to heart failure. When it is necessary to predict a traditional Chinese medicine entity... For a target entity When the system is in operation, it calls a pre-defined dual quaternion interaction operator* to compute the final embedding representation of both. and Interaction score between The score is a real number, and its magnitude is directly proportional to the probability that the two are related.

[0071] To enable the model to make accurate predictions, the entire system minimizes a comprehensive loss function. End-to-end training is performed. The loss function mainly consists of two parts: one part is the pairwise learning loss based on the Bayesian personalized ranking idea. The goal is to ensure that the scores of known positive sample pairs are higher than those of randomly sampled negative sample pairs; another part is the L2 regularization term. This is used to prevent model overfitting. By continuously optimizing the loss function using optimization algorithms such as gradient descent, all parameters of the model (including the dual quaternion embeddings of all entities) are continuously adjusted until the model converges, enabling reliable predictions of unknown traditional Chinese medicine-target pairs.

[0072] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An AI-assisted traditional Chinese medicine drug screening system, characterized in that, Includes the following steps: The data processing module is used to collect and preprocess pre-defined multi-source heterogeneous data; The knowledge graph construction module is used to receive the preprocessed multi-source heterogeneous data and construct a traditional Chinese medicine knowledge graph, wherein the relationships in the traditional Chinese medicine knowledge graph are hierarchical semantic relationships; The knowledge representation learning module is used to perform embedding representation learning on entities and hierarchical semantic relations in the traditional Chinese medicine knowledge graph in the dual quaternion space to obtain the embedding representation of entities and relations. The drug screening prediction module is used to calculate the prediction score between traditional Chinese medicine and heart failure-related targets based on the embedded representation of the entities and relationships.

2. The AI-assisted traditional Chinese medicine drug screening system according to claim 1, characterized in that, The knowledge graph construction module is specifically used for: A multi-agent collaborative framework is used for knowledge extraction and graph construction; the multi-agent collaborative framework includes a relational semantic analysis agent, which is used to define the basic relations extracted from the text as the hierarchical semantic relations based on a preset medical knowledge rule base.

3. The AI-assisted traditional Chinese medicine drug screening system according to claim 1, characterized in that, The knowledge representation learning module employs a dual quaternion graph attention network, which aggregates neighbor node information through a graph attention mechanism to update the entity's own embedding representation.

4. The AI-assisted traditional Chinese medicine drug screening system according to claim 3, characterized in that, The graph attention mechanism calculates attention weights to sum the neighbor entity information in a weighted manner, generating a vector that aggregates neighborhood information. The calculation method is as follows: ;in, For entities The set of first-order neighbors, For entities With its neighboring three groups Normalized attention weights between them For neighboring entities Embedded representation.

5. The AI-assisted traditional Chinese medicine drug screening system according to claim 1, characterized in that, The drug screening and prediction module is specifically used for: The correlation between the final embedding representations of heart failure-related targets and the final embedding representations of traditional Chinese medicines is calculated using dual quaternion interaction operators to obtain the prediction score. The calculation method is as follows: ;in, Targets related to heart failure The final embedding representation, Traditional Chinese medicine The final embedding representation is *, which is the dual quaternion interaction operator.

6. The AI-assisted traditional Chinese medicine drug screening system according to claim 3, characterized in that, The dual quaternion graph attention network is also used to nonlinearly fuse the entity's own embedding representation with a vector that aggregates neighborhood information to generate an updated entity embedding representation.

7. The AI-assisted traditional Chinese medicine drug screening system according to claim 2, characterized in that, The multi-agent collaboration framework also includes a data processing agent, a knowledge extraction agent, and a knowledge verification agent, which work together to automate the construction of the knowledge graph.

8. The AI-assisted traditional Chinese medicine drug screening system according to claim 1, characterized in that, The multi-source heterogeneous data includes at least one of the following: a traditional Chinese medicine knowledge base, a biomedical database, and a scientific literature database.

9. The AI-assisted traditional Chinese medicine drug screening system according to claim 2, characterized in that, The knowledge graph construction module also includes a knowledge verification step, which is used to perform logical verification and correction on the extracted triples according to the predefined ontology library and entity type constraints, so as to ensure the accuracy of the traditional Chinese medicine knowledge graph.

10. A method for an AI-assisted traditional Chinese medicine drug screening system according to any one of claims 1-9, characterized in that, Includes the following steps: The data processing steps involve collecting and preprocessing pre-defined multi-source heterogeneous data. The knowledge graph construction step involves receiving the preprocessed multi-source heterogeneous data and constructing a traditional Chinese medicine knowledge graph, wherein the relationships in the traditional Chinese medicine knowledge graph are hierarchical semantic relationships. The knowledge representation learning steps involve learning the embedding representation of entities and hierarchical semantic relations in the traditional Chinese medicine knowledge graph within the dual quaternion space to obtain the embedding representation of entities and relations. The drug screening prediction step calculates the prediction score between traditional Chinese medicine and heart failure-related targets based on the embedded representation of the entities and relationships.