Treatment target prediction engine system based on omics genetic evidence

By developing a therapeutic target prediction engine system based on omic genetic evidence, integrating multimodal functional genome data and protein interaction networks, the problem of integrating genome and functional genome resources in the existing technology is solved, and more accurate target prediction and drug reuse analysis are achieved.

CN119993255AInactive Publication Date: 2025-05-13RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411875629.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively integrate genomic and functional genomic resources, resulting in multiple challenges in the accurate prediction of genetic targets, including accurate annotation of non-coding region sites, cell type-specific gene regulation, and the identification of distal mechanisms of action.

Method used

Develop a therapeutic target prediction engine system based on omic genetic evidence, which includes a knowledge base related to biomedical ontology and network evidence, a target prediction main engine, a target search main engine and a human-computer interaction module. By integrating multimodal functional genomic data and protein interaction network, it provides more accurate disease association site annotation and target prediction.

Benefits of technology

The system can provide high confidence multiomic regulation and interaction relationships, significantly improve the prediction accuracy of therapeutic targets, cover a wider range of target candidates, support drug reuse potential analysis, and use generative AI to accelerate data integration and feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993255A_ABST
    Figure CN119993255A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bioinformatics data knowledge interpretation, in particular to a therapeutic target prediction engine system based on omics genetic evidence, comprising: a biomedicine ontology and network evidence related knowledge base; the target prediction main engine is used for providing a function of predicting a treatment target based on genome summarized data; the target retrieval main engine is used for providing a treatment target online query function based on genome or genome summarized data; and the man-machine interaction module is used for receiving data input by a user, issuing a specified command to the target sub-engine according to the selected target sub-engine and related parameters, and feeding back a result of receiving online prediction or data retrieval from the target sub-engine. According to the prediction engine, the heredity target prediction and retrieval function can be performed from the beginning, the quantitative recommendation result of the related target and the drug reutilization analysis result of the pathway intersection network are acquired, and a comprehensive and simple method is provided for researchers to perform treatment target prediction based on omics heredity evidence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics data knowledge interpretation, and in particular to a therapeutic target prediction engine system based on omics genetic evidence. Background Art

[0002] Genetic targets are candidate targets obtained by analyzing genetic evidence, and they play a vital role in drug development. Compared with drugs that lack genetic target support, the success rate of drug development with genetic target support is significantly higher, especially those targets that have a causal relationship with the occurrence and development of the disease, which makes this success rate gap more obvious. This shows that target selection and verification based on genetic evidence is one of the key means to improve the success rate of drug research and development.

[0003] Although genetic evidence is now considered an important part of drug development, how to effectively use this information is not a simple process and requires a better understanding of the meaning of specific evidence and the best way to use it, especially for genetic association information from genome-wide association studies (GWAS). GWAS can identify a large number of statistically robust genetic associations, and in recent years, great progress has been made in the functional interpretation of single nucleotide polymorphism sites (SNPs). GWAS provides a huge opportunity for more successful drug development.

[0004] GWAS mainly identifies common SNPs, which usually have small genetic effects. Among them, the number of variants located in protein coding regions is relatively small, but according to the sequence-structure-function association, these variants can have structural and functional effects on proteins, such as TYK2 variants associated with autoimmune diseases; the remaining variants are generally located in non-coding regions, but can also affect gene regulation. Therefore, how to link the disease-related sites generated by GWAS (especially those located in non-coding regions) with the specific genes and pathways responsible for these genetic associations, so as to accurately predict genetic targets, is an extremely complex task. At present, this field still lacks sufficient research tools and data resources, which limits scientists' exploration and breakthroughs in this field.

[0005] Usually, when studying the relationship between non-coding region sites and gene expression, researchers will annotate these sites to adjacent genes. However, this method may ignore the regulatory effects of some sites on distal gene expression, leading to unavoidable false negative results. In addition, non-coding sites can also regulate gene expression by changing the three-dimensional structure of chromatin, and these regulatory methods may be cell type or state specific. In order to overcome the above difficulties and achieve accurate annotation of non-coding region sites, cell type-specific functional genomic data are needed, including but not limited to chromatin interaction data and genetic regulation data. However, these data are often scattered in different studies, with problems such as inconsistent formats and serious batch effects, which increases the difficulty of data integration.

[0006] In recent years, data in the biomedical field has exploded, and both basic research exploring disease mechanisms and clinical research based on patients have generated massive amounts of data. However, with the surge in data volume, the complexity of data analysis has also increased. How to extract useful features and information from them and use them for the prediction of therapeutic targets has become a major challenge.

[0007] Although human genetics has obvious potential in identifying effective drug targets, the breadth and depth of its application have not yet been fully explored. The prediction process of genetic targets mainly relies on the disease genetic association information identified by GWAS, and requires the integration of multi-level functional genomic data and protein interaction data, but this process faces many challenges:

[0008] (1) Associating disease-associated variants with effector genes: Disease-associated variants are mainly located in the non-coding genome. How to accurately link them to effector genes is a major challenge in current research.

[0009] (2) Cell-specific gene regulation: Gene expression is tissue / cell specific. Different types of tissues / cells may produce different regulatory patterns on gene promoters under different conditions due to non-coding mutations.

[0010] (3) Distal effects of non-coding variants: Current studies on the relationship between non-coding variants and gene expression are mostly based on neighboring genes. However, the highly folded nature of chromatin within the genome means that these variants may span multiple genes and have regulatory effects on distal target genes. Therefore, it is necessary to combine remote regulatory genomic datasets for precise identification of target genes.

[0011] (4) Silencing function of non-coding variants: Most existing tools focus on the interaction between non-coding sites and promoters of effector genes (upregulating gene expression), while their silencing effect on genes as silencers has received less attention;

[0012] (5) Heterogeneity among datasets: Techniques such as promoter capture Hi-C, ABC (activity-by-contact), and QTL (quantitative trait loci) mapping have generated a large amount of interaction data, which can be used to identify the regulatory relationship between promoters and distal non-coding sites. However, these data come from different studies, with inconsistent processing methods, different genome versions, certain batch effects, and lack of unified standards;

[0013] (6) Differences between transcript levels and protein abundance: Transcript expression levels cannot fully represent the protein abundance of effector genes, and there are also many variabilities in the protein translation process. Therefore, protein abundance data (pQTL) and protein interaction data (such as from the STRING database) are necessary supplements;

[0014] (7) Analysis of drug repurposing potential: New drug development is a long, complex, time-consuming and labor-intensive process. Therefore, when predicting genetic targets for different diseases, it is crucial to evaluate the therapeutic potential of existing approved drugs for other diseases in treating the target disease. This means that old drugs can be quickly applied to new diseases. Existing tools usually lack corresponding functional support in this regard.

[0015] (8) The potential of generative AI has not yet been fully explored: In recent years, with the introduction of attention mechanisms and the emergence of Transformer deep learning models that can capture contextual relationships, coupled with the rapid development of GPUs, generative AI has demonstrated capabilities that far exceed traditional machine learning models. AI has been deeply integrated into various areas of our lives, such as voice assistants, intelligent conversations, and autonomous driving. These models have excellent feature extraction capabilities and can effectively integrate modeling to generate outputs. Therefore, the use of AI for therapeutic target prediction has great potential.

[0016] At present, there is still a lack of effective tools to integrate genomic and functional genomic resources to predict drug therapeutic targets. How to use the integrated data for knowledge discovery and genetic target prediction is still a key issue that needs to be solved. To this end, it is necessary to develop better algorithms and technologies to process and analyze large-scale genomic and epigenomic data, so as to provide more accurate disease-associated site annotation and target prediction. Summary of the invention

[0017] The purpose of the present invention is to overcome the difficulties of the above-mentioned prior art in integrating genomic and functional genomic resources into effective tools, thereby providing a therapeutic target prediction engine system based on omics genetic evidence.

[0018] The purpose of the present invention can be achieved by the following technical solutions:

[0019] A therapeutic target prediction engine system based on omics genetic evidence, the therapeutic target prediction engine system comprising:

[0020] Biomedical ontology and network evidence related knowledge base: The datasets provided include enhancer-promoter maps, genome-wide promoter interaction capture technology, gene expression / protein abundance, gene / protein interaction networks, biomedical ontology data suitable for knowledge interpretation of omics data, and drug-targets for immune-mediated diseases;

[0021] Target prediction main engine: provides prediction functions for therapeutic targets based on genomic summary data;

[0022] Target search main engine: provides online query function for therapeutic targets based on genome or genome summary data;

[0023] Human-computer interaction module: used to receive data input by the user, issue specified commands to the target sub-engine according to the selected target sub-engine and related parameter types; and feedback the results of data retrieval or online prediction received from the target sub-engine to the user.

[0024] As a preferred technical solution, the biomedical ontology and network evidence related knowledge base includes:

[0025] A knowledge base related to multimodal functional genomics evidence, including enhancer-promoter interaction evidence, including enhancer-promoter maps from the ENCODE and Roadmap databases; promoter capture Hi-C evidence, including manually screened promoter capture Hi-C datasets; and gene expression / protein abundance evidence, including quantitative trait loci data;

[0026] Knowledge sub-base related to gene / protein interaction network evidence, data sources include: STRING database, KEGG database;

[0027] The biomedical ontology knowledge sub-base provides knowledge interpretation of omics summary data applicable to genes, SNPs, intergenic regions and protein domains, covering knowledge areas including: function, pathway, transcriptional regulator, disease, phenotype, drug, domain and evolution;

[0028] The ontological knowledge sub-base related to regulatory interaction gene enrichment is used to perform ontology enrichment analysis of potential regulatory target genes, including function, phenotype, disease and protein domain;

[0029] The drug target knowledge sub-database with clinical theoretical support provides clinically verified target information from sources including: ChEMBL database, DrugBank database, Open Targets database, PubChem database and TCMID database;

[0030] The knowledge base for quantitative recommendation of genetic targets for immune-mediated diseases is derived from the Priority index database. It uses immune genomic datasets from various immune cell types or states, and makes full use of genetic information including disease GWAS, functional immunogenomics, ontology annotations and protein interactions for quantitative recommendation of genetic targets for immune-mediated diseases.

[0031] As a preferred technical solution, the target prediction main engine specifically includes:

[0032] The target prediction sub-engine based on the quantitative recommendation of the priority index is used to receive the user input SNP and its significance of association with the disease, combine the genome proximity, gene expression / protein abundance and the whole genome promoter interaction capture technology knowledge sub-base to define the core genes, and connect the core genes with the peripheral genes to complete the target prediction based on the quantitative recommendation of the priority index;

[0033] The target prediction sub-engine based on the pathway intersection network receives the disease-related SNPs and significance information and specified information input by the user, completes the quantitative recommendation of genes, and generates a pathway intersection network gene list containing a set of high-priority genes;

[0034] The drug repurposing analysis sub-engine based on the pathway intersection network receives the SNPs and their significance of association with the disease input by the user, obtains the pathway intersection network results through the target prediction sub-engine based on the pathway intersection network, receives the specified information, generates the pathway intersection network of the highly prioritized genes, and completes the drug repurposing analysis for the prioritized genes;

[0035] The target prediction engine driven by cross-disease genetic pleiotropy receives the SNPs and their significance of association with diseases input by users, combines the pathway intersection network results obtained by the target prediction engine based on pathway intersection networks, introduces multi-trait scores to quantify the degree to which target genes are highly scored in multiple traits, and uses genes with multi-trait scores to mark the gene network merged by KEGG pathways;

[0036] The target prediction sub-engine based on generative AI receives characteristic trait data input by the user and identifies potential target genes related to it through deep learning algorithms.

[0037] As a preferred technical solution, the target prediction sub-engine based on priority index quantitative recommendation has the following specific workflow:

[0038] Definition of genomic core genes: neighboring genes nGene are defined based on genomic proximity and genomic organization; based on the data of promoter capture Hi-C studies, whether gene promoters physically interact with genomic regions carrying SNPs is identified, and chromatin conformation genes cGene are defined based on chromatin physical interactions; expression-related genes eGene are defined based on summary data from eQTLmapping, and the co-localization method coloc is used to estimate whether a certain SNP plays a causal role in both GWAS and eQTL studies through a Bayesian framework; when the estimated causal probability exceeds the set threshold, the SNP is identified as an expression-related gene eGene, and is scored based on the SNP with the strongest association;

[0039] The scoring of core genes takes into account the following factors: disease genetic association; SNP distance window representing genomic proximity; significance level of genetic association between gene expression in eQTL dataset or protein abundance in pQTL dataset; strength of physical interaction between gene promoter and genomic region containing SNP in PCHi-C dataset; empirical distribution function estimated based on all SNP-gene pairs;

[0040] Using ontology annotations at the gene level, annotation predictors related to immune function / dysfunction were further defined: immune function genes fGene, annotated using Gene Ontology, annotated genes to immune responses with experimental or clinical evidence; disease genes dGene, annotated using the online genetic human genetics database and disease ontology, annotated genes that can cause immune system diseases; immune phenotype genes pGene, annotated abnormal immune system, blood and hematopoietic tissues using the human phenotype ontology, and immune / hematopoietic system phenotypes using the mammalian phenotype ontology;

[0041] The core genes were identified using neighboring genes nGene, chromatin conformation genes cGene and expression-related genes eGene, and immune function genes fGene, disease genes dGene and immune phenotype genes pGene were used to perform functional annotations on the core genes; the output results included a chart reflecting the core genes and their quantitative genetic association scores between diseases and the core genes, as well as a list of evidence showing the data sets used to define the core genes;

[0042] Using core genes as seed nodes, applying knowledge of protein interactions, and using a restarted random walk algorithm to identify peripheral genes under network influence: relying on a knowledge sub-base related to gene / protein interaction network evidence to establish a protein interaction network, incorporating high-confidence interaction relationships in the STRING database, for each type of seed gene, a random walk algorithm is used to identify non-seed genes with network influence; for each proximity, gene expression / protein abundance, and genome-wide promoter interaction capture technology dataset, a predictor containing core and peripheral genes is generated, with an affinity score used to quantify their network connectivity with the core genes; finally, a ranked list of core and peripheral genes based on affinity scores that quantify network connectivity with the input core genes is output;

[0043] The following information is specified in the target prediction sub-engine based on the quantitative recommendation of the priority index: whether to include SNPs in linkage disequilibrium, and if so, based on which population the calculation is based; how to identify core genes based on genomic proximity, gene expression / protein abundance, and genome-wide promoter interaction capture technology; network information for the identified core genes and peripheral genes;

[0044] For each predictor in the gene-predictor matrix, the gene affinity scores were converted to class P values; the class P values ​​were combined between the individual predictors for each gene using Fisher's combination method; the resulting combined class P values ​​were renormalized to priority ranking scores;

[0045] The results are returned to the human-computer interaction module, including input data, priority scores of genes on the genome, gene types, and evidence overview.

[0046] As a preferred technical solution, the target prediction sub-engine based on the pathway intersection network has the following specific workflow:

[0047] Receive the ranked gene list input by the user from the human-computer interaction module, identify a group of genes with high priority and strong correlation with each other, and connect them according to the priority scores;

[0048] The significance of the identified pathway intersection network was evaluated under random conditions by using a degree-preserving permutation test. Multiple permutation tests were performed on the sorted gene list while keeping the degree constant, so that the expected intersection and the intersection to be evaluated had the same / similar number of genes, and the expected intersection was used as the null distribution to estimate the significance of the identified intersection.

[0049] Identify target genes in the prioritized gene list of immune traits, perform gene clustering and visualization, integrate the generated graph with druggability pocket data, and estimate the probability of each hexagon containing a druggable gene; perform enrichment analysis on each clustered gene cluster to identify the enriched approved drug indications and KEGG pathways, and then perform functional annotation on the gene cluster;

[0050] The identified pathway intersection results are returned to the human-computer interaction module in the form of a table of intersection genes and subnetworks.

[0051] As a preferred technical solution, the drug repurposing analysis sub-engine based on the pathway intersection network has the following specific workflow:

[0052] Obtain pathway intersection network results through a target prediction sub-engine based on pathway intersection networks, and receive specified information, including: whether to include SNPs in linkage disequilibrium, and if so, based on which population; how to identify core genes based on genomic proximity, gene expression protein abundance, and genome-wide promoter interaction capture technology; network information used to identify core genes and peripheral genes; confidence of the gene / protein interaction knowledge sub-library used; select the number of genes included in the pathway intersection network and significance information; select the number of genes in the intersection network and the significance information of the identified sub-network; generate pathway intersection networks for high-priority genes, and complete drug repurposing analysis of pathway intersection network genes;

[0053] The identified drug reuse analysis results are returned to the human-computer interaction module in the form of charts.

[0054] As a preferred technical solution, the target prediction sub-engine based on cross-disease genetic pleiotropy drive has the following specific workflow:

[0055] Multi-trait score MRS was used to quantify the extent to which the target gene was highly scored in multiple traits;

[0056]

[0057] In the formula, nTop is the number of times the target gene ranks before the set rank in a certain trait; mRank is the average rank of the target gene; N is the total number of traits included in the analysis;

[0058] Using Fisher's exact test, genes enriched in the intersection were compared with all genes with MRS as the test background to identify individual KEGG pathways;

[0059] Examining the enrichment of pathway intersection network genes at the level of mouse immune-mediated disease phenotypes, drug-interferable gene traits, and immune disease indications by stage and approved therapeutic regimens and druggable pockets;

[0060] Introducing the Multi-Trait Novelty Score (MNS) to quantify how unexplored a target is in most traits:

[0061]

[0062] In the formula, nD represents the number of times the target gene is annotated as a disease gene in a certain trait; nA represents the number of approved drugs for the target gene in these traits; nP represents the number of drugs that the target gene has at least one in the drug development stage in these traits; N represents the total number of traits included in the analysis;

[0063] The identified target prediction results based on cross-disease genetic pleiotropy drivers will be returned to the human-computer interaction module in the form of charts.

[0064] As a preferred technical solution, the target prediction sub-engine based on generative AI has the following specific workflow:

[0065] Based on user-specified information, including the specific type of input characteristic traits and their correlation indicators; selecting the multi-omics data type used to train the generated model; determining the evaluation criteria for the druggability of the target gene; the credibility level of the protein interaction knowledge base used; setting the screening threshold and significance criteria for the target gene; and selecting the number of target genes and related biological pathway information in the output results;

[0066] The generative AI model generates a priority list of candidate targets and predicts the potential biological functions of these targets by learning a multimodal input feature matrix constructed by combining genomic data, network evidence, and ontology annotations at the gene level, and deeply integrating multimodal features. It also identifies target genes that show high clinical application potential in multiple disease contexts by comprehensively analyzing a variety of bioinformatics data. It reveals disease-specific targets for target genes that only appear in specific diseases. It also deeply analyzes the drugability of target genes by combining molecular dynamics simulations.

[0067] The identified target prediction results are fed back to the human-computer interaction module in the form of visual charts.

[0068] As a preferred technical solution, the target search main engine specifically includes:

[0069] The target search sub-engine for immune-mediated diseases receives the user's target search instruction from the human-computer interaction module, and obtains the target priority ranking results of immune-mediated diseases supported by the engine under the navigation tag;

[0070] The target search sub-engine based on the pathway intersection network receives user-specified input information from the human-computer interaction module, and searches and analyzes the targets predicted by multiple traits based on the results of the target prediction sub-engine based on the pathway intersection network;

[0071] The target retrieval sub-engine based on generative AI receives user input from the human-computer interaction module and automatically identifies and predicts potential drug targets based on the generative model.

[0072] As a preferred technical solution, the target search sub-engine for immune-mediated diseases generates target gene ranking results for each disease in two modes, including a complete target gene priority ranking list and an optional number of pathway intersection network gene lists, the modes are: a discovery mode that does not use any known drug target information and a supervision mode that integrates known drug target prediction factors through a random forest algorithm; a downloadable MySQL relational database is provided, as well as detailed documentation on table organization and usage;

[0073] The target search sub-engine based on the pathway intersection network is used to receive user input and specified information from the human-computer interaction module, including different traits for comparison; genes to be evaluated, including all priority genes or all genes identified by the pathway intersection network; choosing to use the discovery mode or the supervision mode for analysis; selecting the index used for priority sorting; based on the results of the target prediction sub-engine based on the pathway intersection network, searching and analyzing the targets predicted by multiple traits;

[0074] The MRS scores of relevant targets in different traits are shown in the figure. Each target provides links to access specific information, including: hyperlinks to GeneCards, Ensembl, HGNC, OMIM, Entrez, and OpenTargets databases to access the target gene sequence, corresponding protein and tertiary structure, related pathways, and disease information; hyperlinks to PDBe-KB and AlphaFold databases to understand the drugability of the target; priority scores of targets in common immune traits; and heat map displays and hyperlinks to gene-related information;

[0075] The target retrieval sub-engine based on generative AI, the target retrieval sub-engine based on generative AI is based on the generative AI large language model, and automatically generates answers according to the user's input information;

[0076] The user fills in the required sub-engine type, input data and expected parameters into the engine's dialog box through text input; the treatment target prediction engine system calls the target retrieval sub-engine based on generative AI for calculation, and feeds the generated result link back to the dialog box.

[0077] Compared with the prior art, the present invention has the following beneficial effects:

[0078] 1) The biomedical ontology and network evidence-related knowledge base of the present invention covers a large number of functional genomic data sets, integrating data of multiple modalities and different batches. These comprehensive and detailed data can provide high-confidence multi-omics regulation and interaction relationships, laying a solid foundation for the prediction of therapeutic targets. In addition, the knowledge base of the present invention is open source, scalable and updated in real time, ensuring the completeness and novelty of the knowledge base.

[0079] 2) The theoretical basis of the target prediction based on the quantitative recommendation of the priority index in the present invention is the whole-genome genetic law model, which believes that complex traits / phenotypes are caused by the joint action of a few core genes and a large number of peripheral genes. Therefore, when identifying therapeutic targets, the present invention not only considers the core genes, but also identifies peripheral genes through the restarted random walk algorithm, and uses ontological annotations at the gene level to further define annotation predictive factors related to immune function / dysfunction, providing more comprehensive and accurate therapeutic target prediction results.

[0080] 3) In the present invention, an algorithm for searching subsets in the gene network obtained by merging all KEGG pathways is developed for target prediction based on pathway intersection networks, and the significance of the identified intersections is estimated. The probability of containing druggable genes in each hexagon is then estimated, and the enriched approved drug indications and KEGG pathways are identified, thereby functionally annotating the gene clusters. The present invention can not only achieve de novo prediction of therapeutic targets based on omics genetic evidence, but also enable users to more deeply explore the pathway information and network information of the target, so that users can explore and analyze personal data in a simpler and more comprehensive way.

[0081] 4) The present invention also provides target prediction driven by cross-disease genetic pleiotropy based on pathway intersection network results, and introduces a multi-trait score to quantify the extent to which a target gene is highly scored in multiple traits, and uses genes with MRS to mark gene networks merged by KEGG pathways; a multi-trait novelty score is introduced to quantify the degree to which a target is unexplored in most traits, in order to analyze the identified pathway intersection network genes and explore the possibility of drug repurposing.

[0082] 5) The present invention combines generative AI to automatically integrate multimodal functional genomic data, identify complex regulatory relationships, and generate accurate hypotheses to identify target genes that show high clinical application potential in multiple disease contexts; reveal disease-specific targets for target genes that only appear in specific diseases; deeply analyze the drugability of target genes in combination with molecular dynamics simulations; and automatically generate answers based on user input information based on the generative AI large language model, and effectively guide users to use the product. Provide innovative technical support for target prediction, query and personalized treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 A schematic diagram of the structure of a therapeutic target prediction engine system based on omics genetic evidence of the present invention;

[0084] Figure 2 Operational flow chart for target prediction main engine. DETAILED DESCRIPTION

[0085] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0086] Example 1

[0087] To solve the problems existing in the prior art, the present invention proposes a therapeutic target prediction engine system based on omics genetic evidence, which is supported by biomedical ontology and network evidence-related knowledge base. The engine can facilitate users to perform de novo genetic target prediction and retrieval functions, and obtain quantitative recommendation results of relevant targets and drug repurposing analysis results of pathway intersection network genes, see Figure 1 .

[0088] Referring to the schematic diagram, a therapeutic target prediction engine system based on omics genetic evidence in this embodiment includes: a biomedical ontology and network evidence related knowledge base, a target prediction main engine, a target retrieval main engine and a human-computer interaction module.

[0089] The biomedical ontology and network evidence related knowledge base is essentially a database system that can provide data sets including enhancer-promoter maps, genome-wide promoter interaction capture technology (PCHi-C), gene expression (eQTL) / protein abundance (pQTL), gene / protein interaction networks, biomedical ontology data suitable for omics data knowledge interpretation, and drug-targets for immune-mediated diseases. This knowledge base is the focus and core content of the present invention, which is collected from public databases and published literature and manually corrected and cleaned.

[0090] It should be noted that the biomedical ontology and network evidence-related knowledge base can be expanded. In the future, the content of the biomedical ontology and network evidence-related knowledge base will continue to be expanded to include more cell state and type-specific data sets, and make the system more comprehensive and the user interface more user-friendly. The knowledge base will be updated every 12 months.

[0091] The knowledge base related to biomedical ontology and network evidence specifically includes the knowledge sub-base related to multimodal functional genomics evidence, the knowledge sub-base related to gene / protein interaction network evidence, the biomedical ontology knowledge sub-base, the ontology knowledge sub-base related to regulatory interaction gene enrichment, the drug target knowledge sub-base with clinical theoretical support, and the knowledge sub-base for quantitative recommendation of genetic targets for immune-mediated diseases.

[0092] The human-computer interaction module is used for data interaction between the user and the target prediction main engine. In this embodiment, it is composed of an interaction layer, a front-end UI, and an access layer. The visual display data is in the form of an HTML web page. Specifically, the human-computer interaction module is used to receive the data type input by the user, and according to the selected target sub-engine and related parameter types, issue a specified command to the target sub-engine. After the target sub-engine executes the corresponding function, the human-computer interaction module receives the results of data retrieval or online prediction from the target sub-engine and feeds it back to the user.

[0093] After receiving the prediction sub-engine type selected by the user from the human-computer interaction module, the target prediction main engine will be directed to the corresponding target prediction sub-engine and display the corresponding data input and parameter selection interface. After the user completes the data input and parameter selection and clicks submit, the interaction layer of the human-computer interaction module predicts the treatment target by calling the Rscript command function in the R language operating environment, and the prediction results will be displayed on the front-end UI layer of the human-computer interaction module.

[0094] After receiving the search sub-engine type selected by the user from the human-computer interaction module, the target search main engine will be directed to the corresponding target search sub-engine and display the corresponding command and parameter selection interface. After the user completes the command and parameter selection and clicks the search, the interaction layer of the human-computer interaction module performs the treatment target search by calling the Rscript command function in the R language operating environment, and the search results will be displayed in the front-end UI layer of the human-computer interaction module in the form of charts.

[0095] Among them, the knowledge base related to biomedical ontology and network evidence includes enhancer-promoter maps, genome-wide promoter interaction capture technology (PCHi-C), gene expression (eQTL) / protein abundance (pQTL), gene / protein interaction network, biomedical ontology data suitable for knowledge interpretation of omics data, drug-target data sets for immune-mediated diseases, etc., specifically including: knowledge sub-base related to multimodal functional genomics evidence, knowledge sub-base related to gene / protein interaction network evidence, biomedical ontology knowledge sub-base, ontology knowledge sub-base related to regulatory interaction gene enrichment, drug target knowledge sub-base with clinical theory support and quantitative recommendation knowledge sub-base of genetic targets for immune-mediated diseases.

[0096] Specifically, the knowledge sub-base related to multimodal functional genomics evidence includes: enhancer-promoter interaction evidence includes enhancer-promoter maps from the ENCODE and Roadmap databases, promoter capture Hi-C evidence includes manually screened promoter capture Hi-C datasets, and gene expression / protein abundance evidence includes 114 quantitative trait loci data (including 113 gene expression data and 1 protein abundance data). By integrating databases of multiple different modalities, it is suitable for systematic and comprehensive target prediction and retrieval.

[0097] A knowledge sub-base related to multimodal functional genomics evidence is set up, which is specifically subdivided into the following: enhancer-promoter interaction evidence, which uses the activity-by-contact (ABC) model to predict the regulatory relationship between non-coding region enhancers and gene promoters, taking into account enhancer activity (Activity) and enhancer-promoter interaction (Contact). The regulatory relationship predicted by the ABC model will be further verified by CRISPR technology to construct an enhancer-promoter interaction map; promoter capture Hi-C evidence, which is based on Hi-C technology and uses special probes to specifically capture the connecting fragments of the promoter region, and then performs high-throughput sequencing to obtain the interaction information between the promoter region and other gene regions. Compared with traditional Hi-C technology, it is more specific and accurate, and can capture genomic regions that interact with specific gene promoters without bias and with high precision, which helps to build a comprehensive gene regulatory network; gene expression / protein abundance evidence, which uses quantitative trait mapping technology to generate genetic regulatory data of gene expression and protein abundance, which can comprehensively analyze the regulatory network from gene genetic variation to gene expression to protein function.

[0098] The knowledge sub-base related to multimodal functional genomics evidence can be applied to target prediction, retrieval and data mining functions of different omics data types such as genes, single nucleotide polymorphism sites (SNPs), and intergenic regions.

[0099] Specifically, the knowledge sub-base related to gene / protein interaction network evidence comes from: 1. STRING database (V12.0). The database sets a confidence score threshold for the interaction relationship (medium confidence is 0.4, high confidence is 0.7, and the highest confidence is 0.9); 2. KEGG database (V112.0 version), which consists of a series of manually drawn KEGG metabolic pathway maps to represent experimental results on metabolism and other cellular and biological functions. Each pathway map includes a network diagram of molecular interactions and reactions, which is designed to connect the genes in the genome and the gene products (mainly proteins) generated by the process shown in the pathway. By default, this engine considers high-confidence interactions, and the corresponding generated network has a total of about 14,000 nodes / genes and about 202,000 interactions / edges.

[0100] The knowledge sub-base related to gene / protein interaction network evidence is set up because this engine follows the "whole gene genetic law model" and considers core genes (small number but large genetic effect) and peripheral genes (large number but small genetic effect) as candidate targets. This engine uses core genes as seeds and uses protein interaction network information and restarted random walk algorithm to identify (non-seed) peripheral genes under the influence of the network. Among them, peripheral genes with higher connectivity with core genes will obtain higher affinity scores, which can expand the screening range of disease-associated genes.

[0101] The knowledge sub-base related to gene / protein interaction network evidence can be used for target prediction, target retrieval and data mining functions of different omics data types such as genes, SNPs, and intergenic regions.

[0102] Specifically, the biomedical ontology knowledge sub-base is used to provide knowledge interpretation for omics summary data such as genes, SNPs, intergenic regions and protein domains, based on the support of multiple ontology knowledge bases. This sub-base covers a wide range of knowledge areas, including functions, pathways, transcriptional regulators, diseases, phenotypes, drugs, domains and evolution.

[0103] The purpose of setting up the biomedical ontology knowledge sub-base is to use the biomedical ontology knowledge sub-base to perform the following ontology analysis:

[0104] 1. Function: Gene Ontology GO (September 2024 version), including the molecular function ontology (GOMF), the cellular component ontology (GOCC) and the biological process ontology (GOBP);

[0105] 2. Pathway: pathway information of KEGG database (version V112.0), REACTOME database (version V90), MSIGDB database (version V2024.1), and MitoPathways pathway information of MitoCarta database (version V3.0);

[0106] 3. Transcriptional regulatory factors: TFs data from the Enrichr database (July 2022 version) and the TRRUST database (April 6, 2018 version);

[0107] 4. Disease: Mondo disease ontology (October 1, 2024 version), human disease ontology HDO (September 2024 version), and EFO experimental factor ontology of GWAS disease characteristics (V3.70.0 version);

[0108] 5. Phenotype: Human Phenotype Ontology HPO (March 2024 version) and Mammalian Phenotype Ontology MPO (September 2024 version);

[0109] 6. Drugs: Drug-gene interaction database DGIdb (V.5.0.7), targeted drug binding pocket ChEMBL database (V34);

[0110] 7. Protein domains and intrinsically disordered proteins: SCOP, Pfam, InterPro and the ontology of intrinsically disordered proteins IDPO;

[0111] 8. Markers and evolution: molecular signatures and genomic phylogenetic stratification in the MSigDB database.

[0112] The biomedical ontology knowledge sub-base is suitable for target prediction, target retrieval and ontology knowledge interpretation and analysis of different omics data types such as genes, SNPs, and intergenic regions.

[0113] Specifically, the ontological knowledge sub-base related to regulatory interaction gene enrichment is used to perform ontological enrichment analysis on potential regulatory target genes, including function, phenotype, disease, protein domain (SCOP superfamily domain, Pfam domain, InterPro domain), etc. The ontological knowledge sub-base related to regulatory interaction gene enrichment currently supports:

[0114] (1) Function: Gene Ontology (GO) (September 2024 version), including the molecular function ontology (GO-MF), the biological process ontology (GO-BP), and the cellular component ontology (GO-CC).

[0115] (2) Pathway: pathway information of KEGG database (version V112.0), REACTOME database (version V90), MSIGDB database (version V2024.1), and MitoPathways pathway information of MitoCarta database (version V3.0).

[0116] (3) Phenotype: Human Phenotype Ontology HPO (April 2024 version) and Mammalian Phenotype Ontology MPO (May 2024 version).

[0117] (4) Disease: Mondo disease ontology (version 1 October 2024), HDO human disease ontology (version 9 September 2024), and GWAS Catalog’s experimental factor ontology EFO (version 3.70.0).

[0118] (5) Protein domains: SCOP superfamily domain, Pfam domain, InterPro domain.

[0119] The ontological knowledge sub-base related to regulatory interaction gene enrichment can be applied to target prediction, target retrieval and ontological knowledge enrichment analysis of omics data types such as genes, SNPs, and intergenic regions.

[0120] Specifically, the drug target knowledge sub-base with clinical theoretical support is an indispensable part of the present invention, which can provide clinically verified target information for the present invention. Based on this information, the target predicted by the present invention can be verified, and the knowledge sub-base can help researchers select more reliable targets and reduce the risk of R&D failure due to improper target selection. In addition, the knowledge base also supports drug reuse, finding new indications for approved drugs, saving development costs and accelerating clinical transformation, and can provide important guidance and support for modern drug research and development.

[0121] Set up a drug target knowledge sub-base with clinical theory support, which mainly includes the following parts:

[0122] 1. ChEMBL database (March 2024 version). This database focuses on collecting data on small molecule compounds and their interactions with biological targets. It contains the structure, biological activity, pharmacology, clinical information, etc. of the compounds, and is widely used in drug discovery, virtual screening, and computational pharmacology. ChEMBL data mainly comes from scientific literature and patents;

[0123] 2. DrugBank database (version 5.1.12). This database integrates a variety of information on small molecule drugs and biologics, including a series of drug chemistry and pharmacology information such as chemical structure, mechanism, metabolic pathway, clinical use, pharmacokinetics and toxicity, as well as drug targets and their relationship with diseases. DrugBank data sources include scientific literature, patents and FDA drug approval information, and are widely used in the fields of drug development, toxicology research and drug repurposing;

[0124] 3. Open Targets Database (July 2024 version), an innovative, large-scale genetic target resource that integrates a variety of public data and uses human genetics (GWAS data) and functional genomics data for systematic drug target identification and prioritization;

[0125] 4. PubChem database (October 2022 version), PubChem is an open chemical database maintained by the National Institutes of Health (NIH) of the United States, which contains a large amount of information about chemical substances. It mainly involves small molecules, but also includes some larger molecules such as nucleotides, carbohydrates, lipids, peptides, and chemically modified macromolecules. The data in PubChem comes from a wide range of sources, covering hundreds of data sources, including government agencies, chemical suppliers, journal publishers, etc. These data include chemical structures, identifiers, physicochemical properties, biological activity, patents, health and safety information, and toxicity data;

[0126] 5. TCMID database (November 2017 version) is an open database focusing on traditional Chinese medicine, aiming to provide multi-dimensional information support for TCM research. It integrates a large amount of information on Chinese medicine ingredients, pharmacological effects, clinical applications and corresponding molecular targets. The construction of TCMID is based on traditional Chinese medicine literature and modern pharmacology research, supports interdisciplinary data analysis, and helps researchers explore the modernization and standardization research path of traditional Chinese medicine.

[0127] The drug target knowledge sub-base with clinical theory support can be applied to precision medicine, drug development and drug repurposing research supported by clinical and genetic evidence.

[0128] Specifically, the quantitative recommendation knowledge sub-base of genetic targets for immune-mediated diseases is derived from the Pi (Priority index) database, which is a comprehensive resource that provides genetic target information for immune-mediated diseases through a genetics-led prioritization strategy.

[0129] The establishment of a quantitative recommendation knowledge sub-base for genetic targets of immune-mediated diseases utilizes immune genomic datasets from various immune cell types or states, and makes full use of genetic information such as disease GWAS, functional immunogenomics, ontology annotations and protein interactions, thereby enhancing the ability to identify and quantitatively recommend drug targets.

[0130] The knowledge sub-base for quantitative recommendation of genetic targets for immune-mediated diseases is suitable for quantitative recommendation of genetic targets for immune-mediated diseases and outputs relevant quantitative recommendation results.

[0131] Among them, the target prediction main engine provides the prediction function of therapeutic targets based on genomic summary data. The target prediction main engine specifically includes: a target prediction sub-engine based on priority index quantitative recommendation, a target prediction sub-engine based on pathway intersection network, a drug repurposing analysis sub-engine based on pathway intersection network, a target prediction sub-engine based on cross-disease genetic pleiotropy drive, and a target prediction sub-engine based on generative AI; the above sub-engines complement each other to provide users with comprehensive, fast and effective prediction functions. .

[0132] Specifically, the target prediction sub-engine based on the quantitative recommendation of the priority index is used to receive the SNP site information and its significance (P-values) associated with the disease input by the user from the human-computer interaction module, and define the core genes by combining the knowledge sub-bases such as genomic proximity, e / pQTL and PCHi-C, and connect the core genes with the peripheral genes, and finally complete the target prediction based on the quantitative recommendation of the priority index. The scoring of the core genes takes into account the following factors: (1) genetic association with the disease (including P-values, significance threshold and linkage disequilibrium R 2 ); (2) SNP distance window representing genomic proximity; (3) significance level of genetic association with gene expression in the eQTL dataset (or protein abundance in the pQTL dataset); (4) strength of physical interaction between gene promoters and genomic regions containing SNPs in the PCHi-C dataset. The empirical distribution function (eCDF) was estimated for all SNP-gene pairs to ensure that the values ​​were within the range of 0-1.

[0133] Implementation: GWAS SNPs were used to define genomic core genes. First, neighboring genes (nGene) were defined based on genomic proximity (located within a certain distance window of the SNP) and genomic organization (located in the same topological association domain as the SNP, using a topological association domain dataset reflecting the GM12878 genomic organization in immune context). The scores for neighboring genes nGene took into account distance (influence range) and were optimized to minimize false positives. Since genes hit by GWAS SNPs are not necessarily the most neighboring genes, chromatin conformation genes (cGene) were defined based on chromatin physical interactions, based on data from promoter capture Hi-C studies to identify whether gene promoters physically interact with genomic regions carrying SNPs. Third, expression-associated genes (eGene) were defined based on summary data from eQTL mapping. In the eQTL–GWAS integration, considering that colocalization analysis can include directional and effect size information of GWAS and eQTL effects in the output, the engine used a common colocalization method, coloc, to estimate whether a SNP plays a causal role in both GWAS and eQTL studies through a Bayesian framework. The default prior probability is: the probability of being associated with a trait is 1×10 -4 , the probability of being associated with two traits at the same time is 1×10 -5 When the estimated causal probability exceeded 0.8, the SNP was identified as an expression-related gene eGene and scored according to the SNP with the strongest association.

[0134] The engine also uses gene-level ontology annotations to further define annotation predictors related to immune function / dysfunction: (1) immune function genes (fGene), which are annotated using Gene Ontology to annotate genes to immune responses with experimental or manual evidence codes; (2) disease genes (dGene), which are annotated using the Online Genetic Human Genetics (OMIM) database and disease ontology to annotate genes that can cause immune system diseases (including primary immunodeficiency diseases); and (3) immune phenotype genes (pGene), which are annotated using the Human Phenotype Ontology to annotate abnormalities of the immune system, blood, and hematopoietic tissues, and the Mammalian Phenotype Ontology to annotate immune / hematopoietic system phenotypes.

[0135] The above nGene, cGene and eGene are the identified core genes, while fGene, dGene and pGene are the functional annotations of the core genes. The output results include a Manhattan diagram and a table, which will be returned to the human-computer interaction module, both of which reflect the core genes and their scores (quantifying the genetic association between diseases and core genes); the results also provide an evidence list showing the data sets used to define the core genes.

[0136] Subsequently, the core genes were used as seed nodes, and the knowledge of protein interactions was applied to identify the peripheral genes (non-seed nodes) under the influence of the network using the restarted random walk algorithm. Specifically: A protein interaction network was established based on the knowledge sub-base related to the gene / protein interaction network evidence, and the interaction relationships defined by the STRING database with an interaction score greater than 0.7 (i.e., high confidence) were included, and marked as "experimental" and "database" (i.e., manual verification), corresponding to approximately 14,000 nodes (genes) and approximately 202,000 edges (interactions). For each type of seed gene, the restarted random walk algorithm was used to identify non-seed genes with network influence.

[0137] For each dataset (i.e., proximity, e / pQTL, and PCHi-C), the above process produced predictors containing core and peripheral genes, with affinity scores quantifying their connectivity to the core gene network. In general, core genes are more likely to receive higher affinity scores than peripheral genes. Nevertheless, if a peripheral gene has high connectivity with most (but not all) core genes, it may also receive a high affinity score, indicating that it is more influenced by the network.

[0138] The method used by this engine to calculate affinity scores ensures comparability between different predictors, and the inclusion of non-seed genes increases the completeness of potential targets. Finally, a ranked list of core genes and peripheral genes is output (ranked by affinity scores that quantify the network connectivity with the input core genes).

[0139] Finally, the following information is specified in the sub-engine: (1) whether to include SNPs in linkage disequilibrium (if included, based on which population the calculation is performed); (2) how to identify core genes based on genomic proximity, e / pQTL, and PCHi-C (such as conformational evidence); and (3) the network information used to identify core genes and peripheral genes. In addition, users can select additional parameters to fine-tune the quantitative recommendation results.

[0140] This engine integrates predictors through a method similar to Fisher's combination meta-analysis. For each predictor in the gene-predictor matrix, we first convert the gene affinity scores into numerical values ​​similar to P-values ​​(Formula 1), and then combine these class P-values ​​between the individual predictors for each gene using Fisher's combination method (Formulas 2 to 4); this method is suitable for cases with moderate or high evidence support. The resulting combined class P-values ​​are re-normalized into priority ranking scores (ranging from 0-5) (Formulas 5 and 6)

[0141]

[0142] in, represents the affinity score of the ith gene on the jth predictor, represents the corresponding transformed class P value, and eCDF is the empirical cumulative density function estimated based on all genes.

[0143]

[0144] x~x 2 (2J)(3)

[0145] CP i =CDF(x) (4)

[0146] where J is the number of predictors, χ 2 (2J) represents a chi-square distribution with 2J degrees of freedom, CP i represents the combined P value of the ith gene, that is, the CDF value of the chi-square distribution at x.

[0147] x i = -log(CP i ) (5)

[0148]

[0149] Among them, PR i is the Pi score of the ith gene, MIN and MAX represent the minimum and maximum values, respectively.

[0150] Finally, the human-computer interaction module will return a summary of this task, including the input data and the running time on the server, display the priority scores of about 14,000 genes on the genome in a Manhattan diagram, and return a table showing the specific priority scores of these genes, gene types (core genes or peripheral genes), and evidence overviews (genomic proximity, e / pQTL, and PCHi-C).

[0151] Specifically, the target prediction sub-engine based on the pathway intersection network is used to receive the disease-related SNPs and significance information (P-values) input by the user from the human-computer interaction module, and the following information needs to be specified in the sub-engine: (1) whether to include SNPs in linkage disequilibrium (if included, based on which population); (2) how to identify core genes based on genomic proximity, e / pQTL and PCHi-C (such as conformational evidence); (3) network information used to identify core genes and peripheral genes; (4) the confidence of the gene / protein interaction knowledge sub-library used; (5) the number of genes included in the pathway intersection network and the significance information. This sub-engine not only completes the quantitative recommendation of genes (consistent with the quantitative target recommendation sub-engine at the gene level), but also generates a pathway intersection containing a set of highly prioritized genes.

[0152] The pathway intersection network receives the user-input sorted gene list from the human-computer interaction module and identifies a set of genes with high priority and strong correlation with each other. This engine develops an algorithm to search for subsets in the gene network (the network obtained by merging all KEGG pathways) so that the generated gene subnetwork (or pathway intersection network) contains a set of high-priority genes that can be connected by a small number of low-priority genes. Subsequently, this engine evaluates the significance (P-values) of the identified subnetwork (pathway intersection network) in a randomized case through a degree-preserving node permutation test. Specifically, we perform 100 permutation tests from the sorted gene list while keeping the degree unchanged, so that the expected intersection identified has the same / similar number of genes as the intersection to be evaluated, and the expected intersection is used as the null distribution to estimate the significance of the intersection identified by this engine.

[0153] This engine identified a total of 878 target genes in the top 1% of the prioritized gene list for 16 immune traits. Gene clustering and visualization can be performed using supra-hexagon. If the generated graph is further integrated with the drugability pocket data, the probability of containing drugability genes in each hexagon can be estimated. For each gene cluster obtained by clustering, enrichment analysis was performed using the R package XGR to identify the enriched approved drug indications (ChEMBL) and KEGG pathways, thereby functionally annotating the gene cluster.

[0154] The final identified pathway intersection network results will be returned to the human-computer interaction module in the form of a table of intersection genes and subnetworks (the color of the nodes / genes depends on the input score).

[0155] Specifically, the drug repurposing analysis sub-engine based on pathway intersection network is used to receive the SNP site information and its significance (P-values) associated with the disease input by the user from the human-computer interaction module. First, the pathway intersection network results are obtained through the target prediction sub-engine based on pathway intersection network. The following information needs to be specified in the sub-engine: (1) whether to include SNPs in linkage disequilibrium (if included, based on which population); (2) how to identify core genes based on genomic proximity, e / pQTL and PCHi-C (such as conformational evidence); (3) network information used to identify core genes and peripheral genes; (4) the confidence of the gene / protein interaction knowledge sub-library used; (5) select the number of genes included in the pathway intersection network and the significance information; (6) select the number of genes in the intersection network and the significance information of the identified sub-network. This sub-engine not only generates a pathway intersection network of highly prioritized genes, but also completes the drug repurposing analysis for the prioritized genes.

[0156] Through the target prediction engine, we can identify target genes associated with specific traits. These target genes not only cover biological importance, but also integrate information from multiple databases, especially data related to drugability. This information includes the characteristics of drug targets, mechanisms of action, and related drug development dynamics. More importantly, the drugability information of target genes is often displayed in the form of three-dimensional structures of proteins, emphasizing the visualization characteristics of drug binding pockets, which provides an intuitive reference for drug design and optimization. In the process of analyzing target genes, target genes with high ratings in multiple diseases deserve further attention. These target genes usually have great clinical application potential and can provide new opportunities for drug repurposing. For example, some target genes may play a role in different types of cancer, metabolic diseases, or neurological diseases, thereby laying the foundation for the re-evaluation and repurposing of existing drugs. At the same time, target genes that only appear in specific diseases may reveal specific targets for the disease. The identification of these targets is crucial for the development of new therapies for specific diseases, which can help researchers deeply understand the disease mechanism and design new drugs in a targeted manner. This combination of target identification and analysis not only provides a foundation for the development of precision medicine, but also promotes innovation and progress in drug development. Through the application of target prediction engines, we can more effectively explore potential opportunities in drug development and promote the implementation of new treatment strategies.

[0157] The final identified drug reuse analysis results will be returned to the human-computer interaction module in the form of charts.

[0158] Specifically, the target prediction sub-engine driven by cross-disease genetic pleiotropy is used to receive the SNP site information and its significance (P-values) associated with the disease input by the user from the human-computer interaction module, combined with the pathway intersection results obtained by the target prediction sub-engine based on the pathway intersection network, and introduced a multi-trait score (MRS) to quantify the degree to which a target gene is highly scored in multiple traits, and use genes with MRS to mark the gene network merged by the KEGG pathway. In order to analyze the pathway composition of the identified intersection (i.e., the participation of individual pathways), Fisher's exact test was used to compare the genes enriched in the intersection with all genes with MRS as the test background to identify a single KEGG pathway. This engine also tested the enrichment of pathway intersection network genes at different levels, such as mouse immune-mediated disease phenotypes (Monarch program), drug interference gene traits (CREEDS), staged and approved treatment regimens indicated by immune diseases (ChEMBL), and druggable pockets. Additionally, we introduced a multi-trait novelty score (MNS) to quantify how unexplored a target is in most traits.

[0159] Among them, MRS (Formula 7) is used to quantify the degree to which the target gene is highly scored in multiple traits. It takes into account the following factors: (1) the number of times the target gene ranks in the top 1% (i.e., top 150) in a certain trait (nTop); (2) the average ranking of the target gene in these traits (mRank); and (3) the total number of traits included in the analysis (N). These three factors are combined through Formula 1, and the value range of MRS is scaled to between 0 and 1. A high MRS indicates that the target gene has received high scores in more traits.

[0160]

[0161] MNS (Formula 8) is used to quantify the extent to which a target gene has not been fully explored in most traits. Novelty is defined from the following aspects: (1) the number of times a target gene is annotated as a disease gene in a trait (nD) (from the DO database); (2) the number of approved drugs for the target gene in these traits (nA) (from the ChEMBL database); (3) the number of drugs for which the target gene has at least one drug in drug development in these traits (nP) (from the ChEMBL database); and (4) the total number of traits included in the analysis (N). These four factors are combined using Formula 2. This formula uses the geometric mean and limits the value range to between 0 and 1. The higher the score, the stronger the novelty of the target gene in the trait. Compared with target genes that only have drugs in the development stage, this formula will increase the penalty score for target genes with approved drugs.

[0162]

[0163] The final identified target prediction results based on cross-disease genetic pleiotropy drivers will be returned to the human-computer interaction module in the form of charts.

[0164] Specifically, the target prediction sub-engine based on generative AI is used to receive the characteristic trait data input by the user from the human-computer interaction module and identify the potential target genes related to it through a deep learning algorithm. The sub-engine first obtains the prediction results of the target gene by comprehensively analyzing a variety of biological information data. In this process, the following information needs to be specified in the sub-engine: (1) the specific type of input characteristic traits and their correlation indicators; (2) the selection of multi-omics data types for training the generative model; (3) the determination of the evaluation criteria for the druggability of the target gene; (4) the credibility level of the protein interaction knowledge base used; (5) setting the screening threshold and significance criteria of the target gene; (6) selecting the number of target genes and related biological pathway information in the output results.

[0165] This sub-engine not only completes the identification of potential target genes, but also evaluates and analyzes the drugability of target genes. Through the powerful information integration, feature extraction and model building capabilities of the generative AI model, researchers use massive biomedical data (such as multi-omics and multi-dimensional data in the knowledge sub-base) for training, which can identify target genes that show high clinical application potential in multiple disease contexts. The pleiotropic characteristics of these target genes make them ideal candidates for drug repurposing, opening up new possibilities for the re-evaluation and application of existing drugs. At the same time, for target genes that only appear in specific diseases, the sub-engine can reveal disease-specific targets and provide valuable target information for new drug research and development. This sub-engine can not only identify potential target genes, but also combine molecular dynamics simulations to deeply analyze the drugability of target genes (similar to AlphaFold), providing an important basis for subsequent drug development.

[0166] Generative AI achieves target prediction through deep learning models (such as generative models based on Transformer architecture). Specific methods include: first, combining genomic data (such as SNPs and significance levels) with network evidence (such as protein-protein interaction network PPI and molecular pathway data) and ontological annotations at the gene level (including immune function genes, disease genes, and immune phenotype genes) to construct a multimodal input feature matrix. The generative AI model generates a priority list of candidate targets and predicts the potential biological functions of these targets through deep fusion of learning of input data and multimodal features. For example, when the input includes the gene expression profile of a disease and its related PPI network, the model can generate a priority ranking of target genes and predict its pleiotropic effects in other diseases in combination with ontological annotations, ultimately providing support for cross-disease target screening. The final identified target prediction results will be fed back to the human-computer interaction module in the form of visual charts for users to conduct further analysis and decision-making.

[0167] Among them, the target retrieval main engine provides an online query function for therapeutic targets based on genome or genome summary data. The target retrieval main engine specifically includes: a target retrieval sub-engine for immune-mediated diseases, a target retrieval sub-engine based on pathway intersection networks, and a target retrieval sub-engine based on generative AI; the above sub-engines complement each other to provide users with comprehensive, fast and effective query functions.

[0168] Specifically, the target retrieval sub-engine for immune-mediated diseases is used to receive user instructions for target retrieval from the human-computer interaction module. Under the "GATEWAY" navigation label, the target priority ranking results for immune-mediated diseases supported by this engine can be obtained. The target gene ranking results for each disease include a complete target gene priority ranking list and an optional number of pathway intersection network gene lists, and can be generated in two modes, namely discovery mode (without using any known drug target information) and supervision mode (integrating predictors of known drug targets, i.e. targets supported by clinical theory, through the random forest algorithm). For most users, especially those who need to find new targets, the discovery mode is recommended, while the supervision mode is more suitable for users who need to explore information related to effective drugs. In addition to the files in each disease-specific page, users can also download the MySQL relational database and detailed documentation on the organization and usage of the table. In this engine, all downloadable files are free to use without any restrictions.

[0169] The above operations can be completed in the graphical interface of the human-computer interaction module.

[0170] Specifically, the target search sub-engine is used to receive user input from the human-computer interaction module, and the following information needs to be specified in the sub-engine: (1) different diseases for comparison; (2) genes to be evaluated, which can be all priority genes or all pathway intersection network genes; (3) whether to use the discovery mode or the supervision mode for analysis; (4) select the indicator used for prioritization, which can be either priority or priority score. Based on the results of the target prediction sub-engine based on the pathway intersection network, the targets predicted by multiple traits can be searched and analyzed.

[0171] These diseases can be easily selected to complete the comparison on the user request page, and the mode can be switched or other options can be selected. Comparable diseases are classified by priority ranking mode (i.e., target information obtained by both supervision and discovery modes, or only discovery mode), and diseases can be added one by one, or all can be selected for comparison. The comparison results are displayed in a table, where the target genes are ranked according to MRS, with corresponding operability and drugability information, and marked with disease-specific ranking results. The MRS scores of relevant targets in different traits are shown in the figure, and each target provides links to access specific information, including: 1. Hyperlinks to databases such as GeneCards, Ensembl, HGNC, OMIM, Entrez, OpenTargets, etc., which can access the sequence of the target gene, the corresponding protein and tertiary structure, related pathways and disease information; 2. Hyperlinks to databases such as PDBe-KB, AlphaFold, etc., can understand the drugability of the target; 3. Priority score (ranking) of the target in common immune traits. Users can easily retrieve information about the target of interest. At the same time, this engine provides an effective way to identify shared target genes or genes with high ratings in specific diseases. Users can explore the potential of these two types of genes in drug reuse through heat map-like displays and hyperlinks to gene-related information (either general or in specific diseases).

[0172] The above results will be returned to the human-computer interaction module in the form of charts.

[0173] Specifically, the target retrieval sub-engine based on generative AI is used to receive user input from the human-computer interaction module, and automatically identify and predict potential drug targets based on the generative model.

[0174] Based on the generative AI large language model, this sub-engine has advanced human-computer interactive dialogue capabilities, can automatically generate answers based on the user's input information, and effectively guide users to use the product. Users can directly fill in the required sub-engine type, input data and expected parameters into the engine's dialog box through simple text input. The system will automatically call the relevant sub-engine for calculation and link the generated results back to the dialog box. In addition, users can also choose to let the system output the calculation results directly in the dialog box, and the output results are consistent with the results obtained when the corresponding sub-engine is directly called, ensuring the consistency and reliability of the information. In this way, users can easily filter and retrieve relevant results using natural language commands, so as to quickly obtain targets of interest and corresponding disease information. This human-computer interaction design not only improves the user experience, but also greatly improves work efficiency, allowing users to more conveniently access and use a large amount of bioinformatics data. Users only need to enter simple text to quickly find diseases, drugs and research data related to specific targets, which promotes the implementation of precision medicine and personalized treatment. With this intelligent sub-engine, researchers and medical workers can explore drug targets more efficiently, providing solid support for new drug development and disease research.

[0175] The above results will be returned to the human-computer interaction module in the form of text, charts, and links.

[0176] The human-computer interaction module is a web page interaction system, which is composed of an interaction layer, a front-end UI, and an access layer. The data displayed visually is in the form of an HTML web page.

[0177] The human-computer interaction module feeds back to the user a self-contained dynamically generated HTML file, which includes all input information, dynamic graphs and tables and other output results. It is an integrated, dynamic, editable and downloadable HTML file.

[0178] In this embodiment, the human-computer interaction module is a web page interaction system, which adopts a three-layer architecture, including an interaction layer, a front-end UI and an access layer.

[0179] The interaction layer is implemented by calling the Rscript command function in the R language running environment.

[0180] The front-end UI is built using UI component libraries, HTML, and CSS technologies.

[0181] The access layer can be accessed through a browser on a personal computer PC or a mobile communication device.

[0182] See also Figure 2 , showing how to operate the target prediction main engine:

[0183] 1) The user selects the sub-engine for relevant target prediction on the web page (human-computer interaction module) and enters relevant data as required;

[0184] 2) Select sub-engine parameters, such as whether to consider SNPs in linkage disequilibrium, genomic proximity, e / pQTL, PCHi-C, the confidence level selected for interaction relationships, the screening threshold of target genes, significance criteria, background traits, etc.;

[0185] 3) The user clicks the SUBMIT button on the web page to call the relevant sub-engine for prediction;

[0186] 4) Based on the input data and selected parameters, the selected sub-engine starts to perform the prediction function and feeds the prediction results back to the front-end layer;

[0187] 5) Users can browse and download prediction results at the access layer.

[0188] The present invention collects a large number of functional genomic data sets and integrates data from multiple modalities and different batches. These comprehensive and detailed data can provide high-confidence multi-omics regulation and interaction relationships, laying a solid foundation for the prediction of therapeutic targets. In addition, the knowledge base of the present invention is open source and scalable, ensuring the completeness and novelty of the knowledge base.

[0189] The theoretical basis of the present invention is the whole-genome genetic law model, which believes that complex traits / phenotypes are caused by the joint action of a few core genes and a large number of peripheral genes. Therefore, when identifying therapeutic targets, the present invention not only takes into account the core genes, but also identifies peripheral genes through a restarted random walk algorithm, providing a more comprehensive and accurate prediction result of therapeutic targets.

[0190] The present invention not only enables de novo prediction of therapeutic targets based on omics genetic evidence, but also enables users to more deeply explore the pathway information and network evidence of the targets, so that users can explore and analyze personal data in a more comprehensive way.

[0191] The present invention can provide a user-friendly interface and come with a detailed user manual providing step-by-step instructions. This enables users to easily browse and effectively utilize resources, promoting integrated and efficient knowledge discovery.

[0192] The present invention combines generative AI to provide innovative technical support for target prediction, query and personalized treatment by automatically integrating multimodal functional genomic data, identifying complex regulatory relationships, and generating accurate hypotheses.

[0193] Based on the above advantages, the present invention can provide researchers with a comprehensive and simplified method to predict therapeutic targets based on omics genetic evidence, promoting more comprehensive and more effective exploration of the prediction of therapeutic targets.

[0194] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A therapeutic target prediction engine system based on omics genetic evidence, characterized in that: The therapeutic target prediction engine system comprises: Biomedical ontology and network evidence related knowledge base: The datasets provided include enhancer-promoter maps, genome-wide promoter interaction capture technology, gene expression / protein abundance, gene / protein interaction networks, drug targets with clinical theoretical support, biomedical ontology data suitable for knowledge interpretation of omics data, and drug-targets for immune-mediated diseases; Target prediction main engine: provides prediction functions for therapeutic targets based on genomic summary data; Target search main engine: provides online query function for therapeutic targets based on genome or genome summary data; Human-computer interaction module: used to receive data input by the user, issue specified commands to the target sub-engine according to the selected target sub-engine and related parameters; and receive the results of online prediction or data retrieval from the target sub-engine and feed them back to the user.

2. A therapeutic target prediction engine system based on omics genetic evidence according to claim 1, characterized in that: The biomedical ontology and network evidence related knowledge base includes: A knowledge base related to multimodal functional genomics evidence, including enhancer-promoter interaction evidence, including enhancer-promoter maps from the ENCODE and Roadmap databases; promoter capture Hi-C evidence, including manually screened promoter capture Hi-C datasets; and gene expression / protein abundance evidence, including quantitative trait loci data; Knowledge sub-base related to gene / protein interaction network evidence, data sources include: STRING database, KEGG database; The biomedical ontology knowledge sub-base provides knowledge interpretation of omics summary data applicable to genes, SNPs, intergenic regions and protein domains, covering knowledge areas including: function, pathway, transcriptional regulator, disease, phenotype, drug, domain and evolution; The ontological knowledge sub-base related to regulatory interaction gene enrichment is used to perform ontology enrichment analysis of potential regulatory target genes, including function, phenotype, disease and protein domain; The drug target knowledge sub-database with clinical theoretical support provides clinically verified target information from sources including: ChEMBL database, DrugBank database, Open Targets database, PubChem database and TCMID database; The knowledge base for quantitative recommendation of genetic targets for immune-mediated diseases is derived from the Priority index database. It uses immune genomic datasets from various immune cell types or states, and makes full use of genetic information including disease GWAS, functional immunogenomics, ontology annotations and protein interactions for quantitative recommendation of genetic targets for immune-mediated diseases.

3. A therapeutic target prediction engine system based on omics genetic evidence according to claim 1, characterized in that: The target prediction main engine specifically includes: The target prediction sub-engine based on the quantitative recommendation of the priority index is used to receive the user input SNP and its significance of association with the disease, combine the genome proximity, gene expression / protein abundance and the whole genome promoter interaction capture technology knowledge sub-base to define the core genes, and connect the core genes with the peripheral genes to complete the target prediction based on the quantitative recommendation of the priority index; The target prediction sub-engine based on the pathway intersection network receives the disease-related SNPs and significance information and specified information input by the user, completes the quantitative recommendation of genes, and generates a pathway intersection network gene list containing a set of high-priority genes; The drug repurposing analysis sub-engine based on the pathway intersection network receives the SNPs and their significance of association with the disease input by the user, obtains the pathway intersection network results through the target prediction sub-engine based on the pathway intersection network, receives the specified information, generates the pathway intersection network of the highly prioritized genes, and completes the drug repurposing analysis for the prioritized genes; The target prediction engine driven by cross-disease genetic pleiotropy receives the SNPs and their significance of association with diseases input by users, combines the pathway intersection network results obtained by the target prediction engine based on pathway intersection networks, introduces multi-trait scores to quantify the degree to which target genes are highly scored in multiple traits, and uses genes with multi-trait scores to mark the gene network merged by KEGG pathways; The target prediction sub-engine based on generative AI receives characteristic trait data input by the user and identifies potential target genes related to it through deep learning algorithms.

4. A therapeutic target prediction engine system based on omics genetic evidence according to claim 3, characterized in that: The target prediction sub-engine based on the priority index quantitative recommendation has the following specific workflow: Define genomic core genes: define neighboring genes nGene based on genomic proximity and genomic organization; identify whether gene promoters physically interact with genomic regions carrying SNPs based on data from promoter capture Hi-C studies, and define chromatin conformation genes cGene based on chromatin physical interactions; define expression-related genes eGene based on summary data from eQTL mapping, and use the co-localization method coloc to estimate whether a certain SNP plays a causal role in both GWAS and eQTL studies through a Bayesian framework; when the estimated causal probability exceeds the set threshold, the SNP is identified as an expression-related gene eGene, and is scored based on the SNP with the strongest association; The scoring of core genes takes into account the following factors: disease genetic association; SNP distance window representing genomic proximity; significance level of genetic association between gene expression in eQTL dataset or protein abundance in pQTL dataset; strength of physical interaction between gene promoter and genomic region containing SNP in PCHi-C dataset; empirical distribution function estimated based on all SNP-gene pairs; Using ontology annotations at the gene level, annotation predictors related to immune function / dysfunction were further defined: immune function genes fGene, annotated using Gene Ontology, annotated genes to immune responses with experimental or clinical evidence; disease genes dGene, annotated using the online genetic human genetics database and disease ontology, annotated genes that can cause immune system diseases; immune phenotype genes pGene, annotated abnormal immune system, blood and hematopoietic tissues using the human phenotype ontology, and immune / hematopoietic system phenotypes using the mammalian phenotype ontology; The core genes were identified using neighboring genes nGene, chromatin conformation genes cGene and expression-related genes eGene, and immune function genes fGene, disease genes dGene and immune phenotype genes pGene were used to perform functional annotations on the core genes; the output results included a chart reflecting the core genes and their quantitative genetic association scores between diseases and the core genes, as well as a list of evidence showing the data sets used to define the core genes; Using core genes as seed nodes, applying knowledge of protein interactions, and using a restarted random walk algorithm to identify peripheral genes under network influence: relying on a knowledge sub-base related to gene / protein interaction network evidence to establish a protein interaction network, incorporating high-confidence interaction relationships in the STRING database, for each type of seed gene, a random walk algorithm is used to identify non-seed genes with network influence; for each proximity, gene expression / protein abundance, and genome-wide promoter interaction capture technology dataset, a predictor containing core and peripheral genes is generated, with an affinity score used to quantify their network connectivity with the core genes; finally, a ranked list of core and peripheral genes based on affinity scores that quantify network connectivity with the input core genes is output; The following information is specified in the target prediction sub-engine based on the quantitative recommendation of the priority index: whether to include SNPs in linkage disequilibrium, and if so, based on which population the calculation is based; how to identify core genes based on genomic proximity, gene expression / protein abundance, and genome-wide promoter interaction capture technology; network information for the identified core genes and peripheral genes; For each predictor in the gene-predictor matrix, the gene affinity scores were converted to class P values; the class P values ​​were combined between the individual predictors for each gene using Fisher's combination method; the resulting combined class P values ​​were renormalized to priority ranking scores; The results are returned to the human-computer interaction module, including input data, priority scores of genes on the genome, gene types, and evidence overview.

5. The therapeutic target prediction engine system based on omics genetic evidence according to claim 3, characterized in that: The target prediction sub-engine based on the pathway intersection network has the following specific workflow: Receive the ranked gene list input by the user from the human-computer interaction module, identify a group of genes with high priority and strong correlation with each other, and connect them according to the priority scores; The significance of the identified pathway intersection network was evaluated under random conditions by using a degree-preserving permutation test. Multiple permutation tests were performed on the sorted gene list while keeping the degree constant, so that the expected intersection and the intersection to be evaluated had the same / similar number of genes, and the expected intersection was used as the null distribution to estimate the significance of the identified intersection. Identify target genes in the prioritized gene list for immune traits, perform gene clustering and visualization, integrate the generated graph with druggability pocket data, and estimate the probability of each hexagon containing a druggable gene; Perform enrichment analysis on each gene cluster obtained by clustering, and then identify the enriched approved drug indications and KEGG pathways, so as to perform functional annotation on the gene cluster; The identified pathway intersection results are returned to the human-computer interaction module in the form of a table of intersection genes and subnetworks.

6. A therapeutic target prediction engine system based on omics genetic evidence according to claim 3, characterized in that: The drug repurposing analysis sub-engine based on the pathway intersection network has the following specific workflow: Obtain pathway intersection network results through a target prediction sub-engine based on pathway intersection networks, and receive specified information, including: whether to include SNPs in linkage disequilibrium, and if so, based on which population; how to identify core genes based on genomic proximity, gene expression protein abundance, and genome-wide promoter interaction capture technology; network information used to identify core genes and peripheral genes; confidence of the gene / protein interaction knowledge sub-library used; select the number of genes included in the pathway intersection network and significance information; select the number of genes in the intersection network and the significance information of the identified sub-network; generate pathway intersection networks for high-priority genes, and complete drug repurposing analysis of pathway intersection network genes; The identified drug reuse analysis results are returned to the human-computer interaction module in the form of charts.

7. The therapeutic target prediction engine system based on omics genetic evidence according to claim 3, characterized in that: The target prediction engine based on cross-disease genetic pleiotropy has the following specific workflow: Multi-trait score MRS was used to quantify the extent to which the target gene was highly scored in multiple traits; In the formula, nTop is the number of times the target gene ranks before the set rank in a certain trait; mRank is the average rank of the target gene; N is the total number of traits included in the analysis; Using Fisher's exact test, genes enriched in the intersection were compared with all genes with MRS as the test background to identify individual KEGG pathways; Examining the enrichment of pathway intersection network genes at the level of mouse immune-mediated disease phenotypes, drug-interferable gene traits, and immune disease indications by stage and approved therapeutic regimens and druggable pockets; Introducing the Multi-Trait Novelty Score (MNS) to quantify how unexplored a target is in most traits: In the formula, nD represents the number of times the target gene is annotated as a disease gene in a certain trait; nA represents the number of approved drugs for the target gene in these traits; nP represents the number of drugs that the target gene has at least one in the drug development stage in these traits; N represents the total number of traits included in the analysis; The identified target prediction results based on cross-disease genetic pleiotropy drivers will be returned to the human-computer interaction module in the form of charts.

8. The therapeutic target prediction engine system based on omics genetic evidence according to claim 3, characterized in that: The target prediction sub-engine based on generative AI has the following specific workflow: Based on user-specified information, including the specific type of input characteristic traits and their correlation indicators; selection of multi-omics data types for training the generative model; determination of the evaluation criteria for the druggability of the target gene; and the credibility level of the protein interaction knowledge base used; Set the screening threshold and significance criteria for target genes; select the number of target genes and related biological pathway information in the output results; The generative AI model generates a priority list of candidate targets and predicts the potential biological functions of these targets by learning a multimodal input feature matrix constructed by combining genomic data, network evidence, and ontology annotations at the gene level, and deeply integrating multimodal features. It also identifies target genes that show high clinical application potential in multiple disease contexts by comprehensively analyzing a variety of bioinformatics data. It reveals disease-specific targets for target genes that only appear in specific diseases. It also deeply analyzes the drugability of target genes by combining molecular dynamics simulations. The identified target prediction results are fed back to the human-computer interaction module in the form of visual charts.

9. The therapeutic target prediction engine system based on omics genetic evidence according to claim 1, characterized in that: The target search main engine specifically includes: The target search sub-engine for immune-mediated diseases receives the user's target search instruction from the human-computer interaction module, and obtains the target priority ranking results of immune-mediated diseases supported by the engine under the navigation tag; The target search sub-engine based on the pathway intersection network receives user-specified input information from the human-computer interaction module, and searches and analyzes the targets predicted by multiple traits based on the results of the target prediction sub-engine based on the pathway intersection network; The target retrieval sub-engine based on generative AI receives user input from the human-computer interaction module and automatically identifies and predicts potential drug targets based on the generative model.

10. A therapeutic target prediction engine system based on omics genetic evidence according to claim 9, characterized in that: The immune-mediated disease target search sub-engine generates target gene ranking results for each disease in two modes, including a complete target gene priority list and a selectable number of pathway intersection network gene lists, namely: a discovery mode that does not use any known drug target information and a supervised mode that integrates known drug target prediction factors through a random forest algorithm; provides a downloadable MySQL relational database and detailed documentation on table organization and usage; The target search sub-engine based on the pathway intersection network is used to receive user input and specified information from the human-computer interaction module, including different traits for comparison; genes to be evaluated, including all priority genes or all genes identified by the pathway intersection network; Choose whether to use the discovery mode or the supervised mode for analysis; select the metric used for prioritization; search and analyze the targets predicted for multiple traits based on the results of the target prediction sub-engine based on the pathway intersection network; The MRS scores of relevant targets in different traits are shown in the figure. Each target provides links to access specific information, including: hyperlinks to GeneCards, Ensembl, HGNC, OMIM, Entrez, and OpenTargets databases to access the target gene sequence, corresponding protein and tertiary structure, related pathways, and disease information; hyperlinks to PDBe-KB and AlphaFold databases to understand the drugability of the target; priority scores of targets in common immune traits; and heat map displays and hyperlinks to gene-related information; The target retrieval sub-engine based on generative AI, the target retrieval sub-engine based on generative AI is based on the generative AI large language model, and automatically generates answers according to the user's input information; The user fills in the engine's dialog box with the required sub-engine type, input data, and expected parameters through text input; The therapeutic target prediction engine system calls the target retrieval sub-engine based on generative AI to perform calculations and feeds the generated result links back to the dialog box.